Data compliance management method and system
The data compliance management system addresses inefficiencies in dataset compliance evaluation by enabling collaborative expert review and AI-driven analysis, ensuring reliable and objective compliance assessments and optimized dataset usage.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- LG MANAGEMENT DEV INST CO LTD
- Filing Date
- 2025-11-20
- Publication Date
- 2026-05-28
Smart Images

Figure KR2025019325_28052026_PF_FP_ABST
Abstract
Description
Data Compliance Management Methods and Systems
[0001] The present invention relates to a method and system for systematically evaluating the compliance of various datasets used for AI model training, managing risks according to user needs, and certifying and reporting compliance results through expert feedback.
[0002] As artificial intelligence technology advances, securing high-quality, large-scale datasets has become a key factor in determining the performance of AI models. These datasets often include vast amounts of data collected from various sources, such as web crawling and open-source purchases.
[0003] However, data involves complex legal issues such as copyright, privacy, and licensing, so the indiscriminate use of non-compliant datasets can lead to serious legal disputes and financial losses.
[0004] Traditionally, legal experts were limited to manually sampling portions of a dataset to review licenses or verifying the inclusion of specific license phrases through simple keyword matching. This approach was not only time-consuming and costly but also had limitations, such as a lack of consistency in results depending on the evaluator's subjectivity.
[0005] Furthermore, even when legal risks were discovered in a dataset, risks were managed in an inefficient manner by discarding the entire dataset to avoid legal disputes, resulting in a massive waste of data resources.
[0006] Therefore, there is an urgent need to invent a standardized collaboration interface that enables multiple legal experts or AI engineers to collaborate to systematically review the complex legal risks of datasets, derive agreed-upon conclusions, and manage the history.
[0007] The present invention was devised to solve the problems of the prior art described above, and aims to provide a data compliance management method and system that provides a systematic collaboration interface in which multiple legal experts mutually review their respective opinions to derive an agreed-upon conclusion.
[0008] In addition, the present invention aims to provide a data compliance management method and system that specifies the compliance level of a finally determined dataset and automatically generates a reliable report authenticated by a digital seal, etc.
[0009] However, the technical problems that the present invention and the embodiments of the present invention aim to solve are not limited to the technical problems described above, and other technical problems may exist.
[0010] A data compliance management method according to an embodiment of the present invention is a method performed on a computer, comprising: a step in which at least one processor analyzes the license compliance of at least one dataset connected to a plurality of entities using at least one artificial intelligence model; a step in which the at least one processor obtains a plurality of evaluation opinions from a plurality of real or virtual agents regarding the analyzed license compliance result; and a step in which the at least one processor controls the expression of a final compliance evaluation of the at least one dataset, which synthesizes the plurality of evaluation opinions, through at least one interface.
[0011] Additionally, the above entity includes a parent dataset used to generate the at least one dataset, or a child dataset derived from the at least one dataset and processed or redistributed.
[0012] In addition, the above-mentioned at least one artificial intelligence model is characterized by including at least one of a first model that searches for data related to the at least one dataset on a web or app, a second model that determines metadata by extracting dependency relationships between a plurality of entities connected to the at least one dataset based on the data searched for the at least one dataset, and a third model that scores risk based on the metadata for the at least one dataset.
[0013] Additionally, the step of analyzing license compliance of at least one dataset further includes the step of recursively performing a license compliance analysis process using the first to third models for lower-level entities derived from each entity among a plurality of entities detected from the at least one dataset.
[0014] In addition, the metadata includes the name, task category, modality, application field, and license type of each of the at least one dataset and the plurality of entities.
[0015] Additionally, the step of analyzing the license compliance includes the step of calculating an individual risk class for the at least one dataset and the step of calculating an integrated risk class for the at least one dataset based on the individual risk classes for a plurality of entities connected to the at least one dataset.
[0016] Additionally, the step of obtaining multiple evaluation opinions from the plurality of actual agents includes the step of expressing the license compliance result through the at least one interface, and the step of receiving an input signal including evaluation opinions input from the plurality of actual agents through the at least one interface.
[0017] Additionally, the step of obtaining multiple evaluation opinions from the plurality of virtual agents comprises: creating a plurality of virtual AI agents predefined to have different evaluation criteria; controlling each of the plurality of virtual AI agents to analyze license compliance for the at least one dataset; and obtaining evaluation opinions on the license compliance results derived for each of the plurality of virtual AI agents.
[0018] In addition, the step of generating the plurality of virtual AI agents is characterized by being trained based on at least one of different resources, model parameters, system prompts, fine-tuning data, a base model, and a database.
[0019] Additionally, the step of obtaining the plurality of evaluation opinions further includes the step of controlling the at least one artificial intelligence model to perform a discussion that cross-examines the evaluation opinions of the plurality of real or virtual agents based on the Delphi technique.
[0020] Additionally, the step of controlling the performance of the above discussion includes the step of correcting the evaluation result derived by the first virtual AI agent by cross-referencing it with at least one other virtual AI agent other than the first virtual AI agent, and the step of controlling the repetition of the correction until a predetermined threshold is satisfied.
[0021] Additionally, a data compliance management method according to an embodiment of the present invention further comprises the steps of: matching the final compliance evaluation with an identifier of the at least one dataset and storing it in at least one memory; and loading at least a portion of the data stored in the at least one memory to dynamically calibrate the internal parameters or weights of at least one of the plurality of virtual AI agents based on the final evaluation.
[0022] In addition, the final compliance evaluation is characterized by being provided as an evaluation report including a final risk class determined by the individual risk class and integrated risk class of at least one dataset, and a digital seal of one of the actual agents.
[0023] Meanwhile, a data compliance management system according to an embodiment of the present invention comprises at least one memory; and at least one processor that executes instructions stored in the at least one memory; wherein the at least one processor receives an input signal from a user through at least one or more means; wherein the input signal is a signal including at least one condition set by the user for at least one dataset to be a service target, and the processor adjusts parameters of a plurality of basic algorithms for evaluating compliance risk for the at least one dataset based on the input signal; and operates according to instructions to provide a user-customized compliance service for the at least one dataset based on an interactive algorithm to which the adjusted parameters are applied.
[0024] In addition, the above at least one is characterized by visualizing the at least one dataset and a plurality of entities connected to the at least one dataset in a hierarchical structure, and visualizing the depth of the plurality of entities connected to the at least one dataset.
[0025] In addition, at least one condition set by the user is characterized by being at least one of user environment information including at least one of the country to which the user belongs, an industry domain, and the purpose of use of the dataset; trend information including the latest information related to a plurality of risk assessment items tracked by crawling or database linkage by at least one artificial intelligence model; and information determining the target risk class of at least one dataset that is the target of the service.
[0026] In addition, the parameters of the above-mentioned plurality of basic algorithms are characterized by including weights or calculation formulas for calculating risk scores.
[0027] In addition, the user-customized compliance service further comprises calculating an integrated risk class for at least one dataset based on an individual risk class for at least one dataset that is the target of the service and individual risk classes for a plurality of entities connected to the at least one dataset, and is characterized in that the risk assessment result calculated by applying the plurality of basic algorithms and the risk assessment result calculated by applying the interactive algorithm are different from each other.
[0028] In addition, the user-customized compliance service further includes performing trimming to delete at least one data record derived from a first entity among the plurality of entities from the at least one dataset, and the integrated risk class of the trimmed at least one dataset satisfies the target risk class.
[0029] In addition, the first entity is characterized by being at least one of an entity in which an individual risk class is less than a preset threshold and an entity in which the depth of the data hierarchy exceeds a preset depth value.
[0030] The data compliance management method and system according to an embodiment of the present invention provides a systematic collaboration interface in which multiple legal experts mutually review their respective opinions to derive an agreed-upon conclusion, thereby excluding the subjectivity of a specific evaluator and having the effect of deriving more objective and reliable compliance evaluation results.
[0031] In addition, the data compliance management method and system according to an embodiment of the present invention has the effect of clearly proving the safety and reliability of the data to the dataset user or consumer by specifying the compliance level of the finally determined dataset and automatically generating a reliable report certified by a digital seal, etc.
[0032] However, the effects obtainable from the present invention are not limited to those mentioned above, and other unmentioned effects can be clearly understood from the description below.
[0033] FIG. 1 illustrates an example of a block diagram of a computing system implementing a data compliance service according to an embodiment of the present invention.
[0034] FIG. 2 illustrates an example of a block diagram of a computing device implementing a data compliance service according to an embodiment of the present invention.
[0035] FIG. 3 illustrates an example of a block diagram in another aspect of a computing device implementing a data compliance service according to one embodiment of the present invention.
[0036] FIG. 4 is a flowchart illustrating a method for providing a data compliance service according to an embodiment of the present invention.
[0037] FIG. 5 is an example of a drawing for explaining a tree-shaped data structure according to an embodiment of the present invention.
[0038] FIG. 6 is an example of a drawing illustrating the role of an artificial intelligence model utilized in a life cycle tracking process according to an embodiment of the present invention.
[0039] FIG. 7 is an example of an entity list displaying metadata of a first dataset according to an embodiment of the present invention.
[0040] FIGS. 8 and 9 are examples of schematic drawings illustrating the detection of a reversal phenomenon using a risk score calculated according to an embodiment of the present invention.
[0041] FIG. 10 is an example of a dynamic interface according to an embodiment of the present invention.
[0042] FIG. 11 is a flowchart illustrating a process for performing risk assessment through a plurality of virtual AI agents according to an embodiment of the present invention.
[0043] FIG. 12 is an example of an evaluation report according to an embodiment of the present invention.
[0044] FIG. 13 is a flowchart illustrating a dynamic risk assessment method for a dataset according to an embodiment of the present invention.
[0045] FIG. 14 is an example of a drawing for explaining a trimming function according to an embodiment of the present invention.
[0046] The present invention is capable of various modifications and may have various embodiments; therefore, specific embodiments are illustrated in the drawings and described in detail in the detailed description. The effects and features of the present invention, and the methods for achieving them, will become clear by referring to the embodiments described in detail below together with the drawings. However, the present invention is not limited to the embodiments disclosed below but can be implemented in various forms. In the following embodiments, terms such as "first," "second," etc., are used not in a limiting sense but for the purpose of distinguishing one component from another. Furthermore, singular expressions include plural expressions unless the context clearly indicates otherwise. Also, terms such as "include" or "have" mean that the features or components described in the specification exist, and do not preclude the possibility that one or more other features or components may be added. Additionally, in the drawings, the size of components may be exaggerated or reduced for convenience of explanation. For example, the size and thickness of each component shown in the drawings are arbitrarily depicted for convenience of explanation, so the present invention is not necessarily limited to what is illustrated.
[0047] Hereinafter, embodiments of the present invention will be described in detail with reference to the attached drawings. When describing with reference to the drawings, identical or corresponding components are given the same reference numerals, and redundant descriptions thereof will be omitted.
[0048]
[0049] FIG. 1 illustrates an example of a block diagram of a computing system performing a data compliance service according to an embodiment of the present invention.
[0050] Referring to FIG. 1, a computing system (1000) for performing a data compliance service according to one embodiment of the present invention includes a user computing device (110), a training computing system (150), and a server computing system (130), and each device and system is connected to communicate through a network (170).
[0051] According to an embodiment of the present invention, 1) a user computing device (110) can perform data compliance services by using a local or / and external machine learning model (120) or by using a machine learning model (140) provided by a server.
[0052] In addition, according to another embodiment of the present invention, 2) a server computing system (130) communicating with a user computing device (110) may provide a data compliance service to the user computing device (110) on an application or / and the web in response to a request from a user through the user computing device (110).
[0053] In addition, according to another embodiment of the present invention, 3) a user computing device (110) and a server computing system (130) may provide data compliance services to a user by performing at least a part of the method of providing data compliance services in conjunction with each other.
[0054] Additionally, according to various embodiments of the present invention, a user computing device (110) and / or a server computing system (130) may learn a machine learning model (120 / 140) that is performed in a method of providing data compliance services through interaction with a training computing system (150) that is communicatedly connected via a network (170). In this case, the training computing system (150) may be separate from the server computing system (130) or may be part of the server computing system (130).
[0055] In some embodiments, the training computing system (150) may be part of the server computing system (130) or part of the user computing device (110).
[0056] In the following description, the user computing device (110) is connected to a server computing system (130) to execute a data compliance service, and the server computing system (130) provides the data compliance service either directly or by using a language model from another server.
[0057] However, it can be understood that cases where part of the process described as being performed in a server computing system (130) is performed in a user computing device (110) are naturally included in the description of the present invention.
[0058] - User Computing Device (110: User Computing Device)
[0059] The user computing device (110) may include all other types of computing devices, such as a smartphone, a mobile phone, a digital broadcasting device, a PDA (personal digital assistants), a PMP (portable multimedia player), a desktop, a wearable device, an embedded computing device and / or a tablet PC.
[0060] Additionally, in the embodiment, the user computing device (110) may further include a predetermined server computing device that provides a data compliance service environment.
[0061] This user computing device (110) includes at least one processor (111) and memory (112).
[0062] Here, the processor (111) may be composed of at least one of a central processing unit (CPU), a graphics processing unit (GPU), ASICs (application specific integrated circuits), DSPs (digital signal processors), DSPDs (digital signal processing devices), PLDs (programmable logic devices), FPGAs (field programmable gate arrays), controllers, microcontrollers, microprocessors, and / or electrical units for performing other functions, or a plurality of electrically connected processors.
[0063] The memory (112) may include one or more non-transient / transient computer-readable storage media such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, and combinations thereof, and may include web storage of a server that performs memory storage functions on the internet. This memory (112) may store data and instructions necessary for the at least one processor (111) to perform the operation of an application for performing data compliance services.
[0064] In one embodiment, the user computing device (110) can perform various deep learning for data compliance services by linking with a deep-learning neural network.
[0065] Here, the deep learning neural network according to the embodiment may include a Convolutional Neural Network (CNN), R-CNN (Regions with CNN features), Fast R-CNN, Faster R-CNN, Mask R-CNN, etc., and may include any deep learning neural network that includes an algorithm capable of performing the embodiments described below, and the embodiments of the present invention do not limit or restrict such deep learning neural networks themselves.
[0066] At this time, according to the embodiment, the deep learning neural network may be installed directly in the server computing system (130) or may operate as a separate device from the server computing system (130) to perform deep learning for the data compliance service.
[0067] Additionally, in one embodiment, the user computing device (110) may store at least one machine learning model (120). For example, the user computing device (110) may be various machine learning models, such as a plurality of neural networks (e.g., deep neural networks) that perform data compliance services based on structured / quantitative data, or other types of machine learning models including non-linear models and / or linear models, and may be configured as a combination thereof.
[0068] For example, machine learning models may include linear regression, decision trees, random forests, gradient boosting pre-trained language models or / and deep learning models. And neural networks may include at least one of feed-forward neural networks, recurrent neural networks (e.g., long short-term memory recurrent neural networks), convolutional neural networks or / and other forms of neural networks.
[0069] Specifically, in the embodiment, the user computing device (110) can store an artificial intelligence agent that performs a life cycle tracking process and a risk optimization process.
[0070] In the embodiment, the artificial intelligence agent can perform a 'life cycle tracking process' that evaluates legal risks by finding multiple upper and lower datasets directly or indirectly connected by sharing a dataset and a predetermined entity, and a 'risk optimization process' that optimizes the dataset with evaluated legal risks into a dataset of a risk class desired by the user by trimming it.
[0071] Such an artificial intelligence agent can perform data compliance services based on at least one machine learning model. In an embodiment, the machine learning model may include a navigation model, a QA (Question-Answer) model, and / or a scoring model.
[0072] Additionally, the user computing device (110) may store a model to be used in each process and a prompt template that serves as the basis for input to the model in order to perform at least part of the process for data compliance services through a large-scale language model (LLM) and / or a large-scale behavior model (LAM).
[0073] For example, the user computing device (110) may store 1) a prompt for generating a query from user input, 2) a prompt for analyzing a domain (or web page), 3) a prompt for generating a trajectory, etc.
[0074] That is, in one embodiment, the user computing device (110) can perform a data compliance service based on received data by requesting the execution of some execution steps in a method of providing a data compliance service through a prompt or the like to a language model of an external server.
[0075] In another embodiment, regarding the method of providing a data compliance service requested through a user computing device (110), the server computing system (130) may perform the data compliance service through at least one machine learning model (140) and a machine learning model of another server to provide data to the user computing device (110).
[0076] Such a user computing device (110) may include at least one input component (121) that detects user input. Specifically, the input component (121) may include a sensor system including an image sensor, a position sensor (IMU), an audio sensor, a distance sensor, a proximity sensor, a contact sensor, etc.
[0077] For example, the user input component (121) may include a touch sensor (e.g., a touch screen or / and a touch pad, etc.) that detects a touch of the user's input medium (e.g., a finger or a stylus), an image sensor that detects the user's motion input, a microphone that detects the user's voice input, a button, a mouse and / or a keyboard, etc.
[0078] Here, the image sensor may include an image processing module. Specifically, the image sensor may process still images or video obtained by an image sensor device (e.g., CMOS or CCD).
[0079] In addition, the image sensor can process a still image or video acquired through the image sensor device using an image recognition process (e.g., OCR, etc.) and / or an image processing module to extract necessary information and transmit the extracted information to a processor.
[0080] Additionally, the input component (121) can receive input from an external controller (e.g., mouse, keyboard, etc.) based on an interface module, and in this case, may include an external output device (e.g., speaker).
[0081] At this time, the interface module may be configured to include at least one of a wired / wireless headset port, an external charger port, a wired / wireless data port, a memory card port, a port for connecting a device equipped with an identification module, an audio I / O (Input / Output) port, a video I / O (Input / Output) port, an earphone port, a power amplifier, an RF circuit, a transceiver, and other communication circuits.
[0082] In addition, the external output device may include a display system that outputs various information related to data compliance services as a graphic image.
[0083] Such a display system may be implemented by including at least one of a liquid crystal display (LCD), a thin film transistor-liquid crystal display (TFT LCD), an organic light-emitting diode (OLED), a flexible display, a 3D display, and an e-ink display.
[0084] Meanwhile, the user computing device (110) including the above-described components may further perform at least some of the functional operations performed by the server computing system (130) described later.
[0085] -Server Computing System (130: Server Computing System)
[0086] The server computing system (130) can perform a series of processes to provide data compliance services.
[0087] In detail, in an embodiment, the server computing system (130) can provide the data compliance service by exchanging data necessary to enable the data compliance service process to run on an external device, such as a user computing device (110), with said external device.
[0088] More specifically, in an embodiment, the server computing system (130) can provide an environment in which an application can run on a user computing device (110).
[0089] To this end, the server computing system (130) may include an application program, data and / or instructions, etc. for the application to operate, and may transmit and receive various data based thereon with the external device.
[0090] Additionally, the server computing system (130) includes at least one processor (131) and memory (132). Here, the processor (131) may be composed of at least one or a plurality of electrically connected processors among a central processing unit (CPU), a graphics processing unit (GPU), ASICs (application specific integrated circuits), DSPs (digital signal processors), DSPDs (digital signal processing devices), PLDs (programmable logic devices), FPGAs (field programmable gate arrays), controllers, microcontrollers, microprocessors, and / or other electrical units for performing functions.
[0091] And the memory (132) may include one or more non-transient / transient computer-readable storage media such as RAM, ROM, EEPROM, EPROM, flash memory device, magnetic disk, etc. and combinations thereof. This memory (132) may store data and instructions required for prompt templates, machine learning models (140), etc., for the processor (131) to perform tasks through the language model of the server computing system (130) or / and the language model of an external server.
[0092] For example, a server computing system (130) may include a neural network or / and other multi-layer non-linear models as a machine learning model (140). Exemplary neural networks may include a feed-forward neural network, a deep neural network, a recurrent neural network, and a convolutional neural network.
[0093] In one embodiment, the server computing system (130) may be implemented to include at least one computing device. For example, the server computing system (130) may be implemented to operate a plurality of computing devices according to a sequential computing architecture, a parallel computing architecture, or a combination thereof. Additionally, the server computing system (130) may include a plurality of computing devices connected via a network.
[0094] In an embodiment, the server computing system (130) may further include a data store computing system (1000) (hereinafter, data store) which is a storage for continuously storing and managing raw data (e.g., log data, demo data, etc.) that forms the basis of a method (service) for providing data compliance services. This data store may include various forms of data storage, ranging from file systems to cloud storage.
[0095] For example, a data store may include at least one database among a relational database that uses a structured query language (SQL) to define and manipulate data, a NoSQL database designed for flexibility and scalability to process unstructured and semi-structured data, a data warehouse optimized for querying and analysis by centralizing large volumes of data from multiple sources as a system used for reporting and data analysis, a data warehouse that stores large volumes of raw data in basic formats such as structured data, semi-structured data, and unstructured data, and a local storage device or Network Attached Storage (NAS) that stores data in files in a format generally accessible by a computer operating system.
[0096] - Training Computing System (150: Training Computing System)
[0097] A training computing system (150) includes at least one processor (151) and a memory (152). Here, the processor (151) may be composed of at least one or a plurality of electrically connected processors, including a central processing unit (CPU), a graphics processing unit (GPU), ASICs (application specific integrated circuits), DSPs (digital signal processors), DSPDs (digital signal processing devices), PLDs (programmable logic devices), FPGAs (field programmable gate arrays), controllers, microcontrollers, microprocessors, and / or other electrical units for performing functions. The memory (152) may include one or more non-transient / transient computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, and combinations thereof. This memory (152) may store data and instructions necessary for the processor (151) to train a machine learning model.
[0098] For example, the training computing system (150) may include a model trainer (160) that trains a machine learning model stored in a user computing device (110) and / or a server computing system (130) using various training or learning techniques, such as back propagation of error.
[0099] For example, the model trainer (160) can perform backpropagation updates to one or more parameters of a machine learning model for data compliance services based on a defined loss function.
[0100] In some embodiments, performing backpropagation of the error may include performing truncated backpropagation through time. The model trainer (160) may perform a number of generalization techniques (e.g., weight decrement, dropout, knowledge distillation, etc.) to improve the generalization ability of the machine learning model being trained.
[0101] And the model trainer (160) includes computer logic utilized to provide the desired function. The model trainer (160) may be implemented as hardware, firmware and / or software that controls a general-purpose processor. For example, in one embodiment, the model trainer (160) includes a program file stored in a storage device, loaded into memory, and executed by one or more processors. In another embodiment, the model trainer (160) includes one or more sets of computer-executable instructions stored in a tangible computer-readable storage medium, such as a RAM hard disk or an optical or magnetic medium.
[0102] Networks (170) include, but are not limited to, 3GPP (3rd Generation Partnership Project) networks, LTE (Long Term Evolution) networks, WIMAX (World Interoperability for Microwave Access) networks, Internet, LAN (Local Area Network), Wireless LAN (Wireless Local Area Network), WAN (Wide Area Network), PAN (Personal Area Network), Bluetooth networks, satellite broadcasting networks, analog broadcasting networks and / or DMB (Digital Multimedia Broadcasting) networks.
[0103] Generally, communication through the network (170) can be performed using any type of wired and / or wireless connection through various communication protocols (e.g., TCP / IP, HTTP, SMTP, FTP), encodings or formats (e.g., HTML, XML), and / or protection schemes (e.g., VPN, Secure HTTP, SSL).
[0104]
[0105] FIG. 2 illustrates an example of a block diagram of a computing device, which is one of the components of a computing system (1000) that performs a data compliance service according to an embodiment of the present invention.
[0106] Referring to FIG. 2, the computing device (100) included in the user computing device (110), server computing system (130), and training computing system (150) includes a plurality of applications (e.g., applications 1 to N). Each application may include a machine learning library.
[0107] For example, applications may include text messaging applications, virtual keyboard applications, browser applications, chatbot applications, etc.
[0108] In an embodiment, the computing device (100) may include a model trainer (160) for training a machine learning model, and may store and operate the machine learning model to perform data compliance services on input data.
[0109] Each application of the computing device (100) can communicate with a number of other components of the computing device, such as, for example, one or more sensors, a context manager, a device state component, and / or additional components. In one embodiment, each application can communicate with each device component using an API (e.g., a public API). In one embodiment, the API used by each application may be specific to that application.
[0110]
[0111] FIG. 3 illustrates an example of a block diagram in another aspect of a computing device, which is one of the components of a computing system (1000) that performs a data compliance service according to an embodiment of the present invention.
[0112] Referring to FIG. 3, the computing device (200) includes a plurality of applications (e.g., Application 1 to Application N). Each application can communicate with a central intelligence layer. For example, the applications may include a virtual keyboard application, a browser application, a chatbot application, etc. In one embodiment, each application can communicate with the central intelligence layer (and a model stored therein) using an API (e.g., a common API across all applications).
[0113] And the central intelligence layer may include prompts using a plurality of machine learning models or / and language models. For example, as illustrated in FIG. 3, each machine learning model and at least some thereof may be provided for each application and managed by the central intelligence layer. In another embodiment, two or more applications may share a single machine learning model. For example, in some embodiment, the central intelligence layer may provide a single model for all applications. In some embodiment, the central intelligence layer may be included within the operating system of the computing device (200) or otherwise implemented.
[0114] The central intelligence layer can communicate with the central device data layer. The central device data layer may be a centralized data store for the computing device (200). As illustrated in FIG. 3, the central device data layer can communicate with a number of other components of the computing device, such as, for example, one or more sensors, a context manager, a device state component, and / or additional components. In some embodiments, the central device data layer can communicate with each device component using an API (e.g., a private API).
[0115] The technology described herein may refer to servers, databases, software applications, and other computer-based systems, as well as information transmitted to or from said systems. It will be recognized that the inherent flexibility of computer-based systems allows for a wide range of possible configurations, combinations, division of tasks, and functionality between and from components. For example, the processes described herein may be implemented using a single device or component or multiple devices or components operating in combination. Databases and applications may be implemented in a single system or in a distributed system across multiple systems. Distributed components may operate sequentially or in parallel.
[0116] - How to provide data compliance services
[0117] Hereinafter, a method for implementing a data compliance service according to one embodiment of the present invention, in which a computing system (1000) systematically evaluates the compliance of various datasets used for AI model training, manages risks according to user needs, and certifies and reports compliance results through expert feedback, is described in detail.
[0118] In the embodiment, the data compliance service may include a 'life cycle tracking process' that identifies multiple upper and lower datasets directly or indirectly connected by sharing a dataset and a predetermined entity to evaluate legal risk, a 'risk optimization process' that trims the dataset with evaluated legal risk to optimize it into a dataset of a risk class desired by the user, and a 'dynamic risk assessment process' that dynamically adjusts the structure of the dataset with evaluated legal risk.
[0119] A data compliance service of a computing system according to one embodiment of the present invention can be performed and provided by an artificial intelligence agent stored in a user computing device (110).
[0120] In an embodiment, the artificial intelligence agent (AA) can perform a lifecycle tracking process that analyzes a given dataset, tracks all nodes connected to the dataset, searches for and records license documents, and evaluates the legal risk that the dataset may have.
[0121] Such an artificial intelligence agent (AA) can perform a life cycle tracking process based on a predetermined machine learning model and / or module.
[0122] Specifically, in the embodiment, the artificial intelligence agent (AA) can perform a lifecycle tracking process based on a navigation model, a QA (Question-Answer) model and / or a scoring model.
[0123] Hereinafter, a method for providing a data compliance service according to one embodiment of the present invention will be described in more detail with reference to the attached drawings.
[0124] FIG. 4 is a flowchart illustrating a method for providing a data compliance service according to an embodiment of the present invention.
[0125] Referring to FIG. 4, a method for providing a data compliance service according to an embodiment of the present invention may include a first dataset acquisition step (S101), a step of identifying a dependent node connected to the acquired first dataset (S103), a step of calculating an individual risk score for the first dataset (S105), a step of calculating an integrated risk score including the dependent node connected to the first dataset (S107), a step of determining an irregular node in which a reversal phenomenon has occurred based on the individual and integrated risk scores (S109), a step of analyzing the dependency of the first dataset for the determined irregular node (S111), and a step of providing a license compliance evaluation result for the first dataset according to the analyzed dependency (S113).
[0126] Specifically, a computing system according to one embodiment of the present invention can acquire a first dataset. (S101)
[0127] In the embodiment, the first dataset refers to a set of data that is the subject of analysis for reviewing license compliance.
[0128] In the embodiment, an entity for the first dataset refers to various objects utilized in the creation of the first dataset. Additionally, detailed information regarding the entity may be matched as metadata for the first dataset.
[0129] Specifically, an entity for the first dataset may include a link containing license-related information for the first dataset, a domain, software / API / tool for operating the first dataset, a source of data constituting the first dataset (e.g., medical images), the type and version of an artificial intelligence model used for labeling that is learned through the first dataset or performed for said learning, a framework or pipeline for learning a machine learning model through the first dataset, and a service platform based on at least one machine learning model including these.
[0130] More specifically, an entity for the first dataset may include a data set utilized prior to the creation of the first dataset (hereinafter referred to as a 'sub dataset' for distinction) and / or a data set utilized for training, services, etc., after the creation of the first dataset (i.e., a sub-dataset). For example, the sub dataset includes the type and version of the artificial intelligence model used for labeling the first dataset, and the sub-dataset includes a service platform provided based on the first dataset.
[0131] Accordingly, in the embodiment, the metadata matched and stored in the first dataset may be 1) data required for web navigation to detect entities for the first dataset and 2) data including detailed information about entities detected for the first dataset.
[0132] Specifically, 1) data required for web navigation to detect entities for the first dataset (hereinafter, the first metadata) may be generated when information about the entity is not found on a certain open source web page and / or community that shares metadata about the entity.
[0133] To search for and determine this first metadata, in the embodiment, the computing system (1000) may perform RAG, internet searching, and / or crawling.
[0134] For example, the first metadata may include the entity name, record name, API, artificial intelligence model name, etc., to be navigated.
[0135] Additionally, 2) data containing detailed information about an entity detected for the first dataset (hereinafter, second metadata) may be generated when information about the entity is retrieved on a certain open source web page and / or community that shares metadata about the entity.
[0136] This second metadata may be determined when detecting detailed information about the entity, or when detecting the result of the scoring process described later.
[0137] For example, the second metadata may include a link where the entity content is disclosed (e.g., a web page that registered the entity), the type of the entity, task category, modality, field of application, risk score, and / or license proof information.
[0138] In the embodiments, the license proof information refers to information proving compliance with or infringement of the license based on the full text of the license for an entity and all usage rights, permissions, regulations, conditions, restrictions, obligations, and / or risk factors derived from the full text. Such license proof information may exist as license terms on a designated website and / or accompanying attachments.
[0139] In addition, license proof information may be defined as a specific license type (e.g., MIT, Apache, etc.). Additionally, it may include a detailed explanation of what regulation the above license type refers to.
[0140] That is, the above license proof information can be detected as an entity and matched and stored as second metadata for the first dataset.
[0141] The first dataset mentioned above is a collection of data that encompasses any data existing on the Internet, and may include, for example, various types of content such as text, images, and videos, software, and / or data crawled from websites.
[0142] In the embodiments, the dataset may not be a simple set of files, but may have a complex tree-shaped data structure in which multiple sub-datasets (nodes) are connected in layers. Hereinafter, 'dataset' may be interchangeably referred to as 'node'.
[0143] FIG. 5 is an example of a drawing for explaining a tree-shaped data structure according to an embodiment of the present invention.
[0144] Specifically, FIG. 5 is a diagram representing a connection graph in which each dataset is viewed as a node and the entities used in the corresponding node (e.g., dataset, software, content, tool, API, AI model, etc.) are set as “Dependency Nodes.” Here, a connected dependency node refers to a node that uses, references, or recreates the first dataset.
[0145] More specifically, the dataset according to the embodiment may have a tree-shaped data structure that starts from a source node (300) and extends to a plurality of sub-nodes that are processed or redistributed n times.
[0146] At this time, the source node (300) is the top-level node and represents the original dataset before redistribution.
[0147] The sub-nodes (310, 320) extending from the source node (300) represent datasets in which the original dataset has been processed or redistributed.
[0148] For convenience of explanation, the sub-node (310) detected in the first stage is referred to as the 'first entity', and the sub-node (320) derived from the first entity and detected in the second stage is referred to as the 'second entity'. In an embodiment of the present invention, the entity may become the nth entity depending on the number of processing and redistributions, and the value of n is not limited. In addition, the value of n may be the processing depth and / or number of layers of the corresponding dataset. Furthermore, the first to nth entities are concepts included in the sub-datasets for the first dataset.
[0149] That is, in the embodiment, the first dataset may be one of the datasets connected to at least one parent node and / or child node (hereinafter, dependent node). The first dataset and dependent node having a tree-shaped data structure may be displayed as visual materials in the license compliance review results provided according to the embodiment of the present invention.
[0150] Additionally, in the embodiment, the computing system (1000) can identify a dependent node connected to the acquired first dataset. (S103)
[0151] In detail, in an embodiment, the computing system (1000) recognizes entities (e.g., datasets, models, software, etc.) detected for a first dataset through web navigation and organizes them into a predetermined data structure (e.g., tree shape and / or graph shape), and can also identify dependent nodes connected to the first dataset by recursively expanding all discovered datasets to finally identify all nodes.
[0152] More specifically, in an embodiment, the computing system (1000) can track at least one dependent node connected to the first dataset based on extensive scalability regarding changes in rights relationships that occur during the process of redistributing and combining datasets.
[0153] To this end, in an embodiment, the computing system (1000) can detect dependent nodes connected to the first dataset (i.e., all nodes located above and / or below the first dataset) based on at least one of a navigation model, a QA model, and / or a scoring model. The navigation model, QA model, and / or scoring model may refer to artificial intelligence models pre-trained to perform a predetermined task.
[0154] FIG. 6 is an example of a drawing illustrating the role of an artificial intelligence model utilized in a life cycle tracking process according to an embodiment of the present invention.
[0155] The navigation model (M1) may be a model that recognizes certain features, including names, types, logos, images, text, etc., for objects (entities) such as datasets, software, and content, and searches for information related to a license (hereinafter, license proof information).
[0156] In the embodiment, the navigation model (M1) can be fine-tuned by repeatedly training in the process of recognizing various types of entities, setting a path for web document navigation for the recognized entities, and analyzing and navigating text included in web documents according to the set path.
[0157] For example, a navigation model (M1) can be pre-trained to search for and provide websites displaying license proof information of a dataset when a dataset is input.
[0158] Accordingly, in the embodiment, the computing system (1000) drives a navigation model (M1) to recognize at least one feature (text, image, etc.) of the first dataset and searches for at least one link related to the first dataset through web navigation, thereby minimizing the search for unnecessary data related to the first dataset and prioritizing the search for only information where license proof information is likely to exist.
[0159] That is, in the embodiment, the computing system (1000) can accurately search for and collect various documents containing license proof information of the corresponding dataset for the first dataset and various upper and lower datasets related thereto (hereinafter collectively referred to as lower datasets) based on the navigation model (M1).
[0160] In another embodiment, the navigation model (M1) may be a model based on a Large-Scale Behavioral Model (LAM) that automatically navigates web pages. Simply put, it is a model that automates the process of verifying information by clicking links one by one through a search engine, and can perform large-scale crawling. In this case, in another embodiment, the navigation model (M1) may be trained to perform web document navigation and text analysis processes using artificially generated synthetic data, even for web pages that are pre-configured to be uncrawlable.
[0161] The QA model (M2) may be a model that analyzes documents related to the first dataset and sub-datasets collected by the navigation model (M1) in a Question-Answer (QA) format to extract license proof information for each dataset and its location, or dependencies between datasets.
[0162] In the embodiment, the QA model (M2) can be fine-tuned using a QA task as training data to extract license proof information and dependency relationships between datasets from documents related to the first dataset and sub-datasets. For example, the QA model (M2) can be pre-trained to output the answer “CC-BY-4.0” to the question “What license does this dataset follow?”
[0163] Accordingly, in the embodiment, the computing system (1000) can flexibly recognize different license phrases, dependency notation methods, license terms and protocols, etc. for the first dataset based on the QA model (M2), and accurately identify license proof information and dependency relationships.
[0164] That is, in the embodiment, the computing system (1000) can extract license proof information and dependency relationships for the first dataset by organizing dependency information and license documents in a structured form based on the QA model (M2), even if licenses, terms and conditions, etc. for the first dataset are scattered throughout the documents.
[0165] To summarize, in the embodiment, the computing system (1000) can identify license proof information and dependency relationships for each dataset for the first dataset and sub-datasets based on the navigation model (M1) and the QA model (M2), and store the identified information by matching it as metadata for the nth entity of each dataset. Accordingly, in the embodiment, the computing system (1000) can define the attributes of a node and identify and display the location within the data layer where the node exists by matching metadata to each dataset for the first dataset and several sub-datasets related thereto, based on the navigation model (M1) and the QA model (M2).
[0166] Additionally, according to another embodiment of the present invention, the first dataset and a plurality of entities therewith may be distributed and stored in a plurality of physically or logically different multi-cloud environments or on-premise environments.
[0167] For example, the original data of the first dataset may be stored in a first private cloud with enhanced security, the first entity processed from the first dataset may exist in a second public cloud, and the second entity, which is an AI model utilizing it, may exist distributed in a third cloud.
[0168] In this case, the computing system (1000) according to an embodiment of the present invention can identify distributed entities through an artificial intelligence agent and / or cloud connector installed in each cloud environment. Specifically, to prevent direct large-scale data transmission between different clouds, the computing system (1000) can perform federated learning and / or federated analysis methods by receiving only license proof information and metadata primarily extracted by the agent within each cloud instead of replicating and retrieving the entire dataset, and calculating an integrated risk score in the central computing system (1000).
[0169] Through this, the present invention provides the effect of seamlessly tracking the entire lifecycle of a dataset and managing compliance in an integrated manner, even in complex enterprise environments where the physical locations of data are dispersed and data movement between storage locations is restricted.
[0170] In various embodiments, a computing system (1000) using a navigation model (M1) and a QA model (M2) can detect dependency nodes for a first dataset based on a method of detecting entities requiring license compliance verification linked to a first dataset and determining only the information containing license proof information among them as metadata, or b) detecting a sub-dataset (i.e., a second entity) connected to an entity determined as metadata for the first dataset n times.
[0171] That is, the computing system (1000) can search for multiple entities utilizing the first dataset on the web based on the above methods, processing and redistributing data, and detect the searched entities to match and store them as metadata.
[0172] The scoring model (M3) may be a model that evaluates the legal risks that each dataset may possess by scoring them according to license evaluation criteria by integrating and reflecting various metadata (e.g., terms of use, entity type, etc.) including license proof information and dependency relationships matched to the entity.
[0173] In the embodiment, the scoring model (M3) can be fine-tuned to assign a risk class or calculate a detailed item score of a dataset when it receives a dataset (e.g., a first dataset, a first entity, a second entity, etc.). To this end, a high-quality license evaluation dataset, in which items requiring complex interpretation are pre-labeled by an expert, can be used as training data. For example, the scoring model (M3) can be pre-trained to output grades such as “A-1” or “C-2” when a dataset is input.
[0174] The scoring model (M3) can output a risk score for the input dataset according to a pre-established prompt, rule, and / or pipeline when it receives a predetermined dataset according to the scoring process described below.
[0175] Accordingly, in the embodiment, the computing system (1000) performs an objective evaluation of the license of the first dataset based on the scoring model (M3) and specific grounds (in the embodiment, license proof information), thereby increasing the accuracy of the legal risk evaluation by applying clear and consistent criteria.
[0176] That is, in the embodiment, the computing system (1000) can evaluate the legal risk of the first dataset by reflecting the legal risk of all dependent nodes (i.e., sub-datasets) affecting the first dataset based on the scoring model (M3).
[0177] In other words, in the embodiment, the computing system (1000) can detect all dependent nodes connected to the first dataset based on at least one model, and search for and map metadata including license proof information for each detected dependent node.
[0178] To summarize, in the embodiment, the computing system (1000) can determine the location of the first dataset by identifying all upper and lower nodes connected to the first dataset based on at least one of the navigation model (M1), the QA model (M2), and / or the scoring model (M3), and map metadata including detected license proof information for all nodes including the first dataset.
[0179] In an embodiment, the computing system (1000) may provide an entity list displaying metadata for a first dataset and / or a sub-dataset.
[0180] FIG. 7 is an example of an entity list displaying metadata of a first dataset according to an embodiment of the present invention.
[0181] Referring to FIG. 7, in an embodiment, the computing system (1000) may provide an entity list (EL) that displays the name, number of sub-datasets, number of sub-dataset layers, task category, modality, application field, license type, and risk score for each of the plurality of datasets.
[0182] Referring again to FIG. 5, the number of lower dataset layers is a value that increases as the number of lower datasets increases, such as 0 when only the source node (300) exists and 1 when the first lower node (310) exists.
[0183] Task categories define the role the dataset performs, including, for example, text classification, question answering, translation, and object detection.
[0184] A modality is the type of data constituting the dataset, including, for example, images, text, numeric, code, and video. Additionally, a single dataset may have multiple modalities.
[0185] The application field specifies whether the dataset is applicable to all fields or, if used only in specific fields, in which fields; in the examples, it is classified into General and Specific.
[0186] The license type specifies the permission, regulation, and conditions of the license, and may be indicated as Unspecific if not defined as one, Custom if defined, or the defined type (e.g., MIT, etc.).
[0187] The risk score refers to the individual risk score and / or integrated risk score determined for each dataset according to the procedure described below.
[0188] In the embodiment, the computing system (1000) may generate and provide individual entity lists for each of the plurality of sub-datasets detected for the first dataset. For example, referring to the illustration, since the number of sub-datasets of Dataset 1 is 5, a total of 5 entities may be displayed in the entity lists of the sub-datasets for Dataset 1.
[0189] Identical content is omitted by applying the above description, and the source and entity type may be displayed in the entity list of the sub-dataset. The source refers to the number of entities derived again from the corresponding sub-dataset, similar to the concept of the second entity described above. In addition, the entity type refers to the entity type of the corresponding sub-dataset constituting the first dataset. As illustrated, for example, Sub-Dataset 1 is another dataset used to construct the above Dataset 1, and Sub-Dataset 5 is software used to construct the above Dataset 1.
[0190] In this way, in the embodiment, the computing system (1000) can detect metadata for the first dataset as well as metadata for all entities used in the creation of the first dataset based on at least one of the navigation model (M1), the QA model (M2), and / or the scoring model (M3), map and store the metadata with the dataset, and provide the contents by displaying them as an entity list.
[0191] Additionally, in the embodiment, the computing system (1000) can calculate individual risk scores for the first dataset. (S105)
[0192] Specifically, in an embodiment, the computing system (1000) can calculate individual risk scores for each dependent node based on at least one license evaluation criterion by reflecting metadata determined per entity.
[0193] The calculation of the individual risk score and integrated risk score (hereinafter, risk score) described below can be performed according to the scoring process of the scoring model (M3) described above.
[0194] Here, the Individual Risk Score according to the embodiment refers to a risk score calculated by scoring each license-related item according to a predetermined license evaluation criterion, reflecting only the license proof information (i.e., the license text and metadata set) of a single entity (in the embodiment, the first dataset).
[0195] In the embodiments, the license evaluation criteria for calculating the risk score follow the first to fourth major categories and the detailed items included in each major category.
[0196] The first major category is a category for risks related to data licenses, and may include detailed items such as 1-1) whether a license exists, 1-2) modification and derivative licenses, 1-3) possibility of disputes regarding output, 1-4) whether rights to the output are granted, and 1-5) obligation to notify of use.
[0197] The second major category is a category for usage period and usage area restrictions, and may include detailed items such as 2-1) data usage period restriction, 2-2) possibility of license revocation, 2-3) AI model service period, and 2-4) usage area restriction.
[0198] The third major category is a category for personal information and data security risks, and may include detailed items such as 3-1) whether personal information is included, 3-2) whether the data subject consents, 3-3) whether pseudonymized information is included, 3-4) possibility of provision to a third party and entrustment, and 3-5) limitation of the scope of users.
[0199] The fourth major category is a category for contract and dispute risks, and may include detailed items such as 4-1) collection process risks, 4-2) existing dispute cases, 4-3) other contract risks, and 4-4) license type classification.
[0200] In the embodiment, the artificial intelligence agent (AA) can generate prompts in advance based on a rule base and determine a score for each detailed item based on the prompts. That is, conditions are set for each subcategory, and it is determined whether the conditions are satisfied.
[0201] For example, regarding 1-1) the existence of a license, 5 points are assigned if it is available for unrestricted commercial use, 3 points if it is explicitly available for internal research purposes, 2 points if no license term is provided, and 1 point if it is explicitly unavailable. Here, rule-based conditions are set for each score. Taking the case of 5 points as an example, the detection of phrases such as 'copyleft' or 'no restriction' can be set as a condition.
[0202] In addition, in the embodiment, the artificial intelligence agent (AA) can pre-set weights for each detailed item (18 in the embodiment). In addition, a predetermined value can be assigned as a score for each detailed item.
[0203] At this time, the above-mentioned predetermined value may be a value assigned as one of the lowest to highest scores according to a pre-set condition, for example, or a value of 0 or 1 assigned according to Y / N.
[0204] In addition, in the embodiment, the artificial intelligence agent (AA) can calculate individual risk scores for the first dataset through Formula 1, which reflects the weight for each item in the value entered for each item.
[0205] [Formula 1]
[0206]
[0207] The above is a detail item of the corresponding entity It refers to the score for, and the above represents the weight for each detailed item. For example, the weight for each detailed item can be pre-matched and stored, such as 15% for “potential for output dispute” and 10% for “permission for modification and derivatives.”
[0208] That is, in the embodiment, the artificial intelligence agent (AA) can determine a score for each detailed item included in the major category and calculate an individual risk score by multiplying the determined score by the weight for each detailed item.
[0209] Additionally, in the embodiment, the artificial intelligence agent (AA) can map the risk class for the calculated individual (integrated) risk score based on the class table. At this time, in the embodiment, the artificial intelligence agent (AA) can also calculate and provide the corresponding risk class when calculating the individual (integrated) risk score. That is, in this embodiment, the individual (integrated) risk score is used to encompass the class corresponding to the individual (integrated) risk score.
[0210] Accordingly, in the embodiment, the artificial intelligence agent (AA) can identify an isolated legal risk level for the first dataset independently of other nodes by calculating individual risk scores of the first dataset.
[0211] Additionally, in the embodiment, the computing system (1000) can calculate an integrated risk score including a dependency node connected to the first dataset. (S107)
[0212] Specifically, in an embodiment, the computing system (1000) can calculate an integrated risk score for the first dataset by reflecting the individual risk scores of each dependent node connected to the first dataset.
[0213] Here, the Aggregate Risk Score according to the embodiment refers to a risk score calculated by considering not only the target entity (the first dataset in the embodiment) but also all parent and child nodes connected to said entity as a single entity, and reflecting the most dangerous item among all dependent nodes according to the license evaluation criteria described above.
[0214] To this end, in the embodiment, the artificial intelligence agent (AA) can calculate an individual risk score for each of all dependent nodes connected to the first dataset.
[0215] Additionally, in the embodiment, the artificial intelligence agent (AA) can extract the dependency node with the highest calculated risk for each detailed item through the calculated individual risk scores. In this embodiment, regarding risk, a higher risk score indicates greater risk, and a lower class indicates greater risk; however, for the sake of convenience of explanation below, it is assumed that a lower score indicates greater risk. That is, the dependency node with the highest calculated risk refers to the dependency node with the lowest calculated risk score.
[0216] Additionally, in the embodiment, the computing system (1000) can calculate an integrated risk score for the first dataset through Equation 2, which reflects the weighted sum of the dependency nodes from which the greatest risk is calculated for each detailed item.
[0217] [Equation 2]
[0218]
[0219] The above ...for each detail item, all nodes It refers to the lowest point selected based on the calculation of the greatest risk (lowest score). This can serve as a mechanism that ensures the risk of the lower nodes is transferred intact to the upper nodes.
[0220] The above means the first dataset and all dependent nodes connected thereto, and the above is, node Details of It means the score for, and the rest of the formula explanation is the same as Formula 1 described above and is applied accordingly.
[0221] Likewise, in the embodiment, the artificial intelligence agent (AA) can map the class for the calculated integrated risk score based on the class table.
[0222] That is, in the embodiment, the artificial intelligence agent (AA) can determine the actual level of legal risk to be exposed by quantifying the maximum legal risk to be finally encountered among the dependent nodes connected to the first dataset by calculating the integrated risk score of the first dataset.
[0223] Additionally, in the embodiment, the artificial intelligence agent (AA) can detect nodes where an inversion phenomenon has occurred in the dataset and all connected dependent nodes by comparing the classes mapped to the individual risk score and the integrated risk score, respectively.
[0224] In the examples, the inversion phenomenon refers to a contradictory situation where a lower node is superficially indicated as safer than a higher node, a phenomenon that occurs due to the practice of looking only at superficial licenses.
[0225] FIGS. 8 and 9 are examples of schematic drawings illustrating the detection of a reversal phenomenon using a risk score calculated according to an embodiment of the present invention.
[0226] Although only the first dataset (DS) is shown in FIG. 8, risk scores can be calculated for all dependent nodes connected to the first dataset.
[0227] In an embodiment, the first dataset (DS) may include a dataset name, URL, type, modality, and / or license proof information.
[0228] Here, modality refers to the type of the first dataset (DS) (e.g., images, videos, etc.), and the license proof information may be simply indicated information that reflects only the fragmentary license of the dataset itself, without reflecting all dependent nodes for the first dataset (DS).
[0229] Referring to FIGS. 8 and 9, in the embodiment, an artificial intelligence agent (AA) can calculate individual risk scores (400) for each of the first dataset (DS) and all dependent nodes connected thereto, according to license evaluation criteria for calculating risk scores. An integrated risk score (500) can be calculated once and applied equally to all nodes.
[0230] For example, as described, when the first dataset (DS) is considered to be the top node, individual risk scores (400) and integrated risk scores (500) can be calculated for each of the sub-nodes connected to the first dataset (DS).
[0231] Additionally, in the embodiment, the computing system (1000) can determine irregular nodes in which an inversion phenomenon has occurred based on the individual and integrated risk scores. (S109)
[0232] To this end, in the embodiment, the artificial intelligence agent (AA) can detect a node (hereinafter referred to as an irregular node) in which the individual risk score (400) is calculated to be higher than the integrated risk score (500).
[0233] In the example, the class determined by the risk score may mean that the closer the alphabet is to A and the closer the number is to 1, the lower the risk.
[0234] Referring to the illustration, for the first dataset (DS), the class according to the individual risk score (400) is A-3 and the class according to the integrated risk score (500) is C-2; thus, it is determined that an inversion phenomenon has occurred, and the first dataset (DS) can be determined as an irregular node. Similarly, the first to third connected nodes (N1, N2, N3) in which an inversion phenomenon has occurred can be determined as irregular nodes.
[0235] That is, in the embodiment, the artificial intelligence agent (AA) can detect and determine at least one irregular node in which a reversal phenomenon has occurred. Accordingly, it provides the advantage of resolving the discrepancy between the nominal license and the actual license risk through a mechanism in which the dispute elements of the lower node are transferred to the upper node.
[0236] Additionally, in the embodiment, the computing system (1000) can analyze the dependency of the first dataset on the determined irregular node. (S111)
[0237] In the embodiments, dependency can be analyzed based on the quantitative proportion of how much a specific problematic node (i.e., irregular node) accounts for in the entire dataset and / or the substantial contribution to how much it substantially contributes to the training of the AI model.
[0238] In an embodiment, the artificial intelligence agent (AA) can calculate a data ratio that estimates the proportion of text and / or image tokens of the irregular node within the entire dataset in order to calculate a quantitative proportion.
[0239] Specifically, in the embodiment, the artificial intelligence agent (AA) records metadata indicating which sub-node each record (e.g., sentence, image, audio segment, etc.) of the dataset originated from, and can calculate the ratio obtained by dividing the number of records of data originating from the irregular node by the total number of records.
[0240] For example, if the number of records originating from the above irregular node is 20,000 out of a total of 100,000 records, 20% of the entire dataset can be interpreted as a license violation risk range.
[0241] Additionally, in the embodiment, the artificial intelligence agent (AA) can evaluate the dependency highly regardless of the calculated quantitative weight when an irregular node is dedicated to a specific domain or provides an essential label. In other words, it can reflect not only the ratio based on the amount of simple data but also the importance of the content.
[0242] In addition, in the embodiment, the artificial intelligence agent (AA) can calculate the model performance contribution by estimating the importance that the irregular node occupies in the AI model performance in order to calculate the substantial contribution.
[0243] To this end, in the embodiment, the artificial intelligence agent (AA) may perform removal experiments and / or contribution analysis experiments.
[0244] An elimination experiment refers to an experiment in which data provided by irregular nodes is excluded during AI model training, and the model is retrained to compare performance. In this case, if performance drops below a pre-set threshold, the contribution of the corresponding irregular node can be estimated and calculated as high.
[0245] A contribution analysis experiment refers to an experiment that compares the similarity between a partial embedding learned solely from random nodes and a full embedding during AI model training. In this case, the higher the similarity, the higher the contribution of the corresponding random node can be estimated and calculated.
[0246] Based on the quantitative weight and substantial contribution calculated above, the artificial intelligence agent (AA) in the embodiment can estimate the dependency of the irregular node as an objectified value. For example, the dependency can be expressed as a predetermined percentage ratio (e.g., 70%).
[0247] That is, in the embodiment, the artificial intelligence agent (AA) can determine the final usability by analyzing the dependency of irregular nodes by estimating the dependency of irregular nodes into an objective value based on quantitative weight and actual contribution.
[0248] Additionally, in an embodiment, the computing system (1000) may provide a license compliance evaluation result for the first dataset according to the analyzed dependency. (S113)
[0249] The results of the above license compliance evaluation may be determined as available or unavailable licenses, or as risk scores for the target (representative) dataset or for each entity constituting the dataset.
[0250] In an embodiment, if the dependency of the irregular node is below a preset threshold, the artificial intelligence agent (AA) can determine the license compliance evaluation result for the first dataset as 'available' by deleting and / or replacing the irregular node.
[0251] Simply put, even if a node is problematic, if the dependency is small, it is possible to support the continued use of the dataset by removing or replacing only a part of it from the dataset.
[0252] Conversely, in the embodiment, if the dependency of the irregular node is greater than or equal to a preset threshold, the artificial intelligence agent (AA) may determine the license compliance evaluation result for the first dataset containing the irregular node as 'unavailable'.
[0253] This is because if the dependency of an irregular node exceeds a preset threshold (e.g., 70%), it means that the entire dataset, including the parent node that has the irregular node as a child node, is nearly unusable.
[0254] That is, in the embodiment, the computing system (1000) can generate a usable identifier indicating that the irregular node is available for use if the analyzed dependency is less than a preset threshold, and can generate an unusable identifier if the analyzed dependency is greater than or equal to the preset threshold.
[0255] Additionally, in an embodiment, the computing system (1000) may provide at least one interface including a dynamic interface and a separate interface so that AI and legal experts review the license compliance evaluation and make a decision regarding final compliance.
[0256] Here, the at least one interface can be implemented in various embodiments.
[0257] In one embodiment, the collaboration process among experts and the final result verification process can both be performed within a single integrated dynamic interface.
[0258] For example, the chat window and edit button may remain active until the evaluation is completed, and once final approval is granted, these features may be locked within the same web page and immediately transitioned into the final result report format.
[0259] In another embodiment, the collaboration process among experts and the final result verification process can be performed individually in a dynamic interface and a separate interface.
[0260] For example, the process of experts collaborating and exchanging opinions is performed through a 'dynamic interface (collaboration tool),' but the finalized compliance assessment results may be provided through a 'separate interface (result report viewer)' that is unmodifiable or has read-only attributes.
[0261] However, in order to explain the features of the invention more clearly, the following description will focus on an embodiment in which a dynamic interface for expert collaboration and a separate interface providing final compliance results operate separately.
[0262] FIG. 10 is an example of a dynamic interface according to an embodiment of the present invention.
[0263] Referring to FIG. 10, the dynamic interface (600) can display the hierarchy structure of the dataset, the risk class (risk score), the basis on which the risk class was determined, the AI model and prompt information used for compliance analysis, and metadata including the aforementioned list of entities.
[0264] In particular, the dynamic interface (600) can provide a visualization library (601) that intuitively visualizes the hierarchical structure of the dataset to show where the risk is found by displaying the depth (D1 to D14) of the dataset.
[0265] The above dynamic interface (600) is provided for multiple users (experts), and each expert user can become an evaluator and input individual evaluation opinions to display on the interface.
[0266] Additionally, in the embodiment, the computing system (1000) may provide a cross-review function in which other evaluators provide opinions on the first evaluator's evaluation opinion and mutually review it based on the dynamic interface (600).
[0267] Multiple evaluators can each present their risk classes and rationale, discuss each other's opinions, and record this history of debate to serve as a reference for the final decision.
[0268] At this time, if the risk class gap among the aforementioned multiple evaluators exceeds a preset threshold, the dataset may be automatically designated as a separate review item.
[0269] Subsequently, in the embodiment, the computing system (1000) can determine the final evaluation by synthesizing the opinions of multiple evaluators. This can be performed based on the Delphi method, in which each expert (evaluator) presents their opinion anonymously, and the judgment is repeatedly modified by referring to each other's opinions to derive an agreed conclusion.
[0270] Once the final evaluation is determined, the first evaluator may finalize the evaluation content based on the agreed content and authenticate it by performing a digital seal or signature.
[0271] Additionally, in the embodiment, the computing system (1000) can generate an evaluation report based on the determined final evaluation and digital seal.
[0272] Meanwhile, the computing system (1000) according to an embodiment of the present invention can perform AI-based risk evaluation by utilizing the Delphi technique performed at the time of the final evaluation decision and driving multiple AI agent models.
[0273] Here, the multiple AI agents according to the embodiment are virtual agents based on an artificial intelligence model, and are created as a concept distinct from the actual agent, the aforementioned expert user.
[0274] This is characterized by controlling the generation of multiple virtual AI agents, each assigned different characteristics, roles, and legal perspectives, rather than having a single AI model evaluate the risk, to perform a multi-faceted risk evaluation on the same first dataset.
[0275] FIG. 11 is a flowchart illustrating a process for performing risk assessment through a plurality of virtual AI agents according to an embodiment of the present invention.
[0276] Referring to FIG. 11, in an embodiment, the computing system (1000) can generate a plurality of virtual AI agents. (S301)
[0277] Specifically, in the embodiment, the virtual AI agent may be an agent predefined to have different evaluation criteria or roles in order to analyze legal risks for the first dataset in a multifaceted and three-dimensional manner.
[0278] More specifically, in the embodiment, the computing system (1000) may create multiple virtual AI agents having different evaluation criteria by allocating resources or model parameters so that each agent has unique analysis characteristics, judgment tendencies, and expertise domains from the time of creation, rather than simply duplicating and calling the same AI model.
[0279] To this end, in the embodiment, the computing system (1000) can set different system prompts for each of the plurality of virtual AI agents. In other words, each virtual AI agent may share the same underlying language model, but may be configured to run a different system prompt when each agent is called.
[0280] In addition, in the embodiment, the computing system (1000) can fine-tune each virtual AI agent by applying different base models and fine-tuning data to each of the plurality of virtual AI agents. Furthermore, inference parameters may be set differentially.
[0281] For example, the computing system (1000) can generate a ‘Strict Lawyer AI Agent’ specialized in conservatively interpreting license agreements and legal provisions. At the same time, it can generate a ‘Flexible Lawyer AI Agent’ specialized in exploring exceptions or permissible ranges by considering the actual usability of the data and industrial practices.
[0282] Additionally, virtual AI agents specialized not only from a legal perspective but also in specific fields (e.g., medical-specialized AI agents) and specific regions (e.g., US lawyer AI agents) can be created.
[0283] In other words, there may be embodiments such as creating a virtual AI agent based on a separate base model fine-tuned with data from a specific field, or creating a virtual AI agent based on a separate external knowledge database that has intensively learned data from a specific region.
[0284] Additionally, in the embodiment, the computing system (1000) may input a first dataset as a subject of discussion to the plurality of virtual AI agents created above. (S303)
[0285] Then, the plurality of virtual AI agents can independently derive risk classes, scores, and / or detailed evaluation grounds for the first dataset and the entities connected thereto, according to the unique interpretation criteria assigned to each. The results derived in this way can be stored as evaluation results of each virtual AI agent.
[0286] Additionally, in an embodiment, the computing system (1000) can control the plurality of virtual AI agents to perform a cross-reference discussion of evaluation results derived for the first dataset. (S305)
[0287] The above discussion may be a process for another agent to correct a specific agent's biased perspective or analytical errors, or to identify logical conflicts between agents to reach a more robust consensus.
[0288] For example, the first agent may generate an opinion refuting the grounds for the first evaluation among the evaluation results derived by the second agent. Accordingly, the second agent may present grounds for defense against the first agent's rebuttal opinion by citing recent case law.
[0289] This process of mutual discussion and rebuttal can be repeated a preset number of times or performed until the evaluation result of each agent satisfies a predetermined threshold.
[0290] Accordingly, in the embodiment, the computing system (1000) can generate a preliminary evaluation result by synthesizing the opinions of a plurality of virtual AI agents through discussion. (S307)
[0291] The above preliminary evaluation results may be the most objective and balanced results, as they are refined or agreed upon through the above interaction and discussion process.
[0292] As a result, potential risks that were not discovered during the initial independent evaluation phase, or conversely, acceptable grounds, can be detected during the discussion process.
[0293] Thus, the actual expert (evaluator) using the dynamic interface (600) can perform the final evaluation based on preliminary evaluation result data from which a multi-faceted legal review has been completed, rather than the biased perspective of a single AI.
[0294] Additionally, in the embodiment, the computing system (1000) may determine a final evaluation based on qualitative feedback regarding the generated preliminary evaluation result. (S309)
[0295] Here, qualitative feedback refers to the subjective and legal opinion of an actual expert (evaluator) based on the dynamic interface (600).
[0296] Specifically, the first evaluator may input qualitative feedback on the discussion history and preliminary evaluation results performed by multiple virtual AI agents on the first dataset to verify the validity of the conclusions derived by the AI and to supplement legal nuances or current contexts that the AI may have overlooked.
[0297] To this end, at least one evaluator may determine the correctness of the first agent's opinion regarding a specific matter and input the basis for the judgment.
[0298] For example, the first evaluator may write a judgment and grounds for judgment such as, “The ‘strict lawyer AI agent’ is correct regarding this matter because the ‘flexible lawyer AI agent’ overlooked the non-commercial exception clause of the latest EU case law.”
[0299] In an embodiment, the computing system (1000) can store input qualitative feedback by matching it with the corresponding dataset. The information stored in this way can be used to dynamically correct the internal parameters and / or weights of the virtual AI agent.
[0300] In other words, by configuring a feedback loop in which human knowledge is transferred to the AI model, the virtual AI agent model itself can be adjusted in real time for similar issues detected in the future, such as temporarily lowering the judgment weight of the first agent and temporarily raising the judgment weight of the second agent.
[0301] Additionally, in the embodiment, the computing system (1000) may provide the determined final evaluation through a dynamic interface or a separate interface. (S311)
[0302] In one embodiment, the computing system (1000) may display the final evaluation result by creating a URL or a separate page independent of the dynamic interface once final approval is granted after the expert discussion is completed. In this case, unlike the dynamic interface, the separate interface may restrict user input or modification functions and statically display only the authenticated result.
[0303] FIG. 12 is an example of an evaluation report according to an embodiment of the present invention.
[0304] Referring to FIG. 12, in the embodiment, the evaluation report (700) may consist of a digital seal (701), a final risk class (702), and risk details (703).
[0305] The above evaluation report (700) may be provided to at least one user (e.g., individual customers, companies and / or experts).
[0306] The digital seal (701) may be in the form of a certificate including the signature of the first evaluator.
[0307] The final risk class (702) may include individual risk information (or score) representing the risk level of the first dataset itself, and integrated risk information (or score) that reflects the potential risk of all dependent nodes connected to the dataset. At this time, when each node is selected, the score for each category of the license evaluation criteria, showing how the corresponding score was calculated, may be displayed.
[0308] This allows users to clearly compare and identify the apparent risk of the dataset with the actual potential risk.
[0309] The risk details (703) can provide a clear conclusion regarding the final usability of the first dataset (e.g., usable, unusable and / or conditionally usable) based on the calculated risk score and the results of the dependency analysis of the irregular nodes.
[0310] This provides the effect of helping users quickly and accurately decide whether to utilize datasets within complex licensing relationships.
[0311] In an embodiment, the computing system (1000) can visualize the finally generated evaluation report in the form of a web page and provide it to the user.
[0312] Additionally, various visualizations may be included to enhance the understanding of the evaluation report. For example, a pie chart can be used to visualize the distribution of sub-datasets within the first dataset, allowing for a quick overview of which domains (e.g., medical, financial, general knowledge, etc.) they belong to. This can aid in the intuitive analysis of the dataset's composition and any bias toward specific domains.
[0313] This web page may include a dynamic user interface (UI) that allows for interaction beyond simply viewing results, enabling users to click on specific nodes or risk items to view details and use filtering functions to select and view only the desired information. This provides users with the effect of intuitively understanding complex analysis results and taking necessary actions quickly.
[0314] Additionally, in the embodiment, the computing system (1000) may build a database of evaluation characteristics for each of the plurality of evaluators and analyze the characteristics to provide feedback to each evaluator.
[0315] To this end, in the embodiment, the computing system (1000) can store all evaluation activities performed by each evaluator at the dynamic interface as log data.
[0316] The log data collected at this time may include the evaluator's initial and final risk classes for a specific dataset, qualitative opinions written as the basis for evaluation and discussions with other evaluators, the degree of agreement and difference with the initial evaluation results presented by the AI, the time taken to reach a final conclusion and / or history of changes in judgment, etc.
[0317] In an embodiment, the computing system (1000) can analyze the collected log data to identify the unique evaluation characteristics of each evaluator and generate a quantified profile.
[0318] For example, a specific evaluator's tendency to consistently assign stricter scores than the average of other experts on the 'privacy' category can be quantified as a 'conservatism index,' or their consistency in evaluating similar cases can be measured as a 'consistency score.' Furthermore, by utilizing Natural Language Processing (NLP) techniques to analyze the text written by evaluators, it is possible to identify the logical reasoning patterns or potential biases regarding specific legal risks.
[0319] Based on the evaluator profile constructed in this way, the computing system (1000) can automatically generate and provide a personalized feedback report to each evaluator.
[0320] The above feedback report may include visualization data comparing the evaluator's evaluation tendency by risk item with the average value of all evaluators, major cases similar to or different from the evaluator's judgment and the grounds for the final decision regarding said cases, and suggestions for maintaining consistency in evaluation.
[0321] Through this feedback structure, individual experts have the opportunity to objectively recognize their potential biases and correct their judgments, thereby providing the effect of increasing the reliability and consistency of the evaluation of the entire expert group, while the computing system (1000) utilizes the accumulated detailed evaluation data of experts to retrain an AI model (e.g., a scoring model (M3)), thereby continuously improving the analysis accuracy of the system itself.
[0322] Furthermore, the computing system (1000) according to an embodiment of the present invention can provide a service through a dynamic risk evaluation process, which does not uniformly apply the same risk evaluation criteria to all users, but rather receives the user's environment and provides customized data compliance optimized for that user.
[0323] FIG. 13 is a flowchart illustrating a dynamic risk assessment method for a dataset according to an embodiment of the present invention.
[0324] Referring to FIG. 13, in the embodiment, the computing system (1000) can obtain user environment information. (S501)
[0325] Here, user environment information refers to context information for interpreting and evaluating the legal risks of a dataset in accordance with the user's specific situation.
[0326] This user environment information can be input from at least one user or obtained from a pre-configured user profile as a parameter for performing a customized assessment based on the user's unique legal compliance obligations and usage scenarios, rather than a fixed risk assessment.
[0327] Specifically, user environment information may include the country to which the user belongs, the industry domain, and the purpose of utilizing the dataset. For example, based on information about the country, the computing system may identify the unique legal restrictions of that country (e.g., the United States, Europe, etc.) (e.g., CCPA (California Consumer Privacy Act), GDPR (General Data Protection Regulation), etc.).
[0328] Since legal restrictions vary by country, this can be used to adjust the algorithm to prioritize the consideration of such legal restrictions during risk assessment.
[0329] As another example, a computing system can identify the characteristics of a relevant industry domain (e.g., healthcare, finance, etc.) based on information about the industry domain.
[0330] This can be utilized to adjust the algorithm so that, in cases where the user's industry domain is a field with high personal information sensitivity, the weight of evaluation items related to personal information protection is set relatively higher compared to other domains during risk assessment.
[0331] As another example, a computing system can set a risk tolerance based on the purpose of utilizing the dataset.
[0332] This can be used to adjust the algorithm to set a higher weight for license-related items when the purpose of use is commercial, depending on whether the dataset is for the development of commercial services provided to the public or for internal research.
[0333] Additionally, in the embodiment, the computing system (1000) can obtain trend information. (S503)
[0334] Although the aforementioned user environment information and trend information are described as being obtained in stages, the two stages can be performed regardless of the order.
[0335] Here, trend information refers to information intended to detect the occurrence of potential risks or changes in existing risks in advance.
[0336] In addition, trend information, as the latest information related to risk assessment items, may include legal precedents, the status of related litigation, and major news articles. This trend information can be continuously tracked by artificial intelligence agents through web crawling, news feed analysis, and integration with legal information databases.
[0337] For example, if an artificial intelligence agent (AA) detects a surge in negative court precedents or articles related to litigation concerning a specific evaluation item (e.g., copyright), the computing system (1000) can identify the item as a main topic and add related information as metadata.
[0338] At the same time, the computing system (1000) can adjust the algorithm to set the risk weight of the corresponding evaluation item higher than that of other items.
[0339] In other words, the above trend information can be obtained to reflect real-time changes in the external legal environment in the risk assessment, rather than a static legal database.
[0340] In this way, in the embodiment, the computing system (1000) can adjust a plurality of basic algorithms used for calculating compliance risk to a form optimized for the user by synthesizing user environment information and trend information. (S505)
[0341] Specifically, in an embodiment, the computing system (1000) can dynamically reconfigure the parameters of a pre-set risk evaluation model (e.g., a scoring model (M3)) using user environment information and trend information as input variables, and adjust them to an optimized form for the user.
[0342] That is, in the embodiment, dynamic risk adjustment means not applying uniform evaluation criteria, but identifying evaluation items that need to respond most sensitively to the user's specific context (e.g., country and / or purpose) and latest trends (e.g., latest case law), and variably adjusting the importance of said items in the algorithm.
[0343] Hereinafter, an algorithm adjusted to be optimized for the above user may be referred to as an 'interactive algorithm'.
[0344] In an embodiment, the computing system (1000) can apply the interactive algorithm as a basic parameter of at least one artificial intelligence model.
[0345] Additionally, in an embodiment, the computing system (1000) can provide a data compliance service to which the interactive algorithm is applied. (S507)
[0346] Specifically, in an embodiment, the computing system (1000) can provide a data compliance service that calculates a user-customized risk score (class) for a first dataset based on the interactive algorithm.
[0347] As a result, even with the same dataset, different evaluation results can be derived depending on the user's unique context.
[0348] For example, assuming the same first dataset is input, the same score is not calculated for all users; instead, the first dataset may be evaluated as A-2 (Safe) grade for a 'research' user in the 'USA', and as C-1 (Risk) grade for a 'commercial' user in 'Europe', with the weight related to commercial use being adjusted upward by the interactive algorithm.
[0349] Additionally, in the embodiment, when the computing system (1000) provides the calculated user-customized risk score as a dynamic interface, it may display the part reflecting user environment information and trend information as a basis for dynamic evaluation.
[0350] For example, detailed grounds such as "Your 'European-Commercial' environment information and 'latest GDPR case law' trend information have been detected, and the weight of the '3-1 Personal Data Contains' item has been increased by 40%" may be displayed and provided.
[0351] This provides the effect of helping experts (evaluators) intuitively understand why the risk was rated high based on the user's specific context, and to make a more accurate final assessment.
[0352]
[0353] Meanwhile, in the embodiment, the artificial intelligence agent (AA) can perform a risk optimization process that provides a trimming function to dynamically adjust the structure of a dataset (e.g., multiple connected entities) in order to increase the utilization of a dataset classified into a risk class lower or higher than a preset threshold.
[0354] FIG. 14 is an example of a drawing for explaining a trimming function according to an embodiment of the present invention.
[0355] Referring to FIG. 14, in an embodiment, an artificial intelligence agent (AA) can detect a selection input for a first entity (EN-1) on a visualization library (601) provided by a dynamic interface (600). This input may be a manual input by a user, but may also be automatically selected according to a predetermined standard.
[0356] The first entity (EN-1) above may be an entity determined to be less than a preset risk class and / or an entity exceeding a preset depth value.
[0357] In the embodiment, the artificial intelligence agent (AA) may physically delete or logically exclude data records (e.g., text sentences, image files) originating from or associated with the first entity (EN-1). That is, trimming refers to a process of increasing the risk class of a dataset by reducing the depth value of an entity with a large depth value.
[0358] For example, for a dataset determined to be unusable due to a low integrated risk score, the risk of the problematic first entity can be blocked from being transferred to the parent node by trimming the first entity.
[0359] More specifically, the artificial intelligence agent (AA) can change part of the answer output from the language model, change the license conditions granted to the language model to a free license, or delete the entire problematic row among the tabular data records.
[0360] Additionally, in the embodiment, the artificial intelligence agent (AA) can re-perform the integrated risk score calculation step (S107) on the trimmed dataset to recalculate the integrated risk score.
[0361] Since the problematic first entity has been removed, the recalculated integrated risk score can be improved compared to the existing integrated risk score prior to trimming.
[0362] Through this, in the embodiment, the computing system (1000) can generate a dataset (hereinafter, clean dataset) consisting only of sources of a risk class (hereinafter, target risk class) that the user allows and provide it to the user.
[0363] Additionally, in the embodiment, the computing system (1000) can perform data augmentation to compensate for the problem of insufficient data that may occur due to trimming.
[0364] To this end, in the embodiment, the computing system (1000) can analyze the attributes of the data records removed due to trimming. For example, the attributes of the data records include label configuration, domain, data format, text semantics, etc.
[0365] Based on the analyzed attributes, in the embodiment, the computing system (1000) can perform content augmentation and / or alternative data sourcing.
[0366] Content augmentation can be performed by generating new data by sampling data with attributes similar to the removed data from the clean dataset remaining in the original dataset.
[0367] Alternative data sourcing can be performed by exploring external alternative data verified as having no licensing risk, targeting the attributes of the removed data to extract data among the alternative data that matches those attributes, and mapping the extracted data to the existing data.
[0368] Through these trimming and augmentation processes, rather than simply discarding datasets with high legal risk, it is possible to dynamically control risk to an acceptable level and provide the effect of maximizing data usability.
[0369] In summary, the data compliance management method and system according to the embodiment of the present invention provides a systematic collaboration interface in which multiple legal experts mutually review their respective opinions to derive an agreed-upon conclusion, thereby excluding the subjectivity of a specific evaluator and having the effect of deriving more objective and reliable compliance evaluation results.
[0370] In addition, the data compliance management method and system according to an embodiment of the present invention has the effect of clearly proving the safety and reliability of the data to the dataset user or consumer by specifying the compliance level of the finally determined dataset and automatically generating a reliable report certified by a digital seal, etc.
[0371]
[0372] The embodiments according to the present invention described above may be implemented in the form of program instructions that can be executed through various computer components and recorded on a computer-readable recording medium. The computer-readable recording medium may include program instructions, data files, data structures, etc., either individually or in combination. The program instructions recorded on the computer-readable recording medium may be those specifically designed and configured for the present invention or those known and available to those skilled in the art of computer software. Examples of computer-readable recording media include magnetic media such as hard disks, floppy disks, and magnetic tapes; optical recording media such as CD-ROMs and DVDs; magneto-optical media such as floptical disks; and hardware devices specifically configured to store and execute program instructions, such as ROM, RAM, and flash memory. Examples of program instructions include machine code, such as that generated by a compiler, as well as high-level language code that can be executed by a computer using an interpreter, etc. Hardware devices may be modified into one or more software modules to perform processing according to the present invention, and vice versa.
[0373] The specific embodiments described in this invention are examples and do not limit the scope of the invention in any way. For the sake of brevity of the specification, descriptions of prior electronic configurations, control systems, software, and other functional aspects of said systems may be omitted. Additionally, the connections of lines or connecting members between components shown in the drawings are illustrative of functional connections and / or physical or circuit connections, and may be replaced or additionally represented as various functional connections, physical connections, or circuit connections in actual devices. Furthermore, unless specifically stated as “essential,” “importantly,” etc., a component may not be strictly necessary for the application of the invention.
[0374] Furthermore, although the detailed description of the present invention has been explained with reference to preferred embodiments of the invention, those skilled in the art or those with ordinary knowledge in the relevant technical field will understand that various modifications and changes can be made to the invention without departing from the spirit and technical scope of the invention as set forth in the claims below. Accordingly, the technical scope of the present invention should not be limited to the contents described in the detailed description of the specification, but should be determined by the claims.
[0375] The data compliance management method and system according to the present invention has industrial applicability in the field of customized compliance management and certification services, in which an AI model development company or a data provider analyzes complex legal risks of a dataset through collaboration between an AI agent and an expert, and further refines the data to match a target risk level set by the user.
Claims
1. As a method performed on a computer, A step in which at least one processor analyzes the license compliance of at least one dataset connected to multiple entities using at least one artificial intelligence model; The above-mentioned at least one processor obtains a plurality of evaluation opinions from a plurality of real or virtual agents regarding the analyzed license compliance result; and The above-mentioned at least one processor controls the final compliance evaluation of at least one dataset, which synthesizes the plurality of evaluation opinions, to be expressed through at least one interface; Data Compliance Management Methods 2. In Paragraph 1, The above entity is, A parent dataset used to generate the above at least one dataset, or a sub-dataset derived from the above at least one dataset and processed or redistributed. Data Compliance Management Methods 3. In Paragraph 2, The above-mentioned at least one artificial intelligence model is, A first model for searching for data related to at least one dataset on the web or app, A second model that determines metadata by extracting dependency relationships between multiple entities connected to the at least one dataset based on data retrieved for the at least one dataset, and Characterized by including at least one of a third model that scores risk based on metadata for at least one dataset. Data Compliance Management Methods 4. In Paragraph 3, The step of analyzing license compliance of at least one dataset above is, A plurality of entities detected from at least one dataset, further comprising the step of recursively performing a license compliance analysis process using the first to third models for lower-level entities derived from each entity. Data Compliance Management Methods 5. In Paragraph 3, The above metadata is, The above at least one dataset and the name, task category, modality, application field, and license type of each of the above plurality of entities Data Compliance Management Methods 6. In Paragraph 1, The step of analyzing the above license compliance is, The step of calculating individual risk classes for at least one dataset, and A step comprising calculating an integrated risk class for at least one dataset based on individual risk classes for a plurality of entities connected to at least one dataset. Data Compliance Management Methods 7. In Paragraph 1, The step of obtaining multiple evaluation opinions from the above-mentioned multiple actual agents is, The step of expressing the above license compliance result through the above at least one interface, and The method includes the step of receiving an input signal including an evaluation opinion input through at least one interface from the plurality of actual agents. Data Compliance Management Methods 8. In Paragraph 1, The step of obtaining multiple evaluation opinions from the above-mentioned multiple virtual agents is, The step of creating a plurality of predefined virtual AI agents having different evaluation criteria, and The step of controlling each of the above plurality of virtual AI agents to analyze license compliance for the above at least one dataset, and A step comprising obtaining an evaluation opinion on the license compliance results derived for each of the plurality of virtual AI agents. Data Compliance Management Methods 9. In Paragraph 8, The step of generating the above plurality of virtual AI agents is, Characterized by being trained based on at least one of different resources, model parameters, system prompts, fine-tuning data, a base model, and a database. Data Compliance Management Methods 10. In Paragraph 1, The step of obtaining the above multiple evaluation opinions is, The method further includes the step of controlling the at least one artificial intelligence model to perform a discussion that cross-examines the evaluation opinions of the plurality of real or virtual agents based on the Delphi technique. Data Compliance Management Methods 11. In Paragraph 10, The step of controlling the execution of the above discussion is, A step of correcting the evaluation result derived by the first virtual AI agent by cross-referencing it with at least one other virtual AI agent other than the first virtual AI agent, and A step of controlling the above correction to be repeated until a predetermined threshold is satisfied. Data Compliance Management Methods 12. In Paragraph 1, The step of matching the above final compliance evaluation with the identifier of the above at least one dataset and storing it in at least one memory, and The method further comprises the step of loading at least a portion of the data stored in the at least one memory and dynamically calibrating the internal parameters or weights of at least one of the plurality of virtual AI agents based on the final evaluation. Data Compliance Management Methods 13. In Paragraph 2, The above final compliance evaluation is, The final risk class determined by the individual risk class and integrated risk class of at least one dataset above, and Characterized by being provided as an evaluation report including the digital seal of any one of the actual agents mentioned above. Data Compliance Management Methods 14. At least one memory; and It includes at least one processor that executes instructions stored in at least one memory; and The above-mentioned at least one processor is, Receiving an input signal from a user through at least one interface; wherein the input signal is a signal including at least one condition set by the user for at least one dataset that is a service target, and Based on the above input signal, the parameters of a plurality of basic algorithms for compliance risk assessment for the above at least one dataset are adjusted; Providing a user-customized compliance service for at least one dataset based on an interactive algorithm to which the above-mentioned adjusted parameters are applied; A data compliance management system that operates according to commands.
15. In Paragraph 14, The above at least one interface is, Visualize the above at least one dataset and a plurality of entities connected to the above at least one dataset in a hierarchical structure, and Characterized by visualizing the depth of multiple entities connected to at least one dataset. Data Compliance Management System.
16. In Paragraph 14, At least one condition set by the above user is, User environment information including at least one of the country to which the user belongs, the industry domain, and the purpose of use of the above dataset, and Trend information including the latest information related to multiple risk assessment items tracked by crawling or database integration by at least one artificial intelligence model, and Characterized as being at least one of the information determining the target risk class of at least one dataset that is the target of the service. Data Compliance Management System.
17. In Paragraph 14, The parameters of the above plurality of basic algorithms are, Characterized by including weights or calculation formulas for calculating risk scores Data Compliance Management System.
18. In Paragraph 14, The above user-customized compliance service is, Individual risk classes for at least one dataset that is the target of the above service, and It further includes calculating an integrated risk class for the at least one dataset based on individual risk classes for a plurality of entities connected to the at least one dataset, and Characterized that the risk assessment result calculated by applying the above-mentioned plurality of basic algorithms and the risk assessment result calculated by applying the above-mentioned interactive algorithm are different from each other. Data Compliance Management System.
19. In Paragraph 16, The above user-customized compliance service is, It further includes performing trimming to delete at least one data record derived from a first entity among the plurality of entities from the at least one dataset. The integrated risk class of at least one trimmed dataset is characterized by satisfying the target risk class. Data Compliance Management System.
20. In Paragraph 19, The above-mentioned first entity is, Entities whose individual risk classes are below a pre-set threshold and Characterized by being at least one of the entities whose depth in the data hierarchy exceeds a preset depth value. Data Compliance Management System.
Citation Information
Patent Citations
Legal knowledge graph construction system
CN112559766A
Battery module
KR1020240055654A
Display apparatus
KR1020250058853A
System and method for regulation compliance
US20130198094A1
Automated document review system combining deterministic and machine learning algorithms for legal document review
US20220004713A1