Data governance method and related device
Patent Information
- Application Number
- PCT/CN2025/139243
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-26
- Filing Date
- 2025-12-02
- Publication Date
- 2026-10-01
Smart Images

Figure CN2025139243_01102026_PF_FP_ABST
Abstract
Description
A data governance method and related equipment
[0001] This application claims priority to Chinese Patent Application No. 202510370690.8, filed with the State Intellectual Property Office of China on March 26, 2025, entitled "A Data Governance Method and Related Equipment", the entire contents of which are incorporated herein by reference. Technical Field
[0002] This application relates to the field of artificial intelligence (AI) technology, and in particular to a data governance method, a data governance system, a computing device cluster, a computer-readable storage medium, and a computer program product. Background Technology
[0003] In recent years, AI models, represented by Large Language Models (LLMs), have achieved remarkable success in many fields and tasks, demonstrating the potential to approach human intelligence. The capabilities of LLMs and other AI models stem from utilizing massive training datasets and a large number of model parameters. Building on this foundation, an increasing number of studies are beginning to use LLMs as core controllers to construct large-scale model applications with human-like decision-making capabilities, such as building large-scale AI agent applications (or simply agents).
[0004] Large-scale production-level model applications typically require high-quality data to improve model performance and effectiveness. Taking code development scenarios as an example, AI code development platforms are usually supported by underlying code data during the R&D process, and rely heavily on the "context" data of the enterprise code repository during application. Therefore, the effort invested in data governance cannot be ignored, and data quality ensures that the final code output meets the expected quality level. At the same time, data quality is also a crucial factor that AI code development platforms cannot ignore in the future as customized solutions.
[0005] To achieve continuous maintenance and enhancement of large-scale production-ready models, it is necessary to continuously collect high-quality data and use this data to continuously optimize the models. A crucial source of high-quality data is real-world user data and feedback from large-scale models deployed in production environments. However, in many scenarios, this data often relies on manual collection by developers, resulting in large-scale models lacking the ability to autonomously iterate prompts, update knowledge, and retrain themselves. Consequently, the performance of these large-scale models (e.g., end-to-end accuracy) declines over time. Summary of the Invention
[0006] This application provides a data governance method that collects source data related to AI applications from data sources, such as real user usage data and feedback data. When triggering conditions are met, a high-quality corpus is constructed based on this source data for autonomous iteration of prompts, autonomous knowledge updates, and autonomous model retraining. This enables continuous optimization of the AI model, preventing a decrease in end-to-end accuracy over time and ensuring the effectiveness of the AI application. This application also provides a data governance system, computing device cluster, computer-readable storage medium, and computer program product corresponding to the above method.
[0007] Firstly, this application provides a data governance method. The data governance method can be executed by a data governance system. The data governance system, also known as a data governance platform, is used to provide an automated triggering mechanism for a data flywheel based on Metrics-Driven Development (MDD) to continuously improve AI applications such as large-scale model applications and prevent the end-to-end accuracy of AI applications from decreasing over time. The data governance system can be a software system, which can be a standalone software system or integrated into other software systems in the form of plug-ins, components, functional modules, mini-programs, services, etc. In some possible implementations, the data governance system may also include a hardware system. For example, the data governance system may include a cluster of computing devices. When the computing device cluster is running, it executes the data governance method of this application.
[0008] Specifically, the data governance system can collect source data related to AI applications from at least one data source. AI applications include those built based on AI models. The data governance system then determines whether a set of triggering conditions related to data governance is met. When a target triggering condition in the set is met, the data governance system constructs a corpus corresponding to the target triggering condition based on the source data. This corpus includes at least one of the following: reordering corpus, embedding representation corpus, knowledge corpus, few-shot cue corpus, or model training corpus. The data governance system performs a quality assessment on the corpus corresponding to the target triggering condition and obtains the assessment result. When the assessment result meets the requirements, the data governance system can update the modules of the AI application based on the corpus corresponding to the target triggering condition. The modules of the AI application include at least one of the following: a reordering model, an embedding representation model, a knowledge base, a cue word center, a model training corpus, and model files.
[0009] This method collects source data related to AI applications from data sources, such as real user usage data and feedback data. When triggering conditions are met, it constructs high-quality corpus based on this source data for autonomous iteration of prompts, autonomous updating of knowledge, and autonomous retraining of the model. This enables continuous optimization of the AI model, preventing the end-to-end accuracy of AI applications (such as large model applications) from decreasing over time and ensuring the effectiveness of AI applications.
[0010] In some possible implementations, the data governance system can also determine the module update mode configured by the user. The module update mode includes a single module update mode, a multi-module update mode, or a full module update mode. Accordingly, the data governance system can determine whether a subset of triggering conditions corresponding to the module update mode is satisfied; this subset is a subset of the set of triggering conditions related to data governance.
[0011] This method supports quantitative optimization and improvement of modules such as the prompt center, knowledge base for Retrieval Augmented Generation (RAG), model training corpus, model files, re-ranking models, and embedding representation models, as well as optimal combination improvement of multiple modules or improvement of all modules, for different business scenarios, offering high flexibility. Specifically, rapid iteration of the training corpus can align with user preferences, alleviate illusion problems, and promote a positive cycle of improved user experience in AI applications such as large-scale model applications. The autonomous iteration of prompts, knowledge, and related models can reduce development and debugging workload, thereby reducing overall costs (such as the customization costs for adapting to new customers).
[0012] In some possible implementations, the set of triggering conditions includes a first type of triggering condition or a second type of triggering condition. The first type of triggering condition includes at least one of quantitative triggering condition or timed triggering condition, and the second type of triggering condition includes triggering conditions related to monitoring indicators.
[0013] This method, by introducing different types of triggering conditions, can lay the foundation for determining triggering conditions in different business scenarios, thereby meeting the needs of different business scenarios.
[0014] In some possible implementations, the monitoring metrics include at least one of the following: retrieval metrics, recall metrics, end-to-end metrics, knowledge deadline, or ranked data volume. Accordingly, the second type of triggering conditions includes one or more of the following: the retrieval metric reaches a first threshold; or the recall metric reaches a second threshold; or the end-to-end metric reaches a third threshold; or the knowledge deadline reaches a fourth threshold; or the ranked data volume reaches a fifth threshold.
[0015] This method introduces monitoring indicators of different dimensions, which can trigger the data flywheel when different triggering conditions are met, thus achieving high availability.
[0016] In some possible implementations, the data governance system can also provide a configuration interface to the user, receiving at least one of the following: rule information for triggering rules or threshold values corresponding to monitoring indicators configured by the user through the configuration interface. The rule information for triggering rules includes the time interval for timed triggering rules or the data volume interval for quantitative triggering rules. The rule information for triggering rules or the threshold values corresponding to monitoring indicators are used to form at least one triggering condition in the set of triggering conditions.
[0017] In this way, trigger conditions can be customized in a unified and standardized manner according to business needs, with low customization costs and the ability to meet individual business requirements. Moreover, this method supports a visual orchestration of the trigger condition judgment process (data flywheel automated triggering mechanism), which helps in subsequent module updates for AI applications through the data flywheel.
[0018] In some possible implementations, when a user enables automated data governance, the data governance system can also provide the user with the changing trends of monitoring indicators for AI applications after enabling automated data governance. This demonstrates the governance effect and provides a reference for users to adjust threshold settings and optimize the data flywheel triggering mechanism.
[0019] In some possible implementations, the data governance system can also determine the actual data distribution in the production environment based on the source data. Then, the data governance system can construct a corpus corresponding to the target triggering conditions based on the actual data distribution in the production environment.
[0020] This method constructs corpora based on real data distribution, which can obtain high-quality corpora and provide a corpus foundation for the autonomous iteration of prompts, autonomous knowledge updates, and autonomous retraining of models.
[0021] In some possible implementations, the data governance system can also acquire the corpus distribution corresponding to the target triggering conditions. The data governance system can then obtain evaluation results based on the corpus distribution and the actual data distribution in the production environment. These evaluation results include information on the differences between the corpus distribution and the actual data distribution.
[0022] This method determines the corpus distribution and evaluates the corpus distribution by combining it with the real data distribution in the production environment. This allows for filtering of corpus data that differs significantly from the real data distribution based on the corpus distribution evaluation results, thus ensuring the quality of the corpus.
[0023] In some possible implementations, the data governance system can also acquire the data source of the corpus corresponding to the target triggering conditions, and then obtain the evaluation result based on the data source. The evaluation result includes the data diversity of the corpus.
[0024] This method determines the data source of the corpus and evaluates the data diversity of the corpus based on that data source. This allows for the filtering of corpus that does not meet the data diversity requirements, thus avoiding overfitting of the corpus and affecting the autonomous iteration of prompts, autonomous updating of knowledge, and autonomous retraining of the model.
[0025] Secondly, this application provides a data governance system. The data governance system includes:
[0026] A data acquisition unit is used to acquire source data related to artificial intelligence (AI) applications from at least one data source, including applications built based on AI models.
[0027] The trigger condition determination unit is used to determine whether the set of trigger conditions related to data governance is met;
[0028] The corpus construction unit is used to construct a corpus corresponding to the target triggering condition based on the source data when the target triggering condition in the set of triggering conditions is met. The corpus corresponding to the target triggering condition includes at least one of the following: reordering corpus, embedding representation corpus, knowledge corpus, few-shot prompting corpus, or model training corpus.
[0029] An evaluation unit is used to perform quality evaluation on the corpus corresponding to the target triggering condition and obtain the evaluation result.
[0030] An update unit is used to update the modules of the AI application according to the corpus corresponding to the target triggering condition when the evaluation result meets the requirements. The modules of the AI application include at least one of a reordering model, an embedding representation model, a knowledge base, a prompt word center, a model training corpus, and a model file.
[0031] In some possible implementations, the system further includes:
[0032] The mode determination unit is used to determine the module update mode configured by the user, wherein the module update mode includes a single module update mode, a multi-module update mode, or a full module update mode.
[0033] The trigger condition determination unit is specifically used for:
[0034] Determine whether the subset of triggering conditions corresponding to the module update mode is satisfied, wherein the subset of triggering conditions is a subset of the set of triggering conditions related to data governance.
[0035] In some possible implementations, the set of triggering conditions includes a first type of triggering conditions or a second type of triggering conditions. The first type of triggering conditions includes at least one of quantitative triggering conditions or timed triggering conditions, and the second type of triggering conditions includes triggering conditions related to monitoring indicators.
[0036] In some possible implementations, the monitoring metrics include at least one of retrieval metrics, recall metrics, end-to-end metrics, knowledge cutoff time, or sorted data volume;
[0037] The second type of triggering conditions includes one or more of the following:
[0038] The search index reaches a first threshold; or...
[0039] The recall metric reaches the second threshold; or...
[0040] The end-to-end metric reaches the third threshold; or...
[0041] The knowledge deadline has reached the fourth threshold; or...
[0042] The amount of sorted data has reached the fifth threshold.
[0043] In some possible implementations, the system further includes:
[0044] A configuration unit is used to provide a configuration interface to the user; receive at least one of the rule information of the triggering rule or the threshold corresponding to the monitoring indicator configured by the user through the configuration interface, wherein the rule information of the triggering rule includes the time interval of the timed triggering rule or the data volume interval of the quantitative triggering rule, and the rule information of the triggering rule or the threshold corresponding to the monitoring indicator is used to form at least one triggering condition in the set of triggering conditions.
[0045] In some possible implementations, the system further includes:
[0046] The governance result providing unit is used to provide the user with the changing trend of the monitoring indicators of the AI application after the user enables automated data governance.
[0047] In some possible implementations, the corpus construction unit is specifically used for:
[0048] Based on the source data, determine the actual data distribution in the production environment;
[0049] Based on the actual data distribution of the production environment, construct a corpus corresponding to the target triggering conditions.
[0050] In some possible implementations, the evaluation unit is specifically used for:
[0051] Obtain the corpus distribution corresponding to the target triggering condition;
[0052] An evaluation result is obtained based on the corpus distribution and the actual data distribution in the production environment. The evaluation result includes information on the differences between the corpus distribution and the actual data distribution.
[0053] In some possible implementations, the evaluation unit is specifically used for:
[0054] Obtain the data source of the corpus corresponding to the target triggering condition;
[0055] Based on the data sources, an evaluation result is obtained, which includes the data diversity of the corpus.
[0056] Thirdly, this application provides a computing device cluster. The computing device cluster includes at least one computing device, and the at least one computing device includes at least one processor and at least one memory. The at least one processor and the at least one memory communicate with each other. The at least one processor is used to execute instructions stored in the at least one memory to cause the computing device or the computing device cluster to perform the data governance method as described in the first aspect or any implementation thereof.
[0057] Fourthly, this application provides a computer-readable storage medium storing instructions that instruct a computing device or a cluster of computing devices to perform the data governance method described in the first aspect or any implementation thereof.
[0058] Fifthly, this application provides a computer program product containing instructions that, when run on a computing device or a cluster of computing devices, causes the computing device or cluster of computing devices to perform the data governance method described in the first aspect or any implementation thereof.
[0059] Based on the implementation methods provided in the above aspects, this application can be further combined to provide more implementation methods. Attached Figure Description
[0060] To more clearly illustrate the technical methods of this application, the accompanying drawings used will be briefly described below.
[0061] Figure 1 is a schematic diagram of the architecture of a data governance system provided in this application;
[0062] Figure 2 is a flowchart of a data governance method provided in this application;
[0063] Figure 3A is a schematic diagram of a configuration interface provided in this application;
[0064] Figure 3B is a schematic diagram of a visual arrangement interface provided in this application;
[0065] Figure 4 is a flowchart of a data distribution assessment and diversity assessment provided in this application;
[0066] Figure 5 is a schematic diagram showing the changing trend of a monitoring indicator provided in this application after the activation of automated governance;
[0067] Figure 6 is a flowchart of a data governance method under a single-module enhancement mode provided in this application;
[0068] Figure 7 is a flowchart of a data governance method under a multi-module enhancement mode provided in this application;
[0069] Figure 8 is a flowchart of a data governance method under a full-module enhancement mode provided in this application;
[0070] Figure 9 is a schematic diagram of the structure of a data governance system provided in this application;
[0071] Figure 10 is a schematic diagram of the structure of a computing device provided in this application;
[0072] Figure 11 is a schematic diagram of the structure of a computing device cluster provided in this application;
[0073] Figure 12 is a schematic diagram of another computing device cluster provided in this application;
[0074] Figure 13 is a schematic diagram of another computing device cluster provided in this application. Detailed Implementation
[0075] The terms "first" and "second" used in the embodiments of this application are for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, a feature defined with "first" and "second" may explicitly or implicitly include one or more of that feature.
[0076] First, some technical terms involved in the embodiments of this application will be introduced.
[0077] Artificial intelligence (AI) is the ability to correctly interpret external data, learn knowledge from that data, and use that learned knowledge to achieve specific goals and tasks. Models built based on AI algorithms are called AI models, or simply models for convenience.
[0078] AI models can include language models (LMs) built based on AI algorithms such as machine learning (ML) or deep learning (DL). Language models can be used to predict the next most likely word based on the input context (e.g., several preceding context words). Language models can also be categorized by parameter size into small language models and large language models (LLMs). Small language models are small-scale language models, or simply small models, while large language models are large-scale language models, or simply large models. Large models can include, but are not limited to, generative pre-trained transformer (GPT) models with a large parameter size.
[0079] Language models, exemplified by LLMs, can serve as core controllers to build intelligent agents or AI agent applications with human-like decision-making capabilities. An agent, or AI agent, is a computer program based on a language model, possessing planning and thinking abilities, memory capabilities, and the ability to use tool functions, enabling it to autonomously complete given tasks. Agents can be used to build applications with human-like decision-making capabilities, such as large-scale production-level AI agent applications.
[0080] Large-scale production-level model applications typically require high-quality data to improve model performance and effectiveness. Taking code development scenarios as an example, AI code development platforms need high-quality data during the R&D process to ensure that the final code output meets expected quality levels. Furthermore, data quality is a crucial factor that AI code development platforms, as customized solutions, cannot ignore in the future.
[0081] To achieve continuous maintenance and enhancement of large-scale production-ready models, it is necessary to continuously collect high-quality data and use this data to continuously optimize the models. A crucial source of high-quality data is real-world user data and feedback from large-scale models deployed in production environments. This data can be used to construct training and knowledge corpora. However, in many scenarios, this data often relies on manual collection by developers, resulting in large-scale models lacking the ability to autonomously iterate prompts, update knowledge, and retrain themselves. Consequently, the performance of these large-scale models (e.g., end-to-end accuracy) declines over time.
[0082] In view of this, this application provides a data governance method. This method can be applied to a data governance system. The data governance system, also known as a data governance platform, provides an automated triggering mechanism for a data flywheel based on Metrics-Driven Development (MDD) to continuously improve AI applications such as large-scale model applications and prevent the end-to-end accuracy of AI applications from decreasing over time. The data flywheel is a data-driven cyclical optimization mechanism that continuously collects, analyzes, and utilizes data to drive continuous optimization and improvement in various business processes, forming a positive cycle. The data governance system can be a software system, which can be an independent software system or integrated into other software systems as plugins, components, functional modules, mini-programs, services, etc. In some possible implementations, the data governance system may also include a hardware system. For example, the data governance system may include a cluster of computing devices. When the computing device cluster is running, it executes the data governance method of this application.
[0083] Specifically, the data governance system can collect source data related to AI applications from at least one data source. AI applications include those built on AI models, such as production-level large-scale model applications. The data governance system can then determine whether a set of triggering conditions related to data governance is met. When the target triggering condition in the set is met, a corpus corresponding to the target triggering condition is constructed based on the source data. This corpus includes at least one of the following: reordering corpus, embedding representation corpus, knowledge corpus, few-shot cue corpus, or model training corpus. The data governance system can determine the value of an evaluation metric based on the evaluation metrics of the corpus corresponding to the target triggering condition. When the evaluation metric value meets the requirements, the data governance system updates the modules of the AI application based on the corpus corresponding to the target triggering condition. The AI application modules include at least one of the following: reordering model, embedding representation model, knowledge base, cue text, model training corpus, and model files.
[0084] This method collects source data related to AI applications from data sources, such as real user usage data and feedback data. When triggering conditions are met, it constructs high-quality corpus based on this source data for autonomous iteration of prompts, autonomous updating of knowledge, and autonomous retraining of the model. This enables continuous optimization of the AI model, preventing the end-to-end accuracy of AI applications (such as large model applications) from decreasing over time and ensuring the effectiveness of AI applications.
[0085] To make the technical solution of this application clearer and easier to understand, the system architecture of this application will be described below with reference to the accompanying drawings.
[0086] Referring to Figure 1, which illustrates the architecture of a data governance system, the data governance system 10 interfaces with an AI application 20. The AI application 20 includes an interactive device 202, a business intelligence agent 204, a prompt word center 206, a knowledge base 208, and an inference service 209. The data governance system 10 includes a data acquisition service 102 and a data development system 104. The data development system 104 is a dedicated hardware and software system for developing datasets and other data, also known as a data development production line. Furthermore, the data governance system 10 may also include a model development system 106 and a repository 108. Similar to the data development system 104, the model development system 106 is a dedicated hardware and software system for developing AI models, also known as a model development production line. The repository 108 can be a data hosting service or a repository provided by a data hosting platform.
[0087] To facilitate understanding, we will first introduce the components and operation process of AI Application 20.
[0088] Interactive device 202, serving as a business entry point, is typically used to obtain user-input queries and agent identifiers (agent_id) and conditions. The agent_id can be specified by the user or determined through intent recognition. Conditions, also known as feature profiles, can include at least one of the following: task type, developer affiliation information, programming language, programming framework, or interactive device type. Task type distinguishes different tasks; for example, task types may include code completion or code testing. Developer affiliation information may include the developer's organization / department. The programming language can be the currently used language, such as C, Java, or JavaScript. A programming framework is an abstract toolkit that provides general functionality, designed to help developers build and maintain applications more easily. Frameworks typically include a set of predefined classes, modules, and functions that developers can extend and customize. The goal of a framework is to provide a structured development approach, reduce repetitive work, and accelerate the application development process. Interactive device 202 may include an integrated development environment (IDE) plugin or a browser plugin; the type of interactive device may include the IDE type. For scenarios where the interactive device 202 is a browser plugin, the conditions may also include domain name and page information. The page information may include, but is not limited to, login information.
[0089] The interactive device 202 is also used to collect the context of the query. For example, when the interactive device 202 is an IDE plugin, it can capture the current position of the cursor and obtain the code snippets before and after the current position. Or, when the interactive device 202 is a browser plugin, it can also capture the domain name, page screenshots, and user actions. These capture operations can be implemented using a Computer-Using Agent (CUA). Furthermore, the interactive device 202 can also obtain cross-file context based on the import statements in the code file where the current position is located. This cross-file context includes the context of other files within the project. As shown in steps ① and ② of Figure 1, the interactive device 202 can receive the user's query input, obtain the agent identifier (agent_id) and conditions, collect the context, and then send the query, agent_id, conditions, and context to the business agent 204.
[0090] As shown in steps ③, ④, ⑤, and ⑥ of Figure 1, the business intelligence agent 204 obtains a prompt template from the prompt word center 206 based on the prompt identifier (prompt_id), and searches the knowledge base 208 based on the search terms and the knowledge base identifier (knowledgebase_id) to obtain search results. These search results can serve as extended context. Referring to steps ⑦ and ⑧ of Figure 1, the business intelligence agent 204 assembles the prompt based on the context, search results, and prompt template to obtain a prompt message. The prompt message and model identifier (model_id) are then sent to the inference service 209. The inference service 209 can call the corresponding AI model to perform inference and return a model-based question answer. It should be noted that the business intelligence agent 204 can also call tools, incorporating the tool call results into the prompt message to optimize its quality. Tools are not shown in Figure 1; however, in practical applications, the AI application 20 may also include tools. Similar to tools, the aforementioned prompt word center 206, knowledge base 208, and reasoning service 209 are optional parts of AI application 20. AI application 20 may also exclude prompt word center 206, knowledge base 208, or reasoning service 209.
[0091] Interactive device 202 is also used to receive model questions and answers returned by business intelligence agent 204 and return model questions and answers to the user, as shown in steps 9 and 10 in Figure 1.
[0092] To facilitate data governance, the tracking data generated by the business intelligence agent 204 can also be retrieved by the data acquisition service 102 for use in the data flywheel, as shown in the steps in Figure 1. As shown. The tracking data can include at least one of front-end tracking data or back-end tracking data. Front-end tracking data includes data collected from front-end pages (such as web pages, mobile application interfaces, etc.) through code, screenshots, or video recordings, which are then parsed to gather user behavior and page-related data. Front-end tracking data can be used to understand how users interact with the front-end interface, such as the buttons clicked, pages viewed, time spent on each page, and scrolling actions. Back-end tracking data includes data collected on the server side (back-end). Back-end tracking data records application running status, business logic execution, and database interaction information. For example, back-end tracking data can track prompts, knowledge bases, AI models, and tool calls. Front-end tracking data can be implemented through code tracking, visual tracking, and inconspicuous tracking. Inconspicuous tracking can utilize the characteristics of browsers or mobile devices to collect data. For example, using browser event monitoring can automatically record user page browsing, scrolling, and other operations. Back-end tracking data can be implemented through tracking in business logic code or using the plugin mechanism of frameworks or middleware.
[0093] Data acquisition service 102 can pull event tracking data from the logs of AI applications. Alternatively, event tracking data can be stored or temporarily stored in object storage service (OBS), distributed message queues (such as Kafka), relational databases (such as MySQL), or distributed search engines (Elastic Search, ES). For real-time data pipeline scenarios such as e-commerce order processing and log analysis, data sources (such as AI application 20) can send event tracking data to Kafka. Kafka Connect or a custom consumer can write the data to MySQL for structured storage and to ES for search / analysis. OBS is used for cold data archiving, and data acquisition service 102 can periodically back up historical data in MySQL / ES to OBS. For high-concurrency read-write separation scenarios, such as social media platform dynamics and news information systems, high-concurrency write requests are buffered through Kafka and asynchronously written to MySQL to avoid directly overwhelming the database. The read-write pattern in this scenario is usually read-heavy and write-light. Therefore, read requests are preferentially retrieved from ES, while OBS stores static files (such as product images) and access is accelerated through a content delivery network (CDN).
[0094] Next, the components of the data governance system 10 and the data governance process will be introduced.
[0095] Data acquisition service 102 is used to collect source data related to AI application 20 from at least one data source. AI application 20 includes applications built based on AI models. As shown in Figure 1, the source data includes business data, such as business knowledge and business annotation data. Business knowledge, also known as domain knowledge, includes specifications / standards, design schemes, and encyclopedia documents. Business annotation data can include at least one of automated annotation data and manually annotated data. The source data can also include tracking data, such as front-end tracking data and back-end tracking data. It should be noted that Figure 1 uses back-end tracking data as an example. In actual applications, data acquisition service 102 can also pull front-end tracking data. For example, the front-end interactive device 202 can send tracking data to the back-end business intelligence agent 204, and data acquisition service 102 can pull front-end tracking data from business intelligence agent 204. Data acquisition service 102 can also receive front-end tracking data pushed by Kafka. For example, the front-end interactive device 202 can interface with the Kafka of the data acquisition service 102. When the data acquisition service 102 subscribes to the front-end tracking data, it can receive the front-end tracking data pushed by Kafka.
[0096] The data development system 104 is used to determine whether a set of triggering conditions related to data governance is met. When a target triggering condition in the set is met, it constructs a corpus corresponding to the target triggering condition based on the source data. This corpus includes at least one of the following: reranking corpus, embedding corpus, knowledge corpus, few-shot cue corpus, or model training corpus. The data development system 104 can obtain source data from the data acquisition service 102 or from the warehouse 108, such as obtaining a dataset from the warehouse 108. The warehouse 108 may include a data warehouse and a model warehouse. The data warehouse may include historical versions of the training set, which can be mixed with the data flywheel dataset for incremental training of AI models. The data warehouse may also include knowledge sets, such as historical knowledge sets, which can be mixed with the data flywheel knowledge set for knowledge storage. The dataset obtained by the data development system 104 from the data warehouse may also include business evaluation sets and standard evaluation sets for AI model evaluation.
[0097] As shown in steps (1), (2a), (2b), (3a), (3b), and (3c) of Figure 1, when a user (such as a model developer) creates a task, the data development system 104 can obtain a dataset from the repository 108 and source data collected by the data collection service 102 through a data flywheel mechanism. The source data collected by the data development system 104 may include, but is not limited to, at least one of the following: question-and-answer request data, business context, operational metrics, tracking data, business annotation data, and business knowledge. The data development system 104 can perform data cleaning and corpus construction, and after construction, perform quality assessment on the corpus corresponding to the target triggering conditions. When the assessment results meet the requirements, the data development system 104 can transfer the corpus. For example, the data development system 104 can provide the dataset to the model development system 106 for model training. Another example is that the data development system 104 can vectorize business knowledge to obtain knowledge corpus and update the knowledge corpus to the knowledge base. For example, the data development system 104 can construct a few-sample prompt corpus based on the processed corpus, which is then used to update the prompt word center 206.
[0098] The data development system 104 is also used to determine the indicator value of the evaluation indicator based on the evaluation indicator of the corpus corresponding to the target triggering condition; when the indicator value of the evaluation indicator meets the requirements, the module of the AI application 20 is updated according to the corpus corresponding to the target triggering condition. The module of the AI application 20 includes at least one of the following: a reordering model, an embedding representation model, a knowledge base 208, a prompt word center, a model training corpus, and a model file. Among them, updating the knowledge base 208 can be updating the knowledge corpus of the knowledge base 208, updating the prompt word center can be updating the prompt text of the prompt word center, and updating the model training corpus can be updating the model training corpus in the model training corpus.
[0099] The reordering model and embedding representation model mentioned above are not shown in Figure 1. The reordering model can be retrained by the model development system 106 based on the reordered corpus, thereby updating the reordering model. Similarly, the embedding representation model can be retrained by the model development system 106 based on the embedding representation corpus, thereby updating the embedding representation model. The model training corpus can be transferred, and the model development system 106 is also used to train the model based on the model training corpus. The model training corpus may include pre-training (PT) corpus, supervised fine-tuning (SFT) corpus, or reinforcement learning (RL) corpus. It should be noted that, as shown in steps (4) and (5) of Figure 1, the model development system 106 can store the model files of the trained AI model, such as checkpoint (ckpt) files, in the model repository for developers to use in secondary development or for fault recovery. Further, as shown in step (6) of Figure 1, the above model files can also be deployed to the inference service 209. The updating of knowledge base 208 may include data development system 104 storing knowledge corpora in vector form. The updating of prompts may include data development system 104 submitting a small sample of prompt corpora to prompt word center 206, so that prompt word center 206 can update prompts based on the small sample of prompt corpora.
[0100] Based on the data governance system 10 shown in Figure 1, this application provides a data governance method. The data governance method of this application will be described in detail below with reference to the accompanying drawings.
[0101] Referring to Figure 2, a flowchart of a data governance method is shown, which includes the following steps:
[0102] S202, Data governance system 10 collects source data related to AI applications from at least one data source.
[0103] AI applications include applications built on AI models. Data sources may include at least one of the following: business data and business intelligence agents.
[0104] Business data can provide business knowledge or business-annotated data. Business knowledge, also known as domain knowledge, includes, but is not limited to, standards, design schemes, and encyclopedic documents. Business-annotated data can include automatically annotated data or manually annotated data. For example, in a code development scenario, manually annotated data can include incremental code or code added to the database.
[0105] A business intelligence agent can provide at least one of the following: question-and-answer request data, business context, operational metrics, and tracking data. Question-and-answer request data includes the original query and the post-processed, user-visible answer. Business context includes the context of user input, the context of the intelligence agent, and the context of the AI model. The intelligence agent's context may include search terms, the name of the retrieved knowledge base, and the search segment results. The model's context may include the complete prompts input into the AI model. Operational metrics (or evaluation metrics) can be divided into end-to-end metrics and retrieval metrics. End-to-end metrics may include, but are not limited to, acceptance rate, retention rate, AI-generated percentage, number of question-and-answer actions, and number / percentage of negative feedback. Retrieval metrics may include, but are not limited to, precision, recall, knowledge deadline, number of Retrieval Augmented Generation (RAG) searches, and number / percentage of clicks to top 1 / top 3 information sources. Tracking data may include user feedback tracking. User feedback tracking can include tracking data for behaviors such as liking, copying all, copying part, inserting, saving as, and clicking on the information source to jump; or tracking data for behaviors such as disliking, stopping answering, and giving negative feedback; or tracking data for behaviors such as canceling liking and canceling disliking.
[0106] S204. Data governance system 10 determines whether the set of triggering conditions related to data governance is met. When the target triggering condition in the set of triggering conditions is met, S206 is executed.
[0107] Specifically, the data governance system 10 can determine the module update mode configured by the user. The module update mode includes a single module update mode (or single module promotion mode), a multi-module update mode (or multi-module promotion mode), or a full module update mode (or full module promotion mode). The triggering conditions for different module update modes can be different. The triggering conditions for each module update mode can be a subset of the set of triggering conditions related to data governance, i.e., a subset of triggering conditions. Based on this, the data governance system 10 can determine whether the subset of triggering conditions corresponding to the module update mode is satisfied.
[0108] In some possible implementations, the set of triggering conditions includes either a first type of triggering condition or a second type of triggering condition. The first type of triggering condition includes at least one of quantitative triggering conditions or timed triggering conditions, while the second type of triggering conditions includes triggering conditions related to monitoring indicators. In some possible implementations, the subset of triggering conditions corresponding to a single-module improvement mode may include the aforementioned first and second types of triggering conditions; the subset of triggering conditions corresponding to a multi-module improvement mode may include the second type of triggering conditions; and the subset of triggering conditions corresponding to a full-module improvement mode may include the first type of triggering conditions.
[0109] The monitoring metrics can include at least one of the following: retrieval metrics, recall metrics, end-to-end metrics, knowledge cutoff time, or sorted data volume. Retrieval metrics can include precision, recall metrics can include recall rate, and end-to-end metrics can include one or more of the following: acceptance rate, retention rate, AI-generated percentage, number of question-and-answer actions, and number / percentage of negative feedback. In practical applications, the data governance system 10 can also expand to include more monitoring metrics based on business needs. For example, the data governance system 10 can provide a configuration interface through which users can customize monitoring metrics.
[0110] Accordingly, the second type of triggering conditions includes one or more of the following: the retrieval metric reaches the first threshold; the recall metric reaches the second threshold; the end-to-end metric reaches the third threshold; the knowledge deadline reaches the fourth threshold; and the sorted data volume reaches the fifth threshold.
[0111] The triggering conditions can be pre-configured. The data governance system 10 can provide a configuration interface to the user, receiving at least one of the following: rule information of the triggering rules or thresholds corresponding to monitoring indicators configured by the user through the configuration interface. The thresholds corresponding to the monitoring indicators can also be referred to as the data flywheel automated triggering thresholds. The rule information of the triggering rules includes the time interval of timed triggering rules or the data volume interval of quantitative triggering rules. The rule information of the triggering rules or the thresholds corresponding to the monitoring indicators are used to form at least one triggering condition in the triggering condition set.
[0112] The following description, in conjunction with the accompanying diagram, explains the configuration of trigger conditions or the configuration of automatic trigger thresholds for the data flywheel.
[0113] Referring to Figure 3A, which illustrates a configuration interface 300, the configuration interface 300 includes a timed trigger rule configuration control 301, a quantitative trigger rule configuration control 302, and multiple monitoring indicator configuration controls. The timed trigger rule configuration control 301 configures the time interval for timed trigger rules, and the quantitative trigger rule control 302 configures the data interval for quantitative trigger rules. The multiple monitoring indicator configuration controls may include an acceptance rate configuration control 303, a retention rate configuration control 304, an AI generation percentage configuration control 305, a negative feedback frequency configuration control 306, a negative feedback percentage configuration control 307, a model training corpus deadline time interval configuration control 308, a knowledge corpus deadline time interval 309, a rank data volume configuration control 310, a retrieval accuracy configuration control 311, a retrieval recall configuration control 312, and a top 3 information source click-through percentage configuration control 313. Furthermore, the configuration interface 300 may also include other reserved indicator configuration controls, such as other indicator configuration controls 314, 315, and 316. Users can extend custom indicators through these other indicator configuration controls according to their needs.
[0114] It should be noted that, in addition to configuring rule information and thresholds, the configuration interface 300 can also be used to orchestrate the order of triggering conditions. Specifically, the configuration interface 300 supports visual orchestration of triggering condition order. For example, when configuring triggering conditions or mechanisms for a single module's boosting mode, users can use drag-and-drop operations through the visual orchestration interface to prioritize triggering conditions such as timed triggering rules or quantitative triggering rules. After the aforementioned source data (such as business context and other business corpora) is cleaned, it can be checked whether the triggering conditions related to retrieval metrics are met, for example, whether the retrieval accuracy is met. If so, it can check whether the knowledge to be retrieved has been added to the database; if so, it can check whether the triggering conditions related to recall metrics are met; if not, it can construct and clean the knowledge corpus. Furthermore, if existing triggering conditions or rules cannot meet the requirements, the configuration interface 300 also supports adding new rules or triggering conditions.
[0115] The following is an example illustration with reference to the accompanying drawings. In the example of Figure 3A, after the user configures the rule information and thresholds and clicks "OK", the configuration interface 300 can jump to the visual orchestration interface. Referring to the schematic diagram of a visual orchestration interface shown in Figure 3B, the visual orchestration interface 400 may include a navigation area 402 and an editing area 404. The navigation area 402 includes trigger conditions generated based on the configured rule information and thresholds, such as trigger condition 1, trigger condition 2... trigger condition N. The user can drag trigger condition 1, trigger condition 3, trigger condition 4, and trigger condition 7 from the navigation area 402 to the work area, and then connect different trigger conditions in the editing area 404 to arrange the order of the trigger conditions and generate a trigger condition judgment process. When the arrangement is complete, the user can click "OK" to submit the trigger condition judgment process.
[0116] S206, Data governance system 10 constructs corpus corresponding to the target triggering conditions based on source data.
[0117] The corpus corresponding to the target triggering condition includes at least one of the following: reranking corpus, embedding corpus, knowledge corpus, few-shot cue corpus, or model training corpus. Among them, the model training corpus includes supervised fine-tuning (SFT) corpus and reinforcement learning (RL) corpus.
[0118] In practice, the data governance system 10 can first determine the actual data distribution in the production environment (also known as the live network distribution) based on the source data, and then construct the corpus corresponding to the target triggering conditions based on the actual data distribution in the production environment. Specifically, the data governance system 10 can construct prompts based on the actual data distribution in the production environment, and construct the corpus corresponding to the target triggering conditions through a language model.
[0119] S208. Data governance system 10 performs quality assessment on the corpus corresponding to the target triggering conditions and obtains the assessment results. When the assessment results meet the requirements, S210 is executed.
[0120] Quality assessment can be conducted on a single data point or on the entire data set. Assessment of a single data point can include at least one of qualitative or quantitative methods. Qualitative assessments may include evaluations of grammatical correctness, formatting, text cleanliness, content effectiveness, and unbiased viewpoints, while quantitative assessments may include quality scores or quality ratings. The evaluation metrics for quality assessments or quality ratings may differ for different data sets. Quality assessment of the entire data set may include an assessment of the total amount of data, as well as distribution and diversity assessments (also known as data source tracing).
[0121] The distribution assessment and diversity assessment are explained in detail below.
[0122] For distribution evaluation, the data governance system 10 can obtain the corpus distribution corresponding to the target triggering conditions, and obtain the evaluation result based on the corpus distribution and the real data distribution in the production environment. The evaluation result includes information on the difference between the corpus distribution and the real data distribution. It should be noted that the difference between the corpus distribution and the real data distribution can be measured using information theory dimensions or distance. Based on this, the difference information can include Kullback-Leibler Divergence, Jensen-Shannon Divergence, and Cross Entropy, or it can include Wasserstein Distance and Maximum Mean Discrepancy (MMD).
[0123] For diversity assessment, the data governance system 10 can acquire the data sources of the corpus corresponding to the target triggering conditions. The data sources can include data source type or data source system. The data governance system 10 can obtain assessment results based on the data sources. The assessment results include the data diversity of the corpus.
[0124] In practical applications, the data governance system 10 can assess the total amount and distribution of data, and perform data source tracing orthogonal determination and diversity assessment to obtain high-quality, diverse corpora. Specifically, when the total amount of data reaches a threshold, the data governance system 10 can determine the difference between the actual data distribution in the production environment and the corpus distribution. When the difference between the corpus distribution and the actual data distribution is less than or equal to a first preset value, the data governance system 10 obtains the data source of the corpus. Then, the data governance system 10 can determine the data diversity based on the data source of the corpus. When the data diversity is greater than or equal to a second preset value, it indicates that the corpus quality meets the requirements, and the data governance system 10 can adopt the corpus.
[0125] To facilitate understanding, an example will be used as an illustration below.
[0126] Referring to Figure 4, a flowchart of data distribution assessment and diversity assessment is provided. The data governance system 10 collects data from various data sources. When the total amount of original or newly added operational data reaches a threshold, the data governance system 10 can obtain the actual distribution of the current network. Specifically, this includes obtaining the actual distribution of "likes / dislikes" in actual user Q&A / retrieval / calls, the actual distribution of "domains / scenes / tasks / entities" in actual user Q&A, the actual distribution of "tool calls" in actual user Q&A, and the actual distribution of "knowledge base retrieval" in actual user Q&A. Then, the data governance system 10 can construct corpora based on the actual distribution of the current network. For example, it can construct SFT / RL corpora based on the actual distribution of "likes / dislikes," SFT corpora based on the actual distribution of "domains / scenes / tasks / entities," SFT corpora based on the actual distribution of "tool calls," and knowledge / embedding corpora based on the actual distribution of "knowledge base retrieval." The data governance system 10 can then compare the corpus distribution with the actual distribution. When the difference is less than or equal to a first preset value, the system can acquire business data source types, business data source systems, and open-source data sources. Business data source types can include handwritten construction, model generation, or data flywheel; business data source systems can include Req, Defect, Design, Report, or Pipeline; and open-source data sources can include communities, forums, and encyclopedia web pages. The data governance system 10 can determine the data diversity of the candidate corpus based on the business data source types, business data source systems, and open-source data sources. When the data diversity is greater than or equal to a second preset value, the data governance system 10 can adopt the aforementioned candidate corpus as the final corpus. It should be noted that Figure 4 uses a DevOps-related business example for illustration; different business types may require different business data source systems.
[0127] S210, the data governance system 10 updates the AI application module based on the corpus corresponding to the target triggering conditions.
[0128] The AI application's modules include at least one of the following: a re-ranking model, an embedding representation model, a knowledge base, prompts, a model training corpus, and model files. The corpus corresponding to the target triggering condition may include re-rank corpus, embedding corpus, knowledge corpus, few-shot prompt corpus, and training corpus. The data governance system 10 can update the re-ranking model based on the re-rank corpus, or update the knowledge base based on the knowledge corpus based on the embedding corpus, update the prompt template or prompt based on the few-shot prompt corpus, update the model training corpus based on the training corpus, or update the model files based on the training corpus.
[0129] Furthermore, when a user enables automated data governance, such as activating a data flywheel, the data governance system 10 can also provide the user with the changing trends of the AI application's monitoring metrics after enabling automated data governance. Enabling automated data governance or the data flywheel can be achieved through a configuration interface. For example, the configuration interface may include an automated data governance switch, which the user can turn on to enable the aforementioned functions. As shown in Figure 5, the data governance system 10 can use an automated trigger indicator dashboard to monitor the real-time status of the AI application's monitoring metrics after enabling automated data governance or the optimization status of the monitoring metrics after the data flywheel is started. Figure 5 illustrates the changing trends of various monitoring metrics over time, including the changing trends of monitoring metrics after enabling automated governance.
[0130] For security and reliability reasons, the data governance system 10 can present monitoring indicators to the user before updating the AI application modules based on the corpus corresponding to the target trigger conditions. Once the user confirms the update, the data governance system 10 can update the AI application modules based on the corpus corresponding to the target trigger conditions.
[0131] To make the technical solution of this application clearer and easier to understand, the data governance methods of single-module promotion mode, multi-module promotion mode and all-module promotion mode are introduced below with reference to the accompanying drawings.
[0132] Referring to Figure 6, a flowchart of a data governance method under a single-module enhancement model is shown. The method includes the following steps:
[0133] S602, Data governance system 10 collects source data related to AI applications from at least one data source.
[0134] In the example in Figure 6, the data source can include business knowledge such as specifications and standards, design schemes, encyclopedia documents, Q&A request data, business context, operational metrics, user feedback points, manually annotated data, and evaluation sets. The evaluation set can include business evaluation sets or standard evaluation sets.
[0135] S604. Data governance system 10 determines whether the timed triggering condition or the quantitative triggering condition is met. If so, data governance system 10 executes S606.
[0136] Specifically, the data governance system 10 can obtain the time difference between the current time and the last time the data flywheel was triggered, or obtain the accumulated data volume between the current time and the last time the data flywheel was triggered. When the time difference meets a pre-configured time interval, it indicates that the timed triggering condition is met. When the data volume meets a pre-configured data volume interval, it indicates that the quantitative triggering condition is met.
[0137] S606, Data Governance System 10 performs corpus cleaning on the collected source data.
[0138] The corpus cleaning process may include at least one of the following: deduplication, compliance, language filtering, removal of private / malicious / sensitive data, replacement of private data, deletion of low-quality data, normalization, and standardization. Deduplication may include hash-based deduplication of code, articles, and paragraphs, and ngram-based deduplication. Compliance may include license compliance. Language filtering may include programming language filtering and internationalized language filtering. In the removal of private / malicious / sensitive data, private data may include personal accounts / passwords, phone numbers, and email addresses; malicious data may include, but is not limited to, malicious code; and sensitive data may include sensitive words. Low-quality data may include low-quality code or low-quality documents. Low-quality code may include excessively short code. Normalization may include the standardization of symbols and spaces. Normalization may include the standardization of business terminology.
[0139] S608. Data governance system 10 determines whether the triggering conditions related to the retrieval indicators are met. If yes, execute S610; otherwise, execute S634.
[0140] The triggering conditions related to the retrieval metrics may include the retrieval metrics reaching a first threshold. In some examples, the data governance system 10 can obtain the retrieval accuracy and determine whether the retrieval accuracy has reached the first threshold.
[0141] S610, Data Governance System 10 determines whether the knowledge to be retrieved has been stored in the database. If yes, execute S612; if no, execute S636.
[0142] S612. Data governance system 10 determines whether the triggering conditions related to the recall indicators are met. If yes, execute S614; if no, execute S624.
[0143] S614, Data Governance System 10 constructs rerank corpus.
[0144] S616 and Data Governance System 10 perform quality assessment of the rerank corpus.
[0145] S618 and Data Governance System 10 are used to train the rerank model.
[0146] The data governance system 10 can train a rerank model using rerank corpora that meet quality requirements. These quality-compliant rerank corpora include corpora whose distribution differs from the true distribution by less than a first preset value and whose data diversity exceeds a second preset value.
[0147] S620 and Data Governance System 10 were evaluated using the Rerank model.
[0148] Specifically, the data governance system 10 can use evaluation sets obtained from data sources, such as business evaluation sets or standard evaluation sets, to evaluate the trained rerank model.
[0149] S622, Data Governance System 10 updates the rerank model.
[0150] Once the evaluation is passed, the data governance system 10 can update the rerank model.
[0151] S624, Data Governance System 10 constructs embedding corpus.
[0152] S626, Data Governance System 10 performs embedding corpus quality assessment.
[0153] S628 and Data Governance System 10 train the embedding model.
[0154] The S630 and Data Governance System 10 were evaluated using an embedding model.
[0155] S632, Data Governance System 10 updates the embedding model.
[0156] The specific implementation of S624 to S632 can be found in the description of rerank-related content in S614 to S622, such as the construction of rerank corpus, quality assessment, and the training, evaluation, and updating of rank model. These details will not be repeated here.
[0157] S634. Data governance system 10 determines whether the knowledge cutoff time interval is greater than the third threshold. If yes, execute S636; if no, execute S642.
[0158] S636, Data Governance System 10 constructs knowledge corpus.
[0159] S638, Data Governance System 10 performs vectorized data entry into the database.
[0160] Specifically, the data governance system 10 can use an embedding model to vectorize the knowledge corpus to obtain knowledge vectors, and then store the knowledge vectors in the database.
[0161] S640 and Data Governance System 10 update their knowledge base.
[0162] S642. Data governance system 10 determines whether the end-to-end metric is greater than the fourth threshold. If not, proceed to S644; if yes, proceed to S650 to S656.
[0163] S644, Data Governance System 10 constructs a few-sample prompt corpus.
[0164] S646, Data Governance System 10 performs quality assessment of the few-sample prompt corpus.
[0165] S648 and Data Governance System 10 updated their prompts.
[0166] S650 and data governance system 10 construct SFT corpus.
[0167] The data governance system 10 can also perform quality assessment, model training, evaluation, and updates on the SFT corpus, as shown in S660 to S666.
[0168] S652, Data Governance System 10 performs corpus dumping.
[0169] S654, Data Governance System 10 updates the model training corpus.
[0170] S656. Data governance system 10 determines whether the amount of sorted data has reached the fifth threshold. If so, execute S658.
[0171] S658, Data Governance System 10 constructs RL corpus.
[0172] The S660 and data governance system 10 were used to evaluate the quality of the training corpus.
[0173] S662 and Data Governance System 10 perform SFT+RL model training.
[0174] S664, Data Governance System 10 conducts SFT+RL model evaluation.
[0175] S666, Data Governance System 10 updates model files.
[0176] The specific implementation of S658 to S666 is described in the relevant content of S614 to S622, such as the construction of rerank corpus, quality assessment and the training, evaluation and updating of rank model, which will not be repeated here.
[0177] The primary objective of the embodiment shown in Figure 6 is to optimize a single bottleneck module that constrains the application experience of large-scale models at a lower cost. Building upon this, and disregarding minimum cost constraints, multiple modules of the AI application can be optimized simultaneously based on user-configured data flywheel thresholds. A detailed description follows with reference to the accompanying drawings.
[0178] Referring to Figure 7, a flowchart of a data governance method under a multi-module enhancement model is shown. The method includes the following steps:
[0179] S702, Data Governance System 10 collects source data related to AI applications from at least one data source.
[0180] S704. Data governance system 10 determines whether the triggering conditions related to the recall indicators are met. If so, execute S714 to S722.
[0181] S706, Data Governance System 10 determines whether the triggering conditions related to the retrieval indicators are met. If so, execute S724 to S732.
[0182] S708, Data Governance System 10 determines whether the knowledge cutoff time interval is greater than the third threshold. If so, execute S734 to S738.
[0183] S710, Data Governance System 10 determines whether the end-to-end metric is greater than the fourth threshold. If not, execute S740 to S744; if yes, execute S746 to S750.
[0184] S712, Data governance system 10 determines whether the amount of sorted data has reached the fifth threshold. If so, execute S752 to S760.
[0185] S714, Data Governance System 10 constructs rerank corpus.
[0186] S716 and Data Governance System 10 perform quality assessment of the rerank corpus.
[0187] S718 and Data Governance System 10 are used to train the rerank model.
[0188] The S720 and Data Governance System 10 were evaluated using the Rerank model.
[0189] S722, Data Governance System 10 performs rerank model updates.
[0190] S724, Data Governance System 10 constructs embedding corpus.
[0191] S726, Data Governance System 10 performs embedding corpus quality assessment.
[0192] S728 and Data Governance System 10 train the embedding model.
[0193] The S730 and Data Governance System 10 were used to evaluate the embedding model.
[0194] S732, Data Governance System 10 updates the embedding model.
[0195] S734, Data Governance System 10 constructs knowledge corpus.
[0196] S736, Data Governance System 10 performs vectorized data entry into the database.
[0197] S738 and Data Governance System 10 update the knowledge base.
[0198] S740 and data governance system 10 construct a few-sample prompt corpus.
[0199] S742, Data Governance System 10 performs quality assessment of the few-sample prompt corpus.
[0200] S744 and Data Governance System 10 updated their prompts.
[0201] S746, Data Governance System 10 constructs SFT corpus.
[0202] S748 and Data Governance System 10 perform corpus dumping.
[0203] S750 and Data Governance System 10 update the model training corpus.
[0204] S752, Data Governance System 10 constructs RL corpus.
[0205] S754 and Data Governance System 10 are used to evaluate the quality of the training corpus.
[0206] S756 and Data Governance System 10 are used to train SFT+RL models.
[0207] S758 and Data Governance System 10 were evaluated using the SFT+RL model.
[0208] S760 and Data Governance System 10 update model files.
[0209] In some possible implementations, the data governance system 10 also supports the enhancement of all modules. This is explained below with reference to the accompanying drawings.
[0210] Referring to Figure 8, which shows a flowchart of a data governance method under a full module enhancement mode, the method includes the following steps:
[0211] S802, Data Governance System 10 collects source data related to AI applications from at least one data source.
[0212] S804. Data governance system 10 determines whether the timed triggering condition or the quantitative triggering condition is met. If so, data governance system 10 executes S806.
[0213] S806, the data governance system 10 performs corpus cleaning on the collected source data.
[0214] S808 and Data Governance System 10 construct the rerank corpus.
[0215] S810 and Data Governance System 10 perform quality assessment of the rerank corpus.
[0216] S812 and Data Governance System 10 are used to train the rerank model.
[0217] S814 and Data Governance System 10 are used for Rerank model evaluation.
[0218] S816 and Data Governance System 10 perform rerank model updates.
[0219] S818 and Data Governance System 10 construct embedding corpora.
[0220] S820 and Data Governance System 10 conduct embedding corpus quality assessment.
[0221] S822 and Data Governance System 10 train the embedding model.
[0222] S824 and Data Governance System 10 conduct embedding model evaluation.
[0223] S826, Data Governance System 10 updates the embedding model.
[0224] S828, Data Governance System 10 constructs knowledge corpus.
[0225] S830 and Data Governance System 10 perform vectorized data entry into the database.
[0226] S832, Data Governance System 10 updates the knowledge base.
[0227] S834, Data Governance System 10 constructs a few-sample prompt corpus.
[0228] S836 and Data Governance System 10 perform quality assessment of the few-sample prompt corpus.
[0229] S838 and Data Governance System 10 updated their prompts.
[0230] S840 and data governance system 10 construct SFT corpus.
[0231] S842, Data Governance System 10 performs corpus dumping.
[0232] S844, Data Governance System 10 updates the model training corpus.
[0233] S846, Data Governance System 10 constructs RL corpus.
[0234] S848 and Data Governance System 10 are used to evaluate the quality of the training corpus.
[0235] S850 and Data Governance System 10 are used to train SFT+RL models.
[0236] S852, Data Governance System 10 conducts SFT+RL model evaluation.
[0237] S854, Data Governance System 10 updates model files.
[0238] The specific implementation of the relevant steps in the embodiments shown in Figures 7 and 8 can be found in the description of the relevant content in the foregoing embodiments, and will not be repeated here.
[0239] Based on the aforementioned data governance methods, this application also provides a data governance system 10. The data governance system provided in this application will be described below from the perspective of functional modularity.
[0240] Referring to Figure 9, which shows a schematic diagram of the structure of a data governance system, the data governance system 10 includes:
[0241] The acquisition unit 902 is used to acquire source data related to artificial intelligence (AI) applications from at least one data source, the AI applications including applications built based on AI models;
[0242] Triggering condition determination unit 904 is used to determine whether the set of triggering conditions related to data governance is met;
[0243] The corpus construction unit 906 is used to construct a corpus corresponding to the target triggering condition based on the source data when the target triggering condition in the set of triggering conditions is met. The corpus corresponding to the target triggering condition includes at least one of the following: reordering corpus, embedding representation corpus, knowledge corpus, few-shot prompting corpus, or model training corpus.
[0244] The evaluation unit 908 is used to perform quality evaluation on the corpus corresponding to the target triggering condition and obtain the evaluation result.
[0245] The update unit 909 is used to update the module of the AI application according to the corpus corresponding to the target triggering condition when the evaluation result meets the requirements. The module of the AI application includes at least one of the following: a reordering model, an embedding representation model, a knowledge base, a prompt word center, a model training corpus, and a model file.
[0246] For example, the acquisition unit 902, trigger condition determination unit 904, corpus construction unit 906, evaluation unit 908, and update unit 909 described above can be implemented in hardware or software. The acquisition unit 902, trigger condition determination unit 904, corpus construction unit 906, evaluation unit 908, and update unit 909 can be units within the data development system 104. In some examples, the update unit 909 can also be a unit within the model development system 106.
[0247] When implemented through software, the acquisition unit 902, trigger condition determination unit 904, corpus construction unit 906, evaluation unit 908, and update unit 909 can be applications running on computing devices, such as computing engines. These applications can also be virtualized and provided to users as virtualization services. Virtualization services can include virtual machine (VM) services, bare metal server (BMS) services, or container services. VM services can be services that use virtualization technology to create virtual machine (VM) resource pools on multiple physical hosts to provide VMs for users to use on demand. BMS services are services that use virtualization technology to create BMS resource pools on multiple physical hosts to provide BMS for users to use on demand. Container services are services that use virtualization technology to create container resource pools on multiple physical hosts to provide containers for users to use on demand. A VM is a simulated virtual computer, that is, a logical computer. A BMS is a scalable, high-performance computing service with computing performance indistinguishable from traditional physical machines and features secure physical isolation. Containers are a kernel virtualization technology that provides lightweight virtualization to isolate user space, processes, and resources. It should be understood that the VM service, BMS service, and container service mentioned above are merely specific examples. In practical applications, virtualization services can also include other lightweight or heavyweight virtualization services, which are not specifically limited here.
[0248] When implemented in hardware, the acquisition unit 902, trigger condition determination unit 904, corpus construction unit 906, evaluation unit 908, and update unit 909 may include at least one computing device, such as a server. Alternatively, the acquisition unit 902, trigger condition determination unit 904, corpus construction unit 906, evaluation unit 908, and update unit 909 may also be implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
[0249] In some possible implementations, the data governance system 10 also includes:
[0250] The mode determination unit 903 is used to determine the module update mode configured by the user, wherein the module update mode includes a single module update mode, a multi-module update mode, or a full module update mode.
[0251] The trigger condition determination unit 904 is specifically used for:
[0252] Determine whether the subset of triggering conditions corresponding to the module update mode is satisfied, wherein the subset of triggering conditions is a subset of the set of triggering conditions related to data governance.
[0253] Similar to the trigger condition determination unit 904, the pattern determination unit 903 can be a unit within the data development system 104. The pattern determination unit 903 can be implemented in software or hardware. When implemented in software, the pattern determination unit 903 can be an application running on a computing device. This application can also be provided to users as a virtualization service such as a BMS, VM, or container. When implemented in hardware, the pattern determination unit 903 can include at least one computing device, such as a server. Alternatively, the pattern determination unit 903 can also be a device implemented using an ASIC or a PLD.
[0254] In some possible implementations, the set of triggering conditions includes a first type of triggering conditions or a second type of triggering conditions. The first type of triggering conditions includes at least one of quantitative triggering conditions or timed triggering conditions, and the second type of triggering conditions includes triggering conditions related to monitoring indicators.
[0255] In some possible implementations, the monitoring metrics include at least one of retrieval metrics, recall metrics, end-to-end metrics, knowledge cutoff time, or sorted data volume;
[0256] The second type of triggering conditions includes one or more of the following:
[0257] The search index reaches a first threshold; or...
[0258] The recall metric reaches the second threshold; or...
[0259] The end-to-end metric reaches the third threshold; or...
[0260] The knowledge deadline has reached the fourth threshold; or...
[0261] The amount of sorted data has reached the fifth threshold.
[0262] In some possible implementations, the data governance system 10 also includes:
[0263] Configuration unit 905 is used to provide a configuration interface to the user; receive at least one of the rule information of the triggering rule or the threshold corresponding to the monitoring indicator configured by the user through the configuration interface, wherein the rule information of the triggering rule includes the time interval of the timed triggering rule or the data volume interval of the quantitative triggering rule, and the rule information of the triggering rule or the threshold corresponding to the monitoring indicator is used to form at least one triggering condition in the set of triggering conditions.
[0264] The configuration unit 905 can be a unit within the data development system 104. The configuration unit 905 can be implemented in software or hardware. When implemented in software, the configuration unit 905 can be an application running on a computing device. This application can also be provided to users as a virtualization service such as a BMS, VM, or container. When implemented in hardware, the configuration unit 905 can include at least one computing device, such as a server. Alternatively, the configuration unit 905 can also be a device implemented using an ASIC or a PLD.
[0265] In some possible implementations, the data governance system 10 also includes:
[0266] The governance result providing unit 907 is used to provide the user with the changing trend of the monitoring indicators of the AI application after the user enables automated data governance.
[0267] The governance result providing unit 907 can be a unit within the data development system 104. The governance result providing unit 907 can be implemented via software or hardware. When implemented via software, the governance result providing unit 907 can be an application running on a computing device. This application can also be provided to users as a virtualization service such as a BMS, VM, or container. When implemented via hardware, the governance result providing unit 907 can include at least one computing device, such as a server. Alternatively, the governance result providing unit 907 can also be a device implemented using an ASIC or a PLD.
[0268] In some possible implementations, the corpus construction unit 906 is specifically used for:
[0269] Based on the source data, determine the actual data distribution in the production environment;
[0270] Based on the actual data distribution of the production environment, construct a corpus corresponding to the target triggering conditions.
[0271] In some possible implementations, the evaluation unit 908 is specifically used for:
[0272] Obtain the corpus distribution corresponding to the target triggering condition;
[0273] An evaluation result is obtained based on the corpus distribution and the actual data distribution in the production environment. The evaluation result includes information on the differences between the corpus distribution and the actual data distribution.
[0274] In some possible implementations, the evaluation unit 908 is specifically used for:
[0275] Obtain the data source of the corpus corresponding to the target triggering condition;
[0276] Based on the data sources, an evaluation result is obtained, which includes the data diversity of the corpus.
[0277] This application also provides a computing device 1000. As shown in FIG10, the computing device 1000 includes: a bus 1002, a processor 1004, a memory 1006, and a communication interface 1008. The processor 1004, the memory 1006, and the communication interface 1008 communicate with each other via the bus 1002. The computing device 1000 can be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in the computing device 1000.
[0278] Bus 1002 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, only one line is used in Figure 10, but this does not imply that there is only one bus or one type of bus. Bus 1002 can include pathways for transmitting information between various components of computing device 1000 (e.g., memory 1006, processor 1004, communication interface 1008).
[0279] The processor 1004 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).
[0280] The memory 1006 may include volatile memory, such as random access memory (RAM). The memory 1006 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid-state drive (SSD). The memory 1006 stores executable program code, which the processor 1004 executes to implement the aforementioned data governance method. Specifically, the memory 1006 stores instructions for the data governance system 10 to execute the data governance method. For example, the memory 1006 may store instructions for implementing the functions of the acquisition unit 902, the trigger condition determination unit 904, the corpus construction unit 906, the evaluation unit 908, and the update unit 909. Furthermore, the memory 1006 may also store instructions for implementing the functions of the pattern determination unit 903, the configuration unit 905, and the governance result providing unit 907.
[0281] The communication interface 1008 uses transceiver modules such as, but not limited to, network interface cards and transceivers to enable communication between the computing device 1000 and other devices or communication networks.
[0282] This application also provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.
[0283] As shown in Figure 11, the computing device cluster includes at least one computing device 1000. The memory 1006 of one or more computing devices 1000 in the computing device cluster may store instructions from the same data governance system 10 for executing data governance methods.
[0284] In some possible implementations, one or more computing devices 1000 in the computing device cluster can also be used to execute some of the instructions used by the data governance system 10 to execute data governance methods. In other words, a combination of one or more computing devices 1000 can jointly execute the instructions used by the data governance system 10 to execute data governance methods.
[0285] It should be noted that the memory 1006 in different computing devices 1000 in the computing device cluster can store different instructions for executing some functions of the data governance system 10.
[0286] Figure 12 illustrates one possible implementation. As shown in Figure 12, two computing devices 1000A and 1000B are connected via a communication interface 1008. The memory in computing device 1000A stores instructions for executing the functions of the acquisition unit 902 and the trigger condition determination unit 904. The memory in computing device 1000B stores instructions for executing the functions of the corpus construction unit 906, the evaluation unit 908, and the update unit 909. Furthermore, the memory of computing device 1000A may also store instructions for executing the functions of the pattern determination unit 903 and the configuration unit 905, and the memory of computing device 1000B may also store instructions for executing the functions of the governance result providing unit 907. In other words, the memories 1006 of computing devices 1000A and 1000B jointly store the instructions used by the data governance system 10 to execute the data governance method.
[0287] The connection method between the computing device clusters shown in Figure 12 can be considered because the data governance method provided in this application requires a lot of resources for corpus construction, governance evaluation, and AI application module updates. Therefore, it is considered that the functions implemented by the corpus construction unit 906, evaluation unit 908, and update unit 909 and the functions implemented by the acquisition unit 902 and trigger condition determination unit 904 are respectively executed by different computing devices. For example, the functions implemented by the acquisition unit 902 and trigger condition determination unit 904 can be executed by computing device 1000A, and the functions implemented by the corpus construction unit 906, evaluation unit 908, and update unit 909 can be executed by computing device 1000B. Similarly, the functions implemented by the pattern determination unit 903 and configuration unit 905 can be executed by computing device 1000A, and the functions implemented by the governance result providing unit 907 can be executed by computing device 1000B.
[0288] It should be understood that the functions of computing device 1000A shown in Figure 12 can also be performed by multiple computing devices 1000. Similarly, the functions of computing device 1000B can also be performed by multiple computing devices 1000.
[0289] In some possible implementations, one or more computing devices in a computing device cluster can be connected via a network. This network can be a wide area network (WAN) or a local area network (LAN), etc. Figure 13 illustrates one possible implementation. As shown in Figure 13, two computing devices 1000C and 1000D are connected via a network. Specifically, they are connected to the network through communication interfaces in each computing device. In this type of possible implementation, the memory 1006 in computing device 1000C stores instructions for executing the functions of the acquisition unit 902 and the trigger condition determination unit 904. Simultaneously, the memory 1006 in computing device 1000D stores instructions for executing the functions of the corpus construction unit 906, the evaluation unit 908, and the update unit 909. Further, the memory of computing device 1000C can also store instructions for executing the functions of the pattern determination unit 903 and the configuration unit 905, and the memory of computing device 1000D can also store instructions for executing the functions of the governance result providing unit 907.
[0290] The connection method between the computing device clusters shown in Figure 13 can be considered in light of the fact that the data governance method provided in this application requires a large amount of resources for corpus construction, quality assessment, and AI application module updates. Therefore, the functions implemented by the acquisition unit 902 and the trigger condition determination unit 904 are considered to be executed by the computing device 1000C, and the functions implemented by the corpus construction unit 906, the evaluation unit 908, and the update unit 909 are executed by the computing device 1000D. In addition, when the data governance system 10 also includes an execution mode determination unit 903, a configuration unit 905, or a governance result providing unit 907, the functions implemented by the mode determination unit 903 and the configuration unit 905 are executed by the computing device 1000C, and the functions implemented by the governance result providing unit 907 are executed by the computing device 1000D.
[0291] It should be understood that the functions of computing device 1000C shown in Figure 13 can also be performed by multiple computing devices 1000. Similarly, the functions of computing device 1000D can also be performed by multiple computing devices 1000.
[0292] This application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that a computing device can store, or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive). The computer-readable storage medium includes instructions that instruct the computing device to execute the data governance method described above in the data governance system 10.
[0293] This application also provides a computer program product containing instructions. The computer program product may be a software or program product containing instructions, capable of running on a computing device or stored on any usable medium. When the computer program product is run on at least one computing device, it causes the at least one computing device to perform the data governance method described above.
[0294] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of this application.
Claims
1. A data governance method, characterized in that, The method includes: Source data related to artificial intelligence (AI) applications are collected from at least one data source, including applications built based on AI models; Determine whether the set of trigger conditions related to data governance has been met; When the target trigger condition in the set of trigger conditions is met, a corpus corresponding to the target trigger condition is constructed based on the source data. The corpus corresponding to the target trigger condition includes at least one of the following: reordering corpus, embedding representation corpus, knowledge corpus, few-shot prompt corpus, or model training corpus. The quality of the corpus corresponding to the target triggering condition is evaluated to obtain the evaluation results; When the evaluation result meets the requirements, the modules of the AI application are updated according to the corpus corresponding to the target triggering condition. The modules of the AI application include at least one of the following: a reordering model, an embedding representation model, a knowledge base, a prompt word center, a model training corpus, and a model file.
2. The method according to claim 1, characterized in that, The method further includes: Determine the module update mode configured by the user, which includes a single module update mode, a multi-module update mode, or a full module update mode; The determination of whether the set of triggering conditions related to data governance is met includes: Determine whether the subset of triggering conditions corresponding to the module update mode is satisfied, wherein the subset of triggering conditions is a subset of the set of triggering conditions related to data governance.
3. The method according to claim 1 or 2, characterized in that, The set of triggering conditions includes a first type of triggering conditions or a second type of triggering conditions. The first type of triggering conditions includes at least one of quantitative triggering conditions or timed triggering conditions, and the second type of triggering conditions includes triggering conditions related to monitoring indicators.
4. The method according to claim 3, characterized in that, The monitoring indicators include at least one of the following: retrieval indicators, recall indicators, end-to-end indicators, knowledge cutoff time, or sorted data volume. The second type of triggering conditions includes one or more of the following: The search index reaches a first threshold; or... The recall metric reaches the second threshold; or... The end-to-end metric reaches the third threshold; or... The knowledge deadline has reached the fourth threshold; or... The amount of sorted data has reached the fifth threshold.
5. The method according to any one of claims 1 to 4, characterized in that, The method further includes: Provide a configuration interface for users; The system receives at least one of the rule information of the triggering rule configured by the user through the configuration interface or the threshold corresponding to the monitoring indicator. The rule information of the triggering rule includes the time interval of the timed triggering rule or the data volume interval of the quantitative triggering rule. The rule information of the triggering rule or the threshold corresponding to the monitoring indicator is used to form at least one triggering condition in the set of triggering conditions.
6. The method according to any one of claims 1 to 5, characterized in that, The method further includes: When a user enables automated data governance, the system provides the user with the changing trends of the monitoring metrics of the AI application after enabling automated data governance.
7. The method according to any one of claims 1 to 6, characterized in that, The step of constructing the corpus corresponding to the target triggering condition based on the source data includes: Based on the source data, determine the actual data distribution in the production environment; Based on the actual data distribution of the production environment, construct a corpus corresponding to the target triggering conditions.
8. The method according to any one of claims 1 to 7, characterized in that, The quality assessment of the corpus corresponding to the target triggering condition, and the obtaining of the assessment results, includes: Obtain the corpus distribution corresponding to the target triggering condition; An evaluation result is obtained based on the corpus distribution and the actual data distribution in the production environment. The evaluation result includes information on the differences between the corpus distribution and the actual data distribution.
9. The method according to any one of claims 1 to 8, characterized in that, The quality assessment of the corpus corresponding to the target triggering condition, and the obtaining of the assessment results, includes: Obtain the data source of the corpus corresponding to the target triggering condition; Based on the data sources, an evaluation result is obtained, which includes the data diversity of the corpus.
10. A data governance system, characterized in that, The data governance system includes: A data acquisition unit is used to acquire source data related to artificial intelligence (AI) applications from at least one data source, including applications built based on AI models. The trigger condition determination unit is used to determine whether the set of trigger conditions related to data governance is met; The corpus construction unit is used to construct a corpus corresponding to the target triggering condition based on the source data when the target triggering condition in the set of triggering conditions is met. The corpus corresponding to the target triggering condition includes at least one of the following: reordering corpus, embedding representation corpus, knowledge corpus, few-shot prompting corpus, or model training corpus. An evaluation unit is used to perform quality evaluation on the corpus corresponding to the target triggering condition and obtain the evaluation result. An update unit is used to update the modules of the AI application according to the corpus corresponding to the target triggering condition when the evaluation result meets the requirements. The modules of the AI application include at least one of a reordering model, an embedding representation model, a knowledge base, a prompt word center, a model training corpus, and a model file.
11. The system according to claim 10, characterized in that, The system also includes: The mode determination unit is used to determine the module update mode configured by the user, wherein the module update mode includes a single module update mode, a multi-module update mode, or a full module update mode. The triggering condition determination unit is specifically used for: Determine whether the subset of triggering conditions corresponding to the module update mode is satisfied, wherein the subset of triggering conditions is a subset of the set of triggering conditions related to data governance.
12. The system according to claim 10 or 11, characterized in that, The set of triggering conditions includes a first type of triggering conditions or a second type of triggering conditions. The first type of triggering conditions includes at least one of quantitative triggering conditions or timed triggering conditions, and the second type of triggering conditions includes triggering conditions related to monitoring indicators.
13. The system according to claim 12, characterized in that, The monitoring indicators include at least one of the following: retrieval indicators, recall indicators, end-to-end indicators, knowledge cutoff time, or sorted data volume. The second type of triggering conditions includes one or more of the following: The search index reaches a first threshold; or... The recall metric reaches the second threshold; or... The end-to-end metric reaches the third threshold; or... The knowledge deadline has reached the fourth threshold; or... The amount of sorted data has reached the fifth threshold.
14. The system according to any one of claims 10 to 13, characterized in that, The system also includes: A configuration unit is used to provide a configuration interface to the user; receive at least one of the rule information of the triggering rule or the threshold corresponding to the monitoring indicator configured by the user through the configuration interface, wherein the rule information of the triggering rule includes the time interval of the timed triggering rule or the data volume interval of the quantitative triggering rule, and the rule information of the triggering rule or the threshold corresponding to the monitoring indicator is used to form at least one triggering condition in the set of triggering conditions.
15. The system according to any one of claims 10 to 14, characterized in that, The system also includes: The governance result providing unit is used to provide the user with the changing trend of the monitoring indicators of the AI application after the user enables automated data governance.
16. The system according to any one of claims 10 to 15, characterized in that, The corpus construction unit is specifically used for: Based on the source data, determine the actual data distribution in the production environment; Based on the actual data distribution of the production environment, construct a corpus corresponding to the target triggering conditions.
17. The system according to any one of claims 10 to 16, characterized in that, The evaluation unit is specifically used for: Obtain the corpus distribution corresponding to the target triggering condition; An evaluation result is obtained based on the corpus distribution and the actual data distribution in the production environment. The evaluation result includes information on the differences between the corpus distribution and the actual data distribution.
18. The system according to any one of claims 10 to 17, characterized in that, The evaluation unit is specifically used for: Obtain the data source of the corpus corresponding to the target triggering condition; Based on the data sources, an evaluation result is obtained, which includes the data diversity of the corpus.
19. A computing device cluster, characterized in that, The computing device cluster includes at least one computing device, the at least one computing device including at least one processor and at least one memory, the at least one memory storing computer-readable instructions; the at least one processor executes the computer-readable instructions to cause the computing device cluster to perform the data governance method as described in any one of claims 1 to 9.
20. A computer-readable storage medium, characterized in that, Includes computer-readable instructions; the computer-readable instructions are used to implement the data governance method according to any one of claims 1 to 9.
21. A computer program product, characterized in that, Includes computer-readable instructions; the computer-readable instructions are used to implement the data governance method according to any one of claims 1 to 9.