Data management method and system based on AI edge device and computing device

By dividing data fields on edge devices and generating data maps, the problem that centralized data management is not suitable for edge data management is solved, and efficient and secure data management and retrieval is achieved, which is suitable for AI edge devices.

CN120448116APending Publication Date: 2025-08-08XFUSION DIGITAL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510541456.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

The existing centralized data management method is not suitable for edge data management, resulting in problems such as latency and low management efficiency.

Method used

Divide the data source into the data domain, and generate a data map for each data domain. Create a data map through metadata association relationships, generate a global data directory, compatible with structured and unstructured data, and is suitable for AI edge devices.

Benefits of technology

It realizes that without centrally storing data, it can effectively manage and retrieve edge data, reduce data transmission delay, improve data management efficiency and security, and is suitable for a variety of data source scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120448116A_ABST
    Figure CN120448116A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a data management method and system based on an AI edge device and a computing device. The method comprises the steps that a first number of data sources are divided into a second number of data fields, and at least one data field comprises an AI edge data source and a non-AI edge data source; extracting metadata for the data field; based on the metadata, a data map corresponding to each data field is generated, and node information of at least one data map is associated with the structured metadata and the unstructured metadata at the same time; and generating a global data directory based on the data maps corresponding to the second number of data fields. According to the embodiment of the invention, one data map is generated for each data field, and one global data directory is generated based on all the data maps, so that the problem that an existing data management method is not suitable for edge equipment can be solved. In addition, the data map is compatible with unstructured data and structured data, and the problem that an existing data management method is not suitable for edge data management can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of data management technology, and in particular to a data management method, system, and computing device based on AI edge devices. Background Art

[0002] Centralized data centers have long been considered the nexus of network connectivity and a fundamental component of the digital economy. Under traditional centralized data center architectures, data management primarily utilizes a data center platform and a unified data lake approach. This approach establishes a centralized data storage warehouse, centrally storing and managing data from various sources, and providing a unified data analysis method.

[0003] With the decentralization of enterprise businesses and the advancement of artificial intelligence (AI) technology, data generation and processing are increasingly reliant on AI architectures, which are also becoming increasingly decentralized. This is leading to an increasing amount of data being generated by AI inference models at the network edge rather than in traditional data centers. Against this backdrop, data management approaches based on a data center and centralized data lakes are unsuitable for edge data management and therefore cannot meet current data management needs. Summary of the Invention

[0004] The embodiments of the present application provide a data management method, system, and computing device based on AI edge devices. Taking into account the network topology characteristics of edge devices, a data map can be generated for each data domain, and a global data directory can be generated based on all data maps. This can effectively solve the problem that the data management method based on the data center and unified data entry is not suitable for edge data management. To achieve the above purpose, the embodiments of the present application adopt the following technical solutions:

[0005] In a first aspect, an embodiment of the present application provides a data management method based on AI edge devices, the method comprising: dividing a first number of data sources into a second number of data domains, wherein at least one data domain includes an AI edge data source and a non-AI edge data source; extracting metadata for each data domain, wherein the metadata is used to describe the attribute information of the data resources in the data domain, and the metadata in the AI edge data source includes unstructured data obtained by reasoning based on the AI model deployed on the edge device; based on the metadata in each data domain, generating a data map corresponding to each data domain, wherein the node information of at least one data map is simultaneously associated with structured metadata and unstructured metadata; based on the data maps corresponding to the second number of data domains, generating a global data directory.

[0006] Based on this solution, the data sources are divided into corresponding data domains, and a data map is generated for each data domain, and then the information in these data maps is summarized to obtain a global data directory. In this way, the data domain where the data is located can be known through the global data directory, and the data information stored in the data source in the data domain can be known through the data map corresponding to the data domain. In this way, the physical dispersion and logical unification of data can be achieved, so that data retrieval and traceability can be achieved based on the global data directory without centrally storing data. The edge data is only stored in the data source, and there is no need to transfer it to the same data warehouse for storage. This can avoid the data transmission delay problem caused by the unified data management method of transferring edge data to the data warehouse.

[0007] On this basis, the data source can include both AI edge data sources and non-AI edge data sources. At this time, the data map generated for the data source can be compatible with structured metadata and unstructured metadata. In this way, the data obtained by AI model reasoning and ordinary business data can be integrated, which is more suitable for edge data management scenarios based on AI architecture.

[0008] In one possible implementation, based on the metadata in each data domain, data maps corresponding to the data domain are generated respectively, including: extracting the association relationship between the metadata in the data source corresponding to the data domain, wherein the association relationship includes at least one of data flow information, data lineage information, and data dependency information; using data resources as nodes and metadata as node information, and establishing a data map based on the association relationship between the metadata.

[0009] This solution extracts metadata corresponding to business data from the data source corresponding to the data domain and creates a data map based on metadata and its data lineage relationships. This allows for a visual display of the relationships and flows of business data within the data domain. Furthermore, metadata preprocessing ensures its quality and accuracy.

[0010] In another possible implementation, in the data source corresponding to the data domain, the association relationship between the metadata is extracted, including: when the data source is an AI edge data source, obtaining a first association relationship between the metadata based on the AI model, wherein the first association relationship is stored in the form of a graph; when the data source is not an AI edge data source, determining a second association relationship between the metadata based on the table structure of the data table.

[0011] This solution captures metadata relationships, enabling interoperability and breaking down data silos. Furthermore, by extracting metadata relationships based on different data sources, it can be applied to scenarios with diverse data sources, offering a wide range of applications and high flexibility.

[0012] In another possible implementation, a global data directory is generated based on data maps corresponding to a second number of data domains, including: creating at least one data product based on each data map, wherein the data product includes at least one metadata in the data map and descriptive information corresponding to the metadata; and publishing the data product to the initial data directory to obtain a global data directory containing the data product.

[0013] This solution uses metadata and corresponding descriptive information to create data products, helping users better understand the metadata. Data products are published to the initial data catalog, resulting in a populated global data catalog. This allows for indexing of data products within the global data catalog, facilitating the search for data products based on actual needs.

[0014] In another possible implementation, the method further includes: obtaining descriptive information input by a first user, wherein the first user is a data owner user corresponding to the data domain where the metadata is located; and / or, when the data source is an AI edge data source, updating the descriptive information based on the AI model in response to trigger information.

[0015] Based on this solution, the data owner can define the descriptive information of the metadata. In this way, the data owner can define the descriptive information in a targeted manner based on actual business needs to ensure the adaptability of the descriptive information to the actual business. For AI edge data sources, the descriptive information can also be automatically generated through the AI model, reducing the workload of the data owner and improving the efficiency of generating descriptive information.

[0016] In another possible implementation, dividing the first number of data sources into a second number of data domains includes: determining the second number of data domains based on the business domains corresponding to each data source, wherein each business domain corresponds to at least one data domain; and dividing the data sources into the data domains based on the correspondence between the data sources and the business domains and the correspondence between the business domains and the data domains.

[0017] This solution divides data domains according to business domains, allowing each data domain to be associated with a specific business domain. This association helps users better understand the business implications of the data domain, leading to more effective data management. Furthermore, by dividing data sources corresponding to business domains into corresponding data domains, this solution further ensures the consistency of business data organization and business logic, improving the accuracy and efficiency of data management.

[0018] In another possible implementation, the method further includes: updating the data map, data products and global data directory corresponding to the data domain corresponding to any data source based on the data deletion instruction corresponding to any data source; and deleting the data resources corresponding to the data deletion instruction in any data source.

[0019] This solution sets constraints for data deletion operations, first updating the data map, data products, and global data catalog, and then deleting the data resources in the data source. This avoids delays in updating the data catalog after data resources are deleted, which can cause users to be unable to retrieve the corresponding data resources when they request data even though they can find data products in the data catalog, thus affecting the user experience.

[0020] In another possible implementation, the method further includes: in response to a search instruction input by a second user, determining a target data product corresponding to the search instruction in a data directory, wherein the second user is a user requesting to use the target data product; returning product identification information corresponding to the target data product to the second user; and in response to a data request instruction input by the second user corresponding to the product identification information, granting the second user access rights to the target data resource corresponding to the data request instruction if the second user authentication is successful.

[0021] Based on this solution, internal enterprise users, also known as secondary users, can query data products through the data catalog and, based on the product ID, obtain corresponding permissions to access the target data resources and then analyze or process them. This allows the global data catalog to be circulated within the enterprise, ensuring that secondary users can easily use and access data resources, maximizing their value.

[0022] In a second aspect, an embodiment of the present application further provides a data management system based on an AI edge device, the system comprising:

[0023] a data domain component, deployed on the first computing device, configured to divide the first number of data sources into a second number of data domains, wherein at least one data domain includes an AI edge data source and a non-AI edge data source;

[0024] A data map component, deployed on a second computing device, is configured to extract metadata for each data domain, wherein the metadata is used to describe attribute information of data resources in the data domain, and the metadata in the AI edge data source includes unstructured data inferred based on the AI model deployed on the edge device; and, based on the metadata in each data domain, generate a data map corresponding to each data domain, wherein each node information of at least one data map is associated with both structured metadata and unstructured metadata;

[0025] A data directory component is deployed on a third computing device and is used to generate a global data directory based on the data maps corresponding to the second number of data domains.

[0026] In one possible implementation, the data map component is used to: extract the association relationship between metadata in the data source corresponding to the data domain, where the association relationship includes at least one of data flow information, data lineage information, and data dependency information; use data resources as nodes and metadata as node information, and establish a data map based on the association relationship between metadata.

[0027] In one possible implementation, the data map component is used to: when the data source is an AI edge data source, obtain a first association relationship between metadata based on an AI model, wherein the first association relationship is stored in the form of a graph; when the data source is not an AI edge data source, determine a second association relationship between metadata based on the table structure of the data table.

[0028] In one possible implementation, the data directory component is used to: create at least one data product based on each data map, wherein the data product includes at least one metadata in the data map and descriptive information corresponding to the metadata; publish the data product to the initial data directory to obtain a global data directory containing the data product.

[0029] In one possible implementation, the data directory component is used to: obtain descriptive information input by a first user, where the first user is the data owner user corresponding to the data domain where the metadata is located; and / or, when the data source is an AI edge data source, update the descriptive information based on the AI model in response to trigger information.

[0030] In one possible implementation, the data domain component is used to: determine a second number of data domains based on the business domains corresponding to each data source, wherein each business domain corresponds to at least one data domain; and divide the data sources into data domains based on the correspondence between the data sources and the business domains and the correspondence between the business domains and the data domains.

[0031] In one possible implementation, the data management system based on AI edge devices also includes a data deletion constraint component, which is used to: update the data map, data products and global data directory corresponding to the data domain corresponding to any data source based on the data deletion instruction corresponding to any data source; and delete the data resources corresponding to the data deletion instruction in any data source.

[0032] In one possible implementation, the data management system based on AI edge devices also includes a data resource access component, which is used to: in response to a search instruction input by a second user, determine the target data product corresponding to the search instruction in the data directory, wherein the second user is the user requesting to use the target data product; return the product identification information corresponding to the target data product to the second user; in response to a data request instruction input by the second user corresponding to the product identification information, if the second user authentication is passed, release access rights to the target data resource corresponding to the data request instruction to the second user.

[0033] In a third aspect, an embodiment of the present application further provides a computing device comprising: a processor and a memory; the processor and the memory are coupled; the memory is used to store program instructions; and the processor is used to execute program instructions to perform any method as described in the first aspect above.

[0034] In a fourth aspect, an embodiment of the present application provides a chip, which is used to execute any method as described in the first aspect above.

[0035] In a fifth aspect, an embodiment of the present application provides a computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are executed by a computer, the method as described in any one of the first aspects is implemented.

[0036] In a sixth aspect, an embodiment of the present application provides a program product, comprising a computer program, which implements any method in the first aspect when executed by a processor. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 This is a simplified architecture diagram of an AI architecture provided in an embodiment of the present application;

[0038] Figure 2 This is a flow chart of a data management method based on AI edge devices provided in an embodiment of the present application;

[0039] Figure 3 This is a flow chart of the data domain partitioning method provided in an embodiment of the present application;

[0040] Figure 4 This is a schematic diagram of the correspondence between the service domain and the data domain provided in an embodiment of the present application;

[0041] Figure 5 This is another schematic diagram of the correspondence between the service domain and the data domain provided in an embodiment of the present application;

[0042] Figure 6 This is a schematic diagram of dividing data sources into data domains provided by an embodiment of the present application;

[0043] Figure 7 This is another schematic diagram of dividing data sources into data domains provided by an embodiment of the present application;

[0044] Figure 8 This is a flow chart for generating a data map provided by an embodiment of the present application;

[0045] Figure 9 This is a flow chart for extracting association relationships provided by an embodiment of the present application;

[0046] Figure 10 This is a flow chart of generating a global directory provided by an embodiment of the present application;

[0047] Figure 11 This is a flowchart of deleting data resources provided by an embodiment of the present application;

[0048] Figure 12 This is a flow chart for obtaining data resources provided by an embodiment of the present application;

[0049] Figure 13 This is an interaction diagram of various components in the process of obtaining data resources provided by an embodiment of the present application;

[0050] Figure 14 This is a schematic diagram of a data management system based on AI edge devices provided in an embodiment of the present application;

[0051] Figure 15 This is a schematic diagram of another data management system based on AI edge devices provided in an embodiment of the present application;

[0052] Figure 16 This is a schematic diagram of a computing device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0053] The following will describe the technical solutions in the embodiments of the present application in conjunction with the accompanying drawings. To facilitate the clear description of the technical solutions in the embodiments of the present application, the first, second, etc. descriptions in the embodiments of the present application are only used for illustration and to distinguish the described objects. There is no order, nor does it represent a special limitation on the number of devices in the embodiments of the present application, and it does not constitute any limitation on the embodiments of the present application.

[0054] The following explains the professional terms mentioned in the embodiments of the present application to facilitate understanding by those skilled in the art.

[0055] Data Map: A data map is a data management tool that helps users understand and manage their data. Data maps typically contain metadata about the data, such as its source, type, and owner.

[0056] Data Catalog: A data catalog is a tool for storing and organizing data. It typically contains one or more data maps and provides search and filtering capabilities, allowing users to easily find the data they need. This article further extracts and packages the information from the data maps into several "data products" and then publishes them to the data catalog.

[0057] AI edge data: AI edge data refers to data generated and processed at the edge of devices (rather than on central servers) and provided in the form of datasets for AI inference or training. These devices may include various IoT devices and mobile devices.

[0058] Unified view: A unified view is a view that integrates information from different sources or formats into a single interface, allowing users to view and manage all relevant information in one place. In this invention, the unified view refers to the global data directory.

[0059] Business domain: A business domain refers to a specific business scope or field within an enterprise or organization, which covers a set of related business processes, functions, rules, and data.

[0060] Data domain: Data domain is the division of data areas, which can be divided from the perspectives of business, data governance, etc.

[0061] The following combination Figures 1 to 13 , which illustrates the process of the data management method based on AI edge devices.

[0062] Figure 1 This is a simplified architecture diagram of an AI architecture provided in an embodiment of the present application.

[0063] like Figure 1 As shown, Figure 1 There are multiple business domains, namely business domain 1 to business domain N, which represent N independent business scopes. For example, an AI architecture based on a medical service enterprise may include sales business domain, medical business domain, financial business domain, etc.

[0064] Figure 1 Each business domain in the data domain corresponds to a data domain, such as the sales business domain corresponds to the sales data domain, the medical business domain corresponds to the medical data domain, and the financial data source corresponds to the financial data domain. Data producers can produce and manage data in the data domain.

[0065] In each data domain, there is at least one data source, which can be a database or other device deployed on a computing or storage device. In some business domains, the device where the data source resides may be an edge device, which can be understood as a device far from the center in the network topology.

[0066] In each data domain, there may also be a computing cluster, which may include multiple servers and other computing devices. In some business domains, the computing devices in the computing cluster may be edge devices, that is, devices far away from the center in the network topology.

[0067] Data management software is deployed on computing devices to manage and analyze data. Data producers can use computing clusters and data management software to generate data resources and store them in data sources. In some business domains, data management software can integrate AI models to generate and process business data through model reasoning.

[0068] Data map software is a type of data management software that can be used to generate data maps corresponding to the business domain, such as Figure 1 The data areas corresponding to each business domain Figure 1 In this way, the information in the data maps of each business domain can be aggregated and published in a unified global data catalog for data users to view.

[0069] It should be noted that Figure 1 This is only a simplified architecture diagram. In some application scenarios corresponding to complex businesses, one business domain may also correspond to multiple data domains; in some application scenarios with large data volumes, one data domain may also have multiple data sources. This application does not limit this.

[0070] Figure 2 This is a flowchart of the data management method based on AI edge devices provided in an embodiment of the present application.

[0071] like Figure 2 As shown, the method includes the following steps S11-S14:

[0072] S11: Divide a first number of data sources into a second number of data domains.

[0073] In step S11, the data sources in the AI architecture are divided into data domains. Each data domain includes at least one data source, and the first and second quantities can be the same or different. In this way, each data domain can serve as a relatively independent data management unit, and different data domains can be isolated from each other to improve the security of data resources in the data source.

[0074] The data source may be an AI edge data source. The computing device or storage device where the AI edge data source resides is located at the edge of the AI architecture network topology, and at least some of the data resources in the AI edge data source are data used and generated during the inference process of the AI model deployed on the edge device.

[0075] For example, at the edge of an AI architecture network topology, a microphone can collect voice chat logs between salespeople and customers. After AI model analysis, the resulting voice analysis results are stored in interaction record data source DS1. Thus, the voice chat logs and voice analysis results become AI edge data, and interaction record data source DS1 becomes the AI edge data source.

[0076] The data source can also be a non-AI edge data source, which can be a non-AI data source located at the edge of the network topology. For example, at the edge of the AI-based network topology, a computing device generates a sales report using traditional non-AI data calculation methods and saves it in the sales report data source DS2. Thus, the sales report is non-AI edge data, specifically non-AI data, and the sales report data source DS2 is a non-AI edge data source, specifically a non-AI data source.

[0077] In addition, non-AI edge data sources can also be AI data sources located at the center of the network topology, or non-AI data sources located at the center of the network topology. The definition method is similar to the aforementioned AI edge data sources and non-AI data sources, and will not be repeated here.

[0078] In this step, at least one data domain includes an AI edge data source and a non-AI edge data source. For example, a sales business domain is located at the edge of the AI architecture network topology, and the data domain includes the aforementioned interaction record data source DS1 and sales report data source DS2.

[0079] The following combination Figure 3 Step S11 is further introduced.

[0080] Figure 3 This is a flow chart of the data domain partitioning method provided in an embodiment of the present application.

[0081] like Figure 3 As shown, in one implementation, step S11 includes the following steps S111-S112:

[0082] S111: Determine a second number of data domains according to the business domains corresponding to the data sources, wherein each business domain corresponds to at least one data domain.

[0083] In step S111, data domains may be divided from a business perspective according to the business domains corresponding to each data source, wherein at least one data domain may be divided for each business, that is, each business domain corresponds to at least one data domain.

[0084] The following combination Figure 4 and Figure 5 This section describes the correspondence between the business domain and the data domain.

[0085] Figure 4 This is a schematic diagram of the correspondence between the business domain and the data domain provided in an embodiment of the present application.

[0086] like Figure 4 As shown, in some scenarios, there is a one-to-one correspondence between business domains and data domains. Specifically, for the aforementioned sales business domain BD1, its corresponding sales data domain DD1 can be determined for sales-related data operations and management. This allows data management and operations within the business domain to be independently completed within the corresponding data domain, eliminating the need for cross-domain operations, improving data processing efficiency and security. Furthermore, this one-to-one correspondence makes the storage and use of data resources more standardized and orderly.

[0087] Figure 5 This is another schematic diagram of the correspondence between the business domain and the data domain provided in an embodiment of the present application.

[0088] like Figure 5 As shown, in other scenarios, such as scenarios with relatively complex businesses, the business domain may correspond to multiple data domains. Specifically, for the medical business domain BD2, if the medical business is too complex, its corresponding health check data domain DD2 and patient file data domain DD3 may be determined. Among them, the health check data domain DD2 can be used for data calculation and management of patient medical examination-related businesses, and the patient file data domain DD3 can be used for data calculation and management of patient file-related businesses. In this way, different business processes in the business domain can be divided into different data domains for management and calculation, thereby realizing the decoupling of complex businesses. Each data domain is only responsible for processing a part of the business logic, making data management more refined and professional.

[0089] It's understandable that if the computing devices deploying the business system are located far apart in the network topology, the aforementioned approach of assigning one business domain to multiple data domains can also be employed. For example, if HR services are used by multiple branch offices, and the branches are network-isolated, a data domain can be set up for each branch, allowing HR services to correspond to multiple data domains.

[0090] It should be noted that in order to highlight the division relationship between the business domain and the data domain, Figure 4 and Figure 5 Omission of data management software, computing clusters and data sources does not mean Figure 4 and Figure 5 The illustrated embodiment does not include data management software, computing clusters, and data sources.

[0091] S112: Divide the data source into the data domain based on the correspondence between the data source and the business domain, and the correspondence between the business domain and the data domain.

[0092] In step S112, after dividing several data domains based on the business domain, the correspondence between the data source and the business domain can be obtained based on the correspondence between the data source and the business domain, and the correspondence between the business domain and the data domain, so as to divide the data source corresponding to each business domain into the data domain.

[0093] The following combination Figure 6 and Figure 7 This section describes how to divide data sources into data domains.

[0094] Figure 6 This is a schematic diagram of dividing data sources into data domains provided in an embodiment of the present application.

[0095] like Figure 6 As shown in , for scenarios where a business domain corresponds to only one data domain, the data source corresponding to the business domain can be directly divided into the data domain corresponding to the business domain. In this case, the data domain and the data source correspond one to one. Figure 6 For the aforementioned sales business domain BD1, the interaction record data source DS1 and sales report data source DS2 corresponding to sales business domain BD1 can be divided into sales data domain DD1. This way, each data source is responsible for storing only a portion of the data resources. This data source division allows for independent backup and recovery operations for each data source, thus resolving the issue of large data resources within a single data source and improving backup and recovery efficiency.

[0096] Figure 7 This is another schematic diagram of dividing data sources into data domains provided in an embodiment of the present application.

[0097] like Figure 7 As shown, for scenarios where a business domain corresponds to multiple data domains, the data source corresponding to the business domain can be dispersed into multiple data domains corresponding to the business domain, with each data domain corresponding to one or more data sources. Figure 7 In the example above, for the medical business domain BD2 corresponding to the health examination data domain DD2 and the patient file data domain DD3, the medical imaging data source DS3 corresponding to the medical business domain BD2 can be divided into the health examination data domain DD2, and the patient personal information data source DS4 can be divided into the patient file data domain DD3. The medical imaging data source DS3 is responsible for storing data resources such as medical imaging data generated during medical examinations, while the patient file data source DS4 is responsible for storing data resources such as personal information data generated during the patient file creation process.

[0098] Based on the corresponding relationship between business domains, data domains, and data sources, the present embodiment divides the data source into corresponding data domains, making the storage of data resources more standardized and intuitive. In addition, since each data domain is independent of each other, the security of the data resources stored in the data source can be improved.

[0099] S12: Extract metadata for each data domain.

[0100] In step S12, metadata is extracted from each data domain based on the data resources stored in the data source. The metadata is used to describe the attribute information of the data resources in the data domain, and may include structured data or unstructured data.

[0101] The following introduces the metadata extraction methods for two types of data resources: structured data and unstructured data.

[0102] For ordinary structured data, metadata can be extracted by extracting field names in the data table.

[0103] For example, in the aforementioned sales data domain DD1, for a non-AI edge data source, such as the sales report data resource stored in the aforementioned sales report data source DS2, the field names in the header position of the sales report, such as product identification (identity document, ID), product name, sales volume, price, etc., which are used to describe the attributes of sales information, can be directly extracted as metadata corresponding to the sales report.

[0104] For unstructured data, metadata can be extracted through AI model reasoning.

[0105] For example, in the aforementioned sales data domain DD1, for a data resource storing voice chat log data, such as the aforementioned interaction record data source DS1, the AI model can be used to parse the voice chat log data and obtain corresponding metadata. For example, the AI model can be used to parse the voice chat log data, tag the voice chat log data, and then obtain metadata corresponding to the voice chat log data based on the tags. Specifically, this metadata may include unstructured metadata such as semantic keywords, contextual intent, semantic recognition results, user demands, and user profiles, as well as structured metadata such as customer ID, session ID, and timestamp.

[0106] In the embodiment of the present application, step S12 is used to determine metadata corresponding to data resources stored in each data source, which is beneficial for analyzing and utilizing data resources and improving the value of data resources.

[0107] S13: Based on the metadata in each data domain, generate a data map corresponding to each data domain.

[0108] In step S13, a data map is generated for each data domain to graphically display the distribution and attribute information of the data resources in the data domain. The data map can be displayed in various forms, including but not limited to tree diagrams, network diagrams, lists, etc.

[0109] The data map generation process can be automatically completed based on the extracted metadata. Specifically, the data map can include metadata and the associations between multiple metadata. Therefore, for a data domain that contains both AI edge data sources and non-AI edge data sources, its data map can include both unstructured metadata extracted from the AI edge data sources and structured metadata extracted from the non-AI edge data sources.

[0110] For example, for the aforementioned sales business domain BD1, the data resources within the sales report data source BS1 include general business data such as sales reports, and the corresponding metadata includes structured data. The data resources within the customer interaction record data source BS2 include chat logs and other inference data derived from AI models, and the corresponding metadata includes unstructured data. Therefore, the data map constructed for the sales business domain BD1 associates both structured and unstructured metadata.

[0111] Furthermore, for a node in the data map, its node information can also be associated with structured metadata and unstructured metadata at the same time.

[0112] For example, for the aforementioned customer interaction record data source BS2, semantic keywords, contextual intent, semantic recognition results, and user profiles are unstructured metadata, while customer ID, session ID, and timestamps are structured metadata. Thus, the node information corresponding to the chat record can be associated with both structured and unstructured metadata.

[0113] Based on this, and considering that node information can be associated with structured metadata and unstructured metadata at the same time, the embodiments of the present application do not limit the storage method of metadata in the data map. In each node of the map data, a suitable storage method can be selected based on the actual format of the metadata. Specifically, unstructured metadata can be stored in formats such as json, and structured metadata can be stored in the form of data tables or key-value pairs. In some complex business scenarios, the extracted metadata can be stored in a nested form. For example, semantic recognition results can be used as parent metadata, and intent recognition results, confidence scores, etc. can be used as child metadata of speech recognition results.

[0114] In addition, before generating the data map, the metadata extracted from the data source can be desensitized to avoid the leakage of sensitive information and ensure the security of data resources.

[0115] In the embodiment of the present application, a data map is generated for each data domain through step S13. This eliminates the need to store edge data uniformly in the data lake and then generate a unified data map. This reduces the data transmission between edge data and the data lake, improving data management efficiency. Therefore, the embodiment of the present application is suitable for edge data management. In addition, in the embodiment of the present application, an independent data map is generated for each data domain, which facilitates more precise control of data access rights and improves the security of data resources.

[0116] The following combination Figure 8 Step S13 is further introduced.

[0117] Figure 8 This is a flow chart for generating a data map provided in an embodiment of the present application.

[0118] like Figure 8 As shown, in one implementation, step S13 includes the following steps S131-S132:

[0119] S131: Extracting association relationships between metadata in a data source corresponding to a data domain.

[0120] In step S131, after extracting metadata from the data source corresponding to each data domain, the association relationship between each metadata can be extracted. The association relationship may include at least one of data flow information, data lineage information, and data dependency information. Data flow information is used to describe the flow path of the data value corresponding to the metadata in the data domain (such as the source, destination, and transmission process of the data value); data lineage information is used to describe the upstream and downstream data of the data value corresponding to the metadata, etc., and may include full-link information on data generation, conversion, and use; data dependency information is used to describe the dependency relationship between the data values corresponding to the metadata, such as whether a data value is referenced by other data values, whether it depends on other data values, etc.

[0121] Furthermore, consider that the same metadata can have multiple different data value sources, data statistical calibers, or analysis methods. For example, the sales metadata in the sales report for the aforementioned sales data domain DD1 may yield different sales values depending on the actual payment collection time, order creation time, or other calibers. This can result in the same metadata having multiple conflicting data values. Therefore, when extracting associations, data quality can be assessed. The assessment results can be used to filter invalid metadata, correct associations, and so on.

[0122] Specifically, based on the data association relationship, the source and statistical processing process of the metadata corresponding to the data value can be determined. By tracing back to the source and the links where deviations occur in the statistical processing process, the more credible data value sources and association relationships can be determined to ensure the validity of data resources.

[0123] The embodiment of the present application extracts the association relationship between metadata through step S131, which can clearly display the logical relationship between each data value in the data resource, thereby facilitating the query, management and use of data.

[0124] The following combination Figure 9 Step S131 is further introduced.

[0125] Figure 9 This is a flow chart for extracting association relationships provided in an embodiment of the present application.

[0126] like Figure 9 As shown, in one implementation, step S131 includes the following steps S1311-S1312:

[0127] S1311: When the data source is an AI edge data source, a first association relationship between metadata is obtained based on the AI model.

[0128] In step S1311, if the data source is an AI edge data source, that is, the data resources in the data source include inference data generated by the AI model deployed on the edge device, the association relationship between the metadata can be obtained by using the AI model reasoning, which is recorded as the first association relationship. Specifically, the label corresponding to the data resource can be obtained by AI model reasoning, and the metadata can be obtained based on the label, and then the first association relationship between the metadata can be obtained based on the label reasoning. For example, for the aforementioned voice chat records, the association relationship between semantic keywords and semantic recognition results, the association relationship between user ID and voiceprint features, etc. can be obtained as the first association relationship.

[0129] The first association relationship includes unstructured data and can therefore be stored in a graph format. Specifically, the first association relationship can be stored in a graph database. For example, nodes in the graph database can represent metadata, edges between nodes can represent associations between metadata, and attributes can represent additional information about the metadata or associations.

[0130] S1312: When the data source is not an AI edge data source, determine a second association relationship between metadata based on the table structure of the data table.

[0131] In step S1312, if the data source is not an AI edge data source, specifically not an AI data source, that is, the data resources in the data source are structured data stored in a data table or other form, then the association relationship between the metadata can be determined based on the table structure of the data table, and recorded as the second association relationship. Specifically, the second association relationship between the metadata can be determined based on the logical relationship between the various data tables, between the various fields within the data tables, and the execution history of SQL statements. For example, for the aforementioned sales report, the association relationship between the promotion period and sales volume, the association relationship between the sales channel and the payment method, etc. can be obtained as the second association relationship.

[0132] The second association relationship may be stored in the form of a data table or a key-value pair, or in a format such as json.

[0133] This embodiment of the present application obtains metadata associations through steps S1311-S1312, enabling metadata interoperability and breaking down data silos. Furthermore, by extracting metadata associations in different ways for different data sources, this approach is applicable to scenarios with diverse data sources, offering a wide range of applications and high flexibility.

[0134] S132: Using data resources as nodes and metadata as node information, a data map is established based on the association relationship between metadata.

[0135] In step S132, for each data domain, a data map corresponding to that data domain is established, using the data resources within the data domain as nodes. Considering that metadata is information used to describe the attributes of data resources, the metadata of each data resource can be used as node information for the node corresponding to each data resource. This allows the node's attributes in various dimensions to be described using node information. Simultaneously, relationships between metadata are added to the data map, linking the data values within the data resources and preventing the formation of data silos.

[0136] It is understood that a data resource can have multiple pieces of metadata, and multiple pieces of metadata can contain both structured and unstructured data. Therefore, a node can have multiple pieces of node information, and multiple pieces of node information for a node can include both structured and unstructured data. Therefore, the embodiments of this application do not limit the storage method of the node information of each node in the data map. In actual application scenarios, the appropriate storage method can be selected based on the format of each node information.

[0137] In this embodiment, steps S131-S132 are used to create a data map based on metadata and its data lineage information, which can intuitively display the relationships and flows of business data within the data domain. On this basis, metadata is pre-processed to ensure metadata quality and accuracy.

[0138] S14: Generate a global data directory based on the data maps corresponding to the second number of data fields.

[0139] In step S14, the data maps corresponding to all data domains are integrated, and information extracted from the data maps is aggregated and published to a global data directory to achieve the integration of data resources from all data sources. For example, summary information such as the sales report and the fields in the table can be extracted and published as a data item to the global directory, thereby forming an index for the sales report in the global directory.

[0140] Among them, before publishing to the global data directory, the information extracted from the data map can be desensitized to avoid the leakage of sensitive information and ensure the security of data resources.

[0141] In step S14, the embodiment of the present application generates a unified global data directory based on the data maps corresponding to each data domain, which can achieve the integration and unified management of resource data. Users can quickly locate the required data resources by searching the global data directory, thereby improving the efficiency of data resource retrieval.

[0142] The following combination Figure 10 Step S14 is further described.

[0143] Figure 10 This is a flowchart for generating a global directory provided by an embodiment of the present application.

[0144] like Figure 10 As shown, in one implementation, step S14 includes the following steps S141-S142:

[0145] S141: Based on each data map, create at least one data product.

[0146] In step S141, based on actual business needs, metadata is extracted from the data map and packaged into data products. It is understood that in the same data map, different metadata can be extracted from the data map and packaged into different data products for different business needs.

[0147] For example, in the sales business domain BD1, to meet the needs of a specific business scenario, such as collecting sales data, the fields that need to be made available externally from several sales-related data resources can be combined into a sales data product. This allows sales personnel to understand product sales through the sales data product.

[0148] In addition, data products can also include descriptive information corresponding to metadata, as well as summary information such as data product descriptions. This allows users to quickly understand the content and purpose of data products based on the descriptive information and select appropriate data products based on actual business needs.

[0149] It should be noted that a data product may include multiple metadata, which may be structured or unstructured. Therefore, the present embodiments do not set a unified standard for the storage format of metadata in data products. In other words, the storage format of each metadata in a data product is determined by the characteristics of the metadata itself.

[0150] In one implementation, the description information may be obtained by manually defining it, and specifically by the following steps:

[0151] Descriptive information input by a first user is obtained, wherein the first user is a data owner user corresponding to the data domain where the metadata is located.

[0152] In this step, the data owner user of each data domain, i.e., the first user corresponding to each data domain, determines the descriptive information for the metadata within that data domain. For example, for the aforementioned sales data domain DD1, the corresponding first user may define descriptive information for metadata such as the product ID, product name, sales volume, price, semantic keywords, contextual intent, and user profile.

[0153] In addition, the first user may also define description information of the data product based on metadata in the data product. For example, for the aforementioned sales data domain DD1, the first user may also define description information corresponding to the sales data product based on metadata in the aforementioned sales data product.

[0154] In the embodiment of the present application, the data owner defines the description information of the metadata. In this way, the data owner can define the description information in a targeted manner based on actual business needs to ensure the adaptability of the description information to the actual business.

[0155] In another implementation, the description information can be obtained through AI model reasoning, which can be obtained through the following steps:

[0156] In the case where the data source is an AI edge data source, the description information is updated based on the AI model in response to the trigger information.

[0157] In this step, if the data source is an AI edge data source, then the data resources in the data source are derived through AI model reasoning, or data generated during the AI model reasoning process. In this case, the description information can be directly derived using AI model reasoning. A timer can be set to periodically trigger the description information update mechanism, using the AI model to update the description information. Alternatively, a manual refresh method can be used, where the first user enters a refresh command to prompt the AI model to update the description information.

[0158] The embodiment of the present application automatically generates description information through an AI model, which can reduce the workload of data owners and improve the efficiency of generating description information.

[0159] S142: Publish the data product to the initial data directory to obtain a global data directory containing the data product.

[0160] In step S142, after creating a data product, it can be published to the initial data catalog. This initial catalog now includes the data product, forming a global data catalog. This allows for external release of the data product, while the global data catalog is accessible to all personnel within the enterprise. This allows everyone within the enterprise to query the data product and obtain publicly available asset information.

[0161] In this embodiment, steps S141-S142 are used to create data products using metadata and corresponding descriptive information, helping users better understand the metadata. Data products are published to the initial data directory, resulting in a populated global data directory. This allows for indexes to be created for data products within the global data directory, facilitating user search for data products based on actual needs.

[0162] It's important to note that permission management and security compliance must be considered throughout the global data catalog generation process. For example, when connecting to data sources to extract metadata, you need to ensure that only authorized users can access the data source; when generating data maps, you need to ensure that all operations comply with relevant regulations; and when publishing data products, you also need to ensure that all operations are secure and compliant.

[0163] This application introduces the generation process of the global data directory through the above embodiments and drawings. Figures 11 to 13 This section describes how to delete data resources in a data source after a global data directory is generated, and how to obtain data resources based on the global data directory.

[0164] Figure 11 This is a flowchart of deleting data resources provided by an embodiment of the present application.

[0165] like Figure 11As shown, in one implementation, step S14 may further include the following steps S15-S16:

[0166] S15: Based on the data deletion instruction corresponding to any data source, update the data map, data products and global data directory corresponding to the data domain corresponding to any data source.

[0167] S16: In any data source, delete the data resource corresponding to the data deletion instruction.

[0168] In steps S15-S16, after the global data directory is generated, if a business need or other application requirement requires the deletion of a data resource A from any data source DS5, a data deletion instruction corresponding to that data source DS5 is first generated. Subsequently, in response to this data deletion instruction, the data map, data products, and global data directory are updated. Specifically, the relevant information about data resource A in the data map, data products, and global directory is removed. After the update is complete, data resource A stored in data source DS5 is deleted.

[0169] During the data map update process, the data domain DD4 corresponding to data source DS5 is first determined. Then, in the data map corresponding to data domain DD4, the data map is updated based on the metadata of data resource A. The updated data map no longer contains metadata for data resource A. Simultaneously, data product B, which contains metadata for data resource A, is removed from the data products corresponding to data domain DD4, and the corresponding index for data product B is deleted from the global data directory. This way, when users query the global data directory, they will not find data product B, and thus will not request data resource A due to a query for data product B.

[0170] The embodiment of the present application sets constraints for the data resource deletion operation through steps S15-S16, which can avoid delays in data directory updates after data resource deletion, resulting in users being able to query data products through the data directory but unable to obtain corresponding data resources when requesting data. This can effectively ensure that the data products queried by users are authentic and available, thereby improving user experience.

[0171] The following combination Figure 12 and Figure 13 , introduces the method of obtaining data resources based on the global data directory after the global data directory is generated.

[0172] Figure 12 This is a flow chart for obtaining data resources provided in an embodiment of the present application.

[0173] like Figure 12 As shown, in one implementation, step S14 may further include the following steps S17-S19:

[0174] S17: In response to the search instruction input by the second user, determine a target data product corresponding to the search instruction in the data catalog.

[0175] The second user is a user who requests to use the target data product.

[0176] In step S17, the second user may enter a search instruction through a unified data resource management view (hereinafter referred to as "unified view"), wherein the unified view may include a search box. After the second user enters the search instruction, the data catalog is searched for a target data product corresponding to the search instruction.

[0177] S18: Return product identification information corresponding to the target data product to the second user.

[0178] In step S18, product identification information corresponding to the target data product, such as product ID, is returned to the second user to inform the second user that the data product corresponding to the search instruction has been found.

[0179] S19: In response to the data request instruction input by the second user corresponding to the product identification information, if the second user passes the authentication, the access right to the target data resource corresponding to the data request instruction is released to the second user.

[0180] In step S19, if the second user enters a data request instruction corresponding to the product identification information, the second user is authenticated. If the second user's authentication is successful, the second user is deemed authorized to access the data resource corresponding to the target data product. Therefore, the second user is granted access to the target data resource corresponding to the data request instruction. The second user can then access the target data resource and perform operations such as data analysis based on the target data resource.

[0181] In the embodiment of the present application, through steps S17-S19, the second user can conveniently use and obtain data resources by querying the global data directory, thereby maximizing the value of the data resources.

[0182] Figure 13 This is an interaction diagram of various components in the process of obtaining data resources provided by an embodiment of the present application.

[0183] like Figure 13 As shown, the second user logs in to the unified view to obtain data resources, wherein the interface of the unified view includes three units: asset management, authorization management, and data supply management.

[0184] Specifically, in step 31, the second user can use the asset management unit to search for a target data product of interest, such as the aforementioned sales data product, through query and filtering operations. In step 32, the asset management unit uses the data directory component to query the global data directory stored in the data asset library and obtain the product ID of the target data product. For example, the asset management unit finds that the product ID of the sales data product is DP001. In step 33, the asset management unit returns the retrieved product ID to the second user. In step 34, the second user sends a data request instruction containing the product ID to the authorization management unit, requesting access to the data resource corresponding to the target data product (hereinafter referred to as the target data resource). During this process, the second user can indicate their interest in the target data product by filling out an application form. For example, by filling out the data request form with information such as user ID, application date, target data product, and purpose, the second user requests access to the target data resource corresponding to the sales data product, i.e., several sales-related data resources. In step 35, based on the data request instruction, the authorization management unit queries the data owner of the target data resource through the authorization management module, i.e., the first user corresponding to the data domain where the target data resource resides. In step 36, after querying the first user, the authorization management module sends an authorization application to the first user, requesting the first user to grant the second user permission to access the target data resources. In step 37, after the first user confirms that the second user has passed the authentication, it sends an authorization license to the authorization management module. Thereafter, the authorization management module grants the second user permission to obtain the target resource data based on the authorization license. In step 38, the authorization module sends a prompt message to the supply management unit to inform the supply management unit that the second user has passed the authentication. In step 39, after confirming that the second user has the permission to obtain the target resource data, the supply management unit provides the target resource data to the second user. For example, data resources related to sales can be provided to the second user. Thereafter, the second user can use the target resource data, such as analyzing data resources related to sales to understand the sales situation of the product.

[0185] It should be understood that the order of execution of the steps in the above embodiments does not necessarily imply a specific order of execution. The order of execution of each process is determined by its function and inherent logic, and does not constitute any limitation on the implementation of the embodiments of the present invention. For example, the steps described in the above embodiments can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in the present invention can be achieved. The present invention does not impose any restrictions on this.

[0186] Corresponding to the aforementioned embodiment of the data management method based on AI edge devices, the present application also provides an embodiment of a data management system based on AI edge devices.

[0187] Figure 14This is a schematic diagram of a data management system based on AI edge devices provided in an embodiment of the present application.

[0188] like Figure 14 As shown, the AI edge device-based data management system 1400 may include a data domain component 1410 , a data map component 1420 , and a data directory component 1430 .

[0189] A data domain component 1410, deployed on a first computing device, is configured to divide a first number of data sources into a second number of data domains, wherein at least one data domain includes an AI edge data source and a non-AI edge data source;

[0190] A data map component 1420, deployed on the second computing device, is configured to extract metadata for each data domain, wherein the metadata is used to describe attribute information of data resources in the data domain, and the metadata in the AI edge data source includes unstructured data inferred based on the AI model deployed on the edge device; and, based on the metadata in each data domain, generate a data map corresponding to the data domain, wherein each node information of at least one data map is associated with both structured metadata and unstructured metadata.

[0191] The data directory component 1430 is deployed on the third computing device and is used to generate a global data directory based on the data maps corresponding to the second number of data domains.

[0192] In one implementation, the data map component 1420 is used to: extract the association relationship between metadata in the data source corresponding to the data domain, where the association relationship includes at least one of data flow information, data lineage information, and data dependency information; use data resources as nodes and metadata as node information, and establish a data map based on the association relationship between metadata.

[0193] In one implementation, the data map component 1420 is used to: when the data source is an AI edge data source, obtain a first association relationship between metadata based on an AI model, wherein the first association relationship is stored in the form of a graph; when the data source is not an AI edge data source, determine a second association relationship between metadata based on the table structure of the data table.

[0194] In one implementation, the data directory component 1430 is used to: create at least one data product based on each data map, wherein the data product includes at least one metadata in the data map and descriptive information corresponding to the metadata; publish the data product to the initial data directory to obtain a global data directory containing the data product.

[0195] In one implementation, the data directory component 1430 is used to: obtain descriptive information input by a first user, where the first user is a data owner user corresponding to the data domain where the metadata is located; and / or, when the data source is an AI edge data source, update the descriptive information based on the AI model in response to trigger information.

[0196] In one implementation, the data domain component 1410 is used to: determine a second number of data domains based on the business domains corresponding to each data source, where each business domain corresponds to at least one data domain; and divide the data sources into data domains based on the correspondence between the data sources and the business domains and the correspondence between the business domains and the data domains.

[0197] In one implementation, the data management system 1400 based on AI edge devices also includes a data deletion constraint component, which is used to: update the data map, data products and global data directory corresponding to the data domain corresponding to any data source based on the data deletion instruction corresponding to any data source; and delete the data resources corresponding to the data deletion instruction in any data source.

[0198] In one implementation, the AI edge device-based data management system 1400 also includes a data resource access component, which is used to: in response to a search instruction input by a second user, determine a target data product corresponding to the search instruction in the data directory, wherein the second user is a user requesting to use the target data product; return product identification information corresponding to the target data product to the second user; in response to a data request instruction input by the second user corresponding to the product identification information, and if the second user authentication is passed, release access rights to the target data resource corresponding to the data request instruction to the second user.

[0199] Figure 15 This is a schematic diagram of another data management system based on AI edge devices provided in an embodiment of the present application.

[0200] like Figure 15 As shown, the AI edge device-based data management system 1400 includes a data map component 1420 and a data directory component 1430.

[0201] exist Figure 15 In the data map, there are data domain A and data domain B, where data domain A includes a data source a and data domain B includes two data sources b1 and b2. The corresponding process of data domain B is similar to that of data domain A, so the diagram of the corresponding steps of data domain B is simplified. Before the data map is generated, first execute Figure 15In step 1, a communication connection is established between each data map component and its corresponding data source, allowing the data map component to extract metadata from the data source. To accommodate multiple data sources, the data source can be connected via a connector component. Then, in step 2, data map component 1420 performs data processing within the data map component's metadata management system. This data processing can include operations such as metadata extraction, metadata enhancement, metadata desensitization and compliance, data quality testing, data lineage extraction, and data problem analysis. Metadata extraction refers to extracting metadata from data sources; metadata enhancement refers to building relationships based on metadata to enrich the metadata; metadata desensitization and compliance refers to desensitizing sensitive information in metadata to ensure data security and compliance; data quality testing refers to testing the quality of the data resources corresponding to the metadata; data lineage extraction refers to extracting the source and flow of the data resources corresponding to the metadata; and data problem analysis refers to analyzing and resolving metadata problems based on the data quality testing results. Data processing can improve the usability and value of metadata. After data processing is complete, step 3 can be executed, generating a data map based on the processed metadata and storing the data map in the metadata information repository. In step 4, the data owners, namely first users A and B, can view data map A and data map B, respectively. After the data map is generated, step 5 can be executed to create data products based on the data map. Finally, step 6 is executed in the data catalog component 1430 to publish the data products and obtain the global data map. The global data map is then stored as a data asset list in the data asset library.

[0202] Figure 16 This is a schematic diagram of a computing device provided in an embodiment of the present application.

[0203] like Figure 16 As shown, the computing device 1600 includes a processor 1601 and a memory 1602. By way of example, the computing device 1600 may further include a communication interface 1603 and a communication bus 1604.

[0204] The processor 1601, the memory 1602, and the communication interface 1603 communicate with each other via a communication bus 1604. The communication interface 1603 may include a system of transmitters and receivers for communicating with other devices or communication networks, and may be a wired interface (port), such as a fiber distributed data interface (FDDI) or a gigabit Ethernet interface (GE).

[0205] In some embodiments, the processor 1601 is configured to execute a program 1605, specifically, to execute the relevant steps in the above-mentioned inference task execution method embodiment. Specifically, the program 1605 may include program code, which includes computer-executable instructions.

[0206] For example, processor 1601 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement some embodiments of the present application. The computing device 1600 may include one or more processors of the same type, such as one or more CPUs, or different types of processors, such as one or more CPUs and one or more ASICs. The CPU may be a single-core CPU (single-CPU) or a multi-core CPU (multi-CPU).

[0207] In some embodiments, the memory 1602 is used to store the program 1605. The memory 1602 may include a high-speed random access memory (RAM) or a non-volatile memory (NVM), such as at least one disk storage.

[0208] The program 1605 may be specifically called by the processor 1601 to enable the computing device 1600 to perform the inference task execution method operation.

[0209] Some embodiments of the present application provide a computer-readable storage medium storing at least one executable instruction. When the executable instruction is executed on a computing device 1600, the computing device 1600 executes the data management method based on the AI edge device in the above embodiment.

[0210] The executable instructions can specifically be used to enable the computing device 1600 to perform data management method operations based on AI edge devices.

[0211] For example, the computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc (CD-ROM), a magnetic tape, a floppy disk, an optical data storage device, and the like.

[0212] Some embodiments of the present application provide a chip system that is applied to a server. The chip system includes one or more interface circuits and one or more processors. The interface circuits and processors are interconnected via circuits. The interface circuits are configured to receive signals from the server's memory and send signals to the processors, the signals including computer instructions stored in the memory. When the server processor executes the computer instructions, the server executes the steps of the data management method based on AI edge devices as described in the above method embodiment.

[0213] The beneficial effects that can be achieved by the readable storage medium provided in some embodiments of the present application can be referred to the beneficial effects of the corresponding AI edge device-based data management method provided above, and will not be repeated here.

[0214] It should be noted that, in the application, relational terms such as first and second, etc. are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or apparatus comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.

[0215] Each embodiment in this specification is described in a related manner. Similar parts between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences between the other embodiments. In particular, the system embodiment is generally similar to the method embodiment, so the description is relatively simple. For related parts, refer to the description of the method embodiment.

[0216] The logic and / or steps represented in the flowchart or otherwise described herein may be considered, for example, as an ordered list of executable instructions for implementing the logical functions, and may be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, system, or device (such as a computer-based system, a system including a processor, or other system that can fetch instructions from and execute instructions on an instruction execution system, system, or device).

[0217] For the purposes of this specification, a "computer-readable medium" can be any system that can contain, store, communicate, propagate, or transport the program for use by or in connection with an instruction execution system, system, or device.

[0218] More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection having one or more wires (electronic systems), a portable computer disk cartridge (magnetic systems), random access memory (RAM), read-only memory (ROM), erasable and programmable read-only memory (EPROM or flash memory), fiber optic systems, and a portable compact disk read-only memory (CDROM).

[0219] In addition, the computer readable medium can even be paper or other suitable medium on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other medium, and then editing, interpreting or processing it in other suitable ways as necessary, and then storing it in the computer memory. It should be understood that various parts of the present application can be implemented in hardware, software, firmware or a combination thereof.

[0220] In the above-described embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one or a combination of the following technologies known in the art can be used: a discrete logic circuit having logic gate circuits for implementing logic functions on data signals, an application-specific integrated circuit having suitable combinational logic gate circuits, a programmable gate array (PGA), a field-programmable gate array (FPGA), etc. The above-described embodiments are merely specific embodiments of the present application and are not intended to limit the scope of protection of the present application. Any modifications, equivalent substitutions, improvements, etc. made based on the technical solutions of the present application shall be included within the scope of protection of the present application.

Claims

1. A data management method based on AI edge devices, characterized in that: The method comprises: Dividing the first number of data sources into a second number of data domains, wherein at least one of the data domains includes an AI edge data source and a non-AI edge data source; Extracting metadata for each of the data domains, wherein the metadata is used to describe attribute information of data resources in the data domain, and the metadata in the AI edge data source includes unstructured data inferred based on an AI model deployed on an edge device; Based on the metadata in each of the data domains, respectively generate a data map corresponding to each of the data domains, wherein node information of at least one of the data maps is associated with both structured metadata and unstructured metadata; A global data directory is generated based on the data maps corresponding to the second number of data fields.

2. The method according to claim 1, characterized in that Generating a data map corresponding to each data domain based on the metadata in each data domain includes: Extracting, from the data source corresponding to the data domain, association relationships between the metadata, wherein the association relationships include at least one of data flow information, data lineage information, and data dependency information; The data map is established by taking the data resources as nodes and the metadata as node information and based on the association relationship between the metadata.

3. The method according to claim 2, characterized in that Extracting the association relationship between the metadata in the data source corresponding to the data domain includes: When the data source is the AI edge data source, obtaining a first association relationship between the metadata based on the AI model, wherein the first association relationship is stored in the form of a graph; In a case where the data source is not the AI edge data source, a second association relationship between the metadata is determined based on a table structure of the data table.

4. The method according to claim 1, wherein Generating a global data directory based on the data map corresponding to the second number of data fields includes: Creating at least one data product based on each of the data maps, wherein the data product includes at least one of the metadata in the data map and description information corresponding to the metadata; The data product is published to an initial data directory to obtain the global data directory containing the data product.

5. The method according to claim 4, characterized in that The method further comprises: Obtaining the description information input by a first user, wherein the first user is a data owner user corresponding to the data domain where the metadata is located; and / or, In a case where the data source is the AI edge data source, the description information is updated based on the AI model in response to trigger information.

6. The method according to claim 1, characterized in that The dividing the first number of data sources into the second number of data domains includes: Determining the second number of data domains according to the business domains corresponding to the data sources, wherein each business domain corresponds to at least one data domain; The data source is divided into the data domain based on the corresponding relationship between the data source and the business domain and the corresponding relationship between the business domain and the data domain.

7. The method according to claim 1, characterized in that The method further comprises: Based on the data deletion instruction corresponding to any data source, updating the data map, the data product and the global data directory corresponding to the data domain corresponding to any data source; In any of the data sources, the data resource corresponding to the data deletion instruction is deleted.

8. The method according to claim 4, characterized in that The method further comprises: In response to a search instruction input by a second user, determining a target data product corresponding to the search instruction in the data catalog, wherein the second user is a user requesting to use the target data product; Returning product identification information corresponding to the target data product to the second user; In response to a data request instruction input by the second user corresponding to the product identification information, if the second user passes authentication, access rights to the target data resource corresponding to the data request instruction are released to the second user.

9. A data management system based on AI edge devices, characterized in that: The system comprises: a data domain component, deployed on the first computing device, configured to divide the first number of data sources into a second number of data domains, wherein at least one of the data domains includes an AI edge data source and a non-AI edge data source; A data map component, deployed on a second computing device, configured to extract metadata for each of the data domains, wherein the metadata is used to describe attribute information of data resources in the data domain, and the metadata in the AI edge data source includes unstructured data inferred based on an AI model deployed on the edge device; and, based on the metadata in each of the data domains, generate a data map corresponding to the data domain, wherein each node information of at least one of the data maps is associated with both structured metadata and unstructured metadata; A data directory component is deployed on a third computing device and is used to generate a global data directory based on the data maps corresponding to the second number of data domains.

10. A computing device, characterized in that The computing device includes a memory and a processor; the memory and the processor are coupled; The memory is used to store computer instructions; The processor is configured to execute the computer instructions so that the computing device performs the data management method based on an AI edge device as described in any one of claims 1 to 8.