Metadata processing method and device based on large model, equipment and storage medium
Through the metadata processing method based on the large language model, an adaptive acquisition strategy is generated, which solves the problem of inefficient metadata acquisition in traditional methods, and realizes efficient and intelligent metadata acquisition, improving data quality and adaptability.
Patent Information
- Application Number
- CN202510399810.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-11
AI Technical Summary
Traditional metadata acquisition methods are difficult to adapt to diversified and isomerized data sources, resulting in a decline in metadata quality and inefficiency.
The metadata processing method based on the large language model is adopted to determine the target characteristics of the metadata acquisition request, generate target prompt information, and use the large language model to generate target acquisition strategies to achieve intelligent and automated collection of metadata.
It improves the efficiency and quality of metadata acquisition, reduces time and labor costs, supports unified collection of multiple data sources, and ensures the effectiveness and reliability of the acquisition strategy.
Smart Images

Figure CN120296074A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of data processing technologies, and in particular to technologies such as artificial intelligence, large models, and big data. Background Art
[0002] Metadata collection is a key link to ensure data quality and promote the effective management and utilization of data. With the in-depth digital transformation, data sources are becoming increasingly diverse and heterogeneous, and traditional methods are difficult to perform effective data collection, seriously affecting the quality of metadata. Therefore, there is an urgent need for an intelligent metadata collection method. Summary of the Invention
[0003] The present disclosure provides a metadata processing method, apparatus, device, and storage medium based on a large model.
[0004] According to one aspect of the present disclosure, there is provided a metadata processing method based on a large model, including:
[0005] Determine a metadata collection request;
[0006] Obtain a target collection feature corresponding to the metadata collection request;
[0007] Based on the target collection feature, obtain a target prompt message corresponding to the metadata collection request;
[0008] According to the target prompt message, and by using a large language model, obtain a target collection strategy corresponding to the metadata collection request, so as to perform metadata collection based on the target collection strategy to obtain target metadata.
[0009] According to another aspect of the present disclosure, there is provided a metadata processing apparatus based on a large model, including:
[0010] An acquisition unit, configured to determine a metadata collection request;
[0011] A feature determination unit, configured to obtain a target collection feature corresponding to the metadata collection request;
[0012] A strategy generation unit, configured to obtain a target prompt message corresponding to the metadata collection request based on the target collection feature; according to the target prompt message, and by using a large language model, obtain a target collection strategy corresponding to the metadata collection request;
[0013] A data collection unit, configured to perform metadata collection based on the target collection strategy to obtain target metadata.
[0014] According to another aspect of the present disclosure, there is provided an electronic device, including:
[0015] At least one processor; and
[0016] A memory communicatively connected to the at least one processor; wherein,
[0017] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute any method in the embodiments of the present disclosure.
[0018] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute any method in the embodiments of the present disclosure.
[0019] According to another aspect of the present disclosure, there is provided a computer program product, including a computer program, which implements any method in the embodiments of the present disclosure when executed by a processor.
[0020] In this way, the solution of the present disclosure can construct target prompt information according to the target collection characteristics corresponding to the metadata collection request, and use the large language model and the target prompt information to obtain a target collection strategy matching the metadata collection request, and then perform metadata collection to obtain target metadata. The above process realizes the intelligence and automation of metadata collection. In this way, a collection strategy matching the given metadata collection request can be customized, saving the time cost and labor cost required in the metadata collection process, and improving the efficiency and quality of metadata collection.
[0021] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0023] Figure 1 is a schematic flowchart of a metadata processing method based on a large model according to an embodiment of the present application Figure 1 ;
[0024] Figure 2 is a schematic flowchart of a metadata processing method based on a large model according to an embodiment of the present application Figure 2 ;
[0025] Figure 3 is a schematic flowchart of a metadata processing method based on a large model according to an embodiment of the present application Figure 3 ;
[0026] Figure 4It is a flowchart of context structure adjustment based on a preset structure template according to an embodiment of the present application;
[0027] Figure 5 It is a schematic diagram of context structure adjustment based on a sliding window method according to an embodiment of the present application;
[0028] Figure 6 It is an architecture diagram of an intelligent metadata acquisition system based on a large model according to an embodiment of the present application;
[0029] Figure 7(a) is a schematic diagram of multi - terminal interaction in a specific embodiment of a metadata processing method based on a large model according to an embodiment of the present application;
[0030] Figure 7(b) is a schematic flowchart of obtaining a target acquisition strategy in a specific embodiment of a metadata processing method based on a large model according to an embodiment of the present application;
[0031] Figure 8 It is a schematic structural diagram of a metadata processing device based on a large model according to an embodiment of the present application;
[0032] Figure 9 It is a block diagram of an electronic device for implementing the metadata processing method based on a large model in an embodiment of the present disclosure. Specific Embodiments
[0033] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to assist in understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for clarity and conciseness, descriptions of well - known functions and structures are omitted below.
[0034] The term "and / or" in this article is merely a description of the association relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The term "at least one" in this article represents any one of multiple types or any combination of at least two of multiple types. For example, including at least one of A, B, and C can represent selecting any one or more elements from the set composed of A, B, and C. The terms "first" and "second" in this article represent referring to multiple similar technical terms and distinguishing them, and do not mean limiting the order, or limiting that there are only two. For example, the first feature and the second feature refer to two types / two features. The first feature can be one or more, and the second feature can also be one or more.
[0035] In addition, for a better illustration of the present disclosure, numerous specific details are provided in the following specific embodiments. Those skilled in the art should understand that the present disclosure can still be implemented without some specific details. In some instances, methods, means, elements, and circuits well-known to those skilled in the art are not described in detail to highlight the gist of the present disclosure.
[0036] The following explains the related technologies of the embodiments of the present disclosure. The following related technologies can be arbitrarily combined with the technical solutions of the embodiments of the present disclosure as alternative solutions, and they all fall within the protection scope of the embodiments of the present disclosure.
[0037] The solution of the present disclosure provides a metadata processing method based on a large model. By leveraging the reasoning ability of a large language model (LLM, also referred to as a large model), it understands the received metadata collection request and finally obtains a collection strategy (corresponding to the target collection strategy) that matches the metadata collection request. In this way, the intelligence and automation of metadata collection are achieved, and the data quality of metadata collection is improved.
[0038] Specifically, Figure 1 is a schematic flowchart of a metadata processing method based on a large model according to an embodiment of the present application Figure 1 This method is optionally applied to an electronic device, such as a personal computer, a server, a server cluster, or other electronic devices.
[0039] Furthermore, this method at least includes at least some of the following content. As Figure 1 shown, it includes:
[0040] Step S101: Determine the metadata collection request.
[0041] Step S102: Obtain the target collection features corresponding to the metadata collection request.
[0042] Step S103: Based on the target collection features, obtain the target prompt information corresponding to the metadata collection request.
[0043] Step S104: According to the target prompt information and using the large language model, obtain the target collection strategy corresponding to the metadata collection request, and perform metadata collection based on the target collection strategy to obtain the target metadata.
[0044] In this way, the disclosed solution can construct a target prompt message according to the target collection feature corresponding to the metadata collection request, and use the large language model and the target prompt message to obtain a target collection strategy that matches the metadata collection request, and then perform metadata collection to obtain target metadata. The above process realizes the intelligence and automation of metadata collection. In this way, a collection strategy that matches the given metadata collection request can be customized, saving the time cost and labor cost required in the metadata collection process, and improving the efficiency and quality of metadata collection.
[0045] In addition, since the target collection strategy used by the disclosed solution for metadata collection is obtained by means of a large model, in other words, it is obtained by using the world knowledge ability and reasoning ability of the large model, the target metadata collected by the disclosed solution better meets the collection requirements of the metadata collection request. Therefore, the data quality of the collected metadata is greatly improved.
[0046] Furthermore, in a specific example, the metadata collection request may further carry collection task description information, which can be used to describe the type of data source to be collected, etc. Further, the obtaining of the target collection feature corresponding to the metadata collection request described above may specifically include:
[0047] Based on the collection task description information carried by the metadata collection request, a target collection feature representing the type of the data set to be collected is obtained. In this way, it is convenient to automatically give a metadata collection strategy adapted to the data source according to the type of the data source, thereby further improving the intelligence degree of metadata collection and further improving the data quality of the collected metadata.
[0048] Furthermore, the disclosed solution does not limit the quantity and type of data sources to be collected. In other words, the disclosed solution can also support batch downloading of multiple data sources. In this way, the problem of unable to uniformly collect metadata caused by data being distributed in different heterogeneous systems (such as distributed in relational databases, non-relational structured query language (NoSQL) databases, file systems, systems based on application programming interfaces (APIs), etc.) and the inconsistent metadata collection standards of each system can be effectively solved. Furthermore, the intelligence degree of metadata collection is further improved, and at the same time, the collection efficiency of metadata is also improved.
[0049] Further, in one example, to verify the effectiveness of the metadata collection strategy and ensure accurate and effective metadata collection, the solution of the present disclosure can also verify the obtained target collection strategy; for example, in one example, after obtaining the target collection strategy, the above-mentioned method further includes:
[0050] Based on the target collection strategy, collect test data;
[0051] Use the collected test data for metadata collection to obtain a collection test result.
[0052] That is to say, after obtaining the target collection strategy, according to the target collection strategy, the collection test data that can apply to the target collection strategy can be determined; then, using the target collection strategy, perform metadata collection to obtain a collection result, that is, a collection test result. In this way, it is convenient to evaluate the effectiveness of the target collection strategy according to the collection test result.
[0053] Here, it should be noted that the verification of the above-mentioned target collection strategy can not only verify the data quality of the collection result, but also verify the logic and grammar accuracy of the target collection strategy, etc. In this way, it provides strong support for subsequent efficient and accurate metadata collection.
[0054] At this time, the above-mentioned metadata collection based on the target collection strategy to obtain target metadata (for example, step S104 above) can specifically include:
[0055] When the collection test result meets the preset requirements, perform metadata collection based on the target collection strategy to obtain target metadata.
[0056] It should be noted that if the collection test result does not meet the preset requirements, the target collection strategy can be readjusted until the collection test result meets the preset requirements. For example, in one example, if the collection test result does not meet the preset requirements, return to step S103, re-call the large language model to update the target collection strategy, and then update the collection test data based on the updated target collection strategy until the collection test result meets the preset requirements.
[0057] In this way, the solution of the present disclosure realizes the intelligent verification of the target collection strategy, ensures the availability and reliability of the target collection strategy, effectively reduces the error rate of metadata collection, and thus provides strong support for effectively improving the collection success rate and the data quality of the collection result.
[0058] Figure 2 It is a schematic flow of a metadata processing method based on a large model according to an embodiment of the present application Figure 2。This method can optionally be applied to electronic devices, such as personal computers, servers, server clusters, and other electronic devices. It can be understood that the relevant content of the above Figure 1 shown method can also be applied to this example, and the relevant associated content will not be elaborated in this example.
[0059] Furthermore, this method includes at least some of the following content. Specifically, as Figure 2 shown, it includes:
[0060] Step S201: Determine a metadata collection request.
[0061] Step S202: Obtain the target collection features corresponding to the metadata collection request.
[0062] Step S203: Based on the target collection features, obtain the target prompt information corresponding to the metadata collection request.
[0063] Step S204: According to the target prompt information and using a large language model, generate an initial collection strategy corresponding to the metadata collection request.
[0064] For example, in one example, when using a large language model and the target prompt information for reasoning, the large language model decomposes and reasons about the collection features, collection tasks, etc. described in the target prompt information in a Chain of Thought manner, such as: (1) Analyze the data source: Analyze the type, structure, and access method of the data source to be collected; (2) Understand the collection requirements: Determine the collection scope and its specification requirements; (3) Determine the collection scheme: According to the characteristics of the data source to be collected, determine a suitable collection method; (4) Generate a collection strategy: Generate a specific collection strategy; (5) Strategy verification: Verify the generated collection strategy to obtain an effective initial collection strategy. In this way, by leveraging a large model to intelligently understand the collection requirements and then quickly provide a matching collection strategy, it provides strong support for improving the collection efficiency and the data quality of the collected metadata. Moreover, the above process does not require manual intervention, so it also saves the labor cost required for the metadata collection process.
[0065] Step S205: Based on the initial collection strategy, obtain a target collection strategy that meets the preset collection rules.
[0066] Step S206: Perform metadata collection based on the target collection strategy to obtain target metadata.
[0067] In this way, the present disclosure provides a specific solution for obtaining a target collection strategy using a large language model, that is, first obtaining an initial collection strategy using a large model, and then obtaining a target collection strategy that meets the preset collection rules based on the initial collection strategy. The above process can intelligently generate a target collection strategy for a metadata collection request without manual intervention. In this way, the time cost and labor cost required for the entire metadata collection process are effectively saved, thereby improving the efficiency and quality of metadata collection. At the same time, it can also flexibly meet the collection requirements of complex data sources.
[0068] Furthermore, in a specific example, the target collection strategy can be obtained in the following manner; specifically, that is, based on the initial collection strategy described above, obtaining a target collection strategy that meets the preset collection rules (for example, step S205 above) can specifically include:
[0069] Step S205-1: Determine a preset rule template that matches the metadata collection request.
[0070] Step S205-2: Based on the initial collection strategy, configure the preset rule template that matches the metadata collection request to obtain a target collection strategy that meets the preset collection rules.
[0071] It should be noted that after obtaining the initial collection strategy using the large language model, in order to improve the effectiveness and reliability of the metadata collection result, the initial collection strategy can also be adjusted and optimized according to the preset rule template corresponding to the metadata collection request. For example, according to the initial collection strategy, parameterize the configuration of the preset rule template. Further, a rule conversion engine can be used to process the parameterized preset rule template, and then a target collection strategy that meets the preset collection rules can be obtained.
[0072] In this way, the present disclosure provides a specific solution for obtaining a target collection strategy, that is, using a preset rule template that matches the metadata collection request to enhance the initial collection strategy to obtain a target collection strategy that meets the preset collection rules. In this way, the effectiveness and reliability of the obtained target collection strategy are further enhanced, thereby laying a foundation for improving the efficiency of subsequent metadata collection and the reliability of the collection result.
[0073] Figure 3 is a schematic flowchart of a metadata processing method based on a large model according to an embodiment of the present application Figure 3 . This method can optionally be applied to electronic devices, such as personal computers, servers, server clusters, and other electronic devices. It can be understood that the relevant content of the above Figure 1 and Figure 2 The relevant content of the shown method can also be applied to this example, and the relevant associated content will not be repeated in this example.
[0074] Further, the method includes at least part of the following content. Specifically, as Figure 3 shown, it includes:
[0075] Step S301: Determine the metadata collection request.
[0076] Step S302: Determine the initial key information in the metadata collection request.
[0077] Step S303: Perform data augmentation processing on the initial key information in the metadata collection request to obtain the initial collection features.
[0078] For example, in an example, the data augmentation processing can be performed in the following manner. Specifically, the above-mentioned data augmentation processing of the initial key information in the metadata collection request to obtain the initial collection features (for example, step S303 above) can specifically include:
[0079] Step S303-1: Based on the similarity, screen out multiple target key information that matches the initial key information in the metadata collection request from the preset knowledge base.
[0080] Step S303-2: Based on the multiple target key information, obtain the initial collection features.
[0081] That is to say, in this data augmentation process, first calculate the similarity between the initial key information in the metadata collection request and each candidate knowledge information in the preset knowledge base, then screen out multiple target key information whose similarity meets the preset similarity requirements from the preset knowledge base to expand the target key information; finally, perform data integration on the multiple target key information, and then obtain the initial collection features. In this way, the data augmentation of the initial key information is effectively realized, the integrity and richness of the key information required for metadata collection are effectively improved, and an effective basis is provided for accurately obtaining the target collection strategy subsequently.
[0082] Step S304: Adjust the context structure of at least part of the initial collection features to obtain target collection features that meet the preset structure requirements.
[0083] Step S305: Based on the target collection features, obtain the target prompt information corresponding to the metadata collection request.
[0084] It should be noted that in one example, after obtaining the target prompt information, the number of tokens of the target prompt information can also be adjusted so that the adjusted target prompt information meets the input requirements of the large language model. For example, a dynamic cropping strategy based on token limitation can be used to adjust the number of tokens of the obtained target prompt information. In this way, while ensuring information integrity, the input scale of the large language model is effectively controlled, providing strong support for quickly obtaining the acquisition strategy subsequently.
[0085] Step S306: According to the target prompt information and using the large language model, obtain the target acquisition strategy corresponding to the metadata acquisition request.
[0086] Step S307: Based on the target acquisition strategy, perform metadata acquisition to obtain the target metadata.
[0087] In this way, in the solution of the present disclosure, since the initial key information in the metadata acquisition request can be effectively data-augmented, the richness of the key information related to the metadata acquisition request is greatly enriched, thereby providing data support for quickly and accurately obtaining the acquisition strategy. Moreover, since at least part of the content in the initial acquisition features obtained after data augmentation is adjusted in the context structure, the context relevance of the acquisition features is effectively optimized, and the information quality of the obtained target acquisition features is effectively improved, providing strong support for the large language model to generate an acquisition strategy applicable to the metadata acquisition request, and thus providing strong support for effectively improving the acquisition success rate and the data quality of the acquisition result.
[0088] Furthermore, in a specific example, the following method can also be used to adjust the context structure of at least part of the initial acquisition features, which specifically includes:
[0089] Method 1: The above-mentioned adjustment of the context structure of at least part of the initial acquisition features to obtain the target acquisition features that meet the preset structure requirements (for example, step S304) specifically includes:
[0090] Step S304-1: Obtain a preset structure template. Here, the preset structure template can represent the required structure relationship between contexts.
[0091] Step S304-2: Configure the preset structure template according to the initial acquisition features to obtain the target acquisition features that meet the preset structure requirements.
[0092] That is to say, as Figure 4As shown, in Method 1, after obtaining a preset structure template that can represent the required structural relationship between contexts, the content of each part in the initial acquisition features can be logically understood first to obtain an understanding result; then, according to this understanding result, the content of each part of the initial acquisition features is configured into the preset structure template to adjust the context structure to ensure the coherence of each part of the content, and thus obtain the target acquisition features with a more compact context structure.
[0093] Method 2: The above-mentioned adjustment of the context structure of at least part of the content in the initial acquisition features to obtain the target acquisition features that meet the preset structure requirements (for example, step S304) specifically includes:
[0094] Using a sliding window method to adjust the context structure of at least part of the content in the initial acquisition features to obtain the target acquisition features that meet the preset structure requirements.
[0095] For example, taking the initial acquisition features including Content 1 to Content 6 as an example, at this time, as Figure 5 shown, the steps of adjusting the context structure using the sliding window method can specifically include:
[0096] Step a: Initialization. For example, set the window length to 3, set the starting position of the window at the head of the initial acquisition features, for example, at Content 1 on the left, and set the sliding direction of the window as from the head to the tail, for example, to the right, and the preset step size to 2.
[0097] It should be noted that the length of the window should be able to contain key context information, and moreover, it should not be too large to avoid including irrelevant information; further, when the window is at the starting position, the content 1, content 2, and content 3 within the window are the current processed context.
[0098] In addition, the division of the content can be based on each paragraph or sentence in the initial acquisition features, and the present disclosure scheme does not limit this.
[0099] Step b: Perform a correlation analysis on the content of each part within the current window to obtain an analysis result.
[0100] Step c: According to the analysis result, adjust the content of each part within the window (such as sorting, merging, splitting, etc.) to obtain a sub-adjustment result. For example, {Content 1, Content 3, Content 2}.
[0101] Step d: According to the sub-adjustment result, update the context structure of each part of the content in the initial acquisition features. For example, the updated current context structure is {Content 1, Content 3, Content 2, Content 4, Content 5, Content 6}.
[0102] Step e: Determine whether the number of steps by which the current window can slide is less than the preset step length; if so, proceed to step f; otherwise, proceed to step h.
[0103] Step f: Determine whether the number of steps by which the current window can slide is greater than 0; if so, proceed to step g; otherwise, output the current context structure of each part in the initial acquisition feature.
[0104] Step g: Slide the window along the preset sliding direction according to the number of steps by which it can slide to obtain a new current window, and return to step b.
[0105] For example, when the number of steps by which it can slide is 1, slide the window 1 step along the preset sliding direction according to the number of steps by which it can slide to obtain a new current window. The three parts of content selected by the new current window are content 5, content 4, and content 6.
[0106] Step h: Slide the window along the preset sliding direction based on the preset step length to obtain a new current window, and return to step b.
[0107] For example, after the first sorting is completed, it is determined that the number of steps by which it can slide (such as 3) is greater than the preset step length. At this time, slide the window along the preset sliding direction based on the preset step length to obtain a new current window. The three parts of content selected by the new current window are content 2, content 4, and content 5.
[0108] Using the above steps, the context structure adjustment of each part in the initial acquisition feature can be completed to obtain the target acquisition feature.
[0109] In this way, the solution of the present disclosure can adjust the context structure of at least part of the content in the initial acquisition feature. In other words, it can achieve the adjustment of the context structure to ensure the coherence of each part of the content, and then obtain the target acquisition feature with a more compact context structure. Thus, it provides strong support for the subsequent large model to capture key information, and further provides strong support for effectively improving the acquisition success rate and the data quality of the acquisition result.
[0110] The following further elaborates on the present disclosure solution with specific examples; the present disclosure solution provides an intelligent metadata collection method based on a large model. Specifically, it can generate target prompt information corresponding to the received metadata collection request, and use the large language model and the target prompt information to obtain an initial collection strategy for the metadata collection request. Furthermore, using the initial collection strategy, it configures a preset rule template pre-determined to match the metadata collection request to obtain a target collection strategy that meets the preset collection rules, so as to use the target collection strategy to collect metadata and obtain target metadata. In this way, the intelligent and automated process of metadata collection is realized, greatly improving the efficiency and data quality of metadata collection.
[0111] Specifically, as Figure 6 shown, the present disclosure solution provides an intelligent metadata collection system based on a large model. The intelligent metadata collection system may specifically include: a data source layer, a core processing layer, and a rule processing layer; among them, the data source layer includes a relational database, a NoSQL database, a file system, and an API interface; the core processing layer includes a knowledge base construction module, a Retrieval-Augmented Generation (RAG) module, and a large model inference module; the rule processing layer includes a collection strategy generation module and a quality verification module. Each module works in coordination to realize the intelligent collection process of metadata.
[0112] Furthermore, the above modules are described as follows:
[0113] (1) Knowledge base construction module: The knowledge base construction module is responsible for constructing and maintaining a professional knowledge base required for metadata collection. The professional knowledge base may specifically include:
[0114] a. Metadata standard library
[0115] The metadata standard library contains standardized information such as metadata format specifications, field definitions, and data type mapping relationships for various data sources.
[0116]
[0117] b. Collection rule library
[0118] The collection rule library is used to store existing collection rule templates (corresponding to the above preset rule templates), best practice cases, and common error handling methods, etc.
[0119]
[0120] c. Data source schema library
[0121] The data source schema library is used to record the access methods and interfaces of different types of data sources
[0122] Schema information such as specifications and certification requirements.
[0123] (2) Retrieval-Augmented Generation Module: This retrieval-augmented generation module is used to perform data augmentation and context recombination on the initial key information in the metadata collection request (corresponding to the above context structure adjustment). The specific processing process is as follows:
[0124] Step 2-a: Using a similarity retrieval method based on a vector database, screen out multiple target key information associated with the initial key information from the professional knowledge base, and perform preliminary integration on the multiple target key information to obtain initial collection features.
[0125]
[0126]
[0127] Step 2-b: Using a sliding window method, perform context recombination on at least part of the content in the initial collection features to obtain target collection features to ensure information coherence.
[0128]
[0129] Step 2-c: Based on the target collection features, obtain target prompt information.
[0130] Step 2-d: Using a dynamic cropping strategy based on token limits, adjust the number of tokens in the target prompt information to ensure information integrity while effectively controlling the input scale of the large model.
[0131]
[0132]
[0133] (3) Large Model Inference Module: This large model inference module performs inference based on the target prompt information in a Chain of Thought manner. The specific inference steps include:
[0134] Step 3-a: Parse the data source: Parse the type, structure, and access method of the data source to be collected.
[0135]
[0136] Step 3-b: Understand the collection requirements: Clearly define the metadata items to be collected and their specification requirements.
[0137]
[0138] Step 3-c: Determine the collection plan: Select an appropriate collection method according to the characteristics of the data source.
[0139]
[0140] Step 3-d: Generate a collection strategy: Generate a specific collection strategy.
[0141] Step 3-e: Verification and evaluation: Verify the effectiveness of the generated acquisition strategy to
[0142] obtain an initial acquisition strategy that passes the verification.
[0143] (4) Acquisition strategy generation module: Based on the templated solution and the initial acquisition strategy, automatically generate the target acquisition strategy. The specific generation steps include:
[0144] Step 4-a: Determine the basic rule template that matches the metadata acquisition request (corresponding to
[0145] the above preset rule templates). For example, select the basic rule template according to the type of data source.
[0146] Step 4-b: Perform parametric configuration on the basic rule template according to the initial acquisition strategy to
[0147] obtain the basic rule template after parametric configuration.
[0148] Step 4-c: Use the rule conversion engine to convert the basic rule template after parametric configuration to
[0149] obtain the target acquisition strategy that meets the preset acquisition rules.
[0150] Here, the solution of the present disclosure supports multiple expression forms. For example, it supports SQL queries, application API calls, file parsing scripts, etc. For example, obtain the target acquisition strategy that meets the SQL query requirements. It can be understood that in actual applications, the appropriate expression form can be selected according to the needs of the actual application scenario. The solution of the present disclosure does not limit this.
[0151]
[0152]
[0153] (5) Quality verification module: This quality verification module mainly verifies the target acquisition strategy through the established multi-level verification mechanism. The specific verification contents include:
[0154] a. Syntax check: Detect the correctness of the syntax of the target acquisition strategy;
[0155] b. Logic verification: Detect the internal consistency of the target acquisition strategy;
[0156] c. Sample test: Verify the effectiveness of the target acquisition strategy using the acquisition test data.
[0157] Here, the acquisition test data is the data selected according to the target acquisition strategy;
[0158] d. Result evaluation: Evaluate the quality of the collection results obtained by using the target collection strategy for metadata collection.
[0159] It should be noted that the solution of the present disclosure can batch process multiple data sources. In other words, it can achieve unified collection of metadata for data sources of different systems. For example, different data sources in the data source layer in this example are distributed in different heterogeneous systems, such as relational databases, NoSQL databases, file systems, API interface-based systems, etc.).
[0160] Furthermore, as shown in FIGS. 7(a) and 7(b), the specific steps of the metadata collection method of the solution of the present disclosure include:
[0161] Step S701: Receive a metadata collection request initiated by a target object.
[0162] Here, the metadata collection request carries task description information for a metadata collection task for one or more data sources.
[0163] Step S702: Parse the task description information in the metadata collection request to extract the initial key information in the metadata collection request.
[0164] Step S703: Use the retrieval enhancement module to retrieve relevant knowledge (corresponding to the above multiple target key information) that matches the initial key information in the metadata collection request from a preset knowledge base (such as the above-mentioned professional knowledge base) based on similarity, and perform preliminary integration on the retrieved relevant knowledge to obtain initial collection features.
[0165] Step S704: Use the retrieval enhancement module to perform context recombination on at least part of the content in the initial collection features. For example, use a sliding window method to perform context recombination on at least part of the content in the initial collection features to obtain target collection features.
[0166] Here, in one example, as shown in FIG. 7(b), the solution of the present disclosure can also first obtain a preset structure template, and then configure the preset structure template according to the initial collection features to perform context structure adjustment on at least part of the content in the initial collection features.
[0167] Step S705: Construct target prompt information corresponding to the metadata collection request based on the target collection features, and then use the large model to perform reasoning in a chain-of-thought manner to obtain the initial collection strategy corresponding to the metadata collection request.
[0168] Step S705: Based on the target collection features, construct the target prompt information corresponding to the metadata collection request, and then use the large model to perform reasoning in a chain-of-thought manner to obtain the initial collection strategy corresponding to the metadata collection request.
[0169] Here, in one example, as shown in FIG. 7(b), after obtaining the target prompt information, the number of tokens of the target prompt information can also be adjusted so that the adjusted target prompt information meets the input requirements of the large model.
[0170] Step S706: Use the collection strategy generation module (such as the collection strategy generator) to select a preset rule template that matches the metadata collection request, and use the initial collection strategy to configure the preset rule template that matches the metadata collection request to obtain a target collection strategy that meets the preset collection rules.
[0171] Step S707: Use the quality verification module to verify the obtained target collection strategy. For example, obtain collection test data that can apply the target collection strategy from the data source, and then use the collection test data to perform metadata collection to obtain a collection test result.
[0172] Step S708: When the collection test result meets the preset requirements, use the target collection strategy to perform metadata collection to obtain target metadata.
[0173] Here, in this example, when the collection test result does not meet the preset requirements, adjust the knowledge retrieval strategy (or adjust the preset structure template, etc.) to re-obtain the target collection strategy.
[0174] Furthermore, the solution of the present disclosure can also be widely applied to the following scenarios:
[0175] (1) Data asset management scenario
[0176] (1.1) Automatically discover and collect data asset information in different systems;
[0177] (1.2) Establish a data asset catalog to support asset evaluation and governance;
[0178] (1.3) Track data asset changes and achieve dynamic updates.
[0179] (2) Data integration scenario
[0180] (2.1) Quickly identify and understand the data structure of the source system;
[0181] (2.2) Automatically generate data mapping rules;
[0182] (2.3) Support data synchronization between heterogeneous systems.
[0183] (3) Data governance scenario
[0184] (3.1) Collect metadata related to data quality;
[0185] (3.2) Support monitoring of data standard execution;
[0186] (3.3) Auxiliary data lineage analysis.
[0187] In summary, compared with the traditional metadata collection method, the disclosed solution has the following advantages:
[0188] First, it improves the collection efficiency. The disclosed solution can utilize the inference ability of the large model to intelligently generate a metadata collection strategy (i.e., the target collection strategy) for the metadata collection request. The above process requires no manual intervention, saving the time cost and labor cost required for the entire metadata collection process, thus significantly improving the metadata collection efficiency.
[0189] Second, it intelligently generates collection strategies. The disclosed solution can, based on the metadata collection request and with the help of the large model, intelligently obtain the metadata collection strategy. The above intelligent generation process can design a collection strategy adapted to different types of data sources. In short, the disclosed solution enhances the adaptability between the collection strategy and the data source. In other words, when a new data source is added or the structure of an existing data source changes, it can be quickly adapted without manually developing a new collector, effectively solving the problem of low efficiency caused by manual adaptation and reducing the maintenance cost required for the metadata collection process. At the same time, it also improves the quality and reliability of the collection results. For example, the target metadata collected is complete and accurate, and...
[0190] Third, it has an intelligent verification mechanism. After obtaining the metadata collection strategy to be used, the disclosed solution can also verify the effectiveness of the metadata collection strategy. In this way, it effectively reduces the error rate of metadata collection and improves the integrity and reliability of the metadata collection results.
[0191] The disclosed solution also provides a metadata processing device based on a large model, as Figure 8 shown, including:
[0192] An acquisition unit 801 for determining a metadata collection request;
[0193] A feature determination unit 802 for obtaining the target collection features corresponding to the metadata collection request;
[0194] A strategy generation unit 803 for, based on the target collection features, obtaining the target prompt information corresponding to the metadata collection request; according to the target prompt information, and using a large language model, obtaining the target collection strategy corresponding to the metadata collection request;
[0195] A data collection unit 804 for performing metadata collection based on the target collection strategy to obtain target metadata.
[0196] In a specific example of the present disclosure solution, the policy generation unit is specifically configured to:
[0197] Generate an initial acquisition policy corresponding to the metadata acquisition request according to the target prompt information and by using a large language model;
[0198] Based on the initial acquisition policy, obtain a target acquisition policy that meets the preset acquisition rules.
[0199] In a specific example of the present disclosure solution, the policy generation unit is specifically configured to:
[0200] Determine a preset rule template that matches the metadata acquisition request;
[0201] Based on the initial acquisition policy, configure the preset rule template that matches the metadata acquisition request to obtain a target acquisition policy that meets the preset acquisition rules.
[0202] In a specific example of the present disclosure solution, the feature determination unit is specifically configured to:
[0203] Determine the initial key information in the metadata acquisition request;
[0204] Perform data augmentation processing on the initial key information in the metadata acquisition request to obtain initial acquisition features;
[0205] Adjust the context structure of at least part of the initial acquisition features to obtain target acquisition features that meet the preset structure requirements.
[0206] In a specific example of the present disclosure solution, the feature determination unit is specifically configured to:
[0207] Obtain a preset structure template, where the preset structure template can represent the required structure relationship between contexts; configure the preset structure template according to the initial acquisition features to obtain target acquisition features that meet the preset structure requirements.
[0208] In a specific example of the present disclosure solution, the feature determination unit is specifically configured to:
[0209] Use a sliding window method to adjust the context structure of at least part of the initial acquisition features to obtain target acquisition features that meet the preset structure requirements.
[0210] In a specific example of the present disclosure solution, the feature determination unit is specifically configured to:
[0211] Based on similarity, screen out multiple target key information that matches the initial key information in the metadata acquisition request from the preset knowledge base;
[0212] Based on multiple target key information, initial acquisition features are obtained.
[0213] In a specific example of the present disclosure solution, it further includes: a verification unit; wherein,
[0214] The verification unit is used to obtain acquisition test data based on a target acquisition strategy; perform metadata acquisition using the acquisition test data to obtain an acquisition test result;
[0215] The data acquisition unit is specifically used to, when the acquisition test result meets the preset requirements, perform metadata acquisition based on the target acquisition strategy to obtain target metadata.
[0216] For the specific functions and example descriptions of the units of the device in the embodiments of the present disclosure, reference may be made to the relevant descriptions of the corresponding steps in the above method embodiments, which will not be elaborated herein.
[0217] In the technical solution of the present disclosure, the acquisition, storage, and application of user personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0218] According to the embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0219] Figure 9 A schematic block diagram of an example electronic device 900 that can be used to implement the embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital assistants, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described herein and / or required.
[0220] As Figure 9 shown, the device 900 includes a computing unit 901, which can execute various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 902 or a computer program loaded from a storage unit 908 into a random access memory (RAM) 903. In the RAM 903, various programs and data required for the operation of the device 900 can also be stored. The computing unit 901, the ROM 902, and the RAM 903 are connected to each other through a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.
[0221] Multiple components in device 900 are connected to I / O interface 905, including: input unit 906, such as a keyboard, mouse, etc.; output unit 907, such as various types of displays, speakers, etc.; storage unit 908, such as a disk, optical disc, etc.; and communication unit 909, such as a network card, modem, wireless communication transceiver, etc. Communication unit 909 allows device 900 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0222] Computing unit 901 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Computing unit 901 executes the various methods and processes described above, such as the metadata processing method based on a large model. For example, in some embodiments, the metadata processing method based on a large model can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed onto device 900 via ROM 902 and / or communication unit 909. When the computer program is loaded into RAM 903 and executed by computing unit 901, one or more steps of the metadata processing method based on a large model described above can be executed. Alternatively, in other embodiments, computing unit 901 can be configured to execute the metadata processing method based on a large model in any other suitable manner (e.g., by means of firmware).
[0223] The various embodiments of the systems and techniques described above in this article can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-chip systems (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: implemented in one or more computer programs, which can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special or general programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0224] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0225] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0226] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).
[0227] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.
[0228] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client-server relationship is generated by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, can also be a server of a distributed system, or a server incorporating a blockchain.
[0229] It should be understood that various forms of the flow shown above can be used, steps can be reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this is not limited herein.
[0230] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of this disclosure shall be included within the protection scope of this disclosure.
Claims
1. A method for metadata processing based on a large model, comprising: Determining a metadata collection request; Obtaining a target collection feature corresponding to the metadata collection request; Based on the target collection feature, obtaining a target prompt message corresponding to the metadata collection request; According to the target prompt message and using a large language model, obtaining a target collection strategy corresponding to the metadata collection request, so as to perform metadata collection based on the target collection strategy to obtain target metadata.
2. The method according to claim 1, wherein The step of obtaining a target collection strategy corresponding to the metadata collection request according to the target prompt message and using a large language model includes: According to the target prompt message and using a large language model, generating an initial collection strategy corresponding to the metadata collection request; Based on the initial collection strategy, obtaining a target collection strategy that meets the preset collection rules.
3. The method according to claim 2, wherein The step of obtaining a target collection strategy that meets the preset collection rules based on the initial collection strategy includes: Determining a preset rule template that matches the metadata collection request; Based on the initial collection strategy, configuring the preset rule template that matches the metadata collection request to obtain a target collection strategy that meets the preset collection rules.
4. The method according to any one of claims 1 to 3, wherein, The step of obtaining a target collection feature corresponding to the metadata collection request includes: Determining initial key information in the metadata collection request; Performing data augmentation processing on the initial key information in the metadata collection request to obtain an initial collection feature; Performing context structure adjustment on at least part of the content in the initial collection feature to obtain a target collection feature that meets the preset structure requirements.
5. The method according to claim 4, wherein The step of performing context structure adjustment on at least part of the content in the initial collection feature to obtain a target collection feature that meets the preset structure requirements includes: Obtaining a preset structure template, and the preset structure template can represent the required structure relationship between contexts; According to the initial collection feature, configuring the preset structure template to obtain a target collection feature that meets the preset structure requirements.
6. The method according to claim 4, wherein The step of performing context structure adjustment on at least part of the content in the initial collection feature to obtain a target collection feature that meets the preset structure requirements includes: Using a sliding window method to perform context structure adjustment on at least part of the content in the initial collection feature to obtain a target collection feature that meets the preset structure requirements.
7. The method according to any one of claims 4-6, wherein The step of performing data augmentation processing on the initial key information in the metadata collection request to obtain an initial collection feature includes: Based on similarity, screening out multiple target key information that matches the initial key information in the metadata collection request from a preset knowledge base; Based on the multiple target key information, obtaining an initial collection feature.
8. The method according to any one of claims 1-7, the method further includes: Based on the target collection strategy, obtaining collection test data; Using the collection test data to perform metadata collection to obtain a collection test result; Wherein, the step of performing metadata collection based on the target collection strategy to obtain target metadata includes: When the collection test result meets the preset requirements, performing metadata collection based on the target collection strategy to obtain target metadata.
9. A metadata processing device based on a large model, comprising: An acquisition unit for determining a metadata collection request; A feature determination unit for obtaining a target collection feature corresponding to a metadata collection request; A policy generation unit for obtaining a target prompt message corresponding to the metadata collection request based on the target collection feature; According to the target prompt message and by using a large language model, obtaining a target collection policy corresponding to the metadata collection request; A data collection unit for collecting metadata based on the target collection policy to obtain target metadata.
10. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1-8.
11. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-8.
12. A computer program product, comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1-8.