A large model-based data resource catalog generation method, device and medium

By establishing a dialogue window between the server and the user, and using a large model to automatically generate a data resource catalog for business database tables, the problem of low efficiency and poor accuracy of manual operation in existing technologies is solved, and efficient and accurate catalog generation is achieved.

CN122173626APending Publication Date: 2026-06-09SHENZHEN TAIJI SOFTWARE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHENZHEN TAIJI SOFTWARE CO LTD
Filing Date
2025-12-29
Publication Date
2026-06-09

AI Technical Summary

Technical Problem

The existing business database table data resource catalog construction relies on manual operation, resulting in long generation time, low efficiency, and susceptibility to differences in staff cognition and negligence.

Method used

By establishing a dialogue window between the server and the user, and utilizing large models to generate a data resource catalog, including keyword extraction, intent modeling, knowledge base querying, and large model generation techniques, a catalog of business database tables is automatically generated.

Benefits of technology

It reduces the generation time of data resource catalogs, improves generation efficiency, and ensures the accuracy of the catalogs, avoiding errors and inconsistencies caused by manual operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122173626A_ABST
    Figure CN122173626A_ABST
Patent Text Reader

Abstract

This application relates to the fields of data management technology and artificial intelligence technology. It discloses a method, device, and medium for generating a data resource catalog based on a large model. The method includes: obtaining the identification information and remarks text of business database tables uploaded by a user terminal device; constructing catalog requirements for business database tables by combining the resource catalog structure, resource catalog list, resource catalog content specifications, core field names, core field types, core field lengths, and core field value ranges; inputting the identification information of the business database tables into a database query engine; retrieving the content of the business database tables through the query engine; inputting the content of the business database tables, the catalog requirements, the identification information, and the remarks text into a large model; and generating a data resource catalog for the business database tables through the large model. This application can improve the efficiency of data resource catalog generation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of data management technology and artificial intelligence technology, and in particular to a method, device and medium for generating a data resource catalog based on a large model. Background Technology

[0002] A business database table is a structured data storage unit designed to meet business needs and is oriented towards business scenarios. The core logic of a business database table is to transform various types of information in business activities into a standardized two-dimensional table structure.

[0003] The construction of data resource catalogs for existing business database tables largely relies on manual operation. This manual approach requires staff to examine each business database table one by one and manually enter relevant information to create the data resource catalog. Therefore, this manual method increases the time required to generate the data resource catalog and hinders efficiency. Thus, how to generate data resource catalogs for business database tables is a pressing technical problem that needs to be solved. Summary of the Invention

[0004] This application provides a method, device, and medium for generating a data resource catalog based on a large model, in order to solve the aforementioned technical problem of how to generate a data resource catalog for a business database table.

[0005] In a first aspect, embodiments of this application provide a method for generating a data resource catalog based on a large model, applied to a server-side device. The data resource catalog generation method includes: Establish a dialog window between the server-side device and the user-side device. Through the dialog window, obtain the identification information and remarks text of the business database table uploaded by the user-side device. Using a keyword extraction algorithm, keywords are extracted from the identifier information and the memo text of the business database table, respectively, to generate keywords for the identifier information and keywords for the memo text of the business database table. Input the keywords of the identification information of the business database table and the keywords of the notes text of the business database table into the intent model. The intent model outputs the intent label of the user terminal device. When the intent label is a query label, the query label, the keywords of the identification information of the business database table and the keywords of the notes text of the business database table are combined to form the intent description information of the user terminal device. The feature vectors of intent description information are queried through the query interface of the knowledge base. The similarity between each vector data in the knowledge base and the feature vector of intent description information is obtained. The vector data with the highest similarity is selected as the target vector data. The storage content associated with the target vector data is extracted. The resource directory structure, resource directory list, resource directory content specifications, core field name, core field type, core field length, and core field value range are obtained from the storage content. The resource directory structure, resource directory list, resource directory content specifications, core field names, core field types, core field lengths, and core field value ranges constitute the directory requirements for the business database tables. The identification information of the business database tables is input into the database query engine, which retrieves the content of the business database tables. The content of the business database tables, the directory requirements of the business database tables, the identification information of the business database tables, and the remarks text of the business database tables are input into the large model, which generates the data resource directory of the business database tables.

[0006] In one possible implementation of the first aspect, establishing a dialog window between the server device and the user device, and obtaining the identification information and remarks text of the business database table uploaded by the user device through the dialog window, includes: Establish a dialogue window between the server device and the user device, and receive interactive information input by the user device through the information input area of ​​the dialogue window; Semantic parsing of the interactive information yields the identifier information and remarks text of the business database tables uploaded by the user device.

[0007] In one possible implementation of the first aspect, the keywords of the identification information of the business database table and the keywords of the remarks text of the business database table are input into the intent model, and the intent model outputs the intent tag of the user terminal device. When the intent tag is a query tag, the query tag, the keywords of the identification information of the business database table, and the keywords of the remarks text of the business database table are combined to form the intent description information of the user terminal device, including: Obtain the model file, load the intent model through the model file, and input the keywords of the identification information of the business database table and the keywords of the notes text of the business database table into the intent model. The intent model outputs the intent label of the user terminal device. When the intent label is a query label, the intent description information of the user terminal device is composed of the query label, the keywords of the identification information of the business database table, and the keywords of the remarks text of the business database table.

[0008] In one possible implementation of the first aspect, the feature vector of the intent description information is queried through the query interface of the knowledge base to obtain the similarity between each vector data in the knowledge base and the feature vector of the intent description information. The vector data with the highest similarity is selected as the target vector data. The storage content associated with the target vector data is extracted, and the resource directory structure, resource directory list, resource directory content specifications, core field names, core field types, core field lengths, and core field value ranges are obtained from the storage content, including: The intent description information is encoded using a feature encoding model to generate a feature vector of the intent description information, and then the feature vector of the intent description information is submitted to the query interface of the knowledge base. The feature vectors of intent description information are queried through the knowledge base query interface. The similarity between each vector data in the knowledge base and the feature vector of intent description information is obtained. The vector data with the highest similarity is selected as the target vector data. The storage content associated with the target vector data is extracted. The resource directory structure, resource directory list, resource directory content specifications, core field names, core field types, core field lengths, and core field value ranges are obtained from the storage content.

[0009] In one possible implementation of the first aspect, the process of constructing a business database table directory requirement by including the resource directory structure, resource directory list, resource directory content specifications, core field names, core field types, core field lengths, and core field value ranges; inputting the business database table identification information into the database query engine; retrieving the business database table content through the query engine; inputting the business database table content, business database table directory requirements, business database table identification information, and business database table comment text into the large model; and generating a data resource directory for the business database table through the large model, including: The resource directory structure, resource directory list, resource directory content specifications, core field names, core field types, core field lengths, and core field value ranges constitute the directory requirements for the business database table. The identification information of the business database table is input into the database query engine, and the content of the business database table is retrieved through the query engine. The content of the business database table includes the business metadata, technical metadata, and data records of the business database table. Using a feature encoding model, the business metadata, technical metadata, data records, directory requirements, identification information, and comment text of the business database tables are encoded respectively, resulting in feature vectors for the business metadata, technical metadata, data records, directory requirements, identification information, and comment text of the business database tables. The feature vectors of the business database table's business metadata, technical metadata, data records, directory requirements, identification information, and remarks text are concatenated to obtain the target feature vector of the business database table. This target feature vector is then input into the large model, which generates the data resource directory for the business database table.

[0010] In one possible implementation of the first aspect, after the process of constructing the business database table directory requirements by including the resource directory structure, resource directory list, resource directory content specifications, core field names, core field types, core field lengths, and core field value ranges, inputting the business database table identification information into the database query engine, retrieving the business database table content through the query engine, inputting the business database table content, business database table directory requirements, business database table identification information, and business database table comment text into the large model, and generating the business database table data resource directory through the large model, the data resource directory generation method includes: The dialog window displays the data resource catalog of the business database tables.

[0011] In one possible implementation of the first aspect, the keyword extraction algorithm is a word frequency inverse document frequency algorithm, and the business database tables include business database tables of government departments and business database tables of enterprises.

[0012] Secondly, embodiments of this application provide a data resource catalog generation device based on a large model, applied to a server-side device, comprising: The acquisition module is used to establish a dialogue window between the server device and the user device. Through the dialogue window, the identification information and remarks text of the business database table uploaded by the user device are obtained. The first generation module is used to extract keywords from the identifier information and the memo text of the business database table using a keyword extraction algorithm, and generate keywords for the identifier information and the memo text of the business database table respectively. The component module is used to input keywords of the identification information of the business database table and keywords of the notes text of the business database table into the intent model, and output the intent label of the user terminal device through the intent model. When the intent label is a query label, the query label, keywords of the identification information of the business database table and keywords of the notes text of the business database table are combined to form the intent description information of the user terminal device. The query module is used to query the feature vector of intent description information through the query interface of the knowledge base, obtain the similarity between each vector data in the knowledge base and the feature vector of intent description information, select the vector data with the highest similarity as the target vector data, extract the storage content associated with the target vector data, and obtain the resource directory structure, resource directory list, resource directory content specifications, core field name, core field type, core field length, and core field value range from the storage content. The second generation module is used to compose the directory requirements of the business database table by combining the resource directory structure, resource directory list, resource directory content specifications, core field names, core field types, core field lengths, and core field value ranges. It inputs the identification information of the business database table into the database query engine, retrieves the content of the business database table through the query engine, and inputs the content of the business database table, the directory requirements of the business database table, the identification information of the business database table, and the remarks text of the business database table into the large model. The large model then generates the data resource directory of the business database table.

[0013] Thirdly, embodiments of this application provide a server device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the data resource catalog generation method described in the first aspect above.

[0014] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the data resource catalog generation method described in the first aspect above.

[0015] Fifthly, embodiments of this application provide a computer program product that, when run on a server device, causes the server device to execute the data resource catalog generation method described in the first aspect.

[0016] The beneficial effects of the embodiments of this application are as follows: Firstly, the resource directory structure, resource directory list, resource directory content specifications, core field names, core field types, core field lengths, and core field value ranges constitute the directory requirements for the business database tables. The identification information of the business database tables is input into the database query engine, which retrieves the content of the business database tables. The content of the business database tables, the directory requirements of the business database tables, the identification information of the business database tables, and the remarks text of the business database tables are input into the large model, which generates the data resource directory of the business database tables. Since no manual operation is required, the generation time of the data resource directory of the business database tables is reduced, which helps to improve the generation efficiency of the data resource directory of the business database tables. Secondly, by generating a data resource catalog for business database tables through a large model, issues such as inconsistent field descriptions, arbitrary attribute determinations, and misuse of coding rules will not arise due to differences in the cognitive abilities, experience limitations, or negligence of the staff. Therefore, this helps to ensure the accuracy of the data resource catalog for database tables. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is an application scenario diagram of the data resource catalog generation method provided in the embodiments of this application; Figure 2 This is a flowchart illustrating the data resource catalog generation method provided in an embodiment of this application; Figure 3 A flowchart illustrating the implementation of S205 provided in this application embodiment; Figure 4 A schematic block diagram of a data resource catalog generation apparatus provided in the embodiments of this application; Figure 5 This is a schematic diagram of the structure of the server device provided in an embodiment of this application. Detailed Implementation

[0019] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application. All other embodiments obtained by those skilled in the art based on the embodiments in this application without inventive effort are within the scope of protection of this application.

[0020] The data resource catalog generation method provided in this application embodiment can be applied to server-side devices, and this application embodiment does not impose any restrictions on the specific type of server-side device.

[0021] Please see Figure 1 , Figure 1 The application scenario diagram of the data resource catalog generation method provided in the embodiments of this application is described in detail below: The server device connects to the user device and establishes a dialogue window between the server device and the user device. Through the information input area of ​​the dialogue window, it receives interactive information input by the user device. Semantic parsing of the interactive information yields the identifier information and remarks text of the business database tables uploaded by the user device.

[0022] In this embodiment, the server device performs semantic parsing on the interactive information to obtain the identification information and remarks text of the business database table uploaded by the user device. Based on the identification information and remarks text of the business database table, the server device can perceive user needs and provide personalized services.

[0023] Please see Figure 2 , Figure 2 This is a flowchart illustrating the data resource catalog generation method provided in this application embodiment, which can be applied to server-side devices.

[0024] like Figure 2 As shown, the data resource catalog generation method provided in this application includes the following steps, detailed below: S201, Establish a dialog window between the server device and the user device, and obtain the identification information and remarks text of the business database table uploaded by the user device through the dialog window. Among them, user-end devices include, but are not limited to, mobile phones, tablets, wearable devices, in-vehicle devices, and laptops.

[0025] S202, using a keyword extraction algorithm, keywords are extracted from the identifier information and the memo text of the business database table, respectively, to generate keywords for the identifier information and keywords for the memo text of the business database table. The keyword extraction algorithm is a term frequency inverse document frequency algorithm.

[0026] Term Frequency-Inverse Document Frequency (TF-IDF) is a weighted statistical algorithm used to evaluate the importance of a word in a specific document.

[0027] S203, input the keywords of the identification information of the business database table and the keywords of the notes text of the business database table into the intent model, and output the intent label of the user terminal device through the intent model. When the intent label is a query label, the query label, the keywords of the identification information of the business database table and the keywords of the notes text of the business database table are combined to form the intent description information of the user terminal device. Specifically, the process involves inputting keywords from the identification information of the business database table and keywords from the notes text of the business database table into the intent model. The intent model then outputs an intent tag for the user terminal device. When the intent tag is a query tag, the query tag, keywords from the identification information of the business database table, and keywords from the notes text of the business database table are combined to form the intent description information for the user terminal device, including: Obtain the model file, load the intent model through the model file, and input the keywords of the identification information of the business database table and the keywords of the notes text of the business database table into the intent model. The intent model outputs the intent label of the user terminal device. When the intent label is a query label, the intent description information of the user terminal device is composed of the query label, the keywords of the identification information of the business database table, and the keywords of the remarks text of the business database table.

[0028] For ease of explanation, the following example is provided: For example, the user terminal device can enter the following in the dialog interface: "Data resource catalog for building citizen social security payment information table, with the note: Includes payment data of insured persons in the city from 2023 to 2024."

[0029] At this point, the intent label is the query label.

[0030] The intent description information for the user's terminal device is: to query the citizen's social security payment information from 2023 to 2024.

[0031] S204. The feature vector of the intent description information is queried through the query interface of the knowledge base. The similarity between each vector data in the knowledge base and the feature vector of the intent description information is obtained. The vector data with the highest similarity is selected as the target vector data. The storage content associated with the target vector data is extracted. The resource directory structure, resource directory list, resource directory content specifications, core field name, core field type, core field length, and core field value range are obtained from the storage content. The knowledge base query interface is a standardized program call interface.

[0032] Among these methods, querying the feature vectors of intent description information through the knowledge base's query interface enables more efficient querying of these feature vectors.

[0033] Specifically, the feature vectors of intent description information are queried through the knowledge base's query interface to obtain the similarity between each vector data in the knowledge base and the feature vectors of intent description information. The vector data with the highest similarity is selected as the target vector data. The associated storage content of the target vector data is extracted, and the resource directory structure, resource directory list, resource directory content specifications, core field names, core field types, core field lengths, and core field value ranges are obtained from the storage content, including: The intent description information is encoded using a feature encoding model to generate a feature vector of the intent description information, and then the feature vector of the intent description information is submitted to the query interface of the knowledge base. The feature vectors of intent description information are queried through the knowledge base query interface. The similarity between each vector data in the knowledge base and the feature vector of intent description information is obtained. The vector data with the highest similarity is selected as the target vector data. The storage content associated with the target vector data is extracted. The resource directory structure, resource directory list, resource directory content specifications, core field names, core field types, core field lengths, and core field value ranges are obtained from the storage content.

[0034] S205: The resource directory structure, resource directory list, resource directory content specifications, core field names, core field types, core field lengths, and core field value ranges constitute the directory requirements for the business database tables. The identification information of the business database tables is input into the database query engine, which retrieves the content of the business database tables. The content of the business database tables, the directory requirements of the business database tables, the identification information of the business database tables, and the remarks text of the business database tables are input into the large model, which generates the data resource directory of the business database tables.

[0035] The data resource directory generation method, after defining the directory requirements for the business database table by assembling the resource directory structure, resource directory list, resource directory content specifications, core field names, core field types, core field lengths, and core field value ranges, inputting the identification information of the business database table into the database query engine, retrieving the content of the business database table through the query engine, inputting the content of the business database table, the directory requirements of the business database table, the identification information of the business database table, and the remark text of the business database table into the large model, and generating the data resource directory of the business database table through the large model, includes: The dialog window displays the data resource catalog of the business database tables.

[0036] The keyword extraction algorithm is a word frequency inverse document frequency algorithm, and the business database tables include business database tables of government departments and business database tables of enterprises.

[0037] For ease of explanation, we will use a business database table from the government department as an example of a table showing citizens' social security payment information: For example, a social security bureau urgently needs to retrieve citizens' social security payment information forms to provide data support for the pension calculation of retirees.

[0038] On the user terminal device, enter the following in the dialog interface: "Data resource directory for building a citizen social security payment information table. Note: Includes payment data of all insured persons in the city from 2023 to 2024."

[0039] The server-side device generates a data resource catalog of the citizen social security payment information table through a large model, and displays the data resource catalog of the citizen social security payment information table through the content display area of ​​the dialog window.

[0040] User devices can directly access the core fields of the citizen social security payment information form, such as the payment base and payment years, through the data resource catalog of the citizen social security payment information form. This can speed up the data retrieval and verification time, which is conducive to reducing the calculation time of pension and improving the calculation efficiency of pension.

[0041] For ease of explanation, let's take the company's business database table as the sales order data table, as an example below: The finance department of a manufacturing company urgently needs to retrieve the company's sales order data table to provide data support for the quarterly operating profit calculation.

[0042] On the user device, enter the following in the dialog interface: the data resource directory for building the enterprise sales order data table, with the note: containing sales order data from 2023 to 2025.

[0043] The server-side device generates a data resource catalog of the enterprise sales order data table through a large model, and displays the data resource catalog of the enterprise sales order data table through the content display area of ​​the dialog window.

[0044] User devices can directly locate core fields of enterprise sales order data tables, such as order amount, transaction time, and payment status, through the data resource directory of the enterprise sales order data table. This eliminates the need to search the enterprise database one by one, thereby reducing the query time for enterprise sales orders and improving the query efficiency.

[0045] The beneficial effects of the embodiments of this application are as follows: Firstly, the resource directory structure, resource directory list, resource directory content specifications, core field names, core field types, core field lengths, and core field value ranges constitute the directory requirements for the business database tables. The identification information of the business database tables is input into the database query engine, which retrieves the content of the business database tables. The content of the business database tables, the directory requirements of the business database tables, the identification information of the business database tables, and the remarks text of the business database tables are input into the large model, which generates the data resource directory of the business database tables. Since no manual operation is required, the generation time of the data resource directory of the business database tables is reduced, which helps to improve the generation efficiency of the data resource directory of the business database tables. Secondly, by generating a data resource catalog for business database tables through a large model, issues such as inconsistent field descriptions, arbitrary attribute determinations, and misuse of coding rules will not arise due to differences in the cognitive abilities, experience limitations, or negligence of the staff. Therefore, this helps to ensure the accuracy of the data resource catalog for database tables.

[0046] Please see Figure 3 , Figure 3 The implementation flowchart of S205 provided in the embodiments of this application is described in detail below: S301, the resource directory structure, resource directory list, resource directory content specifications, core field names, core field types, core field lengths, and core field value ranges are used to form the directory requirements for the business database table. The identification information of the business database table is input into the database query engine, and the content of the business database table is retrieved through the query engine. The content of the business database table includes the business metadata, technical metadata, and data records of the business database table. S302, through the feature encoding model, the business metadata, technical metadata, data records, directory requirements, identification information, and comment text of the business database table are encoded respectively, to obtain the feature vectors of the business metadata, technical metadata, data records, directory requirements, identification information, and comment text of the business database table. S303: The feature vectors of the business metadata of the business database table, the feature vectors of the technical metadata of the business database table, the feature vectors of the data records of the business database table, the feature vectors of the directory requirements of the business database table, the feature vectors of the identification information of the business database table, and the feature vectors of the remarks text of the business database table are concatenated to obtain the target feature vector of the business database table. The target feature vector of the business database table is input into the large model, and the large model generates the data resource catalog of the business database table.

[0047] In this embodiment of the application, the data resource catalog of the business database table is generated by a large model. Compared with manual operation, the generation cycle of the data resource catalog of the database table can be greatly shortened, and the oversights and errors that may occur during manual classification can be effectively avoided, thus greatly improving the reliability of the data resource catalog of the database table.

[0048] For the data resource catalog generation method described in the above embodiments, please refer to [link / reference]. Figure 4 , Figure 4 This is a schematic block diagram of a data resource catalog generation apparatus provided in an embodiment of this application. Figure 4 The data resource catalog generation device 400 shown can be applied to, for example... Figure 1 The application scenario diagram shows the server-side device. The following section uses the server-side device as an example to illustrate this. Figure 4 The data resource catalog generation device 400 shown will be described in detail. The data resource catalog generation device 400 may include an acquisition module 401, a first generation module 402, a composition module 403, a query module 404, and a second generation module 405.

[0049] The acquisition module 401 is used to establish a dialog window between the server device and the user device. Through the dialog window, the identification information of the business database table uploaded by the user device and the remark text of the business database table are obtained. The first generation module 402 is used to extract keywords from the identifier information and the remarks text of the business database table using a keyword extraction algorithm, and generate keywords for the identifier information and the remarks text of the business database table respectively. Module 403 is used to input keywords of the identification information of the business database table and keywords of the notes text of the business database table into the intent model, and output the intent label of the user terminal device through the intent model. When the intent label is a query label, the query label, keywords of the identification information of the business database table and keywords of the notes text of the business database table are combined to form the intent description information of the user terminal device. The query module 404 is used to query the feature vector of intent description information through the query interface of the knowledge base, obtain the similarity between each vector data in the knowledge base and the feature vector of intent description information, select the vector data with the highest similarity as the target vector data, extract the storage content associated with the target vector data, and obtain the resource directory structure, resource directory list, resource directory content specifications, core field name, core field type, core field length, and core field value range from the storage content. The second generation module 405 is used to compose the directory requirements of the business database table by combining the resource directory structure, resource directory list, resource directory content specifications, core field names, core field types, core field lengths, and core field value ranges. It inputs the identification information of the business database table into the database query engine, retrieves the content of the business database table through the query engine, inputs the content of the business database table, the directory requirements of the business database table, the identification information of the business database table, and the remarks text of the business database table into the large model, and generates the data resource directory of the business database table through the large model.

[0050] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0051] The beneficial effects of the embodiments of this application are as follows: Firstly, the resource directory structure, resource directory list, resource directory content specifications, core field names, core field types, core field lengths, and core field value ranges constitute the directory requirements for the business database tables. The identification information of the business database tables is input into the database query engine, which retrieves the content of the business database tables. The content of the business database tables, the directory requirements of the business database tables, the identification information of the business database tables, and the remarks text of the business database tables are input into the large model, which generates the data resource directory of the business database tables. Since no manual operation is required, the generation time of the data resource directory of the business database tables is reduced, which helps to improve the generation efficiency of the data resource directory of the business database tables. Secondly, by generating a data resource catalog for business database tables through a large model, issues such as inconsistent field descriptions, arbitrary attribute determinations, and misuse of coding rules will not arise due to differences in the cognitive abilities, experience limitations, or negligence of the staff. Therefore, this helps to ensure the accuracy of the data resource catalog for database tables.

[0052] Please see Figure 5 , Figure 5 This is a schematic diagram of the structure of the server device provided in an embodiment of this application.

[0053] like Figure 5 As shown, Figure 5 The server device 2 includes: at least one processor 20, a memory 21, and a computer program 22 stored in the memory 21 and executable on the at least one processor 20, wherein the processor 20 executes the computer program 22 to implement the steps in any of the above method embodiments.

[0054] The server-side device 2 may include, but is not limited to, a processor 20 and a memory 21. Those skilled in the art will understand that... Figure 5 This is merely an example of server device 2 and does not constitute a limitation on server device 2. It may include more or fewer components than shown in the figure, or combine certain components, or different components, such as input / output devices, network access devices, etc.

[0055] The processor 20 is used to run a computer program 22 stored in the memory 21, and performs the following steps when executing the computer program 22: Establish a dialog window between the server-side device and the user-side device. Through the dialog window, obtain the identification information and remarks text of the business database table uploaded by the user-side device. Using a keyword extraction algorithm, keywords are extracted from the identifier information and the memo text of the business database table, respectively, to generate keywords for the identifier information and keywords for the memo text of the business database table. Input the keywords of the identification information of the business database table and the keywords of the notes text of the business database table into the intent model. The intent model outputs the intent label of the user terminal device. When the intent label is a query label, the query label, the keywords of the identification information of the business database table and the keywords of the notes text of the business database table are combined to form the intent description information of the user terminal device. The feature vectors of intent description information are queried through the query interface of the knowledge base. The similarity between each vector data in the knowledge base and the feature vector of intent description information is obtained. The vector data with the highest similarity is selected as the target vector data. The storage content associated with the target vector data is extracted. The resource directory structure, resource directory list, resource directory content specifications, core field name, core field type, core field length, and core field value range are obtained from the storage content. The resource directory structure, resource directory list, resource directory content specifications, core field names, core field types, core field lengths, and core field value ranges constitute the directory requirements for the business database tables. The identification information of the business database tables is input into the database query engine, which retrieves the content of the business database tables. The content of the business database tables, the directory requirements of the business database tables, the identification information of the business database tables, and the remarks text of the business database tables are input into the large model, which generates the data resource directory of the business database tables.

[0056] The processor 20 may be a Central Processing Unit (CPU), or it may be other general-purpose processors, digital signal processors, field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.

[0057] In some embodiments, the memory 21 may be an internal storage unit of the server device 2, such as the hard disk or memory of the server device 2.

[0058] This application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps described in the various method embodiments above.

[0059] Since the computer program stored in the computer-readable storage medium can execute any of the large-model-based data resource catalog generation methods provided in the embodiments of this application, the computer-readable storage medium can achieve the beneficial effects that any of the large-model-based data resource catalog generation methods provided in the embodiments of this application can achieve, as detailed in the preceding embodiments, and will not be repeated here.

[0060] This application provides a computer program product that, when run on a server device, causes the server device to execute the aforementioned data resource catalog generation method.

[0061] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium.

[0062] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.

Claims

1. A method for generating a data resource catalog based on a large model, characterized in that, The data resource catalog generation method, applied to server-side devices, includes: Establish a dialog window between the server-side device and the user-side device. Through the dialog window, obtain the identification information and remarks text of the business database table uploaded by the user-side device. Using a keyword extraction algorithm, keywords are extracted from the identifier information and the memo text of the business database table, respectively, to generate keywords for the identifier information and keywords for the memo text of the business database table. Input the keywords of the identification information of the business database table and the keywords of the notes text of the business database table into the intent model. The intent model outputs the intent label of the user terminal device. When the intent label is a query label, the query label, the keywords of the identification information of the business database table and the keywords of the notes text of the business database table are combined to form the intent description information of the user terminal device. The feature vectors of intent description information are queried through the query interface of the knowledge base. The similarity between each vector data in the knowledge base and the feature vector of intent description information is obtained. The vector data with the highest similarity is selected as the target vector data. The storage content associated with the target vector data is extracted. The resource directory structure, resource directory list, resource directory content specifications, core field name, core field type, core field length, and core field value range are obtained from the storage content. The resource directory structure, resource directory list, resource directory content specifications, core field names, core field types, core field lengths, and core field value ranges constitute the directory requirements for the business database tables. The identification information of the business database tables is input into the database query engine, which retrieves the content of the business database tables. The content of the business database tables, the directory requirements of the business database tables, the identification information of the business database tables, and the remarks text of the business database tables are input into the large model, which generates the data resource directory of the business database tables.

2. The data resource catalog generation method according to claim 1, characterized in that, The establishment of a dialog window between the server-side device and the user-side device allows for the acquisition of identification information and remarks text of the business database tables uploaded by the user-side device, including: Establish a dialogue window between the server device and the user device, and receive interactive information input by the user device through the information input area of ​​the dialogue window; Semantic parsing of the interactive information yields the identifier information and remarks text of the business database tables uploaded by the user device.

3. The data resource catalog generation method according to claim 1, characterized in that, The process involves inputting keywords from the identification information of the business database table and keywords from the notes text of the business database table into the intent model. The intent model then outputs an intent tag for the user terminal device. When the intent tag is a query tag, the query tag, keywords from the identification information of the business database table, and keywords from the notes text of the business database table are combined to form the intent description information for the user terminal device, including: Obtain the model file, load the intent model through the model file, and input the keywords of the identification information of the business database table and the keywords of the notes text of the business database table into the intent model. The intent model outputs the intent label of the user terminal device. When the intent label is a query label, the intent description information of the user terminal device is composed of the query label, the keywords of the identification information of the business database table, and the keywords of the remarks text of the business database table.

4. The data resource catalog generation method according to claim 1, characterized in that, The process involves querying the feature vectors of intent description information through the knowledge base's query interface to obtain the similarity between each vector data in the knowledge base and the feature vectors of intent description information. The vector data with the highest similarity is selected as the target vector data. The associated storage content of the target vector data is extracted, and from this storage content, the resource directory structure, resource directory list, resource directory content specifications, core field names, core field types, core field lengths, and core field value ranges are obtained, including: The intent description information is encoded using a feature encoding model to generate a feature vector of the intent description information, and then the feature vector of the intent description information is submitted to the query interface of the knowledge base. The feature vectors of intent description information are queried through the knowledge base query interface. The similarity between each vector data in the knowledge base and the feature vector of intent description information is obtained. The vector data with the highest similarity is selected as the target vector data. The storage content associated with the target vector data is extracted. The resource directory structure, resource directory list, resource directory content specifications, core field names, core field types, core field lengths, and core field value ranges are obtained from the storage content.

5. The data resource catalog generation method according to claim 1, characterized in that, The process involves defining the directory requirements for the business database tables, including the resource directory structure, resource directory list, resource directory content specifications, core field names, core field types, core field lengths, and core field value ranges. The identification information of the business database tables is input into the database query engine, which retrieves the content of the business database tables. The content of the business database tables, the directory requirements, the identification information, and the remarks text are then input into the large model. The large model then generates the data resource directory for the business database tables, including: The resource directory structure, resource directory list, resource directory content specifications, core field names, core field types, core field lengths, and core field value ranges constitute the directory requirements for the business database table. The identification information of the business database table is input into the database query engine, and the content of the business database table is retrieved through the query engine. The content of the business database table includes the business metadata, technical metadata, and data records of the business database table. Using a feature encoding model, the business metadata, technical metadata, data records, directory requirements, identification information, and comment text of the business database tables are encoded respectively, resulting in feature vectors for the business metadata, technical metadata, data records, directory requirements, identification information, and comment text of the business database tables. The feature vectors of the business database table's business metadata, technical metadata, data records, directory requirements, identification information, and remarks text are concatenated to obtain the target feature vector of the business database table. This target feature vector is then input into the large model, which generates the data resource directory for the business database table.

6. The data resource catalog generation method according to claim 1, characterized in that, The data resource directory generation method, after defining the directory requirements for business database tables by assembling the resource directory structure, resource directory list, resource directory content specifications, core field names, core field types, core field lengths, and core field value ranges, inputting the business database table identification information into the database query engine, retrieving the business database table content through the query engine, inputting the business database table content, business database table directory requirements, business database table identification information, and business database table comment text into the large model, and generating the data resource directory for the business database tables through the large model, includes: The dialog window displays the data resource catalog of the business database tables.

7. The data resource catalog generation method according to claim 1, characterized in that, The keyword extraction algorithm is a word frequency inverse document frequency algorithm, and the business database tables include those of government departments and enterprises.

8. A data resource catalog generation device based on a large model, characterized in that, Applied to server-side devices, including: The acquisition module is used to establish a dialogue window between the server device and the user device. Through the dialogue window, the identification information and remarks text of the business database table uploaded by the user device are obtained. The first generation module is used to extract keywords from the identifier information and the memo text of the business database table using a keyword extraction algorithm, and generate keywords for the identifier information and the memo text of the business database table respectively. The component module is used to input keywords of the identification information of the business database table and keywords of the notes text of the business database table into the intent model, and output the intent label of the user terminal device through the intent model. When the intent label is a query label, the query label, keywords of the identification information of the business database table and keywords of the notes text of the business database table are combined to form the intent description information of the user terminal device. The query module is used to query the feature vector of intent description information through the query interface of the knowledge base, obtain the similarity between each vector data in the knowledge base and the feature vector of intent description information, select the vector data with the highest similarity as the target vector data, extract the storage content associated with the target vector data, and obtain the resource directory structure, resource directory list, resource directory content specifications, core field name, core field type, core field length, and core field value range from the storage content. The second generation module is used to compose the directory requirements of the business database table by combining the resource directory structure, resource directory list, resource directory content specifications, core field names, core field types, core field lengths, and core field value ranges. It inputs the identification information of the business database table into the database query engine, retrieves the content of the business database table through the query engine, and inputs the content of the business database table, the directory requirements of the business database table, the identification information of the business database table, and the remarks text of the business database table into the large model. The large model then generates the data resource directory of the business database table.

9. A server-side device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the data resource catalog generation method as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the data resource catalog generation method as described in any one of claims 1 to 7.