Data asset evaluation method, device, equipment and medium based on large model

Through the data asset evaluation method based on large models, the problems of high cost and poor effect of data asset analysis in the existing technology are solved, and efficient and accurate data asset analysis is achieved.

CN117788172BActive Publication Date: 2025-05-20BEIJING XINLIU DATA TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202311871038.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-29
Publication Date
2025-05-20
Estimated Expiration
2043-12-29

AI Technical Summary

Technical Problem

In the prior art, data assets are analyzed manually or by software, resulting in high analysis costs and poor results.

Method used

The data asset evaluation method based on the big model is adopted, and the question information to be used is determined by evaluating question information based on the data identification, and the program to be executed is determined based on the task description information, and the task processing is carried out according to the execution order of the program to be executed, so as to obtain the evaluation results of the data assets.

Benefits of technology

While reducing the cost of data asset analysis, it also improves the accuracy and convenience of data asset analysis, improves the analysis effect, and meets users' needs for data asset analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117788172B_ABST
    Figure CN117788172B_ABST
Patent Text Reader

Abstract

The present invention discloses a data asset evaluation method, device, equipment and medium based on a large model. The method comprises: determining the question information to be used based on the evaluation question information including the data identifier; determining at least one task to be processed corresponding to the question information to be used and the task description information corresponding to each task to be processed; determining the program to be executed corresponding to each task to be processed based on each task description information; performing task processing in sequence according to the execution order of each program to be executed, and obtaining the evaluation result of the data to be evaluated corresponding to the data identifier. The method solves the problem of high analysis cost and poor effect caused by manual or software analysis of data assets in the prior art, reduces the cost of data asset analysis, improves the automation, accuracy and efficiency of data asset analysis, improves the analysis effect, and achieves the effect of meeting the user's demand for data asset analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer processing technologies, and in particular, to a method, apparatus, device, and medium for data asset evaluation based on a large model. Background Art

[0002] With the development of big data and artificial intelligence, the demand for data asset analysis is increasing. Existing data asset analysis methods involve users manually viewing and analyzing data assets, calculating indicators such as data accuracy, record duplication rate, and record filling rate to understand their quality, cost, and potential value.

[0003] This analysis method mainly relies on users' professional skills and experience and requires a large amount of analysis costs. Alternatively, professional data analysis software such as Excel and database management systems is used for data asset analysis. This method requires users to write and run custom scripts or queries to perform analysis tasks. When there are new analysis requirements, the scripts need to be rewritten or modified to perform the analysis tasks, resulting in high costs and poor analysis effects. Summary of the Invention

[0004] The present invention provides a method, apparatus, device, and medium for data asset evaluation based on a large model, so as to reduce the cost of data asset analysis while improving the accuracy and convenience of data asset analysis, enhancing the analysis effect, and achieving the effect of meeting users' data asset analysis requirements.

[0005] According to one aspect of the present invention, there is provided a method for data asset evaluation based on a large model, the method comprising:

[0006] Determining the question information to be used based on the evaluation question information including the data identifier;

[0007] Determining at least one task to be processed corresponding to the question information to be used and task description information corresponding to each of the tasks to be processed;

[0008] Determining an executable program corresponding to each task to be processed based on each of the task description information;

[0009] Performing task processing in sequence according to the execution order of each of the executable programs to obtain an evaluation result of the data to be evaluated corresponding to the data identifier.

[0010] According to another aspect of the present invention, there is provided a data asset evaluation apparatus based on a large model, the apparatus comprising:

[0011] A module for determining the question information to be used, configured to determine the question information to be used based on the evaluation question information including the data identifier;

[0012] A task description information determination module, configured to determine at least one task to be processed corresponding to the to-be-used query information and task description information corresponding to each of the tasks to be processed;

[0013] A to-be-executed program determination module, configured to determine a to-be-executed program corresponding to each task to be processed based on each of the task description information;

[0014] An evaluation result determination module, configured to sequentially perform task processing according to the execution order of each of the to-be-executed programs to obtain an evaluation result of the to-be-evaluated data corresponding to the data identifier.

[0015] According to another aspect of the present invention, there is provided an electronic device, the electronic device includes:

[0016] At least one processor; and

[0017] A memory communicatively connected to the at least one processor; wherein,

[0018] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the data asset evaluation method based on a large model according to any embodiment of the present invention.

[0019] According to another aspect of the present invention, there is provided a computer-readable storage medium, the computer-readable storage medium stores computer instructions, and the computer instructions are used to implement the data asset evaluation method based on a large model according to any embodiment of the present invention when executed by a processor.

[0020] The technical solution of the embodiment of the present invention determines the question information to be used based on the evaluation question information including data identifiers; determines at least one task to be processed corresponding to the question information to be used and the task description information corresponding to each task to be processed; determines the program to be executed corresponding to each task to be processed based on each task description information; and sequentially processes the tasks according to the execution order of each program to be executed to obtain the evaluation result of the data to be evaluated corresponding to the data identifier, solving the problems of high analysis cost and poor effect in the prior art by manually or software analyzing data assets, realizing receiving the evaluation question information input by the user in the form of AI Q&A and processing it to obtain the question information to be used that can more clearly and professionally express the user's evaluation intention, improving the accuracy and effectiveness of the evaluation process, and then selecting at least one task to be processed adapted to the question information to be used and determining the task description information of each task to be processed, determining the program to be executed for the task to be processed according to the task description information, and then sequentially processing the tasks according to the execution order of each program to be executed to obtain the evaluation result of the data to be evaluated corresponding to the data identifier, achieving the effect of reducing the analysis cost of data assets while improving the automation, accuracy and efficiency of data asset analysis, improving the analysis effect, and meeting the user's demand for data asset analysis.

[0021] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings without creative efforts based on these drawings.

[0023] Figure 1 is a flowchart of a method for data asset evaluation based on a large model according to Embodiment 1 of the present invention;

[0024] Figure 2 is a schematic diagram for displaying the evaluation result according to Embodiment 1 of the present invention;

[0025] Figure 3 is a schematic structural diagram of a data asset evaluation device based on a large model according to Embodiment 2 of the present invention;

[0026] Figure 4 is a schematic structural diagram of an electronic device for implementing the method for data asset evaluation based on a large model according to the embodiment of the present invention. Detailed implementation manners

[0027] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0028] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data used in appropriate cases can be interchanged so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0029] Embodiment 1

[0030] Figure 1 is a flowchart of a method for evaluating data assets based on a large model according to Embodiment 1 of the present invention. This embodiment is applicable to the situation of evaluating data assets. This method can be executed by a data asset evaluation device based on a large model. The data asset evaluation device based on a large model can be implemented in the form of hardware and / or software, and the data asset evaluation device based on a large model can be configured in a computing device. As Figure 1 shown, the method includes:

[0031] S110. Determine the question information to be used based on the evaluation question information including the data identifier.

[0032] Among them, the data identifier can be used to uniquely identify the data to be evaluated. The data to be evaluated can refer to the data assets that need to be evaluated, and the data assets can be data resources recorded in a physical or electronic manner. For example, the data to be evaluated can be document materials, electronic data, etc. The evaluation question information includes at least one of text, image, voice, and video.

[0033] In this embodiment, the user can enter the question information corresponding to the data to be evaluated that the user wants to evaluate in the edit box. For example, the entered question information can be a piece of text, such as "Please load dataset 001 and analyze its data accuracy" entered when chatting with AI; it can also be a piece of voice. When receiving the user's question information, it can be considered that the evaluation question information is received, and the semantic analysis can be performed on the evaluation question information to identify the evaluation intention, and the evaluation question information can be rewritten into the text information that conforms to the evaluation expression intention as the question information to be used. The technical solution provided in this embodiment can more clearly and professionally express the question information for evaluation by combining the user's evaluation intention, improving the accuracy and effectiveness of subsequent evaluation processing.

[0034] Exemplarily, the evaluation user provides the evaluation question information in natural language as "Please load dataset 001, analyze its data accuracy, and display the data accuracy score in a bar chart". Identify the intention of the evaluation user, such as the intention to perform accuracy analysis on the data in dataset 001 and display the analysis result in a bar chart. Based on the intention, rewrite the evaluation question information to obtain the question information to be used as "Use dataset number 001, perform data accuracy analysis, and plot the data accuracy score as a bar chart".

[0035] In this embodiment, based on the evaluation question information including the data identifier, determining the question information to be used includes: obtaining the evaluation question information including the data identifier; and determining the question information to be used according to the pre-determined first prompt word template and the evaluation question information.

[0036] Among them, the first prompt word template prompt can be used to guide the model to generate a prompt string that conforms to the evaluation intention. The first prompt word template includes at least one phrase and an evaluation convention corresponding to at least one phrase. For example, the evaluation convention "data accuracy" is related to data quality evaluation.

[0037] Specifically, after obtaining the evaluation question information including the data identifier, natural language analysis technology can be used to understand the vocabulary in the evaluation question information, and based on the meanings of the individual words in the evaluation question information, evaluation key operation recognition and evaluation target recognition are performed. Exemplarily, taking the evaluation question information "Please load dataset 001, analyze its data accuracy, and display the data accuracy score in a bar chart" as an example, analyze the meanings of each word and phrase in the evaluation question information: recognize that "load" means data needs to be obtained, "dataset 001" represents a specific data identifier, "analyze" indicates that data analysis needs to be performed, "data accuracy" represents the specific task of analysis, and "bar chart" represents a specific visualization method; recognize the evaluation key operations as "load" and "analyze"; recognize the evaluation target of the evaluation user as analyzing the data accuracy of dataset 001. Further, the first prompt word template can be used to rewrite the evaluation question information according to the recognized evaluation key operations and evaluation targets, generating the question information to be used, improving the clarity, accuracy, and readability of the question information. For example, according to the prompt words in the first prompt word template, the prompt words in the evaluation question information can be replaced or modified to generate the new question information to be used; or, the prompt words in the evaluation question information can be extracted and filled into the first prompt word template to generate the question information to be used.

[0038] The technical solution provided in this embodiment enables the system to better understand and rewrite the evaluation user's question information through the first prompt word template by setting evaluation common phrases and conventions in the first prompt word template, making the question information more specific and clearer, ensuring that subsequent evaluation processing steps are consistent with the instructions, and improving the accuracy and effectiveness of data asset evaluation.

[0039] S120. Determine at least one task to be processed corresponding to the question information to be used and task description information corresponding to each task to be processed.

[0040] In this embodiment, the question information to be used can be task-decomposed to obtain the task flow to be processed corresponding to the execution of this data evaluation, and the task flow to be processed includes at least one task to be processed. These tasks to be processed may include data acquisition, data query, data preprocessing, data analysis, visualization drawing, etc., and there is also a dependency relationship of execution before and after between the tasks to be processed. Further, according to the task attributes of each task to be processed, its corresponding task description information can be determined. For example, the task description information may include task name, task type, data identifier to be used, data storage location, task operation, etc.

[0041] In this embodiment, determining at least one task to be processed corresponding to the question information to be used and task description information corresponding to each task to be processed includes: retrieving task prompt information; and based on the task prompt information, processing the question information to be used to determine at least one task to be processed and determining task description information corresponding to each task to be processed.

[0042] Among them, the task prompt information includes at least one of description guiding information, information generation format, a set of tasks to be selected of at least one task type, task description word convention, and description information examples. The description guiding information refers to the information used to guide the generation of description information. The description guiding information may include certain process or sequence information. By providing ordered and structured guiding prompts, it helps the model better understand the task requirements and generate corresponding task description information according to the specified process. The information generation format refers to the format of the generated description information. The set of tasks to be selected contains at least one task to be selected, and the task type is related to the operation of the task. The task description can be used to describe the operations performed by the task. The task description word convention refers to the optional words agreed upon when generating the task description. For example, when generating a plot of time series data involving multiple plotting metrics, words such as "obtain sequentially" and "obtain cyclically" can be used. The description information examples may refer to examples of task description information. The task description information includes, but is not limited to, task number, task type, and task description, etc. The task number may refer to the serial number of the task being processed. For example, for the first task to be processed, its number can be task1, and for the second task to be processed, its number can be task2.

[0043] Specifically, the pre-configured task prompt information can be retrieved, and the question information to be used can be processed accordingly through the description guiding information for generating description information in the task prompt information. For example, natural language processing techniques (such as text classification, recognition, etc.) can be used to parse the question information to be used, and the relevant information of the evaluation task can be analyzed. The relevant information may include, but is not limited to, key information such as the goal, input, and output of the evaluation task. According to the relevant information of the task, a task to be selected that is suitable for the evaluation task is selected from the tasks to be selected in the task prompt information as the task to be processed. For example, the matching degree between each step of the evaluation task and the task to be selected is calculated, and the task to be selected with the highest matching degree is selected as the task to be processed for that step. For each task to be processed, the task description information corresponding to each task to be processed can be determined based on the task description word convention, description information examples, information generation format, etc. in the task prompt information, combined with the own attributes of the task to be processed.

[0044] Exemplarily, the system can send the question information to be used as an instruction to the large model. The large model retrieves the pre-configured Task prompt (i.e., task prompt information) and generates task description information for each task to be processed according to the instruction. For example, the Task prompt can be: Please select the most suitable task (task to be processed) according to the given instruction and generate its task description information. The generation format of the description information is: taskn = {'task_type': 'task_instruction'}. There are four types of optional tasks (i.e., tasks to be selected), including [data_selection_task: tasks for selecting data sources, querying data, selecting data columns, etc., data_processing_task: tasks for extracting and processing tasks such as data quality detection, data_analysis_task: tasks for analyzing data asset costs, future revenues, etc., data_visualization_task: tasks for drawing line charts, trend charts or outputting statistical results for one or more charts]; When generating task_instruction (i.e., task description), words such as "acquire sequentially" and "acquire in a loop" can be used for multiple indicators. Line charts are generally drawn for time series data, and bar charts are usually drawn for cross-sectional data. Output in the following format: task1 = {%s: %s}, task2 = {%s: %s}. For prediction tasks, only print the table (i.e., the convention for task description words); {Example of few-shot description information (the example can help the large model generate task description information more accurately and efficiently}; Instruction: Use dataset number 001 to perform data accuracy analysis. The large model can be a large language model, including GPT-3 (Generative Pre-trained Transformer 3), GPT-4, BERT (Bidirectional Encoder Representations from Transformers), RoBERTa, etc. The advantage of using a large model as the core intelligent engine for data asset evaluation is that it can utilize the natural language processing ability of the large model, better understand and execute the natural language instructions proposed by the evaluation user, and improve the evaluation effect and accuracy.

[0045] Exemplarily, the large model can return the following task description information, and the entire task description information can be used to describe the task plan of the entire evaluation task:

[0046] 'task1 = {"data_selection_task": "Select the data source and determine the storage location of the data with dataset number 001."},

[0047] task2 = {"data_selection_task": "Query data and retrieve relevant data from the database."},

[0048] task3 = {"data_processing_task": "Select data columns and pick the data columns that need to be analyzed for precision."},

[0049] task4 = {"data_processing_task": "Calculate according to the instructions and perform precision calculation."},

[0050] task5 = {"data_visualization_task": "Display the precision analysis results in a bar chart.

[0051] "}'

[0052] S130. Based on each task description information, determine the to-be-executed program corresponding to each to-be-processed task.

[0053] Among them, the to-be-executed program can refer to the program code for implementing the task.

[0054] In this embodiment, code generation technology can be used to generate the to-be-executed program corresponding to each to-be-processed task according to the task description information corresponding to each to-be-processed task, so as to perform task processing based on the to-be-executed program.

[0055] In order to improve the accuracy of task processing and the standardization of code generation, during the process of determining the to-be-executed program corresponding to each to-be-processed task based on each task description information, response prompt information can be retrieved; based on the response prompt information and each task description information, determine the to-be-executed program corresponding to each to-be-processed task.

[0056] Among them, the response prompt information includes task guidance information, corresponding examples between description information and interfaces, and interface configuration information. The task guidance information can refer to the information used to guide the generation of task code. The task guidance information can contain certain process or sequence information. By providing ordered and structured guidance prompts, it helps the model better understand the task requirements and generate corresponding outputs according to the specified process. The corresponding example can refer to an example containing the corresponding relationship between task description information and interfaces. The corresponding example can help the model better understand the corresponding relationship between task description information and interfaces and improve the accuracy of interface determination. The interface configuration information includes the interface name, interface comment, input parameters, and output parameters corresponding to at least one to-be-selected interface in the interface tool library. The interface configuration information may not include the function body of the interface. The to-be-selected interface can refer to a pre-encapsulated operation for implementing certain functions, and the interface can include functions, methods, attributes, etc.

[0057] It should be noted that the method for determining the to-be-executed program corresponding to each to-be-processed task is the same. Taking the to-be-executed program of any one of the to-be-processed tasks as an example can be used for introduction.

[0058] In this embodiment, the pre-configured response prompt information can be retrieved, and by combining the interface configuration information in the response prompt information and the task description in the task description information, the interface information capable of executing the task description can be selected from the interface configuration information, and the interface name can be embedded into the to-be-executed code of the to-be-processed task; alternatively, according to the output result of the previous to-be-processed task of the to-be-processed task, the input parameters of the to-be-processed task can be determined, and the to-be-executed code can be updated based on the input parameters; alternatively, the output result of the to-be-processed task can be used as the input result of the next to-be-processed task of the to-be-processed task, the input position of the output result of the to-be-processed task can be determined, and the to-be-executed code can be updated based on the input position to obtain the updated to-be-executed program corresponding to the to-be-processed task.

[0059] Optionally, based on the response prompt information and each task description information, determining the to-be-executed program corresponding to each to-be-processed task includes: based on the task description information, and the interface configuration information and / or corresponding examples in the response prompt information, determining the function parameters corresponding to the corresponding to-be-processed task; based on the task description information, the task guidance information in the response prompt information, and the function parameters corresponding to the to-be-processed task, determining the to-be-executed program corresponding to the to-be-processed task.

[0060] Specifically, the process can be guided by the task guidance information in the response prompt information. According to at least one of the interface configuration information and corresponding examples in the response prompt information, select the function for the task description in the task description information to complete the to-be-processed task, and generate corresponding parameters for the function. Further, the process can continue to be guided by the task guidance information in the response prompt information, generate the corresponding code according to the task description information, and add the function parameters to the code to generate the to-be-executed program corresponding to the to-be-processed task.

[0061] Exemplarily, the large model can retrieve pre-configured response prompt information. For example, the response prompt information includes: Flow prompt (task guidance information): Please complete the task description step by step using the given functions. In each step, only select one or more independent functions from the following function library, and generate corresponding parameters for the functions. The parameter format must strictly follow the function description, and generate corresponding multiple result_i (output results) in parallel. The functions in the subsequent steps use the result output by the previous function as the parameter input, ending with ; Function Library: {interface configuration information, which only includes the annotation, input parameters, and output parameters of the interface, excluding the function body}, {few-shot corresponding examples}. Instruction task description includes: Select the data source, determine the storage location of the data with dataset number 001; Query data, query relevant data from the database; Select data columns, select the data columns that need to perform precision analysis; Calculate according to the instruction, perform precision calculation; Display the precision analysis results in a bar chart.

[0062] Exemplarily, the large model can return the following program to be executed, so that the evaluation task can be implemented according to the program to be executed.

[0063] step1 = {

[0064] "arg1": ["001"],

[0065] "function1": "select_data_source_and_location",

[0066] "output1": "result1",

[0067] "description1": "Selects the data source and determines the data storage location for a given dataset code."

[0068] }, step2 = {

[0069] "arg1": ["result1"],

[0070] "function1": "load_data_to_dataframe",

[0071] "output1": "result2",

[0072] "description1":"Loads data from a specified data storage location into a Pandas DataFrame."

[0073] }, step3 = {

[0074] "arg1": ["result2"],

[0075] "function1": "select_columns",

[0076] "output1": "result3",

[0077] "description1": "Selects specific columns from a DataFrame."

[0078] }, step4 = {

[0079] "arg1": ["result3"],

[0080] "function1": "calculate_column_precision",

[0081] "output1": "result4",

[0082] "description1": "Calculate the precision of each column in a DataFrame."

[0083] }, step5 = {

[0084] "arg1": ["result4", "bar", "precision score"],

[0085] "function1": "plot_dataframe_and_table",

[0086] "output1": "result5",

[0087] "description1": "Plot a DataFrame as a table and a specified type of graph."

[0088] }

[0089] S140. Process tasks in sequence according to the execution order of each program to be executed, and obtain the evaluation result of the data to be evaluated corresponding to the data identifier.

[0090] Among them, the evaluation result can be used to characterize the attributes for evaluating the data to be evaluated. For example, the evaluation result can be data value, data reliability, data integrity, visualization analysis result, etc.

[0091] In this embodiment, each program to be executed represents a corresponding task to be processed. The execution order of each program to be executed is related to the dependency relationship between the tasks to be processed. The program to be executed for the previous task to be processed needs to be executed, and the obtained output result is given to the program to be executed for the current task to be processed, and the program to be executed for the current task to be processed can continue to execute. Of course, the output result of the program to be executed for the current task to be processed also needs to be given to the program to be executed for the next task to be processed as an input parameter. During the process of processing tasks in sequence according to the execution order of each program to be executed, parallel programs to be executed can be put into the thread pool for parallel execution. Using the reflection mechanism of the Python language, multiple tasks other than the visualization task are executed first. After the current program to be executed is completed, the next program to be executed is put into the thread pool for processing. By looping the way of putting the task code, the evaluation result is finally output. This way can effectively accelerate the execution of tasks, improve the task processing efficiency, and ensure the accuracy of the evaluation task processing. At the same time, during the task processing, it is also possible to monitor whether there are potential abnormal situations, and if so, warning prompt information can be generated for feedback.

[0092] Exemplarily, ThreadPoolExecutor can be used to execute tasks in a multi-threaded manner. ThreadPoolExecutor is a thread pool executor in Python. First, the system uses the executor.submit method to submit the code to be executed of the task to be executed to the thread pool.

[0093] The executor.submit method accepts three parameters, namely: the function parameter to be executed (such as parse_and_exe), which represents the task to be performed in each parallel step; the input parameter (such as the call_dict parameter), which refers to the input parameters required by the task, and this parameter is different in each parallel step; the storage parameter (such as result_buffer), which is used to refer to a buffer for storing results. The system uses concurrent.futures.as_completed(futures) to return an iterator that will return the results when the tasks are completed, and collects and processes the results of the pending tasks through a loop. In the loop, the system first uses the future.result() method to try to obtain the results of the pending tasks. If the task is executed successfully, the results will be included in the result variable. If an exception occurs during the execution of the task, the system uses a try-except block to catch and handle the exception, and prints the exception information to the console. It is also possible to output the number of the current parallel step in the successfully executed tasks to track the progress of task execution.

[0094] In this embodiment, the tasks are processed in sequence according to the execution order of each program to be executed, and the evaluation results of the data to be evaluated corresponding to the data identifier are obtained, including: if the program in the program to be executed is a function parameter, the target interface corresponding to the function parameter is obtained from the pre-constructed interface tool library, so as to perform task processing based on the function in the target interface.

[0095] In this embodiment, during the execution of the program to be executed, if there are function parameters in the program, the target interface corresponding to the function parameter can be retrieved from the interface tool library, and each function in the target interface can be executed in parallel to perform task processing, improving the task processing efficiency.

[0096] In this embodiment, the evaluation results can also be visually displayed. For example, a graph corresponding to the evaluation results can be generated through a visualization tool to display the precision analysis results of the data, such as Figure 2 as shown. The specific implementation method can be:

[0097] First, the parse_and_exe method can be used to execute the data visualization task in the previously defined task to be processed by using the reflection mechanism of the python language, that is, to call

[0098] the plot_dataframe_and_table method, and store the results in result_buffer_viz. Specifically, the parse_and_exe method can be called and the call_dict and result_buffer_viz are passed as parameters.

[0099] Next, the results in result_buffer_viz are extracted as finally_output, which includes graphs (plt.Axes), tables (pd.DataFrame), and text descriptions. Among them, plt.Axes is a class in the Matplotlib library used to create and manage graphs. Specifically, the plt.Axes object represents the coordinate axes and subplots in a graph. The coordinate axes are the drawing areas in the graph, which contain the coordinate system of the data and are used to draw charts and graphical elements of the data.

[0100] Then, the extracted results are processed separately. For each element in finally_output, the following operations are performed: If the element is a plt.Axes object, it is displayed as a graph. This includes showing the graph on the screen for the user to view; if the element is a pd.DataFrame object, it is displayed as a table. The table usually contains detailed information about the data; if the element is not a plt.Axes or pd.DataFrame object, it is displayed as a text description. These text descriptions can include data analysis results or other relevant information.

[0101] The advantage of this setting is that by presenting the evaluation results in multiple ways, it is beneficial for the evaluation user to better understand and evaluate the data and analysis results during the evaluation process.

[0102] In this embodiment, a task summary corresponding to the evaluation question information can also be generated based on the evaluation process data. In this way, the evaluation user can check whether the execution of the task conforms to their intention according to the task summary for the evaluation user to review.

[0103] Exemplarily, based on the evaluation process data, a task summary can be generated, which includes the initial instructions proposed by the evaluation user, task decomposition, task plan, summary of the task flow, and execution results of each step. The task summary can be presented in natural language and is used to describe the logical flow of the entire evaluation task. Further, the task summary can be displayed on the display interface. For example, the task summary can be presented to the evaluation user in text form, so that the evaluation user can view the task summary on this interface, understand the overall execution of the evaluation task, and facilitate the evaluation user to check the task summary to verify whether the task plan of the large model is consistent with their original intention, and to see if it meets their intentions and requirements, which helps the evaluation user to complete the data asset evaluation task more quickly and accurately, thereby improving the evaluation accuracy and efficiency.

[0104] The technical solution provided in this embodiment integrates multiple steps, from the understanding of the question instruction, task decomposition, task plan generation, task execution to result visualization, to achieve fast and efficient execution of the data asset evaluation task. At the same time, it supports various types of data asset tasks, such as data accuracy analysis, record filling rate analysis, cost composition analysis, sales revenue prediction, etc. The evaluation user can issue different instructions according to needs, and the system will execute the corresponding tasks according to the instructions and give the evaluation results, improving the convenience and automation of the evaluation and meeting the user's needs for different data asset evaluations.

[0105] In this embodiment, an interface tool library can also be pre-constructed. The implementation method of constructing the interface tool library can be: generating at least one operation instruction based on at least one predetermined seed instruction; analyzing the at least one operation instruction to determine at least one interface tool to be configured; configuring the association information corresponding to each of the interface tools to be configured; generating a function body corresponding to each of the interface tools to be configured according to each piece of the association information to obtain at least one interface to be selected; and constructing an interface tool library based on the at least one interface to be selected.

[0106] Among them, the operation instruction includes at least one of a data acquisition attribute, a data processing attribute, and a visualization attribute.

[0107] In this embodiment, some seed instructions can be pre-configured, such as "load dataset 001 and analyze its record filling rate", and more operation instructions about the dataset can be generated through a large model to obtain richer instructions. These operation instructions can include instructions for operations such as data acquisition, data processing, analysis, and visualization. Further, the large model can be used to analyze these operation instructions to determine which interface tools are needed to implement the requests of these operation instructions. The association information such as the interface function name, input parameters, output parameters, and interface comments of the interface tool to be configured can be configured in a natural language manner. Furthermore, the ability of the large model to generate code can be utilized to generate the specific code for implementing the function of each interface tool to be configured as the function body. These function bodies can be respectively encapsulated to obtain at least one interface to be selected, and these interfaces to be selected constitute the interface tool library, so as to obtain the target interface corresponding to the function parameters from the interface tool library and perform task processing based on the functions in the target interface.

[0108] Exemplarily, the operation instructions include: "Load dataset 001 and analyze its data accuracy"; "Load dataset 001 and check the duplication rate of data records"; "Load dataset 001 and calculate the missing rate of data records"; "Load dataset 001 and detect outliers in the data"; "Load dataset 001 and predict the future sales revenue of the dataset"; "Load dataset 001 and analyze the cost composition of the dataset"; "Load datasets 001 and 002 and compare the historical sales records of dataset 001 and dataset 002"; "Load dataset 001 and visualize the data in the dataset in tabular form", etc. The large model summarizes these instructions to determine the interface tools to be configured and the configuration interface association information. For example, the interface tools to be configured include: a data acquisition interface for extracting data from a data source, with parameters such as the data source name, date range, etc.; a data processing interface for performing operations such as calculation, filtering, and cleaning on the data, with parameters including the data processing operations to be executed; a data display interface for visually presenting the processed data, with parameters such as the chart type, data columns, etc.; a file output interface for further processing the processed data. Further, using the code generation ability of the large model, specific code implementations are generated for each interface tool to obtain an interface tool library.

[0109] The technical solution of this embodiment determines the question information to be used based on the evaluation question information including the data identifier; determines at least one task to be processed corresponding to the question information to be used and the task description information corresponding to each task to be processed; determines the program to be executed corresponding to each task to be processed based on each task description information; and processes the tasks in sequence according to the execution order of each program to be executed to obtain the evaluation result of the data to be evaluated corresponding to the data identifier, solving the problems in the prior art of high analysis cost and poor effect caused by manually or using software to analyze data assets. It realizes receiving the evaluation question information input by the user in the form of AI questions and processing it to obtain the question information to be used that can more clearly and professionally express the user's evaluation intention, improving the accuracy and effectiveness of the evaluation process. Furthermore, at least one task to be processed adapted to the question information to be used is selected, and the task description information of each task to be processed is determined. According to the task description information, the program to be executed for the task to be processed is determined. Then, the tasks are processed in sequence according to the execution order of each program to be executed to obtain the evaluation result of the data to be evaluated corresponding to the data identifier, achieving the effect of reducing the analysis cost of data assets while improving the accuracy and convenience of data asset analysis and enhancing the analysis effect to meet the user's needs for data asset analysis.

[0110] Embodiment 2

[0111] Figure 3 is a schematic structural diagram of a data asset evaluation device based on a large model according to Embodiment 2 of the present invention. As Figure 3As shown in the figure, the device includes: a to-be-used question information determination module 210, a task description information determination module 220, a to-be-executed program determination module 230, and an evaluation result determination module 240.

[0112] Among them, the to-be-used question information determination module 210 is configured to determine the to-be-used question information based on the evaluation question information including the data identifier; the task description information determination module 220 is configured to determine at least one to-be-processed task corresponding to the to-be-used question information and the task description information corresponding to each to-be-processed task; the to-be-executed program determination module 230 is configured to determine the to-be-executed program corresponding to each to-be-processed task based on each task description information; the evaluation result determination module 240 is configured to sequentially perform task processing according to the execution order of each to-be-executed program to obtain the evaluation result of the to-be-evaluated data corresponding to the data identifier.

[0113] The technical solution of this embodiment determines the to-be-used question information based on the evaluation question information including the data identifier; determines at least one to-be-processed task corresponding to the to-be-used question information and the task description information corresponding to each to-be-processed task; determines the to-be-executed program corresponding to each to-be-processed task based on each task description information; sequentially performs task processing according to the execution order of each to-be-executed program to obtain the evaluation result of the to-be-evaluated data corresponding to the data identifier, solving the problems of high analysis cost and poor effect caused by manually or software analyzing data assets in the prior art, realizing receiving and processing the evaluation question information input by the user in the form of AI questions and answers to obtain the to-be-used question information that can more clearly and professionally express the user's evaluation intention, improving the accuracy and effectiveness of evaluation processing, and then selecting at least one to-be-processed task adapted to the to-be-used question information and determining the task description information of each to-be-processed task, determining the to-be-executed program of the to-be-processed task according to the task description information, and then sequentially performing task processing according to the execution order of each to-be-executed program to obtain the evaluation result of the to-be-evaluated data corresponding to the data identifier, achieving the effect of reducing the analysis cost of data assets while improving the accuracy and convenience of data asset analysis and improving the analysis effect to meet the user's demand for data asset analysis.

[0114] Based on the above device, optionally, the to-be-used question information determination module 210 includes an evaluation question information acquisition unit and a to-be-used question information determination unit.

[0115] The evaluation question information acquisition unit is configured to acquire the evaluation question information including the data identifier;

[0116] A question information to be used determination unit, configured to determine question information to be used according to a pre-determined first prompt word template and the evaluation question information; wherein, the first prompt word template includes at least one phrase and an evaluation convention corresponding to the at least one phrase.

[0117] Based on the above device, optionally, the task description information determination module 220 includes a task prompt information determination unit and a task description information determination unit.

[0118] The task prompt information determination unit is configured to retrieve task prompt information; wherein, the task prompt information includes at least one of description guiding information, information generation format, a set of tasks to be selected for at least one task type, a task description word convention, and a description information example.

[0119] The task description information determination unit is configured to process the question information to be used according to the task prompt information, determine at least one task to be processed, and determine task description information corresponding to each of the tasks to be processed; wherein, the task description information includes a task number, a task type, and a task description.

[0120] Based on the above device, optionally, the to-be-executed program determination module 230 includes a response prompt information determination unit and a to-be-executed program determination unit.

[0121] The response prompt information determination unit is configured to retrieve response prompt information; wherein, the response prompt information includes task guiding information, a corresponding example between description information and an interface, and interface configuration information; the interface configuration information includes an interface name, an interface comment, input parameters, and output parameters corresponding to at least one to-be-selected interface in an interface tool library.

[0122] The to-be-executed program determination unit is configured to determine a to-be-executed program corresponding to each of the tasks to be processed based on the response prompt information and each of the task description information.

[0123] Based on the above device, optionally, the to-be-executed program determination unit includes a function parameter determination unit and a to-be-executed program determination subunit.

[0124] The function parameter determination unit is configured to determine function parameters corresponding to a corresponding task to be processed based on the task description information, and the interface configuration information and / or the corresponding example in the response prompt information.

[0125] The to-be-executed program determination unit is configured to determine a to-be-executed program corresponding to the task to be processed based on the task description information, the task guiding information in the response prompt information, and the function parameters corresponding to the corresponding task to be processed.

[0126] Based on the above device, optionally, the evaluation result determination module 240 is configured to, if the program in the to-be-executed program is a function parameter, obtain a target interface corresponding to the function parameter from a pre-constructed interface tool library, so as to perform task processing based on the function in the target interface.

[0127] Based on the above device, optionally, the device further includes an interface tool library construction module, and the interface tool library construction module includes an operation instruction determination unit, a to-be-configured interface tool determination unit, an association information configuration unit, a to-be-selected interface determination unit, and an interface tool library determination unit.

[0128] The operation instruction determination unit is configured to generate at least one operation instruction based on at least one pre-determined seed instruction; wherein, the operation instruction includes at least one of a data acquisition attribute, a data processing attribute, and a visualization attribute;

[0129] The to-be-configured interface tool determination unit is configured to analyze the at least one operation instruction and determine at least one to-be-configured interface tool;

[0130] The association information configuration unit is configured to configure association information corresponding to each of the to-be-configured interface tools;

[0131] The to-be-selected interface determination unit is configured to generate a function body corresponding to each of the to-be-configured interface tools according to the respective association information, so as to obtain at least one to-be-selected interface;

[0132] The interface tool library determination unit is configured to construct an interface tool library based on the at least one to-be-selected interface.

[0133] The data asset evaluation device based on a large model provided by an embodiment of the present invention can execute the data asset evaluation method based on a large model provided by any embodiment of the present invention, and has corresponding functional modules and beneficial effects for executing the method.

[0134] Embodiment IV

[0135] Figure 4It is a schematic structural diagram of an electronic device for implementing the data asset evaluation method based on a large model according to an embodiment of the present invention. The electronic device is intended to represent various forms of digital computers, such as, for example, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, personal digital processors, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0136] As Figure 4 shown, the electronic device 10 includes at least one processor 11, and a memory communicatively connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc. Among them, the memory stores a computer program executable by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0137] Multiple components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0138] The processor 11 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the data asset evaluation method based on a large model.

[0139] In some embodiments, the data asset evaluation method based on large models can be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by the processor 11, one or more steps of the data asset evaluation method based on large models described above can be performed. Alternatively, in other embodiments, the processor 11 can be configured to execute the data asset evaluation method based on large models by any other suitable means (e.g., by means of firmware).

[0140] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a dedicated or general-purpose programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0141] The computer programs for implementing the methods of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to the processor of a general-purpose computer, a dedicated computer, or other programmable data processing devices, such that when the computer programs are executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer programs can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0142] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0143] To provide for interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can also be used to provide for interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0144] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0145] A computing system may include a client and a server. The client and the server are generally far from each other and usually interact via a communication network. The client-server relationship is created by computer programs running on respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.

[0146] It should be understood that various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is imposed herein.

[0147] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A data asset evaluation method based on a large model, characterized in that: include: Determining the question information to be used based on the evaluation question information including the data identifier; Determine at least one task to be processed corresponding to the question information to be used and task description information corresponding to each task to be processed; Based on each of the task description information, determining a program to be executed corresponding to each of the tasks to be processed; If the program in the program to be executed is a function parameter, a target interface corresponding to the function parameter is obtained from a pre-built interface tool library, so as to perform task processing based on the function in the target interface and obtain an evaluation result of the data to be evaluated corresponding to the data identifier; Also includes: Build an interface tool library; among them, The construction interface tool library includes: Based on at least one predetermined seed instruction, at least one operation instruction is generated; wherein the operation instruction includes at least one of a data acquisition attribute, a data processing attribute, and a visualization attribute; Analyzing the at least one operation instruction to determine at least one interface tool to be configured; Configuring association information corresponding to each of the interface tools to be configured; Generate a function body corresponding to each interface tool to be configured according to each of the association information to obtain at least one interface to be selected; Based on the at least one interface to be selected, an interface tool library is constructed.

2. The method according to claim 1, characterized in that The step of determining the question information to be used based on the evaluation question information including the data identifier includes: Obtaining assessment question information including data identification; The question information to be used is determined according to a predetermined first prompt word template and the evaluation question information; wherein the first prompt word template includes at least one phrase and an evaluation convention corresponding to the at least one phrase.

3. The method according to claim 1, characterized in that The determining of at least one to-be-processed task corresponding to the to-be-used question information and task description information corresponding to each to-be-processed task includes: Retrieving task prompt information; wherein the task prompt information includes at least one of description guide information, information generation format, a task set to be selected of at least one task type, task description wording conventions, and description information examples; According to the task prompt information, the question information to be used is processed to determine at least one task to be processed, and task description information corresponding to each task to be processed is determined; wherein the task description information includes a task number, a task type and a task description.

4. The method according to claim 1, characterized in that The step of determining a program to be executed corresponding to each of the tasks to be processed based on the task description information includes: Retrieve response prompt information; wherein the response prompt information includes task guidance information, corresponding examples between description information and interfaces, and interface configuration information; the interface configuration information includes an interface name, interface annotation, input parameters, and output parameters corresponding to at least one interface to be selected in the interface tool library; Based on the response prompt information and each of the task description information, a to-be-executed program corresponding to each of the to-be-processed tasks is determined.

5. The method according to claim 4, characterized in that The step of determining the to-be-executed program corresponding to each of the to-be-processed tasks based on the response prompt information and the task description information includes: Determine the function parameters corresponding to the task to be processed based on the task description information, and the interface configuration information and / or the corresponding example in the response prompt information; Based on the task description information, the task guidance information in the response prompt information, and the function parameters corresponding to the corresponding task to be processed, a program to be executed corresponding to the task to be processed is determined.

6. A data asset evaluation device based on a large model, characterized in that: include: A module for determining question information to be used, used to determine question information to be used based on the evaluation question information including the data identifier; A task description information determination module, used to determine at least one to-be-processed task corresponding to the to-be-used question information and task description information corresponding to each of the to-be-processed tasks; A to-be-executed program determination module, used to determine a to-be-executed program corresponding to each of the to-be-processed tasks based on the task description information; An evaluation result determination module is used for obtaining a target interface corresponding to the function parameter from a pre-built interface tool library if the program in the program to be executed is a function parameter, so as to perform task processing based on the function in the target interface and obtain an evaluation result of the data to be evaluated corresponding to the data identifier; Also includes: An interface tool library construction module, the interface tool library construction module comprising an operation instruction determination unit, a to-be-configured interface tool determination unit, an associated information configuration unit, a to-be-selected interface determination unit, and an interface tool library determination unit; An operation instruction determination unit, configured to generate at least one operation instruction based on at least one predetermined seed instruction; wherein the operation instruction includes at least one of a data acquisition attribute, a data processing attribute, and a visualization attribute; a to-be-configured interface tool determining unit, configured to analyze the at least one operation instruction and determine at least one to-be-configured interface tool; An associated information configuration unit, used to configure associated information corresponding to each of the interface tools to be configured; The interface to be selected determining unit is used to generate a function body corresponding to each interface tool to be configured according to each of the association information, so as to obtain at least one interface to be selected; The interface tool library determining unit is used to construct an interface tool library based on the at least one interface to be selected.

7. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the big model-based data asset assessment method according to any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the large model-based data asset evaluation method according to any one of claims 1 to 5 when executed.

Citation Information

Patent Citations

  • Application integration with a digital assistant

    CN108733438A

  • Asset evaluation method and device and electronic equipment

    CN113298383A

  • Task processing method and device, equipment and storage medium

    CN116956934A