Data analysis method and device, electronic equipment, storage medium and program product
By using the functional modules and analysis functions in the target data analysis model, data analysis problems and data source information are processed, and the execution code is generated and executed, the problem of insufficient complexity of data analysis in the existing technology is solved, and efficient data analysis effect and stability are achieved.
Patent Information
- Application Number
- CN202411970054.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-02
AI Technical Summary
Existing data analysis products are difficult to perform complex data processing and analysis, resulting in poor data analysis results.
Provide a data analysis method, which uses functional modules and analysis functions in the target data analysis model to process analysis problems and data source information, generate execution code, and execute it in the sand table to obtain analysis results.
Adaptive data analysis for different analytical problems and data source information is realized, which improves the complexity and effectiveness of data analysis, and ensures the stability of data analysis by analyzing the function constraint results.
Smart Images

Figure CN119917561A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the fields of data processing technology, text generation technology, large model technology, and large language model technology, and specifically to data analysis methods, devices, electronic devices, storage media, and program products. Background Art
[0002] At present, most data analysis products are single-task solutions based on data query or data processing, which can only be used to assist users in performing data query tasks or data processing tasks, but it is difficult to perform more complex processing and analysis of data, resulting in poor data analysis results. Summary of the invention
[0003] In view of this, the present disclosure provides a data analysis method, device, electronic device, storage medium and program product to solve the problem of poor data analysis effect.
[0004] In a first aspect, the present disclosure provides a data analysis method, the method comprising:
[0005] Obtaining a first analysis question and first data source information;
[0006] Using the functional modules and / or analysis functions corresponding to the data analysis functions in the target data analysis model, the first analysis problem and the first data source information are processed to obtain a first execution code of the first execution content;
[0007] Execute the first execution code in the target sandbox to obtain a first code execution result;
[0008] The first code execution result is processed using the functional modules and / or analysis functions corresponding to the data analysis function in the target data analysis model to obtain a first analysis result of the first analysis problem.
[0009] In a second aspect, the present disclosure provides a data analysis device, the device comprising:
[0010] A data acquisition module, used to acquire a first analysis question and first data source information;
[0011] A first processing module, configured to process the first analysis problem and the first data source information by using a function module and / or an analysis function corresponding to a data analysis function in a target data analysis model, and obtain a first execution code of a first execution content;
[0012] A code execution module, used to execute the first execution code in a target sandbox to obtain a first code execution result;
[0013] The second processing module is used to process the first code execution result by using the functional module and / or analysis function corresponding to the data analysis function in the target data analysis model to obtain the first analysis result of the first analysis problem.
[0014] In a third aspect, the present disclosure provides an electronic device, comprising: a memory and a processor, the memory and the processor are communicatively connected to each other, computer instructions are stored in the memory, and the processor executes the above-mentioned data analysis method by executing the computer instructions.
[0015] In a fourth aspect, the present disclosure provides a computer-readable storage medium having computer instructions stored thereon, the computer instructions being used to enable a computer to execute the above-mentioned data analysis method.
[0016] In a fifth aspect, the present disclosure provides a computer program product, including computer instructions, and the computer instructions are used to enable a computer to execute the above-mentioned data analysis method.
[0017] The data analysis method provided by the embodiment of the present disclosure utilizes the functional modules and / or analysis functions corresponding to the data analysis functions in the target data analysis model to process the first analysis problem and the first data source information to obtain the first execution code of the first execution content. On the one hand, it is possible to adaptively generate the first execution code adapted to the different first analysis problems and the first data source information, and utilize the first execution code to schedule and fuse different functional modules and / or analysis functions to perform complex data processing and analysis. On the other hand, the introduction of the analysis function can constrain the data analysis results of the target data analysis module to ensure the stability of the subsequent data analysis results. Furthermore, the first code execution result of the first execution code is processed using the functional modules and / or analysis functions corresponding to the data analysis functions in the target data analysis model to obtain the first analysis result of the first analysis problem. Therefore, it is possible to adaptively select the appropriate functional modules and / or analysis functions according to the first code execution result of the first execution code to process the first code execution result, thereby effectively improving the data analysis effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the specific embodiments of the present disclosure or the technical solutions in the prior art, the drawings required for use in the specific embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0019] Figure 1 is a flow chart of a data analysis method according to an embodiment of the present disclosure;
[0020] Figure 2 is a flowchart of another data analysis method according to an embodiment of the present disclosure;
[0021] Figure 3 is a data analysis schematic diagram based on a target data analysis model according to an embodiment of the present disclosure;
[0022] Figure 4 is a flow chart of a method for generating a target data distribution model according to an embodiment of the present disclosure;
[0023] Figure 5 is a schematic diagram of a data construction method of sample data and downstream data labels according to an embodiment of the present disclosure;
[0024] Figure 6 is a structural block diagram of a data analysis device according to an embodiment of the present disclosure;
[0025] Figure 7 is a structural block diagram of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0026] In order to make the purpose, technical solution and advantages of the embodiments of the present disclosure clearer, the technical solution in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, rather than all the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present disclosure.
[0027] It is understandable that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, scope of use, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0028] For example, in response to receiving an active request from a user, a prompt message is sent to the user to clearly prompt the user that the operation requested to be performed will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, application, server, or storage medium that performs the operation of the technical solution of the present disclosure according to the prompt message.
[0029] It is determined as an optional but non-limiting implementation method that, in response to receiving an active request from the user, the method of sending the prompt information to the user may be, for example, a pop-up window, in which the prompt information may be presented in text form. In addition, the pop-up window may also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0030] It is understandable that the above notification and the process of obtaining user authorization are merely illustrative and do not constitute a limitation on the implementation of the present disclosure. Other methods that meet the relevant laws and regulations may also be applied to the implementation of the present disclosure.
[0031] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and relevant provisions.
[0032] Data analysis is crucial to the operation of an enterprise. The use of data query and analysis methods in data analysis usually requires complex data processing skills, and the technical threshold for non-technical personnel is high. Therefore, data analysis products for different users are proposed in related technologies to enable users to use data analysis products with a lower threshold and promote data-driven decision-making processes.
[0033] At present, data analysis products in related technologies still have the following problems that need to be solved urgently:
[0034] First, these data analysis products are generally based on the automated conversion of natural language to structured query language (Text to SQL, Text2SQL) to achieve simple data queries. They lack more in-depth processing and analysis of business data, which wastes valuable data and has relatively low value to users.
[0035] Second, the data analysis solutions of these data analysis products are usually divided into two languages: Structured Query Language (SQL) and Python. However, data analysis solutions based on pure Structured Query Language (SQL) and Python have some obvious shortcomings. For example, SQL is more complicated to implement in some primary analysis and it is difficult to implement some high-level analysis functions. Python is difficult to interact directly with the database.
[0036] Third, the data analysis of these data analysis products is mostly concentrated on single tasks such as data query, data analysis, and data summary, making it difficult to complete complex data analysis tasks.
[0037] Fourth, most of these data analysis products use large models to perform data analysis tasks, and the data analysis results generated by large models are usually random. If data analysis is performed purely based on large models, it is difficult to solve some corner cases and it is also difficult to ensure the stability of the results generated by the large models.
[0038] It can be concluded that most of the current data analysis products are single-task solutions based on data query or data processing, and can only be used to assist users in performing data query tasks or data processing tasks, but it is difficult to perform more complex processing and analysis of data. The actual help to users is relatively limited, resulting in poor data analysis results.
[0039] In view of this, according to an embodiment of the present disclosure, a data analysis method embodiment is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0040] In this embodiment, a data analysis method is provided, which can be used in electronic devices such as mobile phones, tablet computers, etc. Figure 1 is a flow chart of a data analysis method according to an embodiment of the present disclosure, such as Figure 1 As shown, the process includes the following steps:
[0041] Step S101, obtaining a first analysis question and first data source information.
[0042] Specifically, the first data source information adopts a lightweight markup language (Markdown).
[0043] Optionally, the first data source information includes a data source name, a field name, a field type, a field description, and an example value of the field. In actual operation, the first data source information may be adjusted according to actual conditions, and is not limited here.
[0044] Exemplarily, the first data source information may be obtained in the following manner: the user configures the first data source information. Alternatively, after the user selects the data source, the metadata of the selected data source is queried through a database such as SQL to obtain the first data source information. Alternatively, after the user uploads the data source, the metadata of the data source is extracted using a reflection mechanism to obtain the first data source information.
[0045] Step S102: Process the first analysis problem and the first data source information using the functional modules and / or analysis functions corresponding to the data analysis functions in the target data analysis model to obtain a first execution code of the first execution content.
[0046] Optionally, the target data analysis model is a natural language processing model (Natural Language Processing, NLP). In addition, other models for data analysis can be selected according to actual conditions, and are not limited here.
[0047] Optionally, the first execution code is Python code.
[0048] Optionally, the data analysis function includes a query function, a processing function, an analysis function, and a summary function. Functional modules corresponding to the data analysis function include a query module, a processing and analysis module, and a summary module. It should be noted that the above data analysis functions and functional modules can be adjusted according to actual conditions and are not limited here.
[0049] Specifically, the output of the query module is a piece of execution code. After querying the data based on the first data source information, the query module stores the query results in a data frame (DataFrame) of a Python data analysis and processing library (Pandas). The input of the built-in function of DataFrame is a string (str) type, corresponding to a query statement of a database. The output of the built-in function is a Pandas DataFrame class, corresponding to the query result of the query statement.
[0050] Specifically, the processing and analysis module is used to generate execution code based on the query results and the target data analysis model to realize data processing and analysis. The processing and analysis module relies on the diversity and freedom of the Python language and can support most of the functions required for data analysis. However, there are many problems in relying solely on the target data analysis model's own capabilities to generate execution code. For example, it is difficult for the target data analysis model to understand the analysis content corresponding to the business needs (i.e., the first analysis problem). The target data analysis model is difficult to handle boundary conditions in real data. Moreover, the analysis content in the target data analysis model is difficult to control. Based on this, some analysis functions can be customized to perform partial data processing and analysis, thereby using these analysis functions to improve the stability and scalability of the target data analysis model.
[0051] Specifically, the input of the summary module mainly includes the first analysis problem, the first execution code, and the first code execution result. The output of the summary module is the first analysis result of the first analysis problem (ie, the conclusion of the summary).
[0052] Optionally, the analysis function includes at least one of a comparative analysis function, a trend analysis function, a concentration analysis function, and an anomaly detection function. Among them, the comparative analysis function is used to analyze the index values (values) of different members in the query results, and to sort and compare the index values of different members. The trend analysis function is used to analyze the index values (values) of different members in the query results, and to perform time series trend analysis on the index values of different members. The concentration analysis function is used to perform concentration analysis or distribution analysis on the index values of different members in the query results. The anomaly detection function is used to group the members in the query results according to a preset grouping dimension, and to perform time series anomaly detection on the index values of the grouped members. It should be noted that in actual operation, the above analysis functions can be adjusted and expanded according to actual conditions.
[0053] Step S103, executing the first execution code in the target sandbox to obtain a first code execution result.
[0054] Specifically, a target sandbox is deployed in the data analysis environment. In addition to supporting general Python execution capabilities, the target sandbox also needs to support the execution results of analysis functions. The execution environment of the target sandbox will inherit the historical execution code and corresponding execution results of the target data analysis model.
[0055] Step S104, using the functional modules and / or analysis functions corresponding to the data analysis functions in the target data analysis model, the first code execution result is processed to obtain a first analysis result of the first analysis problem.
[0056] Specifically, the first code execution result is processed using the functional modules and / or analysis functions corresponding to the data analysis function in the target data analysis model to obtain the second execution content and its operation identifier corresponding to the first code execution result. If the operation identifier represents the generated execution code, a new first execution code is generated. The new first execution code is executed in the target sandbox to obtain a new first code execution result. The functional modules and / or analysis functions corresponding to the data analysis function in the target data analysis model are continued to be used to process the new first code execution result to obtain a new second execution content and its operation identifier. This cycle is repeated until the generated operation identifier represents the data summary. If the operation identifier represents the data summary, the summary module of the target data analysis model is used to summarize the data of all first code execution results to obtain the first analysis result.
[0057] The data analysis method provided in this embodiment uses the functional modules and / or analysis functions corresponding to the data analysis functions in the target data analysis model to process the first analysis problem and the first data source information to obtain the first execution code of the first execution content. On the one hand, it is possible to adaptively generate the first execution code that matches the different first analysis problems and the first data source information, and utilize the first execution code to schedule and integrate different functional modules and / or analysis functions to perform complex data processing and analysis. On the other hand, the introduction of the analysis function can constrain the data analysis results of the target data analysis module to ensure the stability of the subsequent data analysis results. Furthermore, the first code execution result of the first execution code is processed using the functional modules and / or analysis functions corresponding to the data analysis functions in the target data analysis model to obtain the first analysis result of the first analysis problem. Therefore, it is possible to adaptively select appropriate functional modules and / or analysis functions according to the first code execution result of the first execution code to process the first code execution result, thereby effectively improving the data analysis effect.
[0058] In this embodiment, another data analysis method is also provided, which can be used in electronic devices such as mobile phones, tablet computers, etc. Figure 2 is a flow chart of another data analysis method according to an embodiment of the present disclosure, such as Figure 2 As shown, the process includes the following steps:
[0059] Step S201, obtaining a first analysis question and first data source information. See the above step S101 for details, which will not be described in detail here.
[0060] Step S202: Process the first analysis problem and the first data source information using the functional modules and / or analysis functions corresponding to the data analysis functions in the target data analysis model to obtain a first execution code of the first execution content.
[0061] Specifically, the above step S202 includes:
[0062] Step S2021: Observe the first analysis problem and the first data source information using the target data analysis model to obtain first observation information.
[0063] It is worth noting that in a complete data analysis process, the functional modules and analysis functions involved are relatively complex, and some functional modules and analysis functions may not be executed or may be executed multiple times to implement the data analysis function. Therefore, in the data analysis method disclosed in the present invention, the reaction framework and chain of thought (ReAct CoT) paradigm is adopted to call and manage each functional module and analysis function. Among them, a complete ReAct CoT paradigm includes three links: "observation", "thinking" and "execution". The target data analysis module mainly generates related "thinking" and "execution" through the content of "observation", and then generates the observation content of "observation" based on the interaction between "execution" and the environment.
[0064] Specifically, the above step S2021 includes: in the observation phase of the target data analysis model, using the target data analysis model to observe the first analysis problem and the first data source information to obtain first observation information.
[0065] Step S2022: Process the first observation information using the functional modules and / or analysis functions corresponding to the data analysis functions in the target data analysis model to obtain the first execution content and its first operation identifier.
[0066] Specifically, the above step S2022 includes: in the thinking phase of the target data analysis model, using the functional modules and / or analysis functions corresponding to the data analysis functions in the target data analysis model to process the first observation information and obtain the first execution content and its first operation identifier.
[0067] It should be noted that the first execution content includes the execution content that needs to be taken in the current stage inferred based on the first observation information, for example, the functional modules and / or analysis functions that need to be taken.
[0068] Understandably, the introduction of the thinking link can reduce the frequency of hallucinations and errors in the target data analysis model based on CoT, and at the same time, make the data analysis process of the target data analysis model more interpretable, so that it is easier to understand the decision-making basis of the target data analysis model.
[0069] It should be noted that for different data analysis functions, the target data analysis model adopts different thinking processes.
[0070] For example, taking a trend analysis thinking process as an example, the thinking process is "The user's analysis problem is to perform trend analysis on sales in the subcategory dimension, and it needs to be filtered within the time range of the past 6 months. By observing the first data source information, we found that it contains the "sales" indicator and the "subcategory" dimension. We first convert the order date to the preset time type, then filter the data for the past 6 months, and then select the data with subcategories of "office supplies" and "furniture". Then group and aggregate the data for each subcategory by month, and calculate the total sales for each month. Finally, call the "trend analysis function" to perform trend analysis, and splice the results into a DataFrame and output it." According to this thinking process, the first execution content includes the trend analysis function and the indicator values used by the trend analysis function in the trend analysis, such as the preset time type, subcategory dimension, sales indicator, etc.
[0071] Step S2023: If the first operation identifier indicates that an execution code is to be generated, the first execution content is processed using the target data analysis model to generate a first execution code.
[0072] It should be noted that operation identifiers are introduced into the execution content generated in the thinking phase to distinguish between the two operations of generating execution code and data summary, such as Python_code_Sandbox and Summary. Python_code_Sandbox can be used to represent the generated execution code, and Summary can be used to represent the data summary.
[0073] Specifically, the above step S2023 includes: in the execution phase of the target data analysis model, if the first operation identifier represents the generation of an execution code, the first execution content is processed using the target data analysis model to generate a first execution code.
[0074] Further, if the first operation identifier represents data summary, the data is summarized using the target data analysis model and the first execution content to obtain a first analysis result of the first analysis problem.
[0075] Step S203, executing the first execution code in the target sandbox to obtain the first code execution result. See the above step S103 for details, which will not be described in detail here.
[0076] Step S204: Process the first code execution result using the functional modules and / or analysis functions corresponding to the data analysis function in the target data analysis model to obtain a first analysis result of the first analysis problem.
[0077] Specifically, the above step S204 includes:
[0078] Step S2041, using the target data analysis model to observe the first analysis problem, the first data source information and the first code execution result to obtain second observation information.
[0079] Specifically, after obtaining the first code execution result of the first execution code, the first code execution result is transmitted back to the thinking link of the target data analysis model. The above step S2041 includes: in the thinking link of the target data analysis model, using the target data analysis model to observe the first analysis problem, the first data source information and the first code execution result to obtain second observation information.
[0080] It should be noted that the first code execution result includes the code exception information and code output information of the first execution code. Even for the first execution code with execution exception, the first code execution result of the first execution code also needs to be put into the second observation information, so that the target data analysis model can repair the exception in the subsequent thinking and execution stages.
[0081] It can be understood that the observation content of the observation phase of the target data analysis model includes not only the thinking content of the thinking phase that has been carried out (such as the first execution content), but also the execution content of the execution phase that has been carried out (such as the first execution code), as well as the code execution results of the executed execution code.
[0082] Step S2042: Process the second observation information using the functional modules and / or analysis functions corresponding to the data analysis functions in the target data analysis model to obtain the second execution content and its second operation identifier.
[0083] Specifically, the above step S2042 includes: in the thinking phase of the target data analysis model, using the functional modules and / or analysis functions corresponding to the data analysis functions in the target data analysis model to process the second observation information to obtain the second execution content and its second operation identifier.
[0084] It should be noted that the second execution content includes the execution content that needs to be taken in the current stage inferred based on the second observation information, for example, the functional modules and / or analysis functions that need to be taken.
[0085] Step S2043: If the second operation identifier represents a data summary, the target data analysis model and the second execution content are used to perform a data summary on the first code execution result to obtain a first analysis result.
[0086] Optionally, if the second operation identifier represents data summary, the second execution content includes a summary module.
[0087] Specifically, the above step S2043 includes: in the execution link of the target data analysis model, if the second operation identifier represents data summary, then using the target data analysis model and the second execution content to summarize the data of the first code execution result to obtain the first analysis result.
[0088] Specifically, the summary module of the target data analysis model is used to summarize the data of the first code execution result to obtain the first analysis result.
[0089] It is worth noting that when the general target data analysis model adopts data summary, it means that the entire data analysis process of the first analysis problem is completed. At this time, the first code execution result obtained in the entire data analysis process is summarized to give a complete conclusion and obtain the first analysis result of the first analysis problem.
[0090] Furthermore, if the second operation identifier represents the generation of an execution code, the second execution content is processed using the target data analysis model to generate the next first execution code, and the above steps S203 and S204 are returned to execute to process the next first execution code.
[0091] It should be noted that each execution link will inherit the execution content and execution code of the historical execution link. The part of the target data analysis model that generates the execution code corresponds to the query module, processing and analysis module.
[0092] The data analysis method provided in this embodiment uses the target data analysis model to observe the first analysis problem, the first data source information, and the first code execution result of the first execution code generated each time. After obtaining the corresponding observation information, the function module and / or analysis function corresponding to the data analysis function in the target data analysis model is used to process the observation information to obtain the corresponding execution content and its operation identifier, so that different function modules and / or analysis functions can be adaptively called to execute a new round of data analysis functions according to the first analysis problem, the first data source information, and the code execution result each time. If the operation identifier represents the execution code, the current execution content is processed using the target data analysis model to generate the first execution code. If the operation identifier represents the data summary, the first code execution result is summarized using the target data analysis model and the current execution content to obtain the first analysis result. Therefore, the target data analysis model can be used to implement a complete data analysis process to solve complex analysis problems.
[0093] For example, taking the target data analysis model adopting the ReAct CoT paradigm framework as an example, the overall process of the target data analysis model performing data analysis on the first analysis problem and the first data source information is as follows: Figure 3As shown. In the observation phase of the target data analysis model, the first analysis problem, the first data source information, and the first code execution result of the first execution code generated each time in the execution phase are observed to obtain corresponding observation information. In the thinking phase of the target data analysis model, the functional modules and / or analysis functions corresponding to the data analysis functions in the target data analysis model are used to process the current observation information to obtain corresponding execution content. In the execution phase of the target data analysis model, if the operation identifier of the current execution content represents the generation of the execution code, the first execution code of the current execution content is generated using the target data analysis model. The current first execution code is executed using the target sandbox to obtain the first code execution result. In the execution phase of the target data analysis model, if the operation identifier of the current execution content represents the data summary, the target data analysis model and the current execution content are used to summarize the first code execution results of all the first execution codes to obtain the first analysis result of the first analysis problem.
[0094] In some optional embodiments, such as Figure 4 As shown, the target data analysis model is obtained based on the following steps:
[0095] Step S301, obtaining second data source information and sample data, where the sample data includes at least one of a second analysis question, a third execution content, a second execution code, and a second code execution result.
[0096] Specifically, the second analysis result is related to the second data source information.
[0097] It is worth noting that in the ReAct CoT process of the target data analysis model, the observation phase depends on the analysis problem, data source information, and code execution results. The thinking phase and the execution phase depend on the execution content, execution code, and analysis results. Therefore, the data types that the entire target data analysis model depends on include six types of data: analysis problem, data source information, execution content, execution code, code execution results, and analysis results. Different data collection methods can be used for different types of data.
[0098] Among them, the second data source information, the second analysis question, the second execution code and the second analysis result can be collected from known data. For example, the real data source structure and information are used as the second data source information. The second analysis question is obtained from an open source data analysis data set. The existing query and analysis code are converted into the execution code corresponding to the preset data analysis model to obtain the second execution code. The conclusion in the target document (such as the document in the second data source information) is used as the second analysis result.
[0099] Step S302: construct downstream data labels for sample data in a preset data analysis model.
[0100] Specifically, the downstream data type of the sample data in the preset data analysis model is determined, and the downstream data of the sample data in the preset data analysis model is obtained based on the downstream data type to construct the downstream data label of the sample data.
[0101] It should be noted that the downstream data tag includes at least one of an analysis question tag, an execution content tag, an execution code tag, and an analysis result tag.
[0102] Step S303, using the functional modules and / or analysis functions corresponding to the data analysis functions in the preset data analysis model, the second data source information and the sample data are processed to obtain a prediction result, which includes at least one of the prediction execution content, the prediction execution code, and the prediction analysis result.
[0103] Specifically, the second data source information and the functional modules and / or analysis functions corresponding to the data analysis functions in the preset data analysis model are used to predict the downstream data of the sample data to obtain a prediction result.
[0104] It should be noted that the above-mentioned step S303 can refer to the above-mentioned process of processing the first data source information and the first analysis problem by using the functional modules and / or analysis functions corresponding to the data analysis function in the target data analysis model, and the process of summarizing the data of the first code execution result by using the functional modules and / or analysis functions corresponding to the data analysis function in the target data analysis model, which will not be elaborated here.
[0105] Step S304, based on the prediction results and downstream data labels, adjust the parameters of the preset data analysis model to obtain the target data analysis model.
[0106] It should be noted that the data source information, analysis questions and code execution results are not the data predicted by the preset data analysis model. Therefore, the data source information, analysis questions and code execution results used in the preset data analysis model do not participate in the loss calculation. Only the predicted execution content, predicted execution code and predicted analysis results predicted by the preset data analysis model are used to calculate the loss of the preset data analysis model.
[0107] The data analysis method provided in this embodiment constructs downstream data labels of sample data in a preset data analysis model. Then, based on the prediction results of the sample data predicted by the preset data analysis model and the downstream data labels of the sample data, the parameters of the preset data analysis model are adjusted to obtain a target data analysis model. Therefore, the target data analysis model can learn the characteristics of different types of data and their downstream data to improve the data analysis performance of the target data analysis model.
[0108] In some optional implementations, the data analysis of the present disclosure further includes: performing statistical construction on the second data source information to obtain updated second data source information.
[0109] Understandably, considering that most of the data source information does not meet the diversity requirements of the preset data analysis model fine-tuning, after obtaining the second data source information, the obtained second data source information can be statistically constructed to obtain updated second data source information, thereby expanding the diversity of data on the basis of the second data source information. For example, general data source information does not have anomalies, and it is difficult to perform anomaly analysis. Therefore, when detecting anomalies, abnormal data can be added to the second data source information to obtain updated second data source information.
[0110] In some optional implementations, the data analysis of the present disclosure further includes: performing statistical construction on the second code execution result to obtain an updated second code execution result.
[0111] In some optional implementations, the above step S302 includes:
[0112] Step a1, acquiring downstream data of each sample data, where the downstream data of the sample data corresponds to the downstream data type of the sample data in the preset data analysis model.
[0113] Specifically, the downstream data of each sample data can be generated by using external tools (such as a general data analysis model, a data analysis model), or the downstream data of each sample data can be annotated by an annotator through manual annotation.
[0114] Step a2: construct corresponding downstream data labels based on the downstream data of the sample data.
[0115] The data analysis method provided in this embodiment obtains downstream data of each sample data to construct downstream data labels. Therefore, the preset data analysis model can be fine-tuned by utilizing the difference between the downstream data labels and the prediction results of the preset data analysis model for the sample data to ensure that the target data analysis model produces high-quality data.
[0116] In some optional implementations, the above step a1 includes:
[0117] Step a11: if the sample data includes a second analysis question, first downstream data of the second analysis question is obtained, where the first downstream data includes at least one of an execution content, an execution code, and an analysis result generated based on the second analysis question.
[0118] Optionally, an external tool is used to generate the execution content corresponding to the second analysis question. The execution code corresponding to the second analysis question is generated using the external tool and the execution content corresponding to the second analysis question. The execution code corresponding to the second analysis question is executed using the target sandbox to obtain the code execution result corresponding to the second analysis question. The analysis result corresponding to the second analysis question is generated using the external tool and the code execution result corresponding to the second analysis question.
[0119] Alternatively, the annotator annotates the execution content, execution code, and analysis results corresponding to the second analysis question.
[0120] Step a12: if the sample data includes the third execution content, obtain second downstream data of the second execution content, where the second downstream data includes at least one of an execution code generated based on the third execution content and an analysis result.
[0121] Optionally, an execution code corresponding to the third execution content is generated using an external tool. The execution code corresponding to the third execution content is executed using the target sandbox to obtain a code execution result corresponding to the third execution content. An analysis result corresponding to the third execution content is generated using the external tool and the code execution result corresponding to the third execution content.
[0122] Alternatively, the marking personnel mark the execution code and analysis results corresponding to the third execution content.
[0123] Step a13: if the sample data includes the second execution code, then obtain third downstream data of the second execution code, where the third downstream data includes an analysis result generated based on the second execution code.
[0124] Optionally, the target sandbox is used to generate a code execution result corresponding to the second execution code. The external tool and the code execution result corresponding to the second execution code are used to generate an analysis result corresponding to the second execution code.
[0125] Alternatively, a labeling staff labels the analysis result corresponding to the second execution code.
[0126] Step a14: if the sample data includes the second code execution result, fourth downstream data of the second code execution result is obtained, where the fourth downstream data includes an analysis result generated based on the second code execution result.
[0127] Optionally, an external tool and the second code execution result are used to generate an analysis result corresponding to the second code execution result.
[0128] Alternatively, a labeling staff labels the analysis result corresponding to the second code execution result.
[0129] The data analysis method provided in this embodiment obtains corresponding downstream data for different types of sample data, thereby improving the accuracy of downstream data labels.
[0130] In some optional implementations, the above step a2 includes:
[0131] Step a21, in response to a modification instruction for the downstream data of the sample data, modify the downstream data of the sample data to obtain updated downstream data.
[0132] It should be noted that the acquired downstream data may have some defects in details, which are difficult to modify or eliminate in an automated way. Therefore, it is necessary to make some modifications to the downstream data to ensure that the quality of the downstream data meets the expected requirements.
[0133] Step a22, constructing downstream data labels of corresponding sample data based on the updated downstream data.
[0134] The data analysis method provided in this embodiment modifies the downstream data of the sample data to obtain updated downstream data. Then, the downstream data labels of the corresponding sample data are constructed based on the updated downstream data. Therefore, the quality of the downstream data can be improved by modification, so as to improve the quality of the downstream data labels, thereby improving the fine-tuning effect of the preset data analysis model.
[0135] In some optional embodiments, the data analysis method of the present disclosure also includes: performing reverse reasoning on the analysis problem corresponding to the second execution code to obtain a third analysis problem; adding the third analysis problem to the sample data, and constructing a downstream data label for the third analysis problem based on the second execution code.
[0136] It should be noted that, in order to ensure the accuracy of the second execution code, the real execution code may be used on part of the second execution code.
[0137] Specifically, an external tool is used to perform reverse reasoning on the analysis problem corresponding to the second execution code to obtain a third analysis problem.
[0138] It can be understood that the third analysis question is added to the sample data, and the downstream data label of the third analysis question is constructed based on the second execution code. That is, when the third analysis question is used to fine-tune the preset data analysis model, the third analysis question is used as the input of the preset data analysis model, and the second execution code is used as the output of the preset data analysis model.
[0139] The data analysis method provided in this embodiment performs reverse reasoning on the analysis problem corresponding to the second execution code to obtain the third analysis problem and add it to the sample data. Then, the second execution code is constructed as the downstream data label of the third analysis problem. Therefore, the generation quality of the execution code of the target data analysis model can be effectively improved.
[0140] In some optional embodiments, the data analysis method of the present disclosure further includes: in response to a modification instruction for the second analysis problem, obtaining a modification prompt for the second analysis problem; and modifying the second analysis problem based on the modification prompt to obtain a modified second analysis problem.
[0141] It should be noted that most of the second analysis problems are still insufficient in complexity and diversity. Therefore, the second analysis problems can be optimized by using external tools and instruction optimization to obtain modified second analysis problems.
[0142] Specifically, the modification instruction / modification prompt based on the depth adds complex conditions to the second analysis question to obtain the modified second analysis question, so that the modified second analysis question is more complex. Alternatively, the modification instruction / modification prompt based on the breadth adds complex conditions to the second analysis question to obtain the modified second analysis question, so that the modified second analysis question is more diverse.
[0143] The data analysis method provided in this embodiment modifies the second analysis question based on the modification prompt to obtain the modified second analysis question. Therefore, the complexity and / or diversity of the second analysis question can be increased to enrich the number of samples of the second analysis question used for training.
[0144] In some optional implementations, the above step S304 includes:
[0145] Step b1, determining the weight of each prediction result based on the data quality and / or prediction difficulty of each prediction result.
[0146] Specifically, the weight of the predicted execution code is greater than the weight of the predicted execution content, and the weight of the predicted execution content is greater than the weight of the predicted analysis result.
[0147] Step b2, based on the prediction results, downstream data labels and weights, adjust the parameters of the preset data analysis model to obtain the target data analysis model.
[0148] It can be understood that the weight ratio of the entire fine-tuning of the preset data analysis model mainly depends on the data quality and / or prediction difficulty of the prediction results. For example, in data construction, some prediction execution content, prediction execution code and prediction analysis results generated based on the preset data analysis model need to reduce the weight ratio of these prediction results when calculating the loss of the preset data analysis model due to the low overall data quality. At the same time, when the prediction execution content, prediction execution code and prediction analysis results participate in the loss calculation of the preset data analysis model, the weights of the prediction execution content, prediction execution code and prediction analysis results are matched according to the prediction difficulty of the prediction execution code being greater than the prediction difficulty of the prediction execution content, and the prediction difficulty of the prediction execution content being greater than the prediction difficulty of the prediction analysis results.
[0149] The data analysis method provided in this embodiment determines the weight of each prediction result based on the data quality and / or prediction difficulty of each prediction result, and uses the weight to adjust the parameters of the preset data analysis model. Therefore, according to the weight, the model parameters corresponding to the data with lower data quality or greater prediction difficulty are adjusted emphatically to further improve the data quality generated by the target data analysis model and the data analysis quality.
[0150] It is worth noting that the preset data analysis model is often more complex and requires continuous iteration to obtain better results. Therefore, it is necessary to optimize the existing data on the basis of the existing data and add more high-quality sample data to fine-tune the preset data analysis model with high-quality sample data to obtain the target data analysis model.
[0151] As a specific example, fine-tuning a preset data analysis model mainly includes the following steps:
[0152] Step 1: Organize the six types of data: data source information, analysis problems, execution content, execution code, execution results, and analysis results.
[0153] Step 2: Generate new analysis questions, execution content, execution code, and code execution results as sample data by statistically constructing, forward generating, reverse generating, and optimizing instructions for the existing data, and construct downstream data labels for the sample data.
[0154] For example, Figure 5As shown, in the data collection stage, data source information, analysis problems, execution code, and analysis results are obtained. In the statistical construction stage, the data source information is statistically constructed to obtain new data source information, and the code execution results of the execution code are statistically constructed to obtain new code execution results. In the forward generation stage, the analysis problem of the data source information is inferred using external tools to obtain new analysis problems, and the execution content of the analysis problem is generated using external tools, and the code execution results of the execution code are generated using the target sandbox, and the analysis results of the code execution results are generated using external tools. In the reverse generation stage, the execution code is reversely inferred to obtain the analysis problem of the execution code, and the analysis problem is used as sample data, and the corresponding execution code is used as the downstream data label. In the instruction optimization stage, the analysis problem is modified based on the modification prompt to obtain the modified analysis problem. So far, the analysis problem, execution content, execution code, and code execution results of the current stage are used as sample data. In the manual labeling stage, the labeling personnel construct the downstream data label of the sample data based on the downstream data of the labeled sample data. Alternatively, the downstream data of the previously generated sample data is modified to construct the downstream data label of the sample data.
[0155] It should be noted that the entire sample data and downstream data label construction process is not completely based on Figure 5 The process shown in the figure is used for construction. In actual operation, some stages may be skipped for construction. For example, some analysis problems in the forward generation stage can be put into the instruction optimization stage or the manual labeling stage for processing. The data construction methods between different stages are used in conjunction with each other to obtain high-quality sample data and downstream data labels.
[0156] Step 3: Use the sample data and its downstream data labels to fine-tune the preset data analysis model to obtain the target data analysis model.
[0157] As a specific application example, a target application is installed on an electronic device, and the target application is used for data analysis. The target data source information and the target analysis question can be input into the target application. Then, the target application uses the data analysis method disclosed in the present invention to generate a target analysis result corresponding to the target analysis question. The target analysis result is displayed on a target display page of the target application.
[0158] It is worth noting that the data analysis method disclosed in the present invention adopts the design of the ReAct CoT paradigm, calls and integrates different functional modules and analysis functions to execute different data analysis solutions, and can support a variety of data analysis capabilities including query functions, processing functions, analysis functions, and summary functions, so as to effectively mine valuable data in the data. In addition, fine-tuning is performed in combination with the parameters of a unified data analysis model to achieve complete data analysis functions, thereby being able to solve complex data analysis problems. In addition, some built-in analysis functions are defined to support the ability to be compatible with Python and SQL, perform specific analysis functions, and ensure the stability of the analysis results of the data analysis model. After ensuring the stability of the analysis results of the data analysis model, the boundary conditions of the data analysis model can be effectively reduced, further ensuring the stability of the analysis results generated by the data analysis model.
[0159] In this embodiment, a data analysis device is also provided, which is used to implement the above-mentioned embodiments and preferred implementation modes, and the descriptions that have been made will not be repeated. As used below, the term "module" can implement a combination of software and / or hardware of a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceivable.
[0160] This embodiment provides a data analysis device, such as Figure 6 As shown, including:
[0161] The data acquisition module 401 is used to acquire the first analysis question and the first data source information;
[0162] A first processing module 402 is used to process the first analysis problem and the first data source information by using the function module and / or analysis function corresponding to the data analysis function in the target data analysis model to obtain a first execution code of the first execution content;
[0163] The code execution module 403 is used to execute the first execution code in the target sandbox to obtain the first code execution result;
[0164] The second processing module 404 is used to process the first code execution result by using the functional module and / or analysis function corresponding to the data analysis function in the target data analysis model to obtain a first analysis result of the first analysis problem.
[0165] In some optional implementations, the first processing module 402 includes:
[0166] A first observation unit, configured to observe the first analysis problem and the first data source information using a target data analysis model to obtain first observation information;
[0167] A first processing unit, configured to process the first observation information by using a functional module and / or an analysis function corresponding to the data analysis function in the target data analysis model to obtain a first execution content and a first operation identifier thereof;
[0168] The first execution unit is configured to process the first execution content by using a target data analysis model to generate a first execution code if the first operation identifier indicates that an execution code is to be generated.
[0169] In some optional implementations, the second processing module 404 includes:
[0170] A second observation unit is used to observe the first analysis problem, the first data source information and the first code execution result by using the target data analysis model to obtain second observation information;
[0171] A second processing unit, configured to process the second observation information by using a functional module and / or an analysis function corresponding to the data analysis function in the target data analysis model to obtain a second execution content and a second operation identifier thereof;
[0172] The second execution unit is used to summarize the data of the first code execution result by using the target data analysis model and the second execution content to obtain the first analysis result if the second operation identifier represents the data summary.
[0173] In some optional embodiments, the data analysis device of the present disclosure further includes:
[0174] A sample acquisition module, used to acquire second data source information and sample data, where the sample data includes at least one of a second analysis question, a third execution content, a second execution code, and a second code execution result;
[0175] A label construction module is used to construct downstream data labels of sample data in a preset data analysis model;
[0176] A model training module, used to process the second data source information and the sample data using the functional modules and / or analysis functions corresponding to the data analysis functions in the preset data analysis model to obtain a prediction result, wherein the prediction result includes at least one of a prediction execution content, a prediction execution code, and a prediction analysis result;
[0177] The parameter adjustment module is used to adjust the parameters of the preset data analysis model based on the prediction results and downstream data labels to obtain the target data analysis model.
[0178] In some optional embodiments, the label construction module includes:
[0179] A downstream data acquisition unit, used to acquire downstream data of each sample data, wherein the downstream data of the sample data corresponds to the downstream data type of the sample data in the preset data analysis model;
[0180] The label construction unit is used to construct corresponding downstream data labels based on the downstream data of the sample data.
[0181] In some optional implementations, the downstream data acquisition unit includes:
[0182] A first acquisition subunit is configured to acquire first downstream data of the second analysis question if the sample data includes the second analysis question, wherein the first downstream data includes at least one of an execution content, an execution code, and an analysis result generated based on the second analysis question;
[0183] A second acquisition subunit is used to acquire second downstream data of the second execution content if the sample data includes the third execution content, the second downstream data including at least one of an execution code generated based on the third execution content and an analysis result;
[0184] a third acquisition subunit, configured to acquire third downstream data of the second execution code if the sample data includes the second execution code, wherein the third downstream data includes an analysis result generated based on the second execution code;
[0185] The fourth acquisition subunit is used to acquire fourth downstream data of the second code execution result if the sample data includes the second code execution result, wherein the fourth downstream data includes an analysis result generated based on the second code execution result.
[0186] In some optional embodiments, the label construction unit includes:
[0187] A data modification subunit, configured to modify the downstream data of the sample data in response to a modification instruction for the downstream data of the sample data, to obtain updated downstream data;
[0188] The label construction subunit is used to construct the downstream data label of the corresponding sample data based on the updated downstream data.
[0189] In some optional embodiments, the data analysis device of the present disclosure further includes:
[0190] A reverse reasoning module, used for performing reverse reasoning on the analysis problem corresponding to the second execution code to obtain a third analysis problem;
[0191] The data adding module is used to add the third analysis question to the sample data and construct a downstream data label of the third analysis question based on the second execution code.
[0192] In some optional embodiments, the data analysis device of the present disclosure further includes:
[0193] a prompt acquisition module, configured to acquire a modification prompt for the second analysis question in response to a modification instruction for the second analysis question;
[0194] The question modification module is used to modify the second analysis question based on the modification prompt to obtain a modified second analysis question.
[0195] In some optional implementations, the parameter adjustment module includes:
[0196] A weight calculation unit, used to determine the weight of each prediction result based on the data quality and / or prediction difficulty of each prediction result;
[0197] The parameter adjustment unit is used to adjust the parameters of the preset data analysis model based on the prediction results, downstream data labels and weights to obtain the target data analysis model.
[0198] The further functional description of each of the above modules and units is the same as that of the above corresponding embodiments and will not be repeated here.
[0199] The data analysis device in this embodiment is presented in the form of a functional unit, where the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and memory that executes one or more software or fixed programs, and / or other devices that can provide the above functions.
[0200] The present disclosure also provides an electronic device having the above Figure 6 The data analysis device shown.
[0201] See also Figure 7 , Figure 7 is a structural block diagram of an electronic device provided by an optional embodiment of the present disclosure, such as Figure 7 As shown, the electronic device includes: one or more processors 501, a memory 502, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. The various components are connected to each other using different buses for communication, and can be installed on a common mainboard or installed in other ways as needed. The processor can process instructions executed in the electronic device, including instructions stored in or on the memory to display the graphical information of the GUI on an external input / output device (such as a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple electronic devices can be connected, and each device provides some necessary operations (for example, determined as a server array, a group of blade servers, or a multi-processor system). Figure 7A processor 501 is taken as an example.
[0202] The processor 501 may be a central processing unit, a network processor or a combination thereof. The processor 501 may further include a hardware chip. The hardware chip may be a dedicated integrated circuit, a programmable logic device or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic or any combination thereof.
[0203] The memory 502 stores instructions executable by at least one processor 501 , so that the at least one processor 501 executes the method shown in the above embodiment.
[0204] The memory 502 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created according to the use of the electronic device, etc. In addition, the memory 502 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 502 may optionally include a memory remotely arranged relative to the processor 501, and these remote memories may be connected to the electronic device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0205] The memory 502 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid state drive; the memory 502 may also include a combination of the above types of memory.
[0206] The electronic device also includes an input device 503 and an output device 504. The processor 501, the memory 502, the input device 503 and the output device 504 may be connected via a bus or other means. Figure 7 The example of connecting through bus is taken in the following.
[0207] The input device 503 can receive input digital or character information, and generate key signal input related to the user settings and function control of the electronic device, such as a touch screen, a keypad, a mouse, a track pad, a touch pad, an indicator rod, one or more mouse buttons, a trackball, a joystick, etc. The output device 504 may include a display device, an auxiliary lighting device (e.g., an LED) and a tactile feedback device (e.g., a vibration motor), etc. The above-mentioned display device includes but is not limited to a liquid crystal display, a light emitting diode, a display and a plasma display. In some optional embodiments, the display device can be a touch screen.
[0208] The embodiments of the present disclosure also provide a computer-readable storage medium. The above-mentioned method according to the embodiments of the present disclosure can be implemented in hardware, firmware, or can be implemented as a computer code that can be recorded in a storage medium, or can be implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and will be stored in a local storage medium and downloaded through a network, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state drive, etc.; further, the storage medium can also include a combination of the above-mentioned types of memory. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor, or hardware, the method shown in the above embodiment is implemented.
[0209] A part of the present disclosure may be applied as a computer program product, such as a computer program instruction, which, when executed by a computer, can call or provide the method and / or technical solution according to the present disclosure through the operation of the computer. Those skilled in the art should understand that the existence of computer program instructions in computer-readable media includes, but is not limited to, source files, executable files, installation package files, etc., and accordingly, the way in which computer program instructions are executed by a computer includes, but is not limited to: the computer directly executes the instruction, or the computer compiles the instruction and then executes the corresponding compiled program, or the computer reads and executes the instruction, or the computer reads and installs the instruction and then executes the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to the computer.
[0210] Although the embodiments of the present disclosure have been described in conjunction with the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present disclosure, and such modifications and variations are all within the scope defined by the appended claims.
Claims
1. A data analysis method, characterized in that: The method comprises: Obtaining a first analysis question and first data source information; Using the functional modules and / or analysis functions corresponding to the data analysis functions in the target data analysis model, the first analysis problem and the first data source information are processed to obtain a first execution code of the first execution content; Execute the first execution code in the target sandbox to obtain a first code execution result; The first code execution result is processed using the functional modules and / or analysis functions corresponding to the data analysis function in the target data analysis model to obtain a first analysis result of the first analysis problem.
2. The data analysis method according to claim 1, characterized in that: The method of processing the first analysis problem and the first data source information by using the functional module and / or analysis function corresponding to the data analysis function in the target data analysis model to obtain the first execution code of the first execution content includes: Observe the first analysis question and the first data source information using the target data analysis model to obtain first observation information; Processing the first observation information by using a functional module and / or an analysis function corresponding to a data analysis function in the target data analysis model to obtain the first execution content and its first operation identifier; If the first operation identifier indicates generating an execution code, the first execution content is processed using the target data analysis model to generate the first execution code.
3. The data analysis method according to claim 1, characterized in that: The processing of the first code execution result by using the functional module and / or analysis function corresponding to the data analysis function in the target data analysis model to obtain a first analysis result of the first analysis problem includes: Observe the first analysis question, the first data source information, and the first code execution result using the target data analysis model to obtain second observation information; Processing the second observation information by using a functional module and / or an analysis function corresponding to a data analysis function in the target data analysis model to obtain a second execution content and a second operation identifier thereof; If the second operation identifier represents a data summary, the target data analysis model and the second execution content are used to perform a data summary on the first code execution result to obtain the first analysis result.
4. The data analysis method according to claim 1, characterized in that: The target data analysis model is obtained based on the following steps: Acquire second data source information and sample data, where the sample data includes at least one of a second analysis question, a third execution content, a second execution code, and a second code execution result; Constructing downstream data labels for the sample data in a preset data analysis model; Processing the second data source information and the sample data using a functional module and / or an analysis function corresponding to a data analysis function in the preset data analysis model to obtain a prediction result, wherein the prediction result includes at least one of a prediction execution content, a prediction execution code, and a prediction analysis result; Based on the prediction results and the downstream data labels, the parameters of the preset data analysis model are adjusted to obtain the target data analysis model.
5. The data analysis method according to claim 4, characterized in that: The constructing downstream data labels of the sample data in a preset data analysis model includes: Acquire downstream data of each of the sample data, where the downstream data of the sample data corresponds to a downstream data type of the sample data in the preset data analysis model; The downstream data label corresponding to the sample data is constructed based on the downstream data of the sample data.
6. The data analysis method according to claim 5, characterized in that: The obtaining of downstream data of each of the sample data comprises: If the sample data includes a second analysis question, obtaining first downstream data of the second analysis question, wherein the first downstream data includes at least one of an execution content, an execution code, and an analysis result generated based on the second analysis question; If the sample data includes the third execution content, obtaining second downstream data of the second execution content, wherein the second downstream data includes at least one of an execution code generated based on the third execution content and an analysis result; If the sample data includes a second execution code, obtaining third downstream data of the second execution code, wherein the third downstream data includes an analysis result generated based on the second execution code; If the sample data includes a second code execution result, fourth downstream data of the second code execution result is obtained, where the fourth downstream data includes an analysis result generated based on the second code execution result.
7. The data analysis method according to claim 5, characterized in that: The downstream data constructing the corresponding downstream data label based on the sample data includes: In response to a modification instruction for downstream data of the sample data, modify the downstream data of the sample data to obtain updated downstream data; Construct a downstream data label corresponding to the sample data based on the updated downstream data.
8. The data analysis method according to claim 4, characterized in that: The method further comprises: Perform reverse reasoning on the analysis problem corresponding to the second execution code to obtain a third analysis problem; The third analysis question is added to the sample data, and a downstream data tag of the third analysis question is constructed based on the second execution code.
9. The data analysis method according to claim 4, characterized in that: The method further comprises: In response to a modification instruction for the second analysis question, obtaining a modification prompt for the second analysis question; The second analysis question is modified based on the modification prompt to obtain a modified second analysis question.
10. The data analysis method according to claim 4, characterized in that: The parameters of the preset data analysis model are adjusted based on the prediction results and the downstream data labels to obtain the target data analysis model, including Determining the weight of each of the prediction results based on the data quality and / or prediction difficulty of each of the prediction results; Based on the prediction results, the downstream data labels and the weights, the parameters of the preset data analysis model are adjusted to obtain the target data analysis model.
11. A data analysis device, characterized in that: The device comprises: A data acquisition module, used to acquire a first analysis question and first data source information; A first processing module, configured to process the first analysis problem and the first data source information by using a function module and / or an analysis function corresponding to a data analysis function in a target data analysis model, and obtain a first execution code of a first execution content; A code execution module, used to execute the first execution code in a target sandbox to obtain a first code execution result; The second processing module is used to process the first code execution result by using the functional module and / or analysis function corresponding to the data analysis function in the target data analysis model to obtain the first analysis result of the first analysis problem.
12. An electronic device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the data analysis method according to any one of claims 1 to 10 by executing the computer instructions.
13. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the data analysis method according to any one of claims 1 to 10.
14. A computer program product, characterized in that The method comprises computer instructions for causing a computer to execute the data analysis method according to any one of claims 1 to 10.