Model training data generation method and device applied to table, equipment and medium
By acquiring and classifying the logical processing structure and parameter category information of formula operation data, and using a data generation model to reverse-generate model training data, the problem of insufficient accuracy of multidimensional table formula training data in general large model bases is solved. This achieves efficient and accurate training dataset generation, improving model training efficiency and accuracy.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-13
- Publication Date
- 2026-03-17
AI Technical Summary
The lack of knowledge about multidimensional table formulas in general large model bases leads to accuracy problems in directly synthesized training datasets. Furthermore, existing dataset collection methods are costly, inefficient, or have small data volumes, making it impossible to guarantee the consistency of training data and user data distribution.
By acquiring the logical processing structure and parameter category information of the formula operation data, we perform classification processing and sampling, and use the pre-trained data generation model to reverse generate target operation information, operation description information, question information and operation dataset, and synthesize multiple model training data to ensure that their distribution is consistent with the user data.
It generates high-accuracy, low-cost model training data, improving the accuracy and efficiency of model training and helping users automatically generate multidimensional table formulas.
Smart Images

Figure CN120930807B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of large model technology, specifically to methods, apparatus, devices, and media for generating model training data applied to tables. Background Technology
[0002] With the continuous improvement of large-scale model capabilities, synthesizing training datasets using large models has become a common method for constructing training datasets in recent years. However, some general-purpose large-scale model bases, due to their lack of specialized training in certain areas, lack knowledge in those domains, leading to accuracy issues when directly synthesizing training data for those domains. For example, using a general-purpose large-scale model base to synthesize a training dataset of multidimensional table formulas will result in inaccuracies because the general-purpose large-scale model base lacks knowledge related to multidimensional table formulas. Summary of the Invention
[0003] In view of this, this disclosure provides a method, apparatus, device, and medium for generating model training data for tables, in order to solve the problem of inaccurate training data generated directly from a model.
[0004] In a first aspect, this disclosure provides a method for generating model training data for tables, comprising: acquiring the logical processing structure and parameter category information corresponding to each formula operation data, and determining the operation combination corresponding to each logical processing structure and parameter category information; classifying each operation combination to obtain multiple target operation combinations and the weights corresponding to each target operation combination; sampling the logical processing structure and parameter category information corresponding to each target operation combination according to the weights to obtain multiple target logical processing structures and parameter category information; for any target logical processing structure and parameter category information, based on a pre-trained data generation model, obtaining target operation information and operation description information corresponding to the target operation information, and obtaining question information and operation dataset based on the target operation information and operation description information; merging the target operation information, operation description information, question information, and operation dataset into model training data corresponding to the target logical processing structure and parameter category information to obtain multiple model training data.
[0005] Secondly, this disclosure provides a model training data generation device for tables, comprising: an acquisition module for acquiring the logical processing structure and parameter category information corresponding to each formula operation data, and determining the operation combination corresponding to each logical processing structure and parameter category information; a classification processing module for classifying each operation combination to obtain multiple target operation combinations and the weights corresponding to each target operation combination; a sampling module for sampling the logical processing structure and parameter category information corresponding to each target operation combination according to the weights to obtain multiple target logical processing structures and parameter category information; a generation module for, for any target logical processing structure and parameter category information, based on a pre-trained data generation model, obtaining target operation information and operation description information corresponding to the target operation information, and obtaining question information and operation dataset based on the target operation information and operation description information; and a merging module for merging the target operation information, operation description information, question information, and operation dataset into model training data corresponding to the target logical processing structure and parameter category information to obtain multiple model training data.
[0006] Thirdly, this disclosure provides an electronic device, including: a memory and a processor, which are communicatively connected to each other. The memory stores computer instructions, and the processor executes the computer instructions to perform the model training data generation method for tables described in the first aspect or any corresponding embodiment thereof.
[0007] Fourthly, this disclosure provides a computer-readable storage medium storing computer instructions for causing a computer to perform the model training data generation method for tables described in the first aspect or any corresponding embodiment thereof.
[0008] Fifthly, this disclosure provides a computer program product, including computer instructions for causing a computer to execute the model training data generation method for tables applied in any of the first aspects or corresponding embodiments described above.
[0009] The model training data generation method for tables provided in this disclosure classifies the operation combinations corresponding to various logical processing structures and parameter category information to obtain multiple target operation combinations and their corresponding weights. Then, it samples the logical processing structures and parameter category information corresponding to each target operation combination using these weights, ensuring that the distribution of the sampled target logical processing structures and parameter category information matches the distribution of the formula operation data. Furthermore, since the formula operation data is grammatically correct and accurate, using a data generation model to perform step-by-step reverse generation based on each target logical processing structure and parameter category information yields multiple accurate model training data sets with the same distribution as the formula operation data, containing target operation information, operation description information, operation datasets, and question information. This solves the problem of inaccurate training data generated directly from the model. Subsequent model training based on these multiple training data sets improves model accuracy and training efficiency. Attached Figure Description
[0010] To more clearly illustrate the technical solutions in the specific embodiments or related technologies of this disclosure, the accompanying drawings used in the description of the specific embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0011] Figure 1 This is a schematic diagram illustrating an application scenario provided according to embodiments of this disclosure;
[0012] Figure 2 This is a flowchart illustrating a method for generating model training data for tables according to an embodiment of the present disclosure;
[0013] Figure 3 This is a schematic diagram of a multidimensional table interactive page provided according to an embodiment of this disclosure;
[0014] Figure 4 This is a flowchart illustrating another method for generating model training data for tables according to an embodiment of the present disclosure;
[0015] Figure 5 This is a flowchart illustrating another method for generating model training data for tables, according to an embodiment of the present disclosure.
[0016] Figure 6 This is a schematic diagram of a verification page provided according to an embodiment of the present disclosure;
[0017] Figure 7 This is a flowchart illustrating a specific method for generating model training data for tables, according to an embodiment of this disclosure.
[0018] Figure 8 This is a schematic diagram of a model training data generation device for tables provided according to an embodiment of the present disclosure;
[0019] Figure 9 This is a schematic diagram of the hardware structure of an electronic device according to an embodiment of the present disclosure. Detailed Implementation
[0020] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, and not all embodiments. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.
[0021] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0022] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.
[0023] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0024] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.
[0025] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.
[0026] With the continuous improvement of large-scale model capabilities, synthesizing training data using large-scale models has become a common method for constructing training data in recent years. However, some general-purpose large-scale model bases lack knowledge in certain domains due to the lack of specialized training in those areas, which leads to accuracy issues when directly synthesizing training data for those domains.
[0027] For example, multidimensional table formulas are an important method for analyzing data in multidimensional tables. Users can use a wealth of functions to achieve automatic calculation and efficient data processing. However, due to the unique syntax of multidimensional table formulas, writing a correct formula can be challenging for users, especially those who wish to construct complex multidimensional table formulas. They often need to search for relevant knowledge or seek other assistance to achieve the goal of correctly writing multidimensional table formulas.
[0028] Therefore, leveraging the powerful language understanding capabilities of large-scale models to assist users in writing multidimensional table formulas becomes a possibility. However, existing general-purpose large-scale model bases have limited knowledge of formulas, especially unique formula syntax (such as multidimensional table formulas). These general-purpose large-scale model bases may lack any relevant knowledge, thus requiring optimization to improve their performance.
[0029] However, whether fine-tuning a general large model base through prompt engineering or post-training methods such as supervised fine-tuning (SFT) requires a large amount of formula training datasets in the relevant domain to evaluate or train the general large model base in order to improve the performance of the large model on the relevant formulas.
[0030] Because multidimensional table formulas involve user operations and are highly specialized and domain-specific, collecting formula training datasets is a challenging problem. Currently, the following methods are used to collect formula training datasets:
[0031] 1. Use publicly available datasets: Due to the specialized nature of the formulas, publicly available datasets related to formula writing are very scarce.
[0032] 2. Use alternative datasets for conversion: A formula training dataset can be created by converting a text-to-SQL (Structured Query Language, SQL) dataset. However, since the data sources in the NL2SQL dataset are mostly query statements lacking calculation-related formulas, and the NL2SQL dataset mainly consists of single-table formulas with limited cross-table calculation data, the formula training dataset obtained using this approach is not effective for training large models.
[0033] 3. Manual data annotation: This approach has drawbacks such as high cost, low efficiency, slow output, and the need for a large amount of human resources.
[0034] 4. Real-world data collection: This involves inviting users to voluntarily authorize the creation of multidimensional table formulas to form a formula training dataset. However, this approach typically requires user incentives, resulting in a generally small amount of collected data.
[0035] 5. Large Model Synthesis: Using large models for data synthesis has become a common method for constructing training datasets in recent years, offering significant advantages in data production efficiency and cost compared to the aforementioned approaches. However, due to the high precision requirements of multidimensional table formula syntax and the lack of knowledge related to multidimensional table formulas in general large model bases, the accuracy of training datasets using multidimensional table formulas generated from general large model bases is problematic.
[0036] Furthermore, the feature distribution of the model training data should match the user's real data to ensure that the trained model performs consistently in the production environment as it does during offline optimization. However, the model training data generated by the above solutions cannot guarantee consistency with the user data.
[0037] In view of this, this disclosure proposes a method for generating model training data for tables. By sampling the logical processing structure and parameter category information corresponding to the formula operation data, multiple target logical processing structures and parameter category information are obtained. This enables the selection of target logical processing structures and parameter category information from a large amount of logical processing structures and parameter category information, while ensuring that the distribution of the selected target logical processing structures and parameter category information is the same as the distribution of the formula operation data, i.e., user data. Relying on a large model to progressively infer from each target logical processing structure and parameter category information, the target operation information, operation description information, question information, and operation dataset corresponding to each logical processing structure and parameter category information can be obtained. This achieves efficient synthesis of a large amount of high-accuracy training data at a low cost.
[0038] As one optional application scenario of this disclosure embodiment, such as Figure 1 As shown, an electronic device is equipped with a target application containing embedded code and a visual interactive page. Users can edit formula operation data in the editing area of the interactive page. The electronic device responds to the user's editing operations, obtains the user's formula operation data, and executes the formula operation data to generate the corresponding calculation result. After the formula operation data is successfully executed, the electronic device, with user authorization, executes the embedded code to collect the logical processing structure and parameter type information corresponding to the formula operation data used by the user.
[0039] After obtaining the logical processing structure and parameter type information corresponding to each formula operation data, the operation combinations corresponding to each logical processing structure and parameter category information are determined. These operation combinations are then categorized to obtain multiple target operation combinations and their corresponding weights. Subsequently, the logical processing structure and parameter category information corresponding to each target operation combination are sampled using their respective weights, resulting in multiple target logical processing structures and parameter category information. Using a pre-trained data generation model, based on each target logical processing structure and parameter category information, the model generates, step-by-step, reverse-engineering target operation information, operation description information, query information, and operation datasets corresponding to each target logical processing structure and parameter category information, generating operation description information corresponding to the target operation information. Therefore, merging the target operation information, operation description information, query information, and operation datasets corresponding to each target logical processing structure and parameter category information yields multiple sets of model training data.
[0040] Subsequently, the training dataset can be used to train a general large model base to obtain a formula writing assistant. This assistant can automatically translate natural language requirements into multidimensional table formulas, thus helping users write corresponding multidimensional table formulas, i.e., formula manipulation data.
[0041] It should be understood that the tracking code deployed in the target application according to this disclosure can be deployed in the interactive page corresponding to the target application or on the server corresponding to the target application. This disclosure does not restrict the deployment location of the tracking code, and it can be flexibly adjusted according to actual needs. Furthermore, the electronic device will only execute the tracking code after obtaining user authorization, thereby collecting the logical processing structure and parameter type information corresponding to the formula operation data used by the user.
[0042] It should be understood that electronic devices can store and maintain data. Examples of electronic devices may include supercomputers, personal computers, laptops, in-vehicle computing devices, mobile devices (such as smartphones, tablets, etc.), or combinations thereof. It should be understood that the electronic devices described herein are merely exemplary and not limiting; for example, other different types of electronic devices may also be used.
[0043] According to an embodiment of this disclosure, a method for generating model training data applied to tables is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0044] This embodiment provides a method for generating model training data for tables, which can be used in electronic devices. Figure 2 This is a flowchart of a model training data generation method applied to a table according to an embodiment of the present disclosure, such as... Figure 2 As shown, the process includes the following steps:
[0045] Step S201: Obtain the logical processing structure and parameter category information corresponding to each formula operation data, and determine the operation combination corresponding to each logical processing structure and parameter category information.
[0046] Formula operation data can be behavioral data generated when a user interacts with a target application by performing actions in the corresponding interactive areas of the target application's interactive page. The target application can be a communication and collaboration application, a web-based online application, etc., and is not limited here. For example, ... Figure 3 As shown, the target application and the interactive page are set as a multidimensional table. When a user opens the multidimensional table, the multidimensional table page can be displayed on the electronic device. The user edits the multidimensional table formula in the operation editing area corresponding to the target field to be processed, so as to process the corresponding data in the multidimensional table through the multidimensional table formula. The electronic device responds to the user's editing operation under the target field and obtains the multidimensional table formula.
[0047] The logical processing structure can be the combination relationship between functions in the formula operation data. Based on the combination relationship between functions in the formula operation data, we can infer the approximate distribution of user queries, such as calculation dates, statistical data, conditional judgments, etc., and understand the user's formula writing habits. Parameter category information can include the corresponding field information, constant information, and function information in the formula operation data. For example, if the user-input multidimensional table formula is set to IF(ISBLANK([XXX]),0,DAYS([XXX],[YYY])), then the corresponding logical processing structure and parameter category information of the multidimensional table formula can be IF(ISBLANK([field1]),value1,DAYS([field1],[field2])), where XXX and YYY are used to represent the specific parameter values in the multidimensional table formula. Therefore, the logical processing structure and parameter category information can be the function combination relationship, function information, field information, and constant information in the multidimensional table formula, etc.
[0048] As a concrete example, tracking code can be deployed in the target application, and with user authorization, the pre-deployed tracking code can be executed to collect logical processing structure and parameter category information corresponding to the formula operation data used by the user.
[0049] Here, all functions in each logical processing structure and parameter category information can be considered as a set, thus obtaining the operation combination corresponding to that logical processing structure and parameter category information. For example, for a logical processing structure and parameter category information of IF(ISBLANK([field1]), value1, DAYS([field1],[field2])), the corresponding operation combination can be {IF, ISBLANK, DAYS}, that is, one formula operation data corresponds to one operation combination.
[0050] Step S202: Classify each operation combination to obtain multiple target operation combinations and the weights corresponding to each target operation combination.
[0051] Categorization can group operations with the same combination into the same category, thus obtaining multiple target operation combinations. After categorization, one target operation combination can correspond to multiple formula operation data, as well as the logical processing structure and parameter category information corresponding to each formula operation data. Therefore, there is a correspondence between target operation combinations, formula operation data, and the corresponding logical processing structure and parameter category information. For example, if the number of formula operation data is set to 100, the corresponding number of operation combinations is also 100. Categorizing these 100 operation combinations into two main categories, such as {FILTER, SUM} and {FILTER, AND, SUM}, yields two target operation combinations: {FILTER, SUM} and {FILTER, AND, SUM}.
[0052] Weights can be used to characterize the probability of each target operation combination appearing among all operation combinations. Specifically, the weights of each target operation combination can be determined through manual annotation. Alternatively, the weights of the target operation combinations can be determined by statistically analyzing the frequency of their corresponding operation combinations among all operation combinations.
[0053] Step S203: Sample the logic processing structure and parameter category information corresponding to each target operation combination according to the weight to obtain multiple target logic processing structures and parameter category information.
[0054] Here, we can first group the logical processing structures and parameter category information corresponding to each formula operation data according to the target operation combination, resulting in multiple target groups. There is a one-to-one correspondence between the target operation combination and the target group. Then, according to the weight of the target operation combination, a portion of the logical processing structures and parameter category information is randomly selected from its corresponding target group as the target logical processing structures and parameter category information. Following the previous example, let's assume we sample 10 formula operation data from 100 formula operation data, the number of logical processing structures and parameter category information corresponding to the first target operation combination is 30 with a weight of 3 / 10, and the number of logical processing structures and parameter category information corresponding to the second target operation combination is 70 with a weight of 7 / 10. Based on this, we can sample 3 from the 30 logical processing structures and parameter category information corresponding to the first target operation combination, and 7 from the 70 logical processing structures and parameter category information corresponding to the second target operation combination, thus obtaining 10 target logical processing structures and parameter category information.
[0055] Step S204: For any target logic processing structure and parameter category information, based on the pre-trained data generation model, obtain target operation information and operation description information corresponding to the target operation information, and obtain question information and operation dataset based on the target operation information and operation description information.
[0056] The data generation model can generate corresponding target operation information, operation description information, query information, and operation datasets based on the input logical processing structure and parameter category information. This data generation model can be trained on a large language model architecture, a machine learning model architecture, or a combination of multiple model architectures; no specific limitations are made here, as long as it can achieve the function of generating target operation information, operation description information, query information, and operation datasets.
[0057] The target operation information can be derived from the complete operation formula of the user's actual input based on the logical processing structure and parameter category information. Correspondingly, the operation description information is used to explain the target operation information, that is, to interpret the target operation information in natural language, describing the function, logic, or user intent of the operation in human-understandable language.
[0058] Following the previous example, the logical processing structure and parameter category information corresponding to the collected formula operation data are defined as IF(ISBLANK([Field1]), Value1, DAYS(TODAY(),[Field1])). The data generation model generates target operation information based on IF(ISBLANK([Field1]), Value1, DAYS(TODAY(),[Field1])) which can be IF(ISBLANK([Receipt Date]), 0, DAYS(TODAY(),[Receipt Date])). Accordingly, the operation description information can explain the multidimensional table formula and the functions within it. For example, a detailed explanation of ISBLANK([Receipt Date]) and an explanation of the calculation logic and function corresponding to the multidimensional table formula.
[0059] In the process of generating target operation information based on the logical processing structure and parameter category information corresponding to the formula operation data, we can first generate business domain information, namely the scenario positioning information shown later, for the logical processing structure and parameter category information. Then, we can use the data generation model and combine it with the logical processing structure and parameter category information to obtain the corresponding target operation information and operation description information.
[0060] The operation dataset can be the table corresponding to the target operation information. Continuing the previous example, let's define the target operation information as: IF(ISBLANK([Requisition Date]),0,DAYS(TODAY(),[Requisition Date])). Then, the operation dataset can be any table that this multidimensional table formula might involve, such as the office supplies requisition table and the office supplies table. The office supplies requisition table includes employee ID, name, requisition date, and office supplies fields. The office supplies fields are related to the office supplies table, which in turn can include both ID and office supplies fields.
[0061] The question information can be the questions that users input into the formula writing assistant when using the formula writing assistant trained from the model training data to write formula operation data such as multidimensional table formulas.
[0062] Step 205: Merge the target operation information, operation description information, question information, and operation dataset into model training data corresponding to the target logic processing structure and parameter category information, resulting in multiple model training data.
[0063] In scenarios where users use a formula writing assistant to write formula operation data, such as multidimensional table formulas, the operation dataset and question information can be the content that the user inputs into the formula writing assistant, while the target operation information and operation description information can be the content generated by the formula writing assistant based on the user's input.
[0064] As a concrete example, a complete and correct set of model training data can be shown in Table 1 below.
[0065] Table 1 Model Training Data
[0066]
[0067] The model training data generation method for tables provided in this embodiment categorizes the operation combinations corresponding to various logical processing structures and parameter category information to obtain multiple target operation combinations and their corresponding weights. Then, it samples the logical processing structures and parameter category information corresponding to each target operation combination using these weights, ensuring that the distribution of the sampled target logical processing structures and parameter category information matches the distribution of the formula operation data. Furthermore, since the formula operation data is grammatically correct and accurate, the data generation model performs step-by-step reverse generation based on each target logical processing structure and parameter category information. This yields multiple accurate model training data sets with the same distribution as the formula operation data, containing target operation information, operation description information, operation datasets, and question information. This solves the problem of inaccurate training data generated directly from the model. Subsequent model training based on these multiple training data sets improves model accuracy and training efficiency.
[0068] This embodiment provides a method for generating model training data for tables, which can be used in electronic devices. Figure 4 This is a flowchart of a model training data generation method applied to a table according to an embodiment of the present disclosure, such as... Figure 4 As shown, the process includes the following steps:
[0069] Step S401: Obtain the logical processing structure and parameter category information corresponding to each formula operation data, and determine the operation combination corresponding to each logical processing structure and parameter category information. For details, please refer to [link to relevant documentation]. Figure 2 Step S201 of the illustrated embodiment will not be described again here.
[0070] Step S402: Classify the various operation combinations to obtain multiple target operation combinations and the weights corresponding to each target operation combination.
[0071] Specifically, step S402 includes:
[0072] Step S4021: Classify the various operation combinations to obtain multiple target operation combinations. Please refer to the previous text.
[0073] Step S4022: Group multiple logic processing structures and parameter category information according to the target operation combination to obtain multiple target groups, and determine the first quantity of logic processing structures and parameter category information corresponding to each target group.
[0074] Here, the logical processing structure and parameter category information corresponding to formula operation data belonging to the same target operation combination can be grouped into the same group, thus obtaining multiple target groups. Furthermore, it should be understood that since the division is based on the formula operation data corresponding to the target operation combination, the number of target operation combinations is the same as the number of target groups, and one target operation combination corresponds to one target group. A target group can correspond to one or more logical processing structures and parameter category information corresponding to formula operation data.
[0075] Step S4023: Determine the weight corresponding to each target operation combination according to the proportion of each first quantity in the total number of logical processing structures and parameter category information.
[0076] For example, the first logical processing structure and its corresponding parameter category information have the operation combination {FILTER, SUM}; the second logical processing structure and its corresponding parameter category information have the operation combination {FILTER, AND, SUM}; and the third logical processing structure and its corresponding parameter category information have the operation combination {FILTER, SUM}. Classifying these three operation combinations yields two target operation combinations: {FILTER, SUM} and {FILTER, AND, SUM}. Accordingly, based on the target operation combinations, the first logical processing structure and its corresponding parameter category information can be grouped together with the third logical processing structure and its corresponding parameter category information to obtain the first target group (corresponding to the target operation combination {FILTER, SUM}), and the second logical processing structure and its corresponding parameter category information can be grouped together to obtain the second target group (corresponding to the target operation combination {FILTER, AND, SUM}). According to statistics, there are 2 logical processing structures and parameter category information in the first target group and 1 logical processing structure and parameter category information in the second target group. The total number of logical processing structures and parameter category information is 3. Therefore, the weight of the target operation combination {FILTER, SUM} is 2 / 3, and the weight of the target operation combination {FILTER, AND, SUM} is 1 / 3.
[0077] The weight of a target operation combination is determined by the ratio of the number of logical processing structures and key operational information in the target group corresponding to the target operation combination to the total number of such structures. This method can efficiently and quickly determine the weight of the target operation combination, and the weight of each target operation combination can truly reflect the distribution characteristics of the corresponding formula operation data.
[0078] Step S403: Sample the logic processing structure and parameter category information corresponding to each target operation combination according to the weights to obtain multiple target logic processing structures and parameter category information. For details, please refer to [link to relevant documentation]. Figure 2 Step S202 of the illustrated embodiment will not be described again here.
[0079] Step S404: For any target logic processing structure and parameter category information, based on the pre-trained data generation model, obtain target operation information and operation description information corresponding to the target operation information, and obtain question information and operation dataset based on the target operation information and operation description information.
[0080] In the process of generating training data based on a data generation model, there are generally three generation schemes: synchronous generation, forward generation, and backward generation. Synchronous generation involves the data generation model providing a complete set of training data in a single output. Forward generation involves the data generation model first generating the question information, or obtaining the question information from other sources, and then generating corresponding target operation information, such as a multidimensional table formula, based on the question information. Backward generation involves the data generation model first generating the target operation information, such as a multidimensional table formula, or obtaining the target operation information, such as a multidimensional table formula, from other sources, and then using the target operation information, such as the multidimensional table formula, to infer the corresponding question information.
[0081] As mentioned above, synchronous synthesis and forward synthesis require the data generation model to have a thorough understanding of domain knowledge to ensure the accuracy of the generated target operation information, such as multidimensional table formulas. Given that no model currently exists capable of generating the target operation information, such as multidimensional table formulas, required by this disclosure and ensuring accuracy, this disclosure adopts a reverse synthesis approach. Specifically, step S404 further includes:
[0082] Step S4041: Generate the first prompt word based on the logical processing structure and parameter category information.
[0083] The first prompt word is used to present the logical processing structure and parameter category information in the form of instructions, guiding the data generation model to determine the business scenario or operational context to which the logical processing structure and parameter category information belong, thereby obtaining textual description information of the scenario positioning information. As a specific example, the first prompt word template may include an editing area for the logical processing structure and parameter category information. Therefore, the logical processing structure and parameter category information can be filled into the editing area of the first prompt word template to obtain the first prompt word. The first prompt word template can be flexibly set according to specific needs, and this disclosure does not impose any restrictions on the first prompt word template.
[0084] The first prompt can include a few-shot example section. For example, 3-5 examples can be randomly selected from a manually constructed example library each time and added to the first prompt. The first prompt then guides the data generation model's generation process. The inclusion of random examples helps diversify the scene localization information generated by the data generation model.
[0085] Step S4042: Based on the first prompt word, guide the data generation model to generate scene location information corresponding to the logical processing structure and parameter category information.
[0086] Scene location information can be used to define the business domain corresponding to the logical processing structure and parameter category information, as well as the function of the target operation information generated based on the logical processing structure and parameter category information. As a specific example, scene location information could be used in human resource management to calculate an employee's working days.
[0087] Here, one logical processing structure and parameter category information can correspond to one scene positioning information. Of course, multiple scene positioning information can also be generated for one logical processing structure and parameter category information, and this disclosure does not impose any restrictions on this.
[0088] After generating scene location information of a certain scale based on various logical processing structures and parameter category information, statistical analysis can be performed on the generated scene location information to obtain the scene to be completed, i.e., the business domain. Therefore, targeted generation can be performed based on the data generation model for the business domain to be completed, ensuring sample diversity and sample balance. For example, if the number of formula operation data is set to 100, the corresponding number of logical processing structures and parameter category information is also 100. After generating corresponding scene location information for 70 of these logical processing structures and parameter category information, a unified analysis can be performed on the scene location information corresponding to these 70 logical processing structures and parameter category information. If it is found that there is relatively little scene location information in the education domain, then scene location information can be generated specifically for the education domain.
[0089] Step S4043: Generate a second prompt word based on the logical processing structure, parameter category information, and scene positioning information.
[0090] The second prompt word is used to integrate and present logical processing structure, parameter category information, and scene location information in the form of instructions, guiding the data generation model to generate corresponding target operation information based on a clear scene context, and generating textual description information of the operation description information of the target operation information. As a specific example, the second prompt word template may include an editing area for logical processing structure and parameter category information, and an editing area for scene location information. The logical processing structure and parameter category information can be filled into the editing area of the second prompt word template, and the scene location information can be filled into the editing area of the scene location information to obtain the second prompt word. The second prompt word template can be flexibly set according to specific needs, and this disclosure does not impose any restrictions on the second prompt word template.
[0091] Step S4044: Based on the second prompt word, guide the data generation model to generate the target operation information corresponding to the logical processing structure and parameter category information, and generate the operation description information corresponding to the target operation information.
[0092] Here, the second prompt word can guide the data generation model to generate the target field corresponding to the parameter category information. The generated target field is then used to replace the relevant fields in the parameter category information. Finally, the replaced parameter category information is integrated with the logical processing structure to obtain the target operation information.
[0093] For example, for the logical processing structure and parameter category information IF(ISBLANK([field1]), value1, DAYS(TODAY(),[field1])), the second prompt information can guide the data generation model to combine scene location information to generate the requisition date corresponding to field1 and 0 corresponding to value1. After generation, replacing the logical processing structure and parameter category information IF(ISBLANK([field1]), value1, DAYS(TODAY(),[field1])) will yield IF(ISBLANK([requisition date]),0,DAYS(TODAY(),[requisition date])).
[0094] The first prompt word guides the data generation model to generate scene location information. Then, the second prompt word guides the data generation model to combine the scene location information to generate target operation information and operation description information. This ensures that the generated target operation information and operation description information are both grammatically correct and relevant to the actual scene. Subsequent model training based on the model training data can improve the model's ability to understand and generate formulas in different scenarios.
[0095] Specifically, step S404 above also includes:
[0096] Step S4045: Generate a third prompt word based on the target operation information, operation description information, and scene positioning information.
[0097] The third prompt word is used to integrate and present the target operation information, operation description information, and scene location information in the form of instructions, guiding the data generation model to generate text description information of the corresponding operation dataset around the logic, description, and scene context of the operation.
[0098] As a specific example, the third prompt word template may include a target operation information editing area, an operation description information editing area, and a scene location information editing area. Therefore, the target operation information can be filled into the target operation information editing area, the operation description information into the operation description information editing area, and the scene location information into the scene location information editing area of the third prompt word template to obtain the third prompt word. The third prompt word template can be flexibly set according to specific needs, and this disclosure does not impose any restrictions on the third prompt word template.
[0099] Step S4046: Based on the third prompt word, guide the data generation model to generate the operation dataset.
[0100] Here, the third prompt word can guide the data generation model to first extract relevant fields such as the requisition date from the target operation information, then extract the meaning of each field such as the meaning of the requisition date from the operation description information, and extract domain features from the scene location information. Then, the corresponding operation dataset can be generated based on the above information.
[0101] For example, if the generated target operation information is IF(ISBLANK([requisition date]),0,DAYS(TODAY(),[requisition date])), the corresponding operation dataset can be shown in Tables 2 and 3.
[0102] Table 2 Office Supplies Requisition Form
[0103]
[0104] Table 3 Office Supplies List
[0105]
[0106] Step S4047: Generate a fourth prompt word based on the target operation information, operation description information, scene location information, and operation dataset.
[0107] The fourth prompt is a textual description used to integrate target operation information, operation description information, scene location information, and operation dataset in the form of instructions, guiding the data generation model to generate corresponding question information based on the complete operation information. As a specific example, the fourth prompt template may include editing areas for target operation information, operation description information, scene location information, and operation dataset. Therefore, the target operation information, operation description information, scene location information, and operation dataset can be filled into the target operation information editing area, the operation description information, and the operation dataset editing area to obtain the fourth prompt. The fourth prompt template can be flexibly set according to specific needs, and this disclosure does not impose any restrictions on the fourth prompt template.
[0108] Step S4048: Based on the fourth prompt word, guide the data generation model to generate question information.
[0109] The fourth prompt word guides the data generation model to convert the logical operations in the target operation information into business actions; convert the key operation information in the target operation information into business constraints; and convert the operation dataset into query sources. The data generation model then performs semantic understanding on the business actions, business constraints, query sources, scenario location information, and operation description information to obtain the query information.
[0110] By combining target operation information, operation description information, and scenario location information, an operation dataset that can support formula execution, conforms to industry practices, and has a complete structure can be generated. Furthermore, by combining target operation information, operation description information, scenario location information, and operation dataset, user question information can be generated more accurately and more closely to the actual operation scenario.
[0111] Step S405: The target operation information, operation description information, query information, and operation dataset are merged into model training data corresponding to the target logic processing structure and parameter category information, resulting in multiple model training datasets. For details, please refer to [link to relevant documentation]. Figure 2 Step S205 of the illustrated embodiment will not be described again here.
[0112] The model training data generation method for tables provided in this embodiment samples the logical processing structure and parameter category information corresponding to each combination of target operations using weights, ensuring that the distribution characteristics of the generated model training data are the same as those of the user data. Because the syntax of the logical processing structure and parameter category information is precise, the data generation model only needs to synthesize the corresponding fields in the logical processing structure and parameter category information, greatly reducing the model's requirement for understanding formula knowledge and improving the accuracy of the synthesis.
[0113] This embodiment provides a method for generating model training data for tables, which can be used in electronic devices. Figure 5 This is a flowchart of a model training data generation method applied to a table according to an embodiment of the present disclosure, such as... Figure 5 As shown, the process includes the following steps:
[0114] Step S501: Obtain the logical processing structure and parameter category information corresponding to each formula operation data, and determine the operation combination corresponding to each logical processing structure and parameter category information. For details, please refer to [link to relevant documentation]. Figure 2 Step S201 of the illustrated embodiment will not be described again here.
[0115] Step S502 involves classifying the various operation combinations to obtain multiple target operation combinations and their corresponding weights. For details, please refer to [link to relevant documentation]. Figure 2 Step S202 of the illustrated embodiment will not be described again here.
[0116] Step S503: Sample the logic processing structure and parameter category information corresponding to each target operation combination according to the weights to obtain multiple target logic processing structures and parameter category information. For details, please refer to [link to relevant documentation]. Figure 2 Step S203 of the illustrated embodiment will not be described again here.
[0117] Step S504: For any target logic processing structure and parameter category information, based on the pre-trained data generation model, obtain target operation information and corresponding operation description information according to the target logic processing structure and parameter category information. Then, obtain query information and operation dataset based on the target operation information and operation description information. For details, please refer to... Figure 2 Step S204 of the illustrated embodiment will not be described again here.
[0118] Step S505: The target operation information, operation description information, query information, and operation dataset are merged into model training data corresponding to the target logic processing structure and parameter category information, resulting in multiple model training datasets. For details, please refer to [link to relevant documentation]. Figure 2 Step S205 of the illustrated embodiment will not be described again here.
[0119] Step S506: Perform data validation on the model training data and obtain the validation results.
[0120] Data validation can systematically check the generated model training data to verify whether it meets the preset quality standards and usage requirements. This aims to filter out invalid, erroneous, or low-quality data, ensuring the accuracy, consistency, and validity of the data ultimately used for model training, and preventing poor-quality data from affecting the model's learning performance.
[0121] As a concrete example, model training data can be validated from dimensions such as logical consistency, data integrity, and syntactic accuracy. For instance, the logical consistency of the model training data can be determined by verifying whether the operation description information accurately reflects the target operation information. As another example, the data integrity of the model training data can be determined by verifying whether it includes all four elements: target operation information, operation description information, operation dataset, and question information. Furthermore, the syntactic accuracy of the target operation information can be determined by verifying its syntax.
[0122] In an optional implementation, step S506 includes:
[0123] Step S5061: Generate the fifth prompt word using the model training data.
[0124] The fifth prompt is used to present model training data in the form of instructions, guiding each referee model to perform data verification on the training data, thereby obtaining textual descriptions of the metrics information for that model's training data. As a specific example, the fifth prompt template can include a model training data editing area. Therefore, model training data such as target operation information, operation description information, scene positioning information, and operation datasets can be filled into the model training data editing area to obtain the fifth prompt. The fifth prompt template can be flexibly set according to specific needs, and this disclosure does not impose any limitations on the fifth prompt template.
[0125] Step S5062: Use the fifth prompt word to guide the data verification process of each referee model and obtain the indicator information of each referee model output for the model training data.
[0126] The referee model can perform data validation based on the input training data. This referee model can be trained on a large language model architecture, a machine learning model architecture, or a combination of multiple model architectures; no specific limitations are imposed here, as long as it can perform data validation on the training data. Furthermore, different referee models can be different types of large language models. Moreover, the referee model can be the same model as the data generation model, or it can be a different model.
[0127] As a specific example, the indicator information may include the score given by the referee model to the model training data, along with the corresponding reasoning. To ensure the simplicity of the indicator information, the score output by the referee model can be set to 0 or 1. A score of 1 indicates that the model training data has passed the data validation of the corresponding referee model, while a score of 0 indicates that the model training data has not passed the data validation of the corresponding referee model.
[0128] Step S5063: If the information of each indicator corresponding to the model training data meets the preset verification conditions, then the model training data is determined to have passed the data verification, and the verification result is obtained.
[0129] If the various indicator information corresponding to the model training data does not meet the preset verification conditions, then the model training data is determined to have failed the data verification, and the verification result is obtained.
[0130] As a specific example, the number of ratings of 1 corresponding to the model training data can be counted. If the number of ratings of 1 exceeds half of the total number of judge models, it indicates that the model training data has passed the data validation. Conversely, if the number of ratings of 1 does not exceed half of the total number of judge models, it indicates that the model training data has not passed the data validation.
[0131] Since different referee models have slightly different focuses on validating the model training data, using multiple referee models can more comprehensively validate the target training and make the validation results more accurate.
[0132] Step S507: If the verification result indicates that the model training data has failed the verification, in response to the correction operation for the model training data, the model training data is corrected to obtain the corrected model training data.
[0133] For example, such as Figure 6 As shown, the target application may also include a verification page. When the verification result indicates that the model training data has failed the logical consistency verification, the user can click the "Edit" button corresponding to "Logical Consistency" and a verification editing area will pop up. The user can edit in the verification editing area. After editing, the user can click the "OK" button. At this time, the electronic device can respond to the user's editing operation, thereby obtaining the corrected model training data and displaying the corrected model training data.
[0134] The model training data generation method for tables provided in this embodiment performs data verification on the model training data, which can ensure the correctness of the model training data. Subsequent model training based on the model training data can avoid the impact of poor data on the model's learning effect and improve the model's ability to write formulas to manipulate data, thereby further improving the accuracy of the model's formula manipulation of data.
[0135] As a specific application embodiment of this disclosure, such as Figure 7 As shown, the model training data generation method for tables disclosed herein includes steps S701 to S705. Wherein,
[0136] Step S701: Execute the event tracking code with user authorization.
[0137] Step S702: Collect the logical processing structure and parameter category information corresponding to the formula operation data used by the user.
[0138] Step S703: Sample the logical processing structure and parameter category information corresponding to each formula operation data to obtain multiple target logical processing structures and parameter category information.
[0139] Step S704: Using the data generation model, based on the target logic processing structure and parameter category information, the model training data corresponding to the target logic processing structure and parameter category information are generated step by step in reverse.
[0140] This article will take the generation of corresponding model training data based on IF(ISBLANK([field1]), value1, DAYS(TODAY(),[field1])) as an example.
[0141] Step S7041: Generate the corresponding scene location information based on IF(ISBLANK([field1]), value1, DAYS(TODAY(),[field1])). For example, this formula might be used in the field of office supplies requisition to calculate the number of days from the requisition date to the current date.
[0142] Step S7042: Based on IF(ISBLANK([Field 1]), Value 1, DAYS(TODAY(),[Field 1])) and scene location information, and using the data generation model, generate IF(ISBLANK([Receipt Date]), 0, DAYS(TODAY(),[Receipt Date])) and generate operation description information, i.e. formula explanation, for IF(ISBLANK([Receipt Date]), 0, DAYS(TODAY(),[Receipt Date]))).
[0143] Step S7043: Based on IF(ISBLANK([receipt date]),0,DAYS(TODAY(),[receipt date])), operation description information, and scene location information, synthesize the possible operation dataset of IF(ISBLANK([receipt date]),0,DAYS(TODAY(),[receipt date])).
[0144] Step S7044: Based on IF(ISBLANK([receipt date]),0,DAYS(TODAY(),[receipt date])), operation description information, scene location information, and operation dataset, synthesize the question information that the user may ask.
[0145] Step S705: Perform data validation on the training data of each model.
[0146] This section introduces the process of data validation for training data of a model.
[0147] Step S7051: Use each referee model to perform data verification on the model training data to obtain the score value and scoring opinion output by each referee model.
[0148] In step S7052, if the number of scores of 1 is greater than half the total number of judges in the model, the target data is determined to have failed data validation; otherwise, the model training data is determined to have failed data validation. When the model training data fails validation, a manual correction can be requested.
[0149] This embodiment also provides a model training data generation device for tables, which is used to implement the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that performs a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.
[0150] This embodiment provides a model training data generation device for tables, such as... Figure 8 As shown, it includes:
[0151] The acquisition module 801 is used to acquire the logical processing structure and parameter category information corresponding to each formula operation data, and to determine the operation combination corresponding to each logical processing structure and parameter category information.
[0152] The classification processing module 802 is used to classify various operation combinations to obtain multiple target operation combinations and the weights corresponding to each target operation combination.
[0153] The sampling module 803 is used to sample the logical processing structure and parameter category information corresponding to each target operation combination according to the weight, so as to obtain multiple target logical processing structures and parameter category information.
[0154] The generation module 804 is used to generate target operation information and corresponding operation description information based on a pre-trained data generation model for any target logic processing structure and parameter category information, and to generate question information and operation dataset based on the target operation information and operation description information.
[0155] The merging module 805 is used to merge the target operation information, operation description information, question information and operation dataset into model training data corresponding to the target logical processing structure and parameter category information, thereby obtaining multiple model training data.
[0156] In some optional implementations, the classification processing module 802 includes:
[0157] The classification unit is used to classify multiple operation combinations to obtain multiple target operation combinations.
[0158] The grouping unit is used to group multiple logic processing structures and parameter category information according to the target operation combination to obtain multiple target groups, and to determine the first quantity of logic processing structures and parameter category information in each target group.
[0159] The first determining unit is used to determine the weight corresponding to each target operation combination according to the proportion of each first quantity in the total number of logical processing structures and parameter category information.
[0160] In some alternative implementations, the generation module 804 includes:
[0161] The first generation unit is used to generate the first prompt word based on the logical processing structure and parameter category information.
[0162] The second generation unit is used to guide the data generation model based on the first prompt word to obtain scene positioning information corresponding to the logical processing structure and parameter category information.
[0163] The third generation unit is used to generate a second prompt word based on the logical processing structure, parameter category information, and scene positioning information.
[0164] The fourth generation unit is used to guide the data generation model's generation process based on the second prompt word, obtain the target operation information corresponding to the logical processing structure and parameter category information, and generate operation description information corresponding to the target operation information.
[0165] In some alternative implementations, the generation module 804 further includes:
[0166] The fifth generation unit is used to generate a third prompt word based on the target operation information, operation description information, and scene positioning information.
[0167] The sixth generation unit is used to generate the operational dataset by guiding the data generation model based on the third prompt word.
[0168] The seventh generation unit is used to generate the fourth prompt word based on the target operation information, operation description information, scene location information, and operation dataset.
[0169] The eighth generation unit is used to generate question information based on the data generation model guided by the fourth prompt word.
[0170] In some alternative embodiments, the device further includes:
[0171] The verification module is used to verify the model training data and obtain the verification results.
[0172] The correction module is used to correct the model training data if the verification result indicates that the model training data has failed the verification, in response to the correction operation of the model training data, to obtain the corrected model training data.
[0173] In some optional implementations, the verification module includes:
[0174] The ninth generation unit is used to generate the fifth prompt word using the model training data.
[0175] The tenth generation unit is used to guide the data verification process of each judge model using the fifth prompt word, and to obtain the indicator information of each judge model outputting the model training data.
[0176] The second determining unit is used to determine that the model training data has passed the data verification if the various indicator information corresponding to the model training data meets the preset verification conditions, and to obtain the verification result.
[0177] The model training data generation apparatus for tables provided in this disclosure can execute the model training data generation method for tables provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects of the execution method. This solution can obtain multiple sets of accurate model training data with the same distribution as the formula operation data, containing target operation information, operation description information, operation dataset, and question information. This solves the problem of inaccurate training data generated directly from the model. Subsequent model training based on multiple sets of model training data can improve model accuracy and training efficiency. Further functional descriptions of the above modules and units are the same as in the corresponding embodiments described above, and will not be repeated here.
[0178] Figure 9 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure.
[0179] The following is a detailed reference. Figure 9 The diagram illustrates a structural schematic suitable for implementing an electronic device according to embodiments of the present disclosure. The electronic device may include a processor (e.g., a central processing unit, graphics processor, etc.) 901, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 902 or a program loaded from memory 908 into random access memory (RAM) 903. The RAM 903 also stores various programs and data required for the operation of the electronic device. The processor 901, ROM 902, and RAM 903 are interconnected via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.
[0180] Typically, the following devices can be connected to I / O interface 905: input devices 906 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 907 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; memory devices 908 including, for example, magnetic tapes, hard disks, etc.; and communication devices 909. Communication device 909 allows electronic devices to exchange data via wireless or wired communication with other devices. Although Figure 9 Electronic devices with various devices are shown, but it should be understood that it is not required to implement or have all of the devices shown, and more or fewer devices may be implemented or have instead.
[0181] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 909, or installed from a memory 908, or installed from a ROM 902. When the computer program is executed by the processor 901, it performs the functions defined in the method for generating model training data for tables according to embodiments of this disclosure.
[0182] Figure 9 The illustrated electronic device is merely an example and should not be construed as limiting the functionality and scope of the embodiments of this disclosure. This disclosure also provides a computer-readable storage medium in which the methods described above according to embodiments of this disclosure can be implemented in hardware, firmware, or implemented as recordable on a storage medium, or implemented as computer code downloaded over a network and originally stored on a remote storage medium or a non-transitory machine-readable storage medium and subsequently stored on a local storage medium. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium may be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium may also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code. When the software or computer code is accessed and executed by the computer, processor, or hardware, the model training data generation method applied to tables shown in the above embodiments is implemented.
[0183] A portion of this disclosure can be applied to computer program products, such as computer program instructions, which, when executed by a computer, can invoke or provide methods and / or technical solutions according to this disclosure through the operation of the computer. Those skilled in the art will understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, and installation package files. Accordingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executing the instructions; the computer compiling the instructions and then executing the corresponding compiled program; the computer reading and executing the instructions; or the computer reading and installing the instructions and then executing the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to a computer.
[0184] Although embodiments of the present disclosure have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present disclosure, and such modifications and variations all fall within the scope defined by the appended claims.
Claims
1. A method for generating model training data applied to a table, the method comprising: The method comprises: obtaining the logical processing structure and parameter category information corresponding to each formula operation data, and determining the operation combination corresponding to each logical processing structure and parameter category information; classifying each operation combination to obtain a plurality of target operation combinations and the weight corresponding to each target operation combination; sampling the logical processing structure and parameter category information corresponding to each target operation combination according to the weight to obtain a plurality of target logical processing structure and parameter category information; for any one target logical processing structure and parameter category information, obtaining target operation information and operation description information corresponding to the target operation information based on a pre-trained data generation model according to the target logical processing structure and parameter category information, and obtaining question information and an operation data set according to the target operation information and the operation description information, wherein the target operation information is an operation formula obtained based on the target logical processing structure and parameter category information, and the operation data set is a table corresponding to the target operation information; combining the target operation information, the operation description information, the question information and the operation data set as model training data corresponding to the target logical processing structure and parameter category information to obtain a plurality of model training data.
2. The method of claim 1, wherein, The classification of each operation combination to obtain a plurality of target operation combinations and the weight corresponding to each target operation combination comprises: classifying a plurality of operation combinations to obtain a plurality of target operation combinations; grouping a plurality of logical processing structures and parameter category information according to the target operation combination to obtain a plurality of target groups, and determining the first number of logical processing structures and parameter category information corresponding to each target group; determining the weight corresponding to each target operation combination according to the proportion of each first number in the total number of logical processing structures and parameter category information.
3. The method according to any one of claims 1 or 2, characterized in that, The data generation model is pre-trained, and the target operation information and the operation description information corresponding to the target operation information are obtained according to the target logical processing structure and parameter category information, which comprises: generating a first prompt word based on the logical processing structure and parameter category information; guiding the generation process of the data generation model based on the first prompt word to obtain scene positioning information corresponding to the logical processing structure and parameter category information; generating a second prompt word based on the logical processing structure and parameter category information and the scene positioning information; guiding the generation process of the data generation model based on the second prompt word to obtain the target operation information corresponding to the logical processing structure and parameter category information and the operation description information corresponding to the target operation information.
4. The method of claim 3, wherein, The question information and the operation data set are obtained according to the target operation information and the operation description information, which comprises: generating a third prompt word based on the target operation information, the operation description information and the scene positioning information; guiding the generation process of the data generation model based on the third prompt word to generate the operation data set; generate a fourth prompt word based on the target operation information, the operation description information, the scene positioning information, and the operation data set; guide a generation process of the data generation model based on the fourth prompt word, and generate the question information.
5. The method of claim 1, wherein, Further comprising: perform data verification on the model training data to obtain a verification result; if the verification result indicates that the model training data fails the verification, in response to a correction operation on the model training data, correct the model training data to obtain the model training data after correction.
6. The method of claim 5, wherein, The data verification on the model training data to obtain a verification result comprises: generate a fifth prompt word using the model training data; guide a data verification process of each adjudication model using the fifth prompt word to obtain index information output by each adjudication model for the model training data; if each of the index information corresponding to the model training data satisfies a preset verification condition, it is determined that the model training data passes the data verification, and the verification result is obtained. 7.A model training data generation device for a table, comprising: The device comprises: an acquisition module configured to acquire logical processing structures and parameter category information corresponding to each formula operation data, and determine operation combinations corresponding to the logical processing structures and parameter category information; a classification processing module configured to perform classification processing on each of the operation combinations to obtain a plurality of target operation combinations and weights corresponding to each of the target operation combinations; a sampling module configured to sample the logical processing structures and parameter category information corresponding to each of the target operation combinations according to the weights to obtain a plurality of target logical processing structures and parameter category information; a generation module configured to, for any one of the target logical processing structures and parameter category information, based on a data generation model that has been pre-trained, obtain target operation information and operation description information corresponding to the target operation information according to the target logical processing structure and parameter category information, and obtain question information and an operation data set according to the target operation information and the operation description information, the target operation information being an operation formula obtained based on the target logical processing structure and parameter category information, and the operation data set being a table corresponding to the target operation information; a merging module configured to merge the target operation information, the operation description information, the question information, and the operation data set into model training data corresponding to the target logical processing structure and parameter category information to obtain a plurality of the model training data.
8. An electronic device, comprising: Comprise: a memory and a processor, which are communicatively connected, the memory stores computer instructions, and the processor executes the computer instructions to perform the method for generating model training data for a table according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer instructions, and the computer instructions are used to make a computer execute the method for generating model training data for a table according to any one of claims 1 to 6.
10. A computer program product, characterised in that, Computer instructions for causing a computer to perform the model training data generation method applied to a table according to any one of claims 1 to 6 are included.
Citation Information
Patent Citations
Interaction method and device, electronic equipment and computer readable medium
CN115808993A
Data generation method and device, electronic equipment and storage medium
CN118261120A