Method, device, medium and equipment for generating simulation data for evaluating large models

By identifying common source fields and using the same data processing rules to generate simulated data in large model evaluation, the problem that mock data cannot reflect real business scenarios is solved, and the evaluation accuracy is improved while avoiding data leakage.

CN120596930BActive Publication Date: 2025-12-12BEIJING VOLCANO ENGINE TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511094894.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-05
Publication Date
2025-12-12
Estimated Expiration
2045-08-05

AI Technical Summary

Technical Problem

Existing technologies fail to effectively consider the price distribution of different products and the impact of promotional activities on prices when generating mock data. As a result, the generated mock data cannot reflect the real data in the business scenario, affecting the accuracy and reliability of large model evaluation.

Method used

By acquiring business data from real-world business scenarios, identifying fields with the same source, and processing their values ​​using the same data processing rules, simulated field values ​​are generated. This ensures that the simulated field values ​​do not contain business information and have consistent constraints, closely resembling real-world business scenarios.

Benefits of technology

The generated simulated data can provide data that closely resembles real business scenarios for large-scale model evaluation while avoiding data leakage, thereby improving the accuracy of the evaluation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120596930B_ABST
    Figure CN120596930B_ABST
Patent Text Reader

Abstract

A simulation data generation method, device, medium and equipment for evaluating a large model, relating to the technical field of computers and the technical field of large models, the simulation data generation method for evaluating the large model comprising: obtaining a to-be-processed data set, the to-be-processed data set including business data collected under a business scenario; determining homologous fields in the to-be-processed data set, the homologous fields having the same upstream fields; processing field values of the homologous fields using the same data processing rule to obtain simulation field values of the homologous fields, the simulation field values being used to support evaluation of the large model, so that the simulation field values can be ensured to not contain business information, and the field values of the homologous fields having the same upstream fields have the consistency constraint condition, so that the generated simulation field values can be close to the business data under the business scenario, thereby providing data close to the real business scenario for evaluation of the large model while avoiding data leakage.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of computer technology and the technical field of large model, in particular, to a simulation data generation method, device, medium and equipment for evaluating a large model. BACKGROUND

[0002] In the field of data-based evaluation, business data in a business scenario is a key factor to ensure the effectiveness of evaluation. By using business data to simulate real business scenarios, the function and performance of a large model are evaluated to ensure the reliable operation of the large model in actual business applications.

[0003] However, the business data collected in a real business scenario may contain business information, such as user privacy data and confidential data, which should be avoided from being leaked. That is, the business data is not suitable for directly supporting the evaluation of a large model. Therefore, Mock (simulation) data technology is a key technology to solve data leakage. SUMMARY

[0004] This summary is provided to introduce a selection of concepts that are further described below in the detailed description. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter.

[0005] In a first aspect, the present disclosure provides a simulation data generation method for evaluating a large model, comprising:

[0006] obtaining a to-be-processed data set, wherein the to-be-processed data set includes business data collected in a business scenario;

[0007] determining a homologous field in the to-be-processed data set, wherein the homologous field has the same upstream field;

[0008] processing the field values of the homologous field using the same data processing rule to obtain simulation field values of the homologous field, wherein the simulation field values are used to support the evaluation of the large model.

[0009] In a second aspect, the present disclosure provides a simulation data generation device for evaluating a large model, comprising:

[0010] an obtaining module configured to obtain a to-be-processed data set, wherein the to-be-processed data set includes business data collected in a business scenario;

[0011] a first determining module configured to determine a homologous field in the to-be-processed data set, wherein the homologous field has the same upstream field;

[0012] The data processing module is configured to process the field values of the homologous fields by using the same data processing rule to obtain simulated field values of the homologous fields, wherein the simulated field values are used to support evaluation of the large model.

[0013] In a third aspect, the present disclosure provides a computer readable medium having stored thereon a computer program, which, when executed by a processing device, implements the steps of the simulation data generation method of the first aspect.

[0014] In a fourth aspect, the present disclosure provides an electronic device comprising:

[0015] a storage device having stored thereon a computer program;

[0016] a processing device configured to execute the computer program in the storage device to implement the steps of the simulation data generation method of the first aspect.

[0017] In a fifth aspect, the present disclosure provides a computer program product comprising a computer program, which, when executed by a processor, implements the steps of the simulation data generation method of the first aspect.

[0018] By the above technical solution, the field values of the homologous fields are processed by using the same data processing rule to obtain simulated field values of the homologous fields, so that the simulated field values do not contain business information, and the field values of the homologous fields having the same upstream field have the consistency constraint condition, so that the generated simulated field values can be close to the business data under the business scenario, thereby providing data close to the real business scenario for evaluation of the large model in the case of avoiding data leakage, and further providing a data basis for improving the accuracy of the evaluation result of the large model.

[0019] Other features and advantages of the present disclosure will be described in detail in the following detailed description. BRIEF DESCRIPTION OF DRAWINGS

[0020] The above and other features, advantages, and aspects of embodiments of the present disclosure will become more apparent by describing in detail exemplary embodiments thereof with reference to the attached drawings in which:

[0021] Figure 1 is a flowchart of a simulation data generation method for evaluating a large model according to an embodiment of the present disclosure;

[0022] Figure 2 is a schematic diagram of a field of an upstream data set and a field of a downstream data set according to an embodiment of the present disclosure;

[0023] Figure 3 is a flow chart of a simulation data generation method according to an embodiment of the present disclosure;

[0024] Figure 4 is a block diagram of a simulation data generation apparatus for evaluating a large model according to an embodiment of the present disclosure;

[0025] Figure 5 is a structural schematic diagram of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0026] Embodiments of the present disclosure will be described in more detail with reference to the accompanying drawings. While certain embodiments of the present disclosure will be illustrated hereinafter, it is clearly understood that the present disclosure can be carried out in various forms and should not be construed as limited to the embodiments set forth herein, but rather should be construed to cover all alterations, equivalents, and substitutes for the embodiments of the present disclosure. It is understood that the drawings and embodiments are only for illustrative purposes and should not be construed as limiting the scope of the present disclosure.

[0027] It should be understood that each step recited in the method embodiments of the present disclosure can be executed in different order and / or in parallel. In addition, the method embodiments can include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.

[0028] The term “comprising” and variations thereof as used herein are used inclusively, i.e., “comprising but not limited to.” The term “based on” means “based, at least in part, on.” The term “one embodiment” means “at least one embodiment”; the term “another embodiment” means “at least one additional embodiment”; the term “some embodiments” means “at least some embodiments.” Related terms are defined in the description that follows.

[0029] It should be noted that the terms “first”, “second”, and the like in the present disclosure are merely used to distinguish different devices, modules or units, and do not imply the order or interdependence of the functions performed by these devices, modules or units.

[0030] It should be noted that the terms “one”, “multiple” in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that “one or more” should be understood unless otherwise explicitly indicated in the context.

[0031] The names of the messages or information exchanged between the devices in the embodiments of the present disclosure are merely used for illustrative purposes, and are not intended to limit the scope of the messages or information.

[0032] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the type, use range, use scenario, etc. of the personal information involved in the present disclosure should be informed to the user and the authorization of the user should be obtained through appropriate means according to relevant laws and regulations.

[0033] For example, in response to receiving an active request of a user, prompt information is sent to the user to explicitly prompt the user that the operation requested to be performed will require obtaining and using personal information of the user. Thus, the user can voluntarily choose whether to provide personal information to the software or hardware such as an electronic device, an application program, a server or a storage medium, etc. that performs the operation of the technical solutions of the present disclosure according to the prompt information.

[0034] As an optional but non-limiting implementation manner, in response to receiving an active request of a user, the manner of sending prompt information to the user may, for example, be a pop-up window manner, in which the prompt information can be presented in the form of text. In addition, the pop-up window can also carry selection controls for the user to select "agree" or "disagree" to provide personal information to the electronic device.

[0035] It can be understood that the above notification and obtaining of user authorization process is only illustrative and does not limit the implementation manners of the present disclosure, and other manners meeting relevant laws and regulations can also be applied to the implementation manners of the present disclosure.

[0036] At the same time, it can be understood that the data (including but not limited to the data itself, the acquisition or use of the data) involved in the present technical solutions should comply with the requirements of relevant laws and regulations and relevant provisions.

[0037] In the field of data-based evaluation and testing, business data in a business scenario is a key factor to ensure the effectiveness of evaluation and testing. By using business data to simulate real business scenarios, the function and performance of a large model can be evaluated to ensure reliable operation of the large model in actual business applications.

[0038] However, the business data collected in a real business scenario may contain business information such as user privacy data and confidential data, which is required to be avoided from being leaked. That is to say, the business data is not suitable for being directly used to support the evaluation and testing of a large model. Therefore, Mock (simulation) data technology becomes a key technology to solve data leakage.

[0039] In related technologies, only the format and type of data are considered for generating Mock data, such as randomly generating the price of a commodity, without considering the influence of the price distribution of different commodities and preferential activities on the price. This results in that the generated Mock data cannot reflect real data in a business scenario, and further affects the accuracy and reliability of the evaluation and testing of a large model.

[0040] Therefore, the simulation data generation method for evaluating a large model is provided, so that the generated simulation field values can be close to the business data in the business scenario, thereby providing a data basis for improving the accuracy of the evaluation result of the large model.

[0041] The embodiments of the present disclosure are explained and described in conjunction with the accompanying drawings.

[0042] Figure 1 is a flowchart of a simulation data generation method for evaluating a large model according to an embodiment of the present disclosure. The simulation data generation method for evaluating a large model can be applied to an electronic device. And the simulation data generation method for evaluating a large model can be executed by a simulation data generation device for evaluating a large model, wherein the simulation data generation device for evaluating a large model can be implemented by software and / or hardware, and the software and / or hardware can be configured in the electronic device. Refer to Figure 1 The simulation data generation method for evaluating a large model can include steps 110, 120 and 130.

[0043] In step 110, a to-be-processed data set is obtained, wherein the to-be-processed data set includes business data collected in a business scenario.

[0044] The to-be-processed data set includes fields and field values corresponding to the fields, which can be referred to as business data. For example, taking the business scenario as an order scenario, the business data in the order scenario includes fields such as order number, product name, purchase quantity and price, and field values corresponding to each field.

[0045] In step 120, a homologous field in the to-be-processed data set is determined, wherein the homologous field has the same upstream field.

[0046] The to-be-processed data set can be derived based on an upstream data set, and the to-be-processed data set can be referred to as a downstream data set of the upstream data set. The fields in the upstream data set can be referred to as upstream fields of the corresponding fields in the downstream data set, and the fields in the downstream data set can be referred to as downstream fields of the corresponding fields in the upstream data set.

[0047] In this embodiment, the homologous field includes at least two fields, and the corresponding upstream fields in the upstream data set are the same.

[0048] In step 130, the field values of the homologous fields are processed using the same data processing rule to obtain simulation field values of the homologous fields, wherein the simulation field values are used to support the evaluation of the large model.

[0049] It is worth noting that the data processing rule is used to rewrite the field value of the homologous field to obtain the corresponding simulated field value, thereby avoiding the disclosure of real business data.

[0050] Since the homologous fields have the same upstream field, that is, the homologous fields have the same homologous attribute. For example, for the price field of a commodity, since different commodities have different prices, in order to reasonably reflect the price situation of different commodities, it is necessary to impose the same constraint condition on the prices of different commodities, that is, to use the same data processing rule for processing.

[0051] In this way, the same data processing rule is used to process the field values of the homologous fields to obtain the simulated field values of the homologous fields, so that the simulated field values do not contain business information, and the field values of the homologous fields with the same upstream field have consistent constraint conditions, so that the generated simulated field values can be close to the business data in the business scenario, thereby providing data close to the real business scenario for the evaluation of the large model in the case of avoiding data leakage, and further providing a data basis for improving the accuracy of the evaluation result of the large model.

[0052] In some embodiments, for non-homologous fields in the to-be-processed data set, rewriting can be performed according to the type and format of the field value and other requirements to avoid disclosure of business information. It can be understood that the non-homologous field is a field in the to-be-processed data set that does not share an upstream field with other fields.

[0053] In some embodiments, the name of the field in the to-be-processed data set can be rewritten to obtain a field name that does not contain business information, which is referred to as a simulated field in the following.

[0054] In some embodiments, in the generation of simulated data for different business scenarios, by changing the configuration of data processing rules, field relationship tables and other parameters, the generation method of simulated data in a certain business scenario can be migrated to another business scenario for use, thereby being able to quickly generate simulated data that meets the requirements.

[0055] In some embodiments, the above step of determining the homologous fields in the to-be-processed data set can be implemented by: determining the homologous fields in the to-be-processed data set according to the obtained field relationship table, wherein the field relationship table is used to maintain the corresponding field of the field in the downstream data set in the upstream data set.

[0056] Therefore, the homologous fields in the to-be-processed data set can be determined by looking up the table. For example, see the following table:

[0057]

[0058] The above table is a field relationship table. If dataset B is a to-be-processed dataset, based on the above table, since the field a1 and the field a2 in the dataset B have the same upstream field, i.e., the field a, the field a1 and the field a2 are homogenous fields in the dataset B.

[0059] It is worth noting that in the case that the to-be-processed dataset has no upstream dataset, in the field relationship table, the upstream dataset of the to-be-processed dataset can be configured as itself, and the upstream field corresponding to the downstream field is also the upstream field itself.

[0060] In the above manner, the relationship table is used to maintain the fields in the downstream dataset corresponding to the fields in the upstream dataset. In this way, the homogenous fields in the to-be-processed dataset can be determined through table lookup, so that the homogenous fields in the dataset can be quickly determined.

[0061] Figure 2 is a schematic diagram of the fields of an upstream dataset and the fields of a downstream dataset according to an embodiment of the present disclosure. Referring to Figure 2 , the dataset B is a downstream dataset, and the upstream dataset of the dataset B is the dataset A. The upstream field of the field a1 and the field a2 of the dataset B is the field a, and thus the field a1 and the field a2 are homogenous fields and are processed by using the same data processing rule. For the field j of the dataset B, it does not share the upstream field with other fields, and thus the field j is a non-homogenous field in the dataset B.

[0062] In some embodiments, processing by using the same data processing rule can include at least one of the following processing:

[0063] If the field value of the homogenous field is a sensitive word, such as a company name, a dictionary is used to replace the company name with a word in the dictionary; for example, a Chinese name, if the first character hits a preset surname character, a randomly generated Chinese character is used to replace the Chinese name; for example, an English string, at least one English letter in the English string is randomly replaced;

[0064] If the field value of the homogenous field is a number (an integer or a floating point number), the number is disturbed according to a preset variation range, for example, 90%-110%. In this way, the simulated field value has a certain difference from the original field value, but can be maintained at a similar order of magnitude.

[0065] In some embodiments, the simulation data generation method described above can further include the following steps: performing data analysis on the field values corresponding to the fields in the to-be-processed data set to obtain first data analysis results corresponding to the fields, the first data analysis results being used to describe the characteristics of the field values corresponding to the fields; performing data analysis on the simulation field values corresponding to the simulation fields in the simulation data set of the to-be-processed data set to obtain second data analysis results corresponding to the simulation fields, the second data analysis results being used to describe the characteristics of the simulation field values corresponding to the simulation fields; and determining the similarity between the to-be-processed data set and the simulation data set according to the first data analysis results of the fields in the to-be-processed data set and the second data analysis results of the simulation fields in the simulation data set.

[0066] Figure 3 is a flowchart of a simulation data generation method according to an embodiment of the present disclosure, which is combined with Figure 3 The present embodiment is exemplarily described. First, the homologous fields in the to-be-processed data set are determined by table lookup, so that the same data processing rule is used to process the field values of the homologous fields to obtain simulation field values, and the field values of the non-homologous fields are processed by using the related embodiments described above, and for the fields, the fields are rewritten by using the above-mentioned manner to obtain simulation fields, thereby providing a data basis for generating a simulation data set. It can be understood that the simulation fields in the simulation data set are one-to-one corresponding to the fields in the to-be-processed data set, and the simulation field values corresponding to the simulation fields in the simulation data set are the simulation field values of the corresponding fields in the to-be-processed data set. On this basis, the to-be-processed data set and the simulation data set are respectively analyzed to obtain first data analysis results and second data analysis results, so as to calculate the similarity between the to-be-processed data set and the simulation data set.

[0067] It should be noted that in the present embodiment, the simulation field values of the numerical type are analyzed.

[0068] In some embodiments, the data analysis can be performed from at least one of the statistical dimension, the correlation dimension, the trend dimension and the abnormality dimension to obtain the data analysis results, thereby improving the reliability of the similarity calculated subsequently depending on the data analysis results.

[0069] The data analysis result corresponding to the statistical dimension is used to describe a statistical quantity in the field value or the simulated field value, such as at least one of a mean value, a median value, and a variance; the data analysis result corresponding to the correlation dimension is used to describe a correlation between different field values or between different simulated field values, and the correlation can be represented by a Pearson correlation coefficient; the data analysis result corresponding to the trend dimension is used to describe a trend of the data in the field value or the simulated field value, such as at least one of an upward trend, a downward trend, and a periodic fluctuation frequency; and the data analysis result corresponding to the abnormality dimension is used to describe a number of outliers in the field value or the simulated field value. It can be understood that the data analysis result herein includes the first data analysis result or the second data analysis result.

[0070] The similarity is used to measure whether the simulated data in the simulated data set fits the real business scenario, and the higher the similarity is, the higher the fitting degree of the data in the simulated data set to the real business scenario, that is, the higher the quality of the simulated data set.

[0071] In some embodiments, since the data analysis can be performed from multiple dimensions, the corresponding data analysis result includes analysis results in different dimensions. It can be understood that the data analysis result herein includes the first data analysis result or the second data analysis result.

[0072] Further, the step of determining the similarity between the to-be-processed data set and the simulated data set according to the first data analysis result of the field in the to-be-processed data set and the second data analysis result of the corresponding simulated field in the simulated data set can be implemented in the following manner: based on the analysis results in each dimension, determining the analysis result difference between the field in the to-be-processed data set and the corresponding simulated field in the simulated data set in the corresponding dimension; according to the analysis result differences in all dimensions and the first preset weight corresponding to each dimension, determining a fusion difference of the simulated field; and based on the fusion difference of the simulated field, determining the similarity between the to-be-processed data set and the simulated data set.

[0073] For example, the first preset weights of the statistical dimension, the correlation dimension, the trend dimension, and the abnormality dimension are 30%, 20%, 20%, 30%, and 20%, respectively.

[0074] For example, the analysis result difference in each dimension can be determined by using the following formula (1):

[0075] (1);

[0076] In the above formula (1), is the analysis result difference of the dimension, the analysis result of the simulation field in the simulation data set in the dimension, the analysis result of the field in the data set to be processed in the dimension, is a minimum value, avoiding the denominator of the above formula (1) being zero, for example, taking the analysis result of the statistical dimension as the mean value, is the mean value of the corresponding simulation field value in the simulation field in the simulation data set, is the mean value of the corresponding field value of the corresponding field in the data set to be processed. It should be noted that when calculating the analysis result difference using the above formula (1), for the above rising or falling, the value 1 can be used to represent the rising, and the value 0 can be used to represent the falling, to adapt to the calculation of the above formula (1).

[0077] Further, in order to improve the reliability of the similarity, a plurality of different indicators can be configured under one dimension, for example, the mean value, the median value and the variance can be configured under the above statistical dimension, in which case the first preset weight under the corresponding dimension is split to obtain the second preset weight of each indicator, so as to calculate the analysis result difference of the dimension under a plurality of indicators.

[0078] It can be understood that the sum of the second preset weights is equal to the corresponding first preset weight. Taking the above configured example of the first preset weight, for the statistical dimension, three indicators of the mean value, the median value and the variance can be set, and the second preset weight of each indicator is 10% respectively; for example, for the trend dimension, two indicators of the trend (rising or falling) and the number of periodic fluctuations can be set, and the second preset weight of each indicator is 15% respectively.

[0079] The above formula (1) can be used to determine the analysis result difference between the field and the simulation field under a certain dimension and a certain indicator, on the basis of which the can be regarded as the analysis result difference under the indicator, the is the mean value of the corresponding simulation field value in the simulation field in the simulation data set, the is the mean value of the corresponding field value of the corresponding field in the data set to be processed.

[0080] Further, the analysis result difference of the dimension is determined by using the following formula (2):

[0081] (2);

[0082] In the above formula (2), is the analysis result difference of the dimension, L represents the number of indicators under the dimension, represents the th indicator under the dimension, represents the Differences in analysis results under each indicator Indicates the first The second preset weight corresponding to each indicator, where... The calculation method can be referred to formula (1).

[0083] Furthermore, the fusion differences of the simulated fields are determined using the following formula (3):

[0084] (3);

[0085] in, To simulate the differences in field fusion, For the differences in the analysis results of the o-th dimension, Z represents the first preset weight corresponding to the o-th dimension, and Z represents the number of dimensions.

[0086] It is worth noting that the above dataset (simulated dataset or dataset to be processed) includes at least one field (simulated field). For example, a column in the dataset can correspond to a field (simulated field), and the data in each row corresponding to that column is the field value (simulated field value).

[0087] Therefore, to further improve the reliability of similarity, the step of determining the similarity between the dataset to be processed and the simulated dataset based on the fusion difference of simulated fields can be implemented as follows: determine the similarity between the dataset to be processed and the simulated dataset based on the fusion difference of all simulated fields.

[0088] As an example, the similarity between the dataset to be processed and the simulated dataset can be determined using the following formula (4):

[0089] (4);

[0090] In the above formula (4), For similarity, The number of columns in the dataset to be processed, i.e., the number of simulated fields. This represents the fusion difference corresponding to the simulated field in the i-th column.

[0091] In some embodiments, the above-described simulated data generation method may further include the following steps: when the similarity is less than a preset similarity threshold, locating the target simulated field in the simulated dataset, wherein the difference between the second data analysis result corresponding to the target simulated field and the first data analysis result corresponding to the target simulated field is greater than a preset difference; updating the simulated field values ​​corresponding to the target simulated field and the simulated field belonging to the same source field as the target simulated field in the simulated dataset to obtain an updated simulated dataset; and determining the similarity between the updated simulated dataset and the dataset to be processed.

[0092] The similarity between the simulation data set and the to-be-processed data set greater than or equal to a preset similarity threshold can represent that the simulation data set is close to the business data in the business scenario, and the simulation data set can be used for evaluation of the large model. Otherwise, the similarity between the simulation data set and the to-be-processed data set less than the preset similarity threshold can represent that the simulation data set is greatly different from the business data in the business scenario, and the simulation data set cannot be used for evaluation of the large model. In the embodiment, the preset similarity threshold can be set according to actual conditions, which will not be repeated here.

[0093] In the embodiment, the updated simulation field value can be achieved by changing the data processing rule, for example, by reducing the preset variation range to update the simulation field value.

[0094] In the embodiment, the determination manner of the first data analysis result or the second data analysis result and the determination manner of the similarity between the updated simulation data set and the to-be-processed data set can refer to the related embodiments described above, which will not be limited here.

[0095] In the above manner, by positioning the simulation field with the largest difference, the similarity between the simulation data set and the to-be-processed data set can be quickly improved.

[0096] In some embodiments, in the case that the similarity between the updated simulation data set and the to-be-processed data set is still less than the preset similarity threshold, the step of updating the simulation field value is continuously performed until the similarity between the updated simulation data set and the to-be-processed data set is greater than the preset similarity threshold.

[0097] Based on the same concept, the embodiment of the disclosure provides a simulation data generation device for evaluating a large model, Figure 4 is a block diagram of a simulation data generation device for evaluating a large model according to an embodiment of the disclosure, which will be described with reference to Figure 4 The simulation data generation device 400 for evaluating a large model includes:

[0098] The acquisition module 401 is configured to acquire a to-be-processed data set, wherein the to-be-processed data set includes business data collected in a business scenario.

[0099] The first determination module 402 is configured to determine a homologous field in the to-be-processed data set, wherein the homologous field has a same upstream field.

[0100] The data processing module 403 is configured to process field values of the homologous field by using a same data processing rule to obtain simulation field values of the homologous field, wherein the simulation field values are used to support evaluation of the large model.

[0101] Optionally, the first determining module 402 is further configured to determine the homologous fields in the to-be-processed data set according to the obtained field relationship table, where the field relationship table is used to maintain the fields in a downstream data set and the corresponding fields in an upstream data set.

[0102] Optionally, the simulation data generation apparatus 400 for evaluating a large model further includes:

[0103] a first analysis module configured to perform data analysis on the field values corresponding to the fields in the to-be-processed data set to obtain first data analysis results corresponding to the fields, where the first data analysis results are used to describe the characteristics of the field values corresponding to the fields;

[0104] a second analysis module configured to perform data analysis on the simulation field values corresponding to the simulation fields in the simulation data set to obtain second data analysis results corresponding to the simulation fields, where the second data analysis results are used to describe the characteristics of the simulation field values corresponding to the simulation fields, the simulation fields in the simulation data set are in one-to-one correspondence with the fields in the to-be-processed data set, and the simulation field values corresponding to the simulation fields in the simulation data set are simulation field values of the corresponding fields in the to-be-processed data set;

[0105] a second determining module configured to determine the similarity between the to-be-processed data set and the simulation data set according to the first data analysis results of the fields in the to-be-processed data set and the second data analysis results of the simulation fields in the simulation data set.

[0106] Optionally, the simulation data generation apparatus 400 for evaluating a large model further includes:

[0107] a positioning module configured to, in a case where the similarity is less than a preset similarity threshold, position a target simulation field in the simulation data set, where the difference between the second data analysis result corresponding to the target simulation field and the first data analysis result is greater than a preset difference;

[0108] an updating module configured to update the simulation field values corresponding to the target simulation field and the simulation fields belonging to the homologous fields as the target simulation field in the simulation data set to obtain an updated simulation data set;

[0109] a third determining module configured to determine the similarity between the updated simulation data set and the to-be-processed data set.

[0110] Optionally, the data analysis is performed from multiple dimensions, the data analysis results include analysis results in different dimensions, and the second determining module includes:

[0111] a first determining sub-module, configured to determine, based on the analysis result under each dimension, a difference in analysis result between the field in the to-be-processed data set and the corresponding simulation field in the simulation data set under the corresponding dimension;

[0112] a second determining sub-module, configured to determine, according to the difference in analysis result under all the dimensions and the first preset weight corresponding to each dimension, a fusion difference of the simulation field;

[0113] a third determining sub-module, configured to determine, based on the fusion difference of the simulation field, a similarity between the to-be-processed data set and the simulation data set.

[0114] Optionally, the third determining sub-module is further configured to determine, according to the fusion difference of all the simulation fields, the similarity between the to-be-processed data set and the simulation data set.

[0115] The embodiments of the modules in the simulation data generation apparatus 400 for evaluating a large model can refer to the related embodiments of the method, and will not be repeated here.

[0116] The embodiments of the simulation data generation method for evaluating a large model can be implemented by a computer program, and the computer program is stored in a computer readable medium. The computer readable medium can be a volatile memory or a non-volatile memory, or a removable memory or a non-removable memory.

[0117] The embodiments of the simulation data generation method for evaluating a large model can be implemented by a computer program, and the computer program is stored in a computer readable medium. The computer readable medium can be a volatile memory or a non-volatile memory, or a removable memory or a non-removable memory.

[0118] The embodiments of the simulation data generation method for evaluating a large model can be implemented by a computer program, and the computer program is stored in a computer readable medium. The computer readable medium can be a volatile memory or a non-volatile memory, or a removable memory or a non-removable memory.

[0119] a storage device, in which a computer program is stored;

[0120] a processing device, configured to execute the computer program in the storage device to implement the steps of the simulation data generation method for evaluating a large model.

[0121] Reference will be made to the following description of the embodiments of the present disclosure, taken in conjunction with the accompanying drawings, in which Figure 5 which shows a structural schematic diagram of an electronic device 500 suitable for implementing the embodiments of the present disclosure. The terminal device in the embodiments of the present disclosure can include, but is not limited to, mobile terminals such as mobile phones, notebook computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Personal Computers), PMPs (Portable Multimedia Players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), and the like, as well as fixed terminals such as digital TVs, desktop computers, and the like. Figure 5The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0122] like Figure 5 As shown, electronic device 500 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 501, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 502 or a program loaded from storage device 508 into random access memory (RAM) 503. RAM 503 also stores various programs and data required for the operation of electronic device 500. Processing unit 501, ROM 502, and RAM 503 are interconnected via bus 504. Input / output (I / O) interface 505 is also connected to bus 504.

[0123] Typically, the following devices can be connected to I / O interface 505: input devices 506 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 507 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 508 including, for example, magnetic tapes, hard disks, etc.; and communication devices 509. Communication device 509 allows electronic device 500 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 5 An electronic device 500 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0124] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 509, or installed from a storage device 508, or installed from a ROM 502. When the computer program is executed by the processing device 501, it performs the functions defined in the methods of embodiments of this disclosure.

[0125] It is noted that the aforementioned computer-readable medium of the present disclosure can be a computer-readable signal medium or a computer-readable storage medium or any combination thereof. The computer-readable storage medium can be, for example and without limitation, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the computer-readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program used by or in connection with an instruction execution system, apparatus, or device. In the present disclosure, the computer-readable signal medium can include a computer-readable program code transmitted by a computer-readable storage medium or carried by a carrier wave in a baseband or as part of a carrier wave. Such a propagated computer-readable signal medium can take various forms, including but not limited to electro-magnetic, optical, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium that can be used to carry or store a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained in the computer-readable medium can be transmitted by any suitable medium, including but not limited to wire, cable, RF (radio frequency), or the like, or any suitable combination of the foregoing.

[0126] In some embodiments, the electronic device can communicate using any currently known or future developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communications (e.g., a communications network) of any form or medium (e.g., a communications network). Examples of communications networks include local area networks ("LAN"), wide area networks ("WAN"), internetworks (e.g., the Internet), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future developed networks.

[0127] The aforementioned computer-readable medium can be included in the aforementioned electronic device; or can exist separately from the electronic device without being incorporated into the electronic device.

[0128] The computer readable medium described above carries one or more programs, when the one or more programs are executed by the electronic device, cause the electronic device to: acquire a to-be-processed data set, wherein the to-be-processed data set includes service data collected in a service scenario; determine a homologous field in the to-be-processed data set, wherein the homologous field has a same upstream field; and process field values of the homologous field by using a same data processing rule to obtain a simulated field value of the homologous field, wherein the simulated field value is used to support evaluation of the large model.

[0129] Computer program code for carrying out operations of the present disclosure can be written in any one or combination of one or more programming languages, including an object-oriented programming language such as Java, Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).

[0130] The flow diagrams and the block diagrams in the drawings are illustrations of architectures, functionalities, and operations of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flow diagrams or block diagrams can represent a module, a segment, or a portion of code, which comprises one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flow diagrams, and combinations of blocks in the block diagrams and / or flow diagrams, can be implemented by special purpose hardware-based systems that perform the specified functions or operations, or combinations of special purpose hardware and computer instructions.

[0131] The modules described in the embodiments of the present disclosure can be implemented by software or by hardware. In some cases, the name of a module does not constitute a limitation on the module itself. For example, an acquisition module can also be described as an "acquisition to-be-processed data set module".

[0132] The functionality described above in this document can be performed, at least in part, by one or more hardware logic components. For example, and without limitation, non- transitory machine-readable media can include RAM, ROM, programmable ROM (EPROM, EEPROM or flash memory), or any other storage technology, including tangible and / or optical devices. Non- transitory machine-readable media can also include a transmission that can carry the program, and / or a carrier wave comprising the program, depending upon the particular requirements of the operating environment. Any of these media can be included within a single memory device or multiple memory devices.

[0133] In the context of the present disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include one or more lines of a computer program code, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0134] The above description is only preferred embodiments of the present disclosure and the explanation of the technical principles of the application. It should be understood by those skilled in the art that the disclosure range involved in the present disclosure is not limited to the technical solutions formed by the specific combinations of the above technical features, and also covers other technical solutions formed by any combinations of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by replacing the above features with the technical features disclosed in the present disclosure (but not limited to) having similar functions.

[0135] Further, while operations are depicted in a particular order, this should not be understood as requiring the operations to be performed in the particular order shown or in sequential order, and that certain operations can be performed in parallel or in any suitably ordered order. Also, while a number of specific implementation details have been discussed, these should not be construed as limiting the scope of the disclosure. Certain features that are described in the context of separate embodiments can also be implemented in combination. Conversely, various features that are described in the context of a single embodiment can also be implemented or practiced separately or in any suitable subcombination.

[0136] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims. With respect to the devices in the above-described embodiments, in which various modules perform operations, the specific manner in which the various modules perform the operations has been described in detail in the embodiments relating to the method. Here, no detailed explanation will be given.

Claims

1. A method for generating simulation data for evaluating large models, characterized in that, The method comprises: obtaining a to-be-processed data set, wherein the to-be-processed data set comprises business data collected in a business scenario, the to-be-processed data set is derived based on an upstream data set, the to-be-processed data set is referred to as a downstream data set of the upstream data set, and a field in the upstream data set can be referred to as an upstream field of a corresponding field in the downstream data set; determining homologous fields in the to-be-processed data set, wherein the homologous fields have the same upstream fields; applying the same data processing rule to field values of the homologous fields, and rewriting the field values of the homologous fields to obtain simulated field values of the homologous fields, wherein the simulated field values are used to support evaluation of the large model; The method further comprises: performing data analysis on field values corresponding to fields in the to-be-processed data set to obtain first data analysis results corresponding to the fields, wherein the first data analysis results are used to describe characteristics of the field values corresponding to the fields; performing data analysis on simulated field values corresponding to simulated fields in a simulated data set of the to-be-processed data set to obtain second data analysis results corresponding to the simulated fields, wherein the second data analysis results are used to describe characteristics of the simulated field values corresponding to the simulated fields, the simulated fields in the simulated data set correspond one-to-one to the fields in the to-be-processed data set, the simulated field values corresponding to the simulated fields in the simulated data set are simulated field values of corresponding fields in the to-be-processed data set, and the data analysis is performed from at least one of a statistical dimension, a correlation dimension, a trend dimension, and an abnormality dimension to obtain data analysis results; determining a similarity between the to-be-processed data set and the simulated data set according to the first data analysis results of the fields in the to-be-processed data set and the second data analysis results of the simulated fields in the simulated data set; in a case where the similarity is less than a preset similarity threshold, positioning a target simulated field in the simulated data set, wherein a difference between the second data analysis result corresponding to the target simulated field and the first data analysis result is greater than a preset difference; updating simulated field values corresponding to the target simulated field and simulated fields belonging to the homologous fields in the simulated data set to obtain an updated simulated data set; determining a similarity between the updated simulated data set and the to-be-processed data set.

2. The analog data generation method of claim 1, wherein, The method of determining the homologous fields in the to-be-processed data set comprises: determining the homologous fields in the to-be-processed data set according to a field relationship table, wherein the field relationship table is used to maintain fields in a downstream data set in corresponding fields in an upstream data set.

3. The analog data generation method of claim 1, wherein, The data analysis is performed from multiple dimensions, and the data analysis result includes analysis results in different dimensions. The similarity between the to-be-processed data set and the simulation data set is determined according to the first data analysis result of the field in the to-be-processed data set and the second data analysis result of the corresponding simulation field in the simulation data set, and includes: Based on the analysis result in each dimension, the analysis result difference between the field in the to-be-processed data set and the corresponding simulation field in the simulation data set corresponding to the dimension is determined. According to the analysis result difference in all dimensions and the first preset weight corresponding to each dimension, the fusion difference of the simulation field is determined. Based on the fusion difference of the simulation field, the similarity between the to-be-processed data set and the simulation data set is determined.

4. The analog data generation method of claim 3, wherein, The similarity between the to-be-processed data set and the simulation data set is determined based on the fusion difference of the simulation field, and includes: According to the fusion difference of all simulation fields, the similarity between the to-be-processed data set and the simulation data set is determined.

5. A simulation data generation device for evaluating a large model, characterized by, Including: An acquisition module is configured to acquire a to-be-processed data set, wherein the to-be-processed data set includes business data collected in a business scenario, and the to-be-processed data set is derived based on an upstream data set, the to-be-processed data set is referred to as a downstream data set of the upstream data set, and a field in the upstream data set can be referred to as an upstream field of a corresponding field in the downstream data set; A first determination module is configured to determine a homologous field in the to-be-processed data set, wherein the homologous field has the same upstream field; A data processing module is configured to use the same data processing rule to process field values of the homologous field, and rewrite the field values of the homologous field to obtain simulation field values of the homologous field, wherein the simulation field values are used to support evaluation of the large model; The simulation data generation device for evaluating a large model further includes: A first analysis module is configured to perform data analysis on field values corresponding to a field in the to-be-processed data set to obtain a first data analysis result corresponding to the field values, wherein the first data analysis result is used to describe characteristics of the field values corresponding to the field; A second analysis module is configured to perform data analysis on simulation field values corresponding to a simulation field in a simulation data set of the to-be-processed data set to obtain a second data analysis result corresponding to the simulation field values, wherein the second data analysis result is used to describe characteristics of the simulation field values corresponding to the simulation field, the simulation field in the simulation data set is one-to-one corresponding to the field in the to-be-processed data set, the simulation field values corresponding to the simulation field in the simulation data set are simulation field values of the corresponding field in the to-be-processed data set, and the data analysis is performed from at least one of a statistical dimension, a correlation dimension, a trend dimension, and an abnormality dimension to obtain a data analysis result. a second determining module, configured to determine a similarity between the to-be-processed data set and the simulation data set according to the first data analysis result of the field in the to-be-processed data set and the second data analysis result of the simulation field corresponding to the simulation field in the simulation data set; a positioning module, configured to position a target simulation field in the simulation data set when the similarity is less than a preset similarity threshold, wherein a difference between the second data analysis result corresponding to the target simulation field and the first data analysis result is greater than a preset difference; an updating module, configured to update a simulation field value corresponding to the target simulation field and a simulation field belonging to the same homologous field as the target simulation field in the simulation data set to obtain an updated simulation data set; a third determining module, configured to determine a similarity between the updated simulation data set and the to-be-processed data set.

6. A computer readable medium having stored thereon a computer program, characterized in that, The computer program is executed by a processing device to implement the steps of the simulation data generation method in any one of claims 1-4.

7. An electronic device, comprising: comprise: a storage device having a computer program stored thereon; a processing device configured to execute the computer program in the storage device to implement the steps of the simulation data generation method in any one of claims 1-4.

8. A computer program product comprising a computer program, characterized in that, The computer program is executed by a processing device to implement the steps of the simulation data generation method in any one of claims 1-4.

Citation Information

Patent Citations

  • Test method, system and equipment of report system and storage medium

    CN114265780A