Data generation method and device and electronic equipment
Through iterative denoising processing technology, the target timing data is generated using Gaussian noise and time-dependent noise prediction information, which solves the problems of scarcity of time series data and the influence of external factors, and improves the accuracy and reliability of the model.
Patent Information
- Application Number
- CN202510377353.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-06-13
AI Technical Summary
Time series data is scarce in supply chain scenarios and is affected by external factors, resulting in insufficient accuracy and reliability of the model in industrial applications.
By iterative denoising processing based on Gaussian noise and time-dependent noise prediction information, target timing data is generated to characterize the characteristic values in the supply chain scenario.
It improves the accuracy and reliability of data generation, enhances the generalization ability of models in industrial applications, and improves the performance of time series modeling.
Smart Images

Figure CN120144937A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence and data processing, and more particularly, to a data generation method, apparatus, and electronic device. Background Art
[0002] Time series data is continuous data arranged in chronological order that reflects the changes of a certain phenomenon or system state. It is widely used in fields such as finance, meteorology, economy, and manufacturing to describe the characteristics of data changing over time. In the field of time series modeling, insufficient data volume is a major obstacle to building an effective basic model. Compared with large language models, the existing data sources of time series are scarce and highly sensitive, resulting in limited available data volume, thus affecting the performance of the model. In addition, time series data in the industrial field, such as in industries like manufacturing, energy, and finance, are often affected by other external factors such as market fluctuations, and some industrial scenario sequences with important value are scarce, making the models trained based on these data lack the generalization ability for industrial scenarios, and the accuracy and reliability of time series modeling in industrial applications are insufficient. Summary of the Invention
[0003] In view of this, the present disclosure provides a data generation method, apparatus, and electronic device.
[0004] One aspect of the present disclosure provides a data generation method, including: performing iterative denoising processing on the time series noise data based on Gaussian noise prediction information and time-dependent noise prediction information to generate target time series data; wherein the target time series data represents eigenvalue distributions in chronological order in a supply chain scenario, and the eigenvalues are used to represent the object attributes in the supply chain scenario; wherein performing one denoising process includes: obtaining noise prediction information, the noise prediction information including Gaussian noise prediction information and time-dependent noise prediction information; restoring the detailed features of the intermediate time series data based on the Gaussian noise prediction information, and restoring the overall features of the intermediate time series data based on the time-dependent noise prediction information; the intermediate time series data is obtained by performing at least one denoising process on the time series noise data; wherein, as the number of denoising processes increases, the proportion of Gaussian noise prediction information in the noise prediction information decreases, and the proportion of time-dependent noise prediction information increases.
[0005] According to an embodiment of the present disclosure, the noise prediction information obtained in the subsequent denoising process is generated based on the intermediate time series data obtained in the previous denoising process.
[0006] According to an embodiment of the present disclosure, the data generation method further includes obtaining condition information, where the condition information at least characterizes the self - characteristics of the target time - series data; obtaining noise prediction information includes: obtaining noise prediction information according to the condition information, so that when denoising the intermediate time - series data according to the noise prediction information, the process of denoising is constrained by the condition information.
[0007] According to an embodiment of the present disclosure, obtaining noise prediction information includes: predicting the intermediate time - series data to obtain reference time - series data, where the reference time - series data at least characterizes the data distribution prediction information of the denoised time - series data; obtaining noise prediction information according to the reference time - series data, so that when denoising the intermediate time - series data according to the noise prediction information, the process of denoising is constrained by the reference time - series data.
[0008] According to an embodiment of the present disclosure, the data generation method further includes: obtaining a target matrix, where the target matrix at least characterizes the target change trend of the data within the target time range; according to the target matrix, adjusting the change trend of the data within the target time range in the target time - series data to be closer to the target change trend.
[0009] According to an embodiment of the present disclosure, the noise prediction information is generated by a first model, and the training process of the first model includes: randomly selecting a time step from a set of sequentially arranged time steps; obtaining sample noise corresponding to the time step, where the sample noise includes sample Gaussian noise and sample time - dependent noise, and as the time step increases, the proportion of sample Gaussian noise in the sample noise increases, and the proportion of sample time - dependent noise decreases; adding noise to the first sample time - series data based on the sample noise to obtain the noise - added sample time - series data; performing noise prediction on the noise - added sample time - series data according to the first model to obtain sample noise prediction information; calculating a target loss according to the sample noise prediction information and the sample noise; adjusting the parameters of the first model according to the target loss; repeating the above process until the number of repetitions reaches a threshold or the first model converges.
[0010] According to an embodiment of the present disclosure, performing noise prediction on the noise - added sample time - series data according to the first model to obtain sample noise prediction information includes: obtaining sample condition information corresponding to the first sample time - series data, where the sample condition information at least characterizes the self - characteristics of the first sample time - series data; the first model performs noise prediction on the noise - added sample time - series data according to the sample condition information to generate sample noise prediction information.
[0011] According to an embodiment of the present disclosure, the time - dependent noise is generated by the following operations: constructing a plurality of generation methods, each generation method corresponding to at least one time - series feature; generating at least one time - dependent noise based on the plurality of generation methods, where the time - dependent noise at least partially has the time - series features corresponding to the respective plurality of generation methods.
[0012] Another aspect of the present disclosure provides a data generation device, including: a denoising module configured to perform iterative denoising processing on time-series noise data based on Gaussian noise prediction information and time-dependent noise prediction information to generate target time-series data; wherein the target time-series data represents eigenvalue distributed in chronological order in a supply chain scenario, and the eigenvalue is used to represent object attributes in the supply chain scenario; wherein the denoising module is specifically configured to: obtain noise prediction information, where the noise prediction information includes Gaussian noise prediction information and time-dependent noise prediction information; restore the detail features of intermediate time-series data based on the Gaussian noise prediction information, and restore the overall features of the intermediate time-series data based on the time-dependent noise prediction information; the intermediate time-series data is obtained by performing at least one denoising process on the time-series noise data; wherein, as the number of denoising processes increases, the proportion of Gaussian noise prediction information in the noise prediction information decreases, and the proportion of time-dependent noise prediction information increases.
[0013] Another aspect of the present disclosure provides an electronic device, including: at least one processor; and a memory connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the following operations: perform iterative denoising processing on time-series noise data based on Gaussian noise prediction information and time-dependent noise prediction information to generate target time-series data; wherein the target time-series data represents eigenvalue distributed in chronological order in a supply chain scenario, and the eigenvalue is used to represent object attributes in the supply chain scenario; wherein performing one denoising process includes: obtaining noise prediction information, where the noise prediction information includes Gaussian noise prediction information and time-dependent noise prediction information; restoring the detail features of intermediate time-series data based on the Gaussian noise prediction information, and restoring the overall features of the intermediate time-series data based on the time-dependent noise prediction information; the intermediate time-series data is obtained by performing at least one denoising process on the time-series noise data; wherein, as the number of denoising processes increases, the proportion of Gaussian noise prediction information in the noise prediction information decreases, and the proportion of time-dependent noise prediction information increases.
[0014] Another aspect of the present disclosure provides a computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the data generation method according to any one of the foregoing embodiments.
[0015] Another aspect of the present disclosure provides a computer program product, including computer programs / instructions, characterized in that when the computer programs / instructions are executed by a processor, the operations of the data generation method according to any one of the foregoing embodiments are implemented. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Through the following description of the embodiments of the present disclosure with reference to the accompanying drawings, the above and other objects, features, and advantages of the present disclosure will become clearer. In the drawings:
[0017] Figure 1 Schematically shows a flowchart of a data generation method according to an embodiment of the present disclosure;
[0018] Figure 2 Schematically shows a flowchart of obtaining noise prediction information in the data generation method according to an embodiment of the present disclosure;
[0019] Figure 3 Schematically shows another flowchart of obtaining noise prediction information in the data generation method according to an embodiment of the present disclosure;
[0020] Figure 4 Schematically shows another flowchart of the data generation method according to an embodiment of the present disclosure;
[0021] Figure 5 Schematically shows a comparison diagram after processing target time-series data by a target matrix in the data generation method according to an embodiment of the present disclosure;
[0022] Figure 6 Schematically shows a flowchart of training a first model in the data generation method according to an embodiment of the present disclosure;
[0023] Figure 7 Schematically shows a flowchart of obtaining sample noise prediction information in the data generation method according to an embodiment of the present disclosure;
[0024] Figure 8 Schematically shows a flowchart of generating time-dependent noise in the data generation method according to an embodiment of the present disclosure;
[0025] Figure 9 Schematically shows a structural block diagram of a data processing method according to an embodiment of the present disclosure;
[0026] Figure 10 Schematically shows a block diagram of a data generation device according to an embodiment of the present disclosure; and
[0027] Figure 11 Schematically shows a block diagram of an electronic device suitable for implementing the method described above according to an embodiment of the present disclosure. Detailed implementation manners
[0028] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present disclosure. In the following detailed description, for the purpose of explanation, numerous specific details are set forth in order to provide a thorough understanding of the embodiments of the present disclosure. However, it is obvious that one or more embodiments can also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily obscuring the concepts of the present disclosure.
[0029] The terms used herein are merely for describing specific embodiments and are not intended to limit the present disclosure. The terms "including", "comprising", etc. used herein indicate the presence of the described features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0030] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0031] In cases where expressions similar to "at least one of A, B, and C, etc." are used, generally, it should be interpreted according to the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include, but not be limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.).
[0032] In the embodiments of the present disclosure, in terms of the collection, update, analysis, processing, use, transmission, provision, disclosure, storage, etc. of the data involved (for example, including but not limited to user personal information), they all comply with the provisions of relevant laws and regulations, are used for legal purposes, and do not violate public order and good customs. In particular, necessary measures are taken for user personal information to prevent illegal access to user personal information data and to safeguard the security of user personal information, network security, and national security.
[0033] Embodiments of the present disclosure provide a data generation method, including: performing iterative denoising processing on time series noise data based on Gaussian noise prediction information and time-dependent noise prediction information to generate target time series data; wherein the target time series data represents eigenvalue distributed in chronological order in a supply chain scenario, and the eigenvalue is used to represent object attributes in the supply chain scenario; wherein performing one denoising process includes: obtaining noise prediction information, the noise prediction information including Gaussian noise prediction information and time-dependent noise prediction information; restoring the detailed features of the intermediate time series data based on the Gaussian noise prediction information, and restoring the overall features of the intermediate time series data based on the time-dependent noise prediction information; the intermediate time series data is obtained by performing at least one denoising process on the time series noise data; wherein, as the number of denoising processes increases, the proportion of the Gaussian noise prediction information in the noise prediction information decreases, and the proportion of the time-dependent noise prediction information increases.
[0034] Figure 1 Schematically shows a flowchart of the data generation method according to an embodiment of the present disclosure.
[0035] As Figure 1 shown, the data generation method may at least include operation S110.
[0036] In operation S110, iterative denoising processing is performed on the time series noise data based on the Gaussian noise prediction information and the time-dependent noise prediction information to generate target time series data; wherein the target time series data represents eigenvalue distributed in chronological order in a supply chain scenario, and the eigenvalue is used to represent object attributes in the supply chain scenario.
[0037] Time series data is a set of data points arranged in chronological order, while time series noise data is time series data that at least contains noise, which may be obtained by adding noise to specific time series data, or may be pure noise data generated randomly or in a specific manner according to the format of time series data.
[0038] Gaussian noise prediction information is the prediction of Gaussian noise in time series noise data. Gaussian noise is a typical random noise, and its distribution is usually assumed to be a normal distribution.
[0039] Time dependence refers to a basic property, that is, the state of a certain variable or data is not completely independent in the time dimension, but has some association with previous or subsequent time points. At the simplest level, it can be understood that there is some regularity or correlation between data points over time, rather than a random independent distribution. For example, in a supply chain scenario, the inventory level of a certain commodity may change due to the sales situation of the previous day, and this front-back association is an embodiment of time dependence.
[0040] Time dependence is not limited to the sequential relationship of time points. It can also be quantified by the evolution of statistical characteristics (such as mean, variance, or covariance) over time. In time series analysis, this characteristic often manifests as trends (long-term increases or decreases), seasonality (periodic repetitions), or local smooth changes (short-term fluctuations).
[0041] Time-dependent noise can be regarded as a special type of noise for time series data, with its core lying in its time-varying dependence. Time-dependent noise is different from traditional Gaussian noise, which assumes that the noise at each time point is independently and identically distributed, while time-dependent noise has time-correlated statistical characteristics. For example, its covariance may decrease as the time interval increases, reflecting a structure similar to seasonality or trends. The time-dependent noise prediction information is a prediction of the time-dependent noise present in the time series noise data.
[0042] The target time series data can refer to a sequence of characteristic values related to a specific object arranged in chronological order in a supply chain scenario. The characteristic values in the sequence characterize the attributes of some important objects in the supply chain and change over time. For example, the object can be the spare parts inventory level, product sales volume, production cycle, logistics delivery time, etc.
[0043] Specifically, for example, if the object is the spare parts inventory level, then the characteristic value is the inventory quantity of spare parts at each time point. The target time series data is a sequence describing the change of the spare parts inventory level over time. For example, as demand fluctuates and replenishment situations change, the inventory level of spare parts may increase or decrease, and the target time series data can help analyze the change trend of the inventory and the replenishment timing.
[0044] Specifically, for example, if the object is the product sales volume, then the characteristic value is the product sales quantity at each time point. The target time series data records the change of the product sales volume over time. For example, during holidays or promotional periods, the sales volume may increase sharply, while in the off-season, the sales volume may decrease. This time series data can help predict the future change trend of sales volume and optimize production and inventory management.
[0045] Specifically, for example, if the object is the logistics delivery time, then the characteristic value is the time length of each delivery. The target time series data is a sequence describing the change of the logistics delivery time over time, which can help optimize the transportation route and method and improve the delivery efficiency.
[0046] Performing iterative denoising processing on the time series noise data can take the time series noise data as the starting state of the iterative process. In each round of iteration, the noise components are gradually removed based on the current state, and after multiple rounds of iteration, the target time series data closer to the true time series distribution is gradually restored.
[0047] According to an embodiment of the present disclosure, operation S110 may include operations S111 to S112.
[0048] In operation S111, noise prediction information is obtained, and the noise prediction information includes Gaussian noise prediction information and time-dependent noise prediction information.
[0049] In operation S112, the detailed features of the intermediate time-series data are restored based on the Gaussian noise prediction information, and the overall features of the intermediate time-series data are restored based on the time-dependent noise prediction information; the intermediate time-series data is obtained by performing at least one denoising process on the time-series noise data.
[0050] It should be noted that in the embodiment of the present disclosure, before restoring the detailed features of the intermediate time-series data based on the Gaussian noise prediction information and restoring the overall features of the intermediate time-series data based on the time-dependent noise prediction information, it is necessary to perform a denoising on the initial time-series noise data to obtain the intermediate time-series data. That is, before operation S112, there is also a denoising process on the time-series noise data, and the process includes: obtaining noise prediction information, where the noise prediction information includes Gaussian noise prediction information and time-dependent noise prediction information. Restoring the detailed features of the time-series noise data based on the Gaussian noise prediction information and restoring the overall features of the time-series noise data based on the time-dependent noise prediction information to obtain the intermediate time-series data.
[0051] In the iterative denoising process, by predicting and removing noise in each iteration, the generated result is the intermediate time-series data. These intermediate data are not the final result, but are the phased data generated after one or more processes in the denoising process. They are close to the target data but have not completely removed the noise. As the number of iterations increases, the structure of the original time-series data will be restored more and more accurately, the quality of the intermediate time-series data will gradually improve, and the noise component will gradually decrease, and finally the output data (i.e., the target time-series data) will be obtained.
[0052] According to an embodiment of the present disclosure, as the number of denoising processes increases, the proportion of the Gaussian noise prediction information in the noise prediction information decreases, and the proportion of the time-dependent noise prediction information increases. In the iterative denoising process, the composition of the noise prediction information may include two types of noise components: Gaussian noise prediction information and time-dependent noise prediction information. As the denoising process progresses, the structure of the noise prediction information will change, specifically manifested as the proportion of the Gaussian noise prediction information gradually decreasing, while the proportion of the time-dependent noise prediction information gradually increasing.
[0053] Gaussian noise has the property of independent and identically distributed. Each noise is independent and follows the same distribution. It affects each part of the data in a relatively uniform way. Then, the detailed parts are easily submerged by the noise, while the overall trend is not easily submerged by the noise. Since Gaussian noise is independent, random fluctuations usually occur at each time point, and these fluctuations have a greater impact on the short-term changes and detailed parts in the data. Therefore, it is easy to cause the details of the time series data to be submerged by the noise.
[0054] Time-dependent noise has a dependence relationship in time. The noise at the current time is related to the previous noise value. Therefore, extreme fluctuations do not occur at each time point, but it is more inclined to produce continuous and smooth perturbations. This makes its impact on the high-frequency details in the data weaker, while having a greater impact on the overall interference. Due to its time dependence, time-dependent noise usually shows long-term periodic fluctuations or trend changes, affecting the long-time scale characteristics in the data.
[0055] Therefore, as the number of denoising processes increases, the proportion of Gaussian noise prediction information in the noise prediction information decreases, and the proportion of time-dependent noise prediction information increases. This can make the denoising process gradually focus on removing the long-term trends and periodic fluctuations in the data and retain the overall pattern of the data. Specifically, as Gaussian noise is gradually removed, the short-term random fluctuations in the time series data are smoothed, thus avoiding the interference of these fluctuations on trend analysis and prediction. The increase in the proportion of time-dependent noise makes the denoising process more focused on the adjustment of the long-term structure of the data, ultimately ensuring that the denoised data can more accurately reflect the true trends and periodic characteristics of the data.
[0056] For example, when processing a sales data, a large number of short-term promotional fluctuations (Gaussian noise) will be removed during the initial denoising, making the periodic sales fluctuations and long-term growth trends in the data gradually emerge. As the number of denoising times increases, the focus of denoising begins to shift to removing seasonal fluctuations or periodic fluctuations (time-dependent noise). The finally obtained data is more in line with the actual sales trends and fluctuation rules, which helps with more accurate demand forecasting and inventory management.
[0057] According to the embodiments of the present disclosure, the noise prediction information obtained from the subsequent denoising process is generated based on the intermediate time series data obtained from the previous denoising process. The process of performing iterative denoising can be as follows.
[0058] In the first denoising process, noise prediction is performed based on the initial time series noise data to obtain noise prediction information. The detailed features of the time series noise data are restored based on the Gaussian noise prediction information, and the overall features of the time series noise data are restored based on the time-dependent noise prediction information to generate the first intermediate time series data. In subsequent iterations, each iteration will perform further denoising based on the intermediate time series data generated by the previous denoising process. As the iteration proceeds, each iteration will predict the noise information in the intermediate time series data obtained by the previous denoising to obtain Gaussian noise prediction information and time-dependent noise prediction information. In the early stage of the iteration, the Gaussian noise prediction information in the noise prediction information accounts for a large proportion of the total noise prediction information, and the time-dependent noise accounts for a small proportion of the total noise prediction information. In the early stage of the denoising process, the Gaussian noise component will be removed first because it has a greater impact on the details and is the main source of interference in the early stage of the data. By removing Gaussian noise, the high-frequency random fluctuations in the data are gradually reduced, the details of the data become smoother, and the interference of these noises on the overall trend of the data is reduced.
[0059] As the iteration proceeds, the proportion of Gaussian noise prediction information in the noise prediction information in the total noise prediction information is small, because most of the short-term noise has been removed, and the remaining noise components are more manifested as long-term, time-related fluctuations. At this time, the proportion of time-dependent noise prediction information in the total noise prediction information gradually increases. As the proportion of time-dependent noise increases, the denoising process gradually focuses more on removing long-term noise such as long-term trend fluctuations and periodic fluctuations in the data. Since these noises are caused by time dependencies, they usually appear as continuous and smooth disturbances, and will not produce sudden and violent fluctuations at each time point, but gradually affect the entire data sequence. In the later stage of the denoising process, with the removal of Gaussian noise, the remaining noise is mostly time-dependent noise, and the goal of the denoising process turns to further recovering the long-term trend and periodic characteristics of the data. At this time, in the process of removing time-dependent noise, the overall trend of the data begins to become more obvious, the long-term pattern is restored, and the periodic fluctuations are better eliminated. After multiple rounds of iterative denoising, the high-frequency noise in the time series data almost completely disappears, and important features such as long-term trends and periodic fluctuations in the data gradually become clear. The denoised data can more truly reflect the actual change patterns of the target time series data.
[0060] Figure 2 The flowchart of obtaining noise prediction information in the data generation method according to an embodiment of the present disclosure is schematically shown.
[0061] like Figure 2As shown, based on the foregoing embodiments, the data generation method may include operation S210 of obtaining condition information, where the condition information at least characterizes the self - characteristics of the target time - series data. The condition information may be, for example, the historical values of the data, trend information, periodic characteristics, seasonal variations, etc. These condition information can provide important background and attributes regarding the target time - series data, enabling more precise control of the data generation characteristics during the generation process. For example, specifying to generate a sequence with Seasonalstrength of 0.8, Entropy of 0.5, and Linearity of 0.9 means that the generated time - series data should have strong seasonality (Seasonal strength of 0.8), relatively stable fluctuations (Entropy of 0.5), and an obvious linear trend (Linearity of 0.9). Through these condition information, the generated time - series data can be more consistent with the patterns and laws of the actual target data.
[0062] In operation S111, according to the condition information, noise prediction information is obtained so that when denoising the intermediate time - series data based on the noise prediction information, the process constraint of the condition information on the denoising process is realized. Specifically, the condition information, as a constraint condition, ensures that the long - term trend and periodic characteristics of the data are maintained during the denoising process, while removing the noise components that do not conform to the rules of the target data. During the denoising process, the generation of the noise prediction information is not only based on the noise characteristics of the current data but also combines the condition information to predict the structure of the noise.. The condition information provides an effective constraint on the denoising process, making the denoised data not only able to remove noise but also retain the important structural characteristics related to the target time - series data, avoiding distorting the self - characteristics of the obtained data and ensuring the authenticity and regularity of the generated data.
[0063] For example, assume that the daily temperature data of a certain city is being processed. The condition information may include the historical temperature data of the city (such as the seasonal fluctuation patterns in the past few years) and the current month's climate trend (such as whether it is in summer or winter). During the denoising process, the noise prediction information is not only used to remove the random fluctuations in the data but also ensures that the removed noise does not affect the seasonal change pattern of the temperature or the long - term climate change trend. In this way, the generated denoised data will better retain these seasonal and long - term trend characteristics, making the denoised data more consistent with the real temperature fluctuations.
[0064] Figure 3 Another flowchart showing the obtaining of the noise prediction information in the data generation method according to an embodiment of the present disclosure is schematically shown.
[0065] As Figure 3 shown, based on the foregoing embodiments, S111 may include operations S310 - S320.
[0066] In operation S310, the intermediate time series data is predicted to obtain reference time series data, which at least characterizes the data distribution prediction information of the denoised time series data. The reference time series data is not a fixed template, but is regenerated according to the current intermediate time series data before each iteration, aiming to provide an ideal data distribution prediction for the current denoising step as the reference answer for the denoised data. This reference data not only includes the long-term patterns such as the trends and seasonal characteristics that the denoised data should possess, but also can guide the way of noise removal during the current denoising process to ensure the rationality of the data structure after noise removal.
[0067] In operation S320, according to the reference time series data, noise prediction information is obtained so that when the intermediate time series data is denoised according to the noise prediction information, the process constraint of the reference time series data on the denoising process can be achieved. Specifically, the role of the reference time series data is to help accurately predict the distribution characteristics of the noise in the current denoising step, ensuring that while removing the noise, the long-term trends, periodic characteristics and overall structure of the data can still be retained. As the iteration progresses, each new intermediate time series data will generate new reference time series data before denoising to guide the current denoising process, making the denoised data gradually converge to the ideal state described by the reference time series data.
[0068] During the denoising process of the time series data, especially in the application of non-recursive models, the boundary inharmony problem usually appears in the starting and ending parts of the data. Because the noise information in these regions usually cannot be self-corrected through recursive relationships, it is easy to cause discontinuity and abrupt transitions generated in the boundary regions, affecting the overall quality of the data. Therefore, during the denoising process, the reference time series data plays an important role in alleviating this problem. By providing an ideal target distribution for each iteration, it helps the model maintain a smooth transition when processing the boundary and avoid such inharmonious boundaries.
[0069] According to the embodiments of the present disclosure, predicting the intermediate time series data to obtain the reference time series data may include predicting the intermediate time series data according to the hybrid AR model to obtain the reference time series data. Specifically, the hybrid AR model combines the advantages of the autoregressive (AR) model and other time series modeling methods, and can predict the long-term trends and seasonal characteristics of the data at multiple time steps and generate an ideal distribution of the denoised time series data.
[0070] The hybrid AR model can include an autoregressive (AR) part that predicts current and future data values by using past time-series data points (e.g., data from the past few days, weeks, or months). The AR model is based on the idea of linear regression and predicts the data at the current moment by calculating a linear combination of past moments. During the denoising process, the AR model can help capture long-term trends and seasonal variations in time-series data. For example, in sales data, the AR model can help capture the regular patterns of annual cyclic changes, or in temperature data, capture seasonal fluctuations.
[0071] The hybrid AR model can include a hybrid part that combines a moving average (MA) model or other forms of time-series modeling methods to better handle noise components and random fluctuations in the data. The hybrid AR model can also introduce external factors (such as holiday effects, market fluctuations, etc.) as additional inputs in some cases to better generate reference time-series data.
[0072] Figure 4 Another flowchart of the data generation method according to an embodiment of the present disclosure is schematically shown.
[0073] As Figure 4 shown, based on the foregoing embodiments, the data generation method may include operations S410 to S420.
[0074] In operation S410, a target matrix is obtained. The target matrix at least characterizes the target change trend of the data within the target time range. Specifically, the target matrix is used to quantify the "significance" of the target time-series data, and by designing an Augment matrix M that characterizes the significance of the time-series position, it helps capture the change trend of the time-series data at key positions. The target matrix defines the change trends at different time positions to optimize the generated time-series data and make it more conform to the regularity of the target time series.
[0075] In operation S420, according to the target matrix, the change trend of the data within the target time range in the target time-series data is adjusted to approach the target change trend. Specifically, after iterative denoising processing, the generated time-series data often has certain local deviations, especially at key time points, and may not fully match the target change trend. At this time, the target matrix M comes into play. By adjusting the generated time-series data, the change trend of the data is gradually guided to the target trend, correcting the deviation in the generated result. In this way, the finally generated data not only approximates the real data in the overall structure but also shows the expected change trend at critical moments and key positions.
[0076] According to an embodiment of the present disclosure, the process of obtaining the target matrix may include the following operations.
[0077] Initialize the target matrix; the target matrix needs to be initialized first, and the values in the initial matrix can be set according to certain predefined rules or empirical values.
[0078] Loop and execute the following operations until the stop condition is met:
[0079] Perform weighted processing on the second sample time series data according to the target matrix to obtain the third sample time series data; in each iteration, the target matrix will perform weighted processing on the second sample time series data. The purpose of weighted processing is to adjust the data so that the data can show specific change trends at key time points or important positions, such as generating obvious peaks or valleys at certain moments. Through the transformation of the target matrix, the third sample time series data will reflect the data characteristics after being weighted by the matrix.
[0080] Use the second sample time series data as the input and the third sample time series data as the verification data to train the second model; the third sample time series data after being transformed by the target matrix will be used as the verification data for training a new prediction model. At this time, the second model can identify important trend characteristics in the time series data, such as peaks, valleys, periodic changes, etc., by learning these transformed data samples.
[0081] Obtain the sample input time series data and the sample output time series data. The sample output time series data has a sample change trend within the sample time range relative to the sample input time series data; the sample input time series data and the sample output time series data can be used as a key verification set. For example, in the scenario of generating sales data, the sample output time series data may show characteristics such as sales peaks and seasonal fluctuations, and these trends will be key factors for model verification.
[0082] Based on the second model, generate output information according to the sample input time series data;
[0083] In response to the error between the output information and the sample output time series data being greater than the error threshold, adjust the values in the target matrix; the output information predicted by the model will be compared with the verification data (i.e., the sample output time series data). If the error between the predicted result and the actual data is greater than the predetermined error threshold, it means that the weighted effect of the target matrix fails to correctly adjust the change trend of the data. In this case, some values in the target matrix need to be adjusted to optimize the weighted processing so that the data can better reflect the target change trend (such as peak characteristics).
[0084] The stop condition includes that the error is less than or equal to the error threshold. When the error between the predicted output of the model and the sample output data drops below the predetermined threshold, the training process will stop and the adjustment of the target matrix is also completed. At this time, the matrix has been optimized so that the generated time series data can accurately show specific change trends and key characteristics (such as peaks, valleys, etc.).
[0085] Figure 5 Schematically shows a comparison graph after processing target time series data by a target matrix in the data generation method according to an embodiment of the present disclosure.
[0086] Reference Figure 5 , the data waveform of the target time series data in the time range of t0 to t1 is as Figure 5 shown on the left half, while in the time range of t0 to t1, it should be a trend of rising gently first and then dropping suddenly, and the obtained target time series data does not well reflect this trend. After processing the target time series data by the target matrix, the data of the target time series data in the time range of t0 to t1 is weighted and transformed to change its change trend, and the time series data as shown in Figure 5 the right side is obtained as the final target time series data.
[0087] Figure 6 Schematically shows a flowchart of training a first model in the data generation method according to an embodiment of the present disclosure.
[0088] As Figure 6 shown, on the basis of the foregoing embodiment, the noise prediction information is generated by the first model. The first model can be, for example, a diffusion model, and the diffusion model is a generative model based on a step-by-step denoising process, and performs denoising of time series data by simulating the process of data gradually recovering from a pure noise state to a real data state.
[0089] In each round of denoising iteration, the first model will perform noise prediction according to the time series noise data or the intermediate time series data, and then perform denoising on the time series noise data or the intermediate time series data according to the obtained noise prediction information to restore the overall characteristics and detailed characteristics of the time series data.
[0090] According to an embodiment of the present disclosure, the training process of the first model may include operations S610 to S660.
[0091] In operation S610, a time step is randomly selected from a set of sequentially arranged time steps. The set of time steps can be a set of integers, and the integers in the set are arranged in ascending order of numerical value. For example, it can be [0, T], and the data in the set is an arithmetic sequence with a difference of 1. T is the number of integers.
[0092] In operation S620, sample noise corresponding to the time step is obtained, where the sample noise includes sample Gaussian noise and sample time-dependent noise. As the time step increases, the proportion of sample Gaussian noise in the sample noise increases, and the proportion of sample time-dependent noise decreases.
[0093] For example, an interpolation formula can be set: Σ = ωtΣr + (1 - ωt)Σg.
[0094] Among them, Σ is the noise covariance matrix at the current time step, Σr is the covariance related to time-dependent noise, Σg is the covariance of Gaussian noise, ωt is the interpolation coefficient controlling the proportion of the two, and t is the time step selected in operation S610. ωt is dynamically adjusted with the change of the time step t. Initially (when t is small), the interpolation coefficient ωt is large, resulting in time-dependent noise dominating. As the time step increases, ωt gradually decreases, leading to an increasing proportion of Gaussian noise. In this way, at the initial stage of adding noise, more time-correlated noise is introduced, and in the later stage, more Gaussian noise is gradually introduced during the noise addition process.
[0095] During the training process of the diffusion model, choosing to introduce more time-correlated noise at the initial stage of adding noise and gradually introducing more Gaussian noise in the later stage is to better simulate and recover the key features in time series data, especially the dependence and stationarity of time series data. At the initial stage of the data, the model needs to pay more attention to the local patterns and dependencies in the data. If Gaussian noise is introduced too early, it may cause the model to ignore the internal relationships in time series data, resulting in unsatisfactory denoising effects. Introducing time-correlated noise can enable the model to better understand the variation law of the data in the time dimension and provide a better starting point for subsequent denoising.
[0096] In operation S630, based on the sample noise, noise is added to the first sample time series data to obtain the noise-added sample time series data. The noise generated by the dynamic interpolation mechanism can be adjusted according to ωt at each time step to ensure the appropriate proportion of noise in different stages, thereby helping the model to gradually denoise.
[0097] In operation S640, according to the first model, noise prediction is performed on the noise-added sample time series data to obtain sample noise prediction information.
[0098] In operation S650, the target loss is calculated according to the sample noise prediction information and the sample noise.
[0099] In operation S660, the parameters of the first model are adjusted according to the target loss.
[0100] Repeat operations S610 to S660 until the number of repetitions reaches a threshold or the first model converges. Reaching the threshold of the number of repetitions is to ensure the sufficiency of the number of training times, rather than just that the loss function is small enough. Although the decrease in the loss function indicates that the model has a good noise prediction effect at a certain time step, this only reflects the local performance of the model in a single training. In fact, in a single training, a small loss does not necessarily mean that the model performs equally well at all time steps. Therefore, it is necessary to ensure a sufficient number of training times to ensure that the model has a balanced learning effect for each time step. Different time steps may have different noise characteristics and data distributions. If only the minimization of the loss at a certain time step is concerned, it may lead to poor prediction effects of the model at some time steps, or even insufficient noise processing ability at specific time points. Therefore, through a sufficiently large number of training times, the model can be gradually optimized in multiple time steps to ensure good noise prediction effects at different time steps and avoid overfitting of the model to certain specific time steps.
[0101] In some embodiments, it can be determined whether to stop iterative training by whether the first model converges, that is, when the loss of the model tends to be stable and cannot be further significantly reduced in multiple training iterations, the training should also be stopped. The convergence criterion is that the training loss or validation loss of the model has reached a stable minimum value after several rounds of iteration and hardly changes in subsequent training. At this time, even if the number of training times is continuously increased, the model performance will not be significantly improved, so it can be considered that the model is good enough.
[0102] Figure 7 Schematically shows a flowchart of obtaining sample noise prediction information in the data generation method according to an embodiment of the present disclosure.
[0103] As Figure 7 shown, on the basis of the foregoing embodiment, S640 may include operations S710 to S720.
[0104] In operation S710, obtain sample conditional information corresponding to the first sample time series data, and the sample conditional information at least characterizes the self-characteristics of the first sample time series data. During the training process, we not only consider the noise of the sample time series data itself, but also need to introduce conditional information to assist the learning process of the denoising network. Specifically, the conditional information can provide additional context information for the denoising process, helping the network better understand the internal characteristics of the data, such as the trend, periodicity, peaks or valleys at specific time points of the time series data.
[0105] During training, for each time series in the first sample time series data, a corresponding conditional vector can be obtained. This conditional vector serves as a guidance for the denoising network, helping the network to more accurately remove noise and restore the true features of the data. Specifically, the conditional vector characterizes the inherent features of the first sample time series data, which can be, for example, statistical features (such as mean, variance), periodic information, or other conditions set according to specific task requirements (such as seasonal fluctuations, external interferences, etc.).
[0106] In operation S720, the first model makes a noise prediction on the noisy sample time series data according to the sample conditional information, generating sample noise prediction information. In this process, the sample conditional information, as one of the inputs, guides the noise prediction of the denoising process, realizing the process constraint of the sample conditional information on the denoising process. This also means that the first model not only learns how to denoise, but also learns how to denoise under specific context conditions (conditional information).
[0107] Figure 8 Schematically shows a flowchart for generating time-dependent noise in the data generation method according to an embodiment of the present disclosure.
[0108] In the training process of traditional diffusion models, standard Gaussian noise is usually used for noise addition. However, in time series data generation tasks, relying solely on Gaussian noise may not be able to well simulate the time series characteristics of real data. Especially in the denoising process, Gaussian noise may not be able to effectively explain the frequency information reconstructed by the denoising network, resulting in the generated time series data lacking reasonable temporal correlation. Therefore, it is necessary to generate a kind of time-dependent noise to improve this problem.
[0109] As Figure 8 shown, based on the foregoing embodiments, the generation process of time-dependent noise may include operations S810~S820.
[0110] In operation S810, multiple generation methods are constructed, and each generation method corresponds to at least one time series feature. The generation method can be, for example, a kernel function, and each kernel function corresponds to at least one time series feature. For example, kernel function 1 is a linear function corresponding to linear features, and kernel function 2 is a periodic function corresponding to periodic features, etc.
[0111] Specifically, for example, a kernel function pool can be constructed, and different kernel functions correspond to different types of time series features: Linear Kernel: mainly used to represent long-term trends and is suitable for time series with global change patterns. Radial Basis Function (RBF) Kernel: used to model local smooth changes and is suitable for stationary time series with local fluctuations. Periodic Kernel: used to model seasonal changes and is suitable for time series with periodic patterns, etc.
[0112] In operation S820, at least one time-dependent noise is generated based on multiple generation methods, where the time-dependent noise at least partially has the timing characteristics corresponding to each of the multiple generation methods. For example, several kernel functions can be randomly selected from a kernel function pool and synthesized using different combination methods (such as addition or multiplication) to construct a more complex time-dependent pattern. A large number of time series samples are generated based on these kernel functions using a Gaussian Process, and their covariance matrices are calculated. Finally, these covariance matrices are averaged to obtain a final time-dependent noise covariance matrix for generating time-dependent noise that conforms to the timing characteristics.
[0113] Figure 9 A structural block diagram of a data processing method according to an embodiment of the present disclosure is schematically shown.
[0114] Combined with the foregoing Figures 1 - 8 shown embodiments, the overall processing flow of the data processing method provided by the present disclosure can be as shown in the block diagram Figure 9 shown.
[0115] Referring to Figure 9 , Figure 9 The part represented by block 910 in is the prerequisite content in the data processing method, that is, it can be the training part of a first model (such as a diffusion model). When the first model is a diffusion model, this process can also be called the forward process or the forward process of the diffusion model. Among them, the part shown in 911 is the process of generating noise data during the training process of the first model.
[0116] As shown in part 911 of Figure 9 , multiple generation methods can be prepared in advance, for example, a function pool composed of a linear kernel function, a radial basis kernel (RBF) function, and a periodic kernel (Periodic) function. Then, multiple of these kernel functions are randomly selected and combined into a new combined function by multiplying or adding. Then, based on the combined function, at least one time series sample is generated using a Gaussian Process, and their covariance matrices are calculated. These covariance matrices are averaged to obtain a final time-dependent noise covariance matrix. Using this time-dependent noise covariance matrix and Gaussian noise, through the interpolation formula, Σ = ωtΣr+(1 - ωt)Σg, the noise information corresponding to each time step is generated.
[0117] Among them, Σ is the noise covariance matrix of the current time step, Σr is the covariance related to the time-dependent noise, Σg is the covariance of the Gaussian noise, ωt is the interpolation coefficient controlling the ratio of the two, and t is the time step.
[0118] Subsequently, based on the obtained noise information and the first sample time-series data, the first model is trained. During the training of the first model, the conditional information corresponding to the first sample time series can also be generated based on the first sample time-series data, and the first model is trained according to the conditional information and the noise information, so that the noise prediction network in the first model can achieve good prediction results for the noise information at each time step, and it can learn that after constraining the prediction of the noise information through the conditional information, the first model can be applied to the inference process, that is, the process of generating the target time-series data.
[0119] Reference Figure 9 As shown in part 920 of the reference, in the process of generating the target time-series data, first, noise data (such as pure noise data in time-series format, or noise time-series data obtained by adding noise to specific time-series data at least once), and / or the conditional information of the target time-series data are input into the first model. After the noise prediction network in the first model predicts the noise in the noise data, the noise data is denoised according to the predicted noise prediction result. After iteratively executing multiple times, the target time-series data is obtained. When the input information of the first model also includes conditional information, the conditional information is also input into the first model. For example, the specified seasonality is 0.8, the volatility is 0.5, and the linearity is 0.9.
[0120] Reference Figure 9 As shown in part 921 of the reference, the denoising network can include an encoder and a decoder. The encoder generates at least one spatial representation vector, and these spatial representation vectors are used to characterize the specific noise part in the input noise data. Then, through the decoder and these spatial representation vectors, the noise of the data is removed to obtain the denoised intermediate time-series data.
[0121] According to the embodiments of the present disclosure, during a denoising process, a hybrid AR model can also be used to make a prediction of the output of the model during the denoising process. First, a reference time-series data is given, and then this reference time-series data is also used as one of the inputs of the denoising network to constrain the denoising process of the denoising network, so that the data distribution of the generated intermediate time-series data approaches the data distribution of the reference time-series data.
[0122] Figure 9 The data flow shown in part 921 of the reference is a process of one denoising. Through an iterative method, after performing denoising multiple times, the target time-series data is obtained. The iteration can stop when a fixed number of iterations is reached, the output target time-series data converges, etc.
[0123] Reference Figure 9For the part shown in 930, after generating the target time-series data, the target time-series data can be further adjusted by the target matrix. Among them, the part shown in 931 is the obtaining process of the target matrix. After the second sample time-series data is transformed by the target matrix, the third sample time-series data is obtained. The second model is trained according to the third sample time-series data. The second model generates output information based on the sample input time-series data and compares it with the sample output time-series data (true value). If the prediction effect is poor, then the value of the target matrix is adjusted, and the whole process is repeated until the prediction effect of the second model is good, which means that the value of the target matrix is already relatively reasonable.
[0124] Figure 10 Schematically shows a block diagram of a data generation device according to an embodiment of the present disclosure.
[0125] As Figure 10 shown, the data generation device 1000 may include a denoising module 1010.
[0126] The denoising module 1010 is configured to perform iterative denoising processing on the time-series noise data based on the Gaussian noise prediction information and the time-dependent noise prediction information to generate target time-series data; wherein, the target time-series data represents the eigenvalue distributed in chronological order in the supply chain scenario, and the eigenvalue is used to represent the object attribute in the supply chain scenario. In some embodiments, the denoising module 1010 may be configured to perform the operation S110 in the above data generation method, which will not be elaborated here.
[0127] According to an embodiment of the present disclosure, the denoising module 1010 is specifically configured to obtain noise prediction information, and the denoising module may include a first prediction module 1011 and a first recovery module 1012.
[0128] The first prediction module is configured to obtain noise prediction information, and the noise prediction information includes Gaussian noise prediction information and time-dependent noise prediction information. In some embodiments, the first prediction module 1011 may be configured to perform the operation S111 in the above data generation method, which will not be elaborated here.
[0129] The first recovery module 1012 is configured to restore the detailed features of the intermediate time-series data based on the Gaussian noise prediction information and restore the overall features of the intermediate time-series data based on the time-dependent noise prediction information; the intermediate time-series data is obtained by performing at least one denoising process on the time-series noise data; wherein, as the number of denoising processes increases, the proportion of the Gaussian noise prediction information in the noise prediction information decreases, and the proportion of the time-dependent noise prediction information increases. In some embodiments, the first recovery module 1012 may be configured to perform the operation S112 in the above data generation method, which will not be elaborated here.
[0130] According to an embodiment of the present disclosure, the data generation device may include a first acquisition module.
[0131] The first acquisition module is used to acquire condition information, and the condition information at least characterizes the self-characteristics of the target time-series data. In some embodiments, the first acquisition module may be used to perform operation S210 in the above data generation method, which will not be elaborated here.
[0132] On this premise, the denoising module may include a first denoising sub-module, which is used to obtain noise prediction information according to the condition information, so that when the intermediate time-series data is denoised according to the noise prediction information, the process constraint of the condition information on the denoising process is realized. In some embodiments, the first denoising sub-module may be used to perform operation S111 in the above data generation method, which will not be elaborated here.
[0133] According to an embodiment of the present disclosure, the denoising module may include a first prediction module and a first acquisition module.
[0134] The first prediction module is used to predict the intermediate time-series data to obtain reference time-series data, and the reference time-series data at least characterizes the data distribution prediction information of the denoised time-series data. In some embodiments, the first prediction module may be used to perform operation S310 in the above data generation method, which will not be elaborated here.
[0135] The first acquisition module is used to obtain noise prediction information according to the reference time-series data, so that when the intermediate time-series data is denoised according to the noise prediction information, the process constraint of the reference time-series data on the denoising process is realized. In some embodiments, the first acquisition module may be used to perform operation S320 in the above data generation method, which will not be elaborated here.
[0136] According to an embodiment of the present disclosure, the data generation device may further include a first acquisition module and a first adjustment module.
[0137] The first acquisition module is used to acquire a target matrix, and the target matrix at least characterizes the target change trend of the data within the target time range. In some embodiments, the first acquisition module may be used to perform operation S410 in the above data generation method, which will not be elaborated here.
[0138] The first adjustment module is used to adjust the change trend of the data within the target time range in the target time-series data to be close to the target change trend according to the target matrix. In some embodiments, the first adjustment module may be used to perform operation S420 in the above data generation method, which will not be elaborated here.
[0139] According to an embodiment of the present disclosure, the noise prediction information is generated by a first model. The present disclosure also provides a training device for training the first model. The training device may include a first selection module, a second acquisition module, a first noise addition module, a second prediction module, a first calculation module, and a second adjustment module.
[0140] The first selection module is configured to randomly select a time step from a set of sequentially arranged time steps. In some embodiments, the first selection module may be configured to perform operation S610 in the above data generation method, which will not be elaborated herein.
[0141] The second acquisition module is configured to acquire sample noise corresponding to the time step, where the sample noise includes sample Gaussian noise and sample time-dependent noise. As the time step increases, the proportion of sample Gaussian noise in the sample noise increases, and the proportion of sample time-dependent noise decreases. In some embodiments, the second acquisition module may be configured to perform operation S620 in the above data generation method, which will not be elaborated herein.
[0142] The first noise addition module is configured to add noise to the first sample time series data based on the sample noise to obtain the noise-added sample time series data. In some embodiments, the first noise addition module may be configured to perform operation S630 in the above data generation method, which will not be elaborated herein.
[0143] The second prediction module is configured to perform noise prediction on the noise-added sample time series data according to the first model to obtain sample noise prediction information. In some embodiments, the second prediction module may be configured to perform operation S640 in the above data generation method, which will not be elaborated herein.
[0144] The first calculation module is configured to calculate a target loss according to the sample noise prediction information and the sample noise. In some embodiments, the first calculation module may be configured to perform operation S650 in the above data generation method, which will not be elaborated herein.
[0145] The second adjustment module is configured to adjust the parameters of the first model according to the target loss. In some embodiments, the second adjustment module may be configured to perform operation S660 in the above data generation method, which will not be elaborated herein.
[0146] According to an embodiment of the present disclosure, the training device may include a third acquisition module and a third prediction module.
[0147] The third acquisition module is configured to acquire sample condition information corresponding to the first sample time series data, and the sample condition information at least characterizes the self-characteristics of the first sample time series data. In some embodiments, the third acquisition module may be configured to perform operation S710 in the above data generation method, which will not be elaborated herein.
[0148] The third prediction module is used for the first model to perform noise prediction on the noisy sample time series data according to the sample condition information, and generate sample noise prediction information. In some embodiments, the third prediction module can be used to execute operation S720 in the above data generation method, which will not be elaborated here.
[0149] According to an embodiment of the present disclosure, the training device may include a time-dependent noise generation module for generating time-dependent noise, and the time-dependent noise generation module may include a fourth acquisition module and a fourth prediction module.
[0150] The fourth acquisition module is used to construct multiple generation methods, and each generation method corresponds to at least one time series feature. In some embodiments, the fourth acquisition module can be used to execute operation S810 in the above data generation method, which will not be elaborated here.
[0151] The fourth prediction module is used to generate at least one time-dependent noise based on multiple generation methods, wherein the time-dependent noise at least partially has the time series features corresponding to the respective multiple generation methods. In some embodiments, the fourth prediction module can be used to execute operation S820 in the above data generation method, which will not be elaborated here.
[0152] According to the embodiments of the present disclosure, any plurality of modules, sub-modules, units, and sub-units, or at least part of the functions of any plurality of them can be implemented in one module. Any one or more of the modules, sub-modules, units, and sub-units according to the embodiments of the present disclosure can be split into multiple modules for implementation. Any one or more of the modules, sub-modules, units, and sub-units according to the embodiments of the present disclosure can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on substrate, a system on package, an application specific integrated circuit (ASIC), or can be implemented by any other reasonable way of integrating or packaging circuits in hardware or firmware, or implemented in any one of the three implementation ways of software, hardware, and firmware, or in an appropriate combination of any several of them. Alternatively, one or more of the modules, sub-modules, units, and sub-units according to the embodiments of the present disclosure can be at least partially implemented as a computer program module, and when the computer program module is run, it can execute the corresponding functions.
[0153] For example, any number of the first prediction module 1011 and the first recovery module 1012 can be combined and implemented in one module / unit / sub-unit, or any one of the modules / units / sub-units can be split into multiple modules / units / sub-units. Alternatively, at least part of the functions of one or more of these modules / units / sub-units can be combined with at least part of the functions of other modules / units / sub-units and implemented in one module / unit / sub-unit. According to an embodiment of the present disclosure, at least one of the first prediction module 1011 and the first recovery module 1012 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on substrate, a system on package, an application specific integrated circuit (ASIC), or any other reasonable manner of integrating or packaging circuits, etc., implemented by hardware or firmware, or implemented in any one of the three implementation manners of software, hardware, and firmware, or in an appropriate combination of any several of them. Alternatively, at least one of the first prediction module 1011 and the first recovery module 1012 can be at least partially implemented as a computer program module, and when the computer program module is run, the corresponding functions can be executed.
[0154] It should be noted that the data processing system part in the embodiments of the present disclosure corresponds to the data processing method part in the embodiments of the present disclosure. For the description of the data processing system part, please refer to the data processing method part specifically, and details will not be repeated here.
[0155] Figure 11 A block diagram of an electronic device suitable for implementing the method described above according to an embodiment of the present disclosure is schematically shown. Figure 11 The electronic device shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present disclosure.
[0156] As Figure 11 shown, the electronic device 1100 according to an embodiment of the present disclosure includes a processor 1101, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1102 or a program loaded from a storage section 1108 into a random access memory (RAM) 1103. The processor 1101 can include, for example, a general microprocessor (such as a CPU), an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (such as an application specific integrated circuit (ASIC)), etc. The processor 1101 can also include on-board memory for caching purposes. The processor 1101 can include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present disclosure.
[0157] In the RAM 1103, various programs and data required for the operation of the electronic device 1100 are stored. The processor 1101, the ROM 1102, and the RAM 1103 are connected to each other via a bus 1104. The processor 1101 performs various operations of the method flow according to the embodiments of the present disclosure by executing programs in the ROM 1102 and / or the RAM 1103. It should be noted that the programs may also be stored in one or more memories other than the ROM 1102 and the RAM 1103. The processor 1101 may also perform various operations of the method flow according to the embodiments of the present disclosure by executing programs stored in the one or more memories.
[0158] According to an embodiment of the present disclosure, the electronic device 1100 may further include an input / output (I / O) interface 1105, and the input / output (I / O) interface 1105 is also connected to the bus 1104. The electronic device 1100 may further include one or more of the following components connected to the input / output (I / O) interface 1105: an input portion 1106 including a keyboard, a mouse, etc.; an output portion 1107 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage portion 1108 including a hard disk, etc.; and a communication portion 1109 including a network interface card such as a LAN card, a modem, etc. The communication portion 1109 performs communication processing via a network such as the Internet. A drive 1110 is also connected to the input / output (I / O) interface 1105 as needed. A removable medium 1111, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1110 as needed so that a computer program read from it can be installed into the storage portion 1108 as needed.
[0159] According to an embodiment of the present disclosure, the method flow according to the embodiments of the present disclosure may be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable storage medium, and the computer program includes program codes for performing the method shown in the flowchart. In such an embodiment, the computer program may be downloaded and installed from a network via the communication portion 1109, and / or installed from the removable medium 1111. When the computer program is executed by the processor 1101, the above functions defined in the system according to the embodiments of the present disclosure are executed. According to an embodiment of the present disclosure, the above-described system, device, apparatus, module, unit, etc. may be implemented by computer program modules.
[0160] The present disclosure also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the foregoing embodiments; or may exist alone without being assembled into the device / apparatus / system. The foregoing computer-readable storage medium carries one or more programs, and when the one or more programs are executed, the method according to the embodiments of the present disclosure is implemented.
[0161] According to an embodiment of the present disclosure, the computer-readable storage medium may be a non-volatile computer-readable storage medium. For example, it may include but is not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing. In the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program may be used by or in combination with an instruction execution system, apparatus, or device.
[0162] For example, according to an embodiment of the present disclosure, the computer-readable storage medium may include one or more memories other than the foregoing ROM 1102 and / or RAM 1103 and / or ROM 1102 and RAM 1103.
[0163] Embodiments of the present disclosure further include a computer program product, which includes a computer program. The computer program contains program codes for executing the method provided by the embodiments of the present disclosure. When the computer program product runs on an electronic device, the program codes are used to cause the electronic device to implement the control method provided by the embodiments of the present disclosure.
[0164] When the computer program is executed by the processor 1101, the above functions defined in the system / apparatus of the embodiments of the present disclosure are executed. According to an embodiment of the present disclosure, the foregoing systems, apparatuses, modules, units, etc. may be implemented by computer program modules.
[0165] In one embodiment, the computer program may rely on tangible storage media such as optical storage devices and magnetic storage devices. In another embodiment, the computer program may also be transmitted and distributed in the form of signals on a network medium, and downloaded and installed through the communication part 1109, and / or installed from the removable medium 1111. The program code included in the computer program can be transmitted using any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above. According to the embodiments of the present disclosure, the program code for executing the computer program provided by the embodiments of the present disclosure can be written in any combination of one or more programming languages. Specifically, these computing programs can be implemented using high-level procedures and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include but are not limited to, such as Java, C++, Python, the "C" language, or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (for example, by using an Internet service provider to connect through the Internet).
[0166] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the above-mentioned module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions. Those skilled in the art can understand that the features described in the various embodiments of the present disclosure can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in the present disclosure. In particular, without departing from the spirit and teachings of the present disclosure, the features described in the various embodiments of the present disclosure can be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present disclosure.
[0167] The embodiments of the present disclosure have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present disclosure. Although the embodiments have been described separately above, this does not mean that the measures in each embodiment cannot be used advantageously in combination. Without departing from the scope of the present disclosure, those skilled in the art can make various substitutions and modifications, and all such substitutions and modifications should fall within the scope of the present disclosure.
Claims
1. A data generation method, comprising: Performing iterative denoising on the time series noise data based on Gaussian noise prediction information and time-dependent noise prediction information to generate target time series data; wherein the target time series data represents characteristic values distributed in time order in the supply chain scenario, and the characteristic values are used to represent object attributes in the supply chain scenario; The denoising process is performed once, including: Obtaining noise prediction information, wherein the noise prediction information includes Gaussian noise prediction information and time-dependent noise prediction information; Restoring detailed features of the intermediate time series data based on the Gaussian noise prediction information, and restoring overall features of the intermediate time series data based on the time-dependent noise prediction information; the intermediate time series data is obtained by performing at least one denoising process on the time series noise data; Among them, as the number of denoising processes increases, the proportion of the Gaussian noise prediction information in the noise prediction information decreases, and the proportion of the time-dependent noise prediction information increases.
2. According to the method of claim 1, the noise prediction information obtained by the subsequent denoising process is generated based on the intermediate time series data obtained by the previous denoising process.
3. The method according to claim 1, further comprising: Acquire condition information, where the condition information at least represents the intrinsic characteristics of the target time series data; The obtaining of noise prediction information comprises: The noise prediction information is obtained according to the condition information, so that when the intermediate time series data is subjected to denoising processing according to the noise prediction information, the process constraint of the denoising processing by the condition information is implemented.
4. The method according to claim 1, wherein obtaining noise prediction information comprises: Predicting the intermediate time series data to obtain reference time series data, wherein the reference time series data at least represents data distribution prediction information of the denoised time series data; The noise prediction information is obtained according to the reference time series data, so that when the intermediate time series data is subjected to denoising processing according to the noise prediction information, the process constraint of the reference time series data on the denoising processing is implemented.
5. The method according to claim 1, further comprising: Acquire a target matrix, wherein the target matrix at least represents a target change trend of data within a target time range; According to the target matrix, the change trend of the data within the target time range in the target time series data is adjusted to be close to the target change trend.
6. The method according to claim 1, wherein the noise prediction information is generated by a first model, and the training process of the first model comprises: Randomly select a time step from the sequentially ordered set of time steps; Acquire a sample noise corresponding to the time step, wherein the sample noise includes sample Gaussian noise and sample time-dependent noise, and as the time step increases, the proportion of the sample Gaussian noise in the sample noise increases, and the proportion of the sample time-dependent noise decreases; Based on the sample noise, adding noise to the first sample time series data to obtain the noisy sample time series data; According to the first model, performing noise prediction on the sample time series data after the noise is added to obtain sample noise prediction information; Calculating a target loss according to the sample noise prediction information and the sample noise; adjusting parameters of the first model according to the target loss; The above process is repeated until the number of repetitions reaches a threshold, or the first model converges.
7. The method according to claim 6, wherein the step of performing noise prediction on the noisy sample time series data according to the first model to obtain sample noise prediction information comprises: Acquire sample condition information corresponding to the first sample time series data, where the sample condition information at least represents a characteristic of the first sample time series data; The first model performs noise prediction on the noisy sample time series data according to the sample condition information to generate the sample noise prediction information.
8. The method according to claim 6, wherein the time-dependent noise is generated by: Constructing multiple generation methods, each of the generation methods corresponds to at least one time series feature; Based on the multiple generation modes, at least one time-dependent noise is generated, wherein: The time-dependent noise at least partially has the timing characteristics corresponding to each of the multiple generation modes.
9. A data generating device, comprising: A denoising module, used to perform iterative denoising on the time series noise data based on Gaussian noise prediction information and time-dependent noise prediction information to generate target time series data; wherein the target time series data represents characteristic values distributed in time order in the supply chain scenario, and the characteristic values are used to represent object attributes in the supply chain scenario; Wherein, the denoising module is specifically used for: Obtaining noise prediction information, wherein the noise prediction information includes Gaussian noise prediction information and time-dependent noise prediction information; The detailed features of the intermediate time series data are restored based on the Gaussian noise prediction information, and the overall features of the intermediate time series data are restored based on the time-dependent noise prediction information; the intermediate time series data is obtained by performing at least one denoising process on the time series noise data; Among them, as the number of denoising processes increases, the proportion of the Gaussian noise prediction information in the noise prediction information decreases, and the proportion of the time-dependent noise prediction information increases.
10. An electronic device, comprising: at least one processor; as well as A memory connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the following operations: performing iterative denoising processing on the time series noise data based on Gaussian noise prediction information and time-dependent noise prediction information to generate target time series data; wherein the target time series data represents feature values distributed in time order in the supply chain scenario, and the feature values are used to represent object attributes in the supply chain scenario; The denoising process is performed once, including: Obtaining noise prediction information, wherein the noise prediction information includes Gaussian noise prediction information and time-dependent noise prediction information; The detailed features of the intermediate time series data are restored based on the Gaussian noise prediction information, and the overall features of the intermediate time series data are restored based on the time-dependent noise prediction information; the intermediate time series data is obtained by performing at least one denoising process on the time series noise data; Among them, as the number of denoising processes increases, the proportion of the Gaussian noise prediction information in the noise prediction information decreases, and the proportion of the time-dependent noise prediction information increases.