Dataset construction method and apparatus, and computer, storage medium and program product
By splitting the dataset and constructing a set mapping model, the problem of increased complexity in robust optimization in data center temperature control is solved, and efficient and accurate dataset construction and temperature control are achieved.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- TENCENT TECHNOLOGY (SHENZHEN) CO LTD
- Filing Date
- 2025-11-11
- Publication Date
- 2026-05-21
AI Technical Summary
Existing technologies for temperature control in data centers, which construct elliptic uncertain sets, increase the complexity of robust optimization problems, increase the amount of computation required, and reduce the efficiency of solving them.
The first business forecast dataset is split into a first dataset to be processed and a second dataset to be processed. A data uncertainty set is constructed based on the first dataset to be processed. Data poles are obtained from the data uncertainty set. A set mapping model is constructed to obtain data mapping values. The second business forecast dataset is constructed based on the set size parameter and converted into a linear constraint for robust optimization problem solving.
This reduces the complexity and resource consumption of solving robust optimization problems, improves the accuracy and efficiency of dataset construction, and ensures that robust optimization results meet feasibility constraints.
Smart Images

Figure CN2025134157_21052026_PF_FP_ABST
Abstract
Description
Dataset construction methods, devices, computers, storage media, and program products
[0001] This application claims priority to Chinese Patent Application No. 2024116152188, filed on November 12, 2024, entitled “Dataset Construction Method, Apparatus, Computer, Storage Medium and Program Product”, the entire contents of which are incorporated herein by reference. Technical Field
[0002] This application relates to the field of computer technology, and in particular to a dataset construction method, apparatus, computer, storage medium, and program product. Background Technology
[0003] In data center energy management, the temperature is typically expected to be controlled within a range below a preset value. This process involves uncertainties used to control the temperature. Currently, the common approach is to obtain historical datasets, construct an elliptic uncertainty set based on these datasets, and then constrain the uncertainties using this elliptic uncertainty set. However, when using this method to solve robust optimization problems, the elliptic uncertainty set is transformed into a second-order cone constraint, increasing the complexity of the robust optimization problem, the computational load, and reducing the efficiency of the solution. Summary of the Invention
[0004] This application provides a dataset construction method, apparatus, computer, storage medium, and program product, which can reduce resource consumption and improve the efficiency of dataset construction.
[0005] This application provides a dataset construction method, which includes:
[0006] The first business forecast dataset is split into a first dataset to be processed and a second dataset to be processed, and a data uncertainty set is constructed based on the first dataset to be processed.
[0007] Obtain data poles from the uncertain data set and form a pole vector from the data poles;
[0008] A set mapping model is constructed based on the pole vector, and the set mapping model is used to obtain the data mapping value corresponding to the second dataset to be processed;
[0009] Obtain the set size parameter from the data mapping value corresponding to the second dataset to be processed, and construct the second business prediction dataset based on the set size parameter; the second business prediction dataset is used for business prediction.
[0010] One embodiment of this application provides a dataset construction apparatus, which includes:
[0011] The data splitting module is used to split the first business forecast dataset into a first dataset to be processed and a second dataset to be processed, and to construct a data uncertainty set based on the first dataset to be processed.
[0012] The pole processing module is used to obtain data poles from the data uncertainty set and form a pole vector from the data poles;
[0013] The data mapping module is used to construct a set mapping model based on the pole vectors, and to obtain the data mapping value corresponding to the second dataset to be processed using the set mapping model;
[0014] The data construction module is used to obtain the set size parameter from the data mapping value corresponding to the second dataset to be processed, and to construct the second business prediction dataset based on the set size parameter; the second business prediction dataset is used for business prediction.
[0015] Specifically, when splitting the first business forecast dataset into a first dataset to be processed and a second dataset to be processed, this data splitting module can be used for:
[0016] Obtain the data splitting parameters, perform logarithmic processing on the data splitting parameters, and obtain the business quantity threshold;
[0017] Based on the business quantity threshold and the number of first business data included in the first business prediction dataset, the first business prediction dataset is split into a first dataset to be processed and a second dataset to be processed; the number of first business data included in the second dataset to be processed is greater than or equal to the business quantity threshold.
[0018] Specifically, when constructing an uncertain data set based on the first dataset to be processed, this data splitting module can be used for:
[0019] Based on the data distribution of the first dataset to be processed, an initial uncertain set is constructed;
[0020] The center of the initial uncertainty set is rotated to the origin to obtain the data uncertainty set; the origin refers to the origin corresponding to the feature axis used to describe the data distribution of the first business prediction dataset.
[0021] Specifically, when constructing the initial uncertain set based on the data distribution of the first dataset to be processed, this data splitting module can be used for:
[0022] Obtain the covariance matrix and data mean of the first dataset to be processed, and combine the covariance matrix and data mean with the elliptic function to obtain the initial uncertain set;
[0023] When rotating the center of the initial uncertain set to the origin to obtain the data uncertain set, this data splitting module can be used for:
[0024] Eigenvalue decomposition is performed on the covariance matrix to obtain the eigenvalue matrix and the data orthogonal matrix;
[0025] By using the eigenvalue matrix and the data orthogonal matrix, the center of the initial uncertainty set is rotated to the origin, thus obtaining the data uncertainty set.
[0026] Specifically, when obtaining data poles from an uncertain data set and assembling them into a pole vector, this pole processing module can be used for:
[0027] Obtain N eigenvalues from the eigenvalue matrix, and perform pole coordinate transformation on each of the N eigenvalues to obtain N positive data poles and N negative data poles; N is a positive integer; the N positive data poles and N negative data poles correspond to N characteristic axes, and each characteristic axis includes one positive data pole and one negative data pole;
[0028] Combine N positive data poles with N negative data poles to form 2 N There are N pole matrices; each pole matrix contains N data poles, and the N data poles in each pole matrix belong to different characteristic axes;
[0029] 2 N Performing vector transformations on the pole matrices, we obtain 2 N A pole vector.
[0030] Specifically, when constructing a set mapping model based on pole vectors, this data mapping module can be used for:
[0031] Obtain the data orthogonal matrix and the data average value, and based on the rotation method of the initial uncertain set, form a mapping parameter from the data orthogonal matrix and the data average value;
[0032] A set mapping model is constructed based on pole vectors and mapping parameters.
[0033] Specifically, when retrieving the set size parameter from the data mapping value corresponding to the second dataset to be processed, this data construction module can be used for:
[0034] Sort the data mapping values corresponding to the second dataset to be processed to obtain a sequence of data mapping values;
[0035] Based on the amount of first business data included in the second dataset to be processed and the data splitting parameters, a limit determination model is constructed, and the limit determination model is parsed to obtain the data location;
[0036] The data mapping value located at the data position in the data mapping value sequence is determined as the set size parameter.
[0037] Specifically, when constructing the second business forecast dataset based on the set size parameter, this data construction module can be used for:
[0038] Obtain the pole vector and mapping parameters from the set mapping model;
[0039] Based on the set size parameter, the product of the pole vector and the mapping parameter is constrained to obtain the second business prediction dataset.
[0040] The second business prediction dataset is used to represent the error value of business prediction for business scenarios.
[0041] The device also includes:
[0042] The business forecasting module is used to perform business forecasting based on business scenarios and obtain the first business forecast result.
[0043] The business forecasting module is also used to obtain the business forecasting error value from the second business forecasting dataset, and to determine the sum of the first business forecasting result and the business forecasting error value as the second business forecasting result for the business scenario.
[0044] The business scenario is a data center temperature control scenario; the second business prediction result is used to represent the temperature relative to the outside of the data center.
[0045] The device also includes:
[0046] This business forecasting module is also used to obtain data center temperature control scenarios, including the data center temperature value and temperature change at the first moment, and to obtain latency parameters.
[0047] The business forecasting module is also used to determine the temperature parameter by summing the second business forecast result with the temperature change.
[0048] This business forecasting module is also used to weight the data center temperature value at the first moment and the temperature parameter using a delay parameter to obtain the data center temperature value at the second moment; the temperature at the first moment is less than that at the second moment.
[0049] In acquiring latency parameters, this business prediction module is also used for:
[0050] Obtain the time slot length between the first and second moments, and obtain the thermal capacity and thermal resistance of the data center;
[0051] Based on the time slot length, heat capacity, and thermal resistance, determine the temperature effect data;
[0052] The temperature effect data is processed exponentially to obtain the delay parameter.
[0053] One embodiment of this application provides a computer device, including a processor, a memory, and an input / output interface;
[0054] The processor is connected to a memory and an input / output interface, respectively. The input / output interface is used to receive and output data, the memory is used to store computer programs, and the processor is used to call the computer programs so that the computer device containing the processor executes the dataset construction method in one aspect of the embodiments of this application.
[0055] One aspect of this application provides a computer-readable storage medium storing a computer program adapted to be loaded and executed by a processor, such that a computer device having the processor performs the dataset construction method of one aspect of this application.
[0056] One aspect of this application provides a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in various optional embodiments of this application. In other words, when the computer instructions are executed by the processor, they implement the methods provided in various optional embodiments of this application.
[0057] Implementing the embodiments of this application will have the following beneficial effects:
[0058] In this embodiment, the first business prediction dataset can be split into a first dataset to be processed and a second dataset to be processed. An uncertain data set is constructed based on the first dataset to be processed. Data poles are obtained from the uncertain data set and arranged into a pole vector. A set mapping model is constructed based on the pole vector, and the set mapping model is used to obtain the data mapping value corresponding to the second dataset to be processed. A set size parameter is obtained from the data mapping value corresponding to the second dataset to be processed, and a second business prediction dataset is constructed based on the set size parameter. The second business prediction dataset is used for business prediction. Through this method, after constructing the uncertain data set, it is not directly used for solving the robust optimization problem. Instead, after splitting the initial first business prediction dataset into two independent datasets, an uncertain set is constructed using one dataset (i.e., the first dataset to be processed), and then a transformation is performed based on the data poles in the uncertain set. This transforms the robust optimization problem solution from a second-order cone constraint to a linear constraint, thereby reducing the complexity and resource consumption of solving the robust optimization problem. Meanwhile, by using the set mapping model, the final determined second business prediction dataset can conform to the distribution of the first business prediction dataset, which means that the robust optimization results meet the feasibility constraints, thereby improving the accuracy of dataset construction. Attached Figure Description
[0059] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0060] Figure 1 is a network interaction architecture diagram for dataset construction provided in an embodiment of this application;
[0061] Figure 2 is a schematic diagram of a dataset construction scenario provided in an embodiment of this application;
[0062] Figure 3 is a flowchart of a dataset construction method provided in an embodiment of this application;
[0063] Figure 4 is a schematic diagram of a specific implementation scenario of dataset construction provided in an embodiment of this application;
[0064] Figure 5 is a flowchart of a business processing method provided in an embodiment of this application;
[0065] Figure 6 is a schematic diagram of the total scheduling cost provided in an embodiment of this application;
[0066] Figure 7 is a schematic diagram showing the average solution time of two uncertain sets and the percentage difference between them provided in an embodiment of this application;
[0067] Figure 8 is a schematic diagram of a dataset construction device provided in an embodiment of this application;
[0068] Figure 9 is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation
[0069] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0070] If this application requires the collection of object data (such as user data), a prompt interface or pop-up window will be displayed before and during the collection process. This prompt interface or pop-up window is used to inform the user that certain data is being collected. The data acquisition steps will only begin after the user confirms the prompt interface or pop-up window; otherwise, the process will end. Furthermore, the acquired user data will be used in reasonable and legal scenarios or for reasonable purposes. Optionally, in scenarios where user data needs to be used but user authorization has not been obtained, authorization can be requested from the user, and the user data will be used only after authorization is granted. In other words, the use of user data in this application complies with relevant laws and regulations, meaning that the user data will be used within a reasonable and legal scope.
[0071] In this embodiment of the application, please refer to Figure 1. Figure 1 is a network interaction architecture diagram for dataset construction provided in this embodiment of the application. As shown in Figure 1, the computer device 101 can obtain a first business prediction dataset, which includes historical prediction data under the business scenario, that is, data that has already been generated. Therefore, it can be considered that the data distribution of the first business prediction dataset conforms to the actual business prediction situation. On this basis, an uncertain set is constructed to make the uncertain set conform to the actual situation, which can ensure the accuracy of the dataset construction to a certain extent. Specifically, the computer device 101 can split the first business prediction dataset into a first dataset to be processed and a second dataset to be processed. An initial uncertain set (i.e., a data uncertain set) is constructed based on the first dataset to be processed. The center of the data uncertain set is the origin, which refers to the origin corresponding to the feature axis describing the data distribution of the first business prediction dataset. The initial uncertain set is converted into linear constraints based on the second dataset to be processed, thereby reducing the computational amount of solving the robust optimization problem and improving the efficiency of dataset construction. The computer device 101 can obtain the first business forecast dataset from locally stored historical forecast data, or from business devices (such as business devices 102a, 102b, or 102c), or it can generate the first business forecast dataset based on locally stored historical forecast data and historical forecast data obtained from business devices, etc., without any restrictions. The uncertainty set is used to describe and handle uncertainties in the model, and is used to define the possible range of uncertain parameters (or uncertain quantities).
[0072] Specifically, please refer to Figure 2, which is a schematic diagram of a dataset construction scenario provided by an embodiment of this application. As shown in Figure 2, the computer device can split the first business prediction dataset 201 (which can be denoted as ψ) into a first dataset to be processed 2011 (which can be denoted as ψ1) and a second dataset to be processed 2012 (which can be denoted as ψ2). The computer device can construct a data uncertainty set 202 based on the first dataset to be processed 2011, obtain data poles from the data uncertainty set 202, and form a pole vector from the data poles. A set mapping model is constructed based on the pole vector. The set mapping model is used to obtain the data mapping value corresponding to the second dataset to be processed. Data mapping is realized based on the data poles and the set mapping model, thereby transforming the solution of the robust optimization problem into a linear constraint. Furthermore, the set size parameter is obtained from the data mapping value corresponding to the second dataset to be processed. The second business prediction dataset 203 is constructed based on the set size parameter to map the position of the second business prediction dataset 203 to the position of the first business prediction dataset. This makes the second business prediction dataset 203 conform to the data distribution of the first business prediction dataset, thereby reducing the computational cost of solving the robust optimization problem while ensuring the accuracy of the dataset construction.
[0073] Robust optimization is a type of stochastic optimization problem. Its goal is to find a solution that satisfies the constraints of the business scenario for all possible scenarios and obtains the optimal solution for the objective function.
[0074] It is understood that the computer equipment mentioned in the embodiments of this application includes, but is not limited to, terminal equipment or servers. In other words, the computer equipment can be a server or a terminal equipment, or a system composed of a server and a terminal equipment. Among them, the terminal equipment mentioned above can be an electronic device, including but not limited to mobile phones, tablets, desktop computers, laptops, handheld computers, in-vehicle equipment, augmented reality / virtual reality (AR / VR) devices, head-mounted displays, smart TVs, wearable devices, smart speakers, digital cameras, webcams and other mobile internet devices (MIDs) with network access capabilities, or terminal equipment in scenarios such as trains, ships, and flights. As shown in Figure 1, the terminal equipment can be a laptop (as shown in business device 102b), a mobile phone (as shown in business device 102c), or an in-vehicle equipment (as shown in business device 102a), etc. Figure 1 only lists some of the devices. Optionally, business device 102a refers to the device located in the vehicle 103. Business device 102a can be used to manage data 1021 (such as historical prediction data). The servers mentioned above can be independent physical servers, server clusters or distributed systems composed of multiple physical servers, or cloud servers that provide basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, vehicle-road cooperation, content delivery networks (CDN), and big data and artificial intelligence platforms.
[0075] Optionally, the data involved in the embodiments of this application may be stored in a computer device, or may be stored based on cloud storage technology or a blockchain network, without limitation.
[0076] Further, please refer to Figure 3, which is a flowchart of a dataset construction method provided in an embodiment of this application. As shown in Figure 3, the dataset construction process includes the following steps:
[0077] Step S301: The first business forecast dataset is split into a first dataset to be processed and a second dataset to be processed, and a data uncertainty set is constructed based on the first dataset to be processed.
[0078] In this embodiment, a computer device can acquire a first business prediction dataset, which includes historical prediction data for a business scenario, where a business scenario refers to a scenario using an uncertain set. Specifically, the computer device can acquire a business scenario, acquire historical prediction data associated with the business scenario, and assemble the historical prediction data into the first business prediction dataset. For example, if the business scenario is a data center temperature control scenario, the historical prediction data corresponding to this scenario is the external temperature prediction error relative to the data center. This external temperature prediction error refers to the error between the historically predicted temperature and the historically actual temperature when historically predicting the temperature outside the data center. Similarly, if the business scenario is a building temperature management scenario, the historical prediction data corresponding to this scenario is the building temperature prediction error, etc. In other words, this application can be applied to any business scenario that can use an uncertain set, without limitation. When applying this application, the corresponding historical prediction data can be acquired based on the applied business scenario for subsequent generation of uncertain sets.
[0079] Furthermore, the computer device can split the first business prediction dataset into a first dataset to be processed and a second dataset to be processed. For example, refer to Figure 4, which is a schematic diagram of a specific implementation scenario of dataset construction provided by an embodiment of this application. As shown in Figure 4, the computer device can obtain the first business prediction dataset 401 and split it into a first dataset to be processed 4011 and a second dataset to be processed 4012, which are two independent datasets.
[0080] Specifically, the computer equipment can acquire data splitting parameters, perform logarithmic processing on these parameters, and obtain a business quantity threshold. The data splitting parameters are parameters used to limit the range of data quantity in the datasets after splitting the first business prediction dataset. They can be considered preset parameters, used to determine the quantity of first business data included in the datasets after splitting the first business prediction dataset. Logarithmic processing refers to the process of converting the data splitting parameters into logarithmic form. Optionally, the data splitting parameters may include a first splitting parameter (which can be denoted as ε) and a second splitting parameter (which can be denoted as δ). Logarithmic processing is performed on the first and second splitting parameters to obtain the business quantity threshold, which can be “logδ / log(1-ε)”. That is, the difference between the logarithmic threshold (e.g., “1” in “logδ / log(1-ε)”) and the first splitting parameter is logarithmically processed to obtain the first parameter logarithm. The second splitting parameter is then logarithmically processed to obtain the second parameter logarithm. The quotient of the second parameter logarithm and the first parameter logarithm is determined as the business quantity threshold. The first and second splitting parameters can be preset parameters determined based on the actual operation of the business scenario. Both the first and second splitting parameters can fall within the range of the first parameter. For example, if the range of the first parameter is (0, 1), it can be denoted as ε,δ∈(0,1), indicating that both the first and second splitting parameters are between 0 and 1. Since the logarithmic function passes through the point (1, 0), the logarithmic processing threshold can be 1. Furthermore, based on the business quantity threshold and the number of first business data included in the first business prediction dataset, the first business prediction dataset can be split into a first dataset to be processed and a second dataset to be processed. The number of first business data included in the second dataset to be processed is greater than or equal to the business quantity threshold, which can be denoted as n2≥logδ / log(1-ε). The number of first business data included in the first business prediction dataset can be denoted as the total historical data, represented by n; the number of first business data included in the first dataset to be processed can be denoted as the first historical quantity, represented by n1; and the number of first business data included in the second dataset to be processed can be denoted as the second historical quantity, represented by n2. Where n1 = n - n2.In other words, a second quantity range can be determined based on a business quantity threshold, and a first quantity range can be determined based on the total historical data and the second quantity range. Based on the first and second quantity ranges, a first historical quantity and a second historical quantity are determined, and the sum of the first and second historical quantities is the total historical data. The first historical quantity belongs to the first quantity range, and the second historical quantity belongs to the second quantity range. Using the first and second historical quantities, the first business prediction dataset is randomly split into a first dataset to be processed and a second dataset to be processed. The number of first business data included in the first dataset to be processed is the first historical quantity, and the number of first business data included in the second dataset to be processed is the second historical quantity.
[0081] Furthermore, the computer device can construct a data uncertainty set based on the first dataset to be processed. This means the first dataset to be processed can be constructed as a data uncertainty set centered at the origin. The origin refers to the origin corresponding to the characteristic axis used to describe the data distribution of the first business prediction dataset. Specifically, as shown in Figure 4, the computer device can construct an initial uncertainty set 402 corresponding to the first dataset to be processed. This initial uncertainty set 402 is an elliptical uncertainty set. Specifically, it can be constructed based on the data distribution of the first dataset to be processed, as shown in Figure 4. The distribution of this initial uncertainty set 402 is approximately the same as the data distribution of the first dataset to be processed, and can be used to represent the data location distribution of the first business data in the first dataset to be processed. Rotating this initial uncertainty set 402 yields a data uncertainty set 403. This process rotates the initial uncertainty set 402 to the origin, that is, it rotates the center of the initial uncertainty set 402 to the origin, resulting in the data uncertainty set 403. Here, the origin refers to the origin corresponding to the characteristic axis used to describe the data distribution of the first business prediction dataset. Specifically, the computer device can obtain the covariance matrix Σ and the data mean μ of the first dataset ψ1 to be processed, and construct an initial uncertainty set based on the covariance matrix Σ and the data mean μ. This initial uncertainty set can be found in formula (1): (y-μ) T Σ -1 (y-μ)≤1 (1)
[0082] As shown in formula (1), the superscript T is used to denote the transpose of the matrix, the superscript "-1" is used to denote the inverse matrix, and y is used to denote the first business data in the first dataset to be processed. The computer device can combine the covariance matrix Σ and the data mean μ with an elliptic function to obtain the initial uncertain set.
[0083] Furthermore, computer equipment can perform eigenvalue decomposition on the covariance matrix to obtain the eigenvalue matrix and the data orthogonal matrix. Specifically, this involves performing eigenvalue decomposition on the inverse matrix of the covariance matrix, a process that can be denoted as Σ.-1 =P T ΛP, where P represents Σ -1 The matrix composed of eigenvectors can be denoted as the data orthogonal matrix. Λ is used to represent the eigenvalue matrix, and can be written as Λ=diag[λ1,λ2,…,λ]. N In the eigenvalue matrix, diag[] represents the diagonal matrix. For example, the contents of diag[] here are the values of the elements on the diagonal of the eigenvalue matrix. N is the dimension of the eigenvalue matrix. N is a positive integer and can be used to represent the number of eigenaxis. As shown in Figure 4, there are two eigenaxis, D1 and D2, and N is 2. Eigenvalue decomposition refers to the process of decomposing a matrix into the product of a set of eigenvalues (i.e., the eigenvalue matrix) and an eigenvector matrix (i.e., the data orthogonal matrix). The data orthogonal matrix is a matrix composed of multiple linearly independent eigenvectors.
[0084] Furthermore, the initial uncertain set can be rotated using eigenvalue matrices and data orthogonal matrices to obtain the data uncertain set. Specifically, the center of the initial uncertain set can be rotated to the origin using eigenvalue matrices and data orthogonal matrices to obtain the data uncertain set. Alternatively, the initial uncertain set can be parameterized using an elliptic function at the origin to obtain the data uncertain set. Specifically, the parameters to be merged in the initial uncertain set can be determined based on the elliptic function at the origin, and these parameters can be used as unknowns to generate the data uncertain set. The elliptic function at the origin refers to an elliptic function centered at the origin. Specifically, the parameters to be merged, "P(y-μ)", can be obtained and used as unknowns y', which can be written as y'=P(y-μ) to obtain the data uncertain set, which can be seen in formula (2): y ′T Λy′≤1 (2)
[0085] As shown in formula (2), y' represents the data in the uncertain data set, which can be seen in the uncertain data set 403 in Figure 4. In simple terms, the computer device can obtain the parameters to be merged from the initial uncertain set based on the origin elliptic function, merge these parameters into the origin transformation parameters (i.e., y'), and update the parameters to be merged in the initial uncertain set to the origin transformation parameters to obtain the uncertain data set. At this point, the processing of the uncertain set is transformed to the origin, which simplifies the subsequent linear transformation of the uncertain set, thereby improving the efficiency of dataset construction.
[0086] Step S302: Obtain data poles from the data uncertainty set and form a pole vector from the data poles.
[0087] In this embodiment, a computer device can obtain data poles from a data uncertainty set and form a pole vector from these data poles. A data pole refers to a point located on a characteristic axis within the boundary of the data uncertainty set. For example, if the number of characteristic axes is N (where N is a positive integer), then the data pole refers to a point located on N characteristic axes within the boundary of the data uncertainty set. Each characteristic axis has a positive half-axis and a negative half-axis. Each characteristic axis has one positive data pole on its positive half-axis and one negative data pole on its negative half-axis, resulting in 2N data poles.
[0088] Specifically, the computer device can obtain N eigenvalues from the eigenvalue matrix, perform pole coordinate transformation on the N eigenvalues respectively, and obtain N positive data poles and N negative data poles; N is a positive integer; the N positive data poles and N negative data poles correspond to N characteristic axes, and each characteristic axis includes one positive data pole and one negative data pole. Among them, the N positive data poles can be referred to as formula (3):
[0089] As shown in formula (3), V + This is used to represent N positive data poles. The N negative data poles can be found in formula (4):
[0090] As shown in formula (4), V - Used to represent N negative data poles. Where V i + The data poles on the i-th positive half-axis can be denoted as: V i - The data poles on the i-th negative half-axis can be denoted as:
[0091] Furthermore, N positive data poles and N negative data poles can be combined to form 2 N There are N pole matrices, each consisting of a data pole from each of the N characteristic axes. That is, each pole matrix consists of any one data pole included in each of the N characteristic axes, meaning each pole matrix contains N data poles belonging to different characteristic axes. The j-th pole matrix can be found in formula (5).
[0092] As shown in formula (5), j is a positive integer, j = 1, 2, ..., 2 N For example, assuming N is 3, there are 6 data poles, namely... It can form 2 3 The pole matrices are respectively... and
[0093] Furthermore, we can consider 2 N Each of the pole matrices is transformed into a vector, resulting in 2 N There are j-th pole vectors. The j-th pole vector can be found in equation (6):
[0094] In this way, 2 can be obtained. N The two extreme vectors, N The pole vectors form the edges connecting N positive data poles and N negative data poles, which constitute the first polyhedral uncertainty set, as shown in Figure 4, the first polyhedral uncertainty set 404. This achieves a linear transformation of the uncertainty set, converting the robust optimization problem solution into a linear constraint, thereby reducing the computational cost of solving the robust optimization problem and improving the efficiency of dataset construction.
[0095] Step S303: Construct a set mapping model based on the pole vector, and use the set mapping model to obtain the data mapping value corresponding to the second dataset to be processed.
[0096] In this embodiment, the computer device can construct a set mapping model based on pole vectors. This set mapping model is used to restore the data position distribution indicated by the uncertain data set to the data position distribution indicated by the first business prediction dataset. In other words, the set mapping model refers to a model that converts the position of data located in the uncertain data set into its position within the data position distribution indicated by the first business prediction dataset. Specifically, the data orthogonal matrix and data average value can be obtained. Based on the rotation method of the initial uncertain set's position, the data orthogonal matrix and data average value are used to form mapping parameters. The set mapping model is constructed based on the pole vectors and mapping parameters. Specifically, the pole vectors and mapping parameters can form a set mapping model. This set mapping model is used to restore the position of the uncertain set, that is, to restore the rotation of the initial uncertain set, so that the set mapping model can indicate the second polyhedral uncertain set, as shown in Figure 4, the second polyhedral uncertain set 405. This set mapping model can be denoted as... This mapping parameter can be denoted as P(y-μ), and can be used to represent the size boundary of the second polyhedron uncertainty set 405, where, The vector of the j-th pole is used to represent the data orthogonal matrix. Since y' = P(y-μ), by constructing this set mapping model, the first polyhedral uncertain set can be rotated to the second polyhedral uncertain set, thereby restoring the position of the uncertain set to the data distribution of the first business prediction dataset.
[0097] Furthermore, the computer device can use a set mapping model to obtain the data mapping value corresponding to the second dataset to be processed; that is, for Obtain the data mapping value corresponding to the second dataset to be processed, where, ψ2 is used to represent all, and ψ2 is used to represent the second dataset to be processed. For example, the first business data in the second dataset to be processed can be substituted into the set mapping model to obtain the data mapping value of the first business data in the second dataset to be processed.
[0098] Step S304: Obtain the set size parameter from the data mapping value corresponding to the second dataset to be processed, and construct the second business prediction dataset based on the set size parameter; the second business prediction dataset is used for business prediction.
[0099] In this embodiment, the second business forecast dataset is used for business forecasting. Specifically, the computer device can obtain the set size parameter from the data mapping values corresponding to the second dataset to be processed. Specifically, the computer device can sort the data mapping values corresponding to the second dataset to be processed to obtain a data mapping value sequence, which can be seen in formula (7):
[0100] As shown in formula (7), the subscript “(x)” is used to represent the xth data mapping value in the data mapping value sequence, and n2 is used to represent the second historical number of the first business data included in the second dataset to be processed.
[0101] Furthermore, a limit determination model can be constructed based on the quantity of first business data included in the second dataset to be processed and the data splitting parameters. For example, the quantity of first business data included in the second dataset to be processed and the data splitting parameters can be used to form a limit determination model; the limit determination model can be parsed to obtain the data location. The limit determination model can be found in formula (8).
[0102] As shown in formula (8), Used to represent combination operations, it can be written as By determining the model through analytical limits, the data location l can be obtained. * This data location can be used to indicate the location value when applying statistical feasibility constraints. The data mapping value located at the data location within the data mapping value sequence can be defined as the set size parameter, which can be denoted as... It can be used to represent the boundary values of the second dataset to be processed.
[0103] Furthermore, the computer equipment can construct a second business prediction dataset based on the set size parameter. Specifically, the set size parameter is used to constrain the dataset to obtain the second business prediction dataset. Specifically, the pole vector and mapping parameters can be obtained from the set mapping model; based on the set size parameter, the product of the pole vector and the mapping parameters is constrained to obtain the second business prediction dataset, as shown in Figure 4, which is the second business prediction dataset 406. The second business prediction dataset can be referred to in formula (9):
[0104] As shown in formula (9), U P This is used to represent the second business forecast dataset, where ρ is a value obtained based on the set mapping model. The second business forecast dataset is constructed using the set mapping model, making it a linearly constrained set. That is, the second business forecast dataset is a polyhedral uncertain set, which can be considered as an uncertain set construction method based on statistically guaranteed vertex link (SGVL). This makes the solution to the robust optimization problem linearly constrained, thereby reducing the computational load, reducing resource consumption, and improving the solution efficiency of the robust optimization problem.
[0105] In this embodiment, the first business prediction dataset can be split into a first dataset to be processed and a second dataset to be processed. An uncertain data set is constructed based on the first dataset to be processed. Data poles are obtained from the uncertain data set and arranged into a pole vector. A set mapping model is constructed based on the pole vector, and the set mapping model is used to obtain the data mapping value corresponding to the second dataset to be processed. A set size parameter is obtained from the data mapping value corresponding to the second dataset to be processed, and a second business prediction dataset is constructed based on the set size parameter. The second business prediction dataset is used for business prediction. Through this method, after constructing the uncertain data set, it is not directly used for solving the robust optimization problem. Instead, after splitting the initial first business prediction dataset into two independent datasets, an uncertain set is constructed using one dataset (i.e., the first dataset to be processed), and then a transformation is performed based on the data poles in the uncertain set. This transforms the robust optimization problem solution from a second-order cone constraint to a linear constraint, thereby reducing the complexity and resource consumption of solving the robust optimization problem. Meanwhile, by using the set mapping model, the final determined second business prediction dataset can conform to the distribution of the first business prediction dataset, which means that the robust optimization results meet the feasibility constraints, thereby improving the accuracy of dataset construction.
[0106] Furthermore, this application can be used in any business scenario that utilizes an uncertain set. For example, please refer to Figure 5, which is a flowchart of a business processing method provided by an embodiment of this application. As shown in Figure 5, the business processing includes the following steps:
[0107] Step S501: The first business forecast dataset is split into a first dataset to be processed and a second dataset to be processed, and an uncertain data set is constructed based on the first dataset to be processed.
[0108] In the embodiments of this application, the process can be referred to the relevant description in step S301 of FIG3, and will not be repeated here.
[0109] Step S502: Obtain data poles from the data uncertainty set and form a pole vector from the data poles.
[0110] In the embodiments of this application, the process can be referred to the relevant description in step S302 of FIG3, and will not be repeated here.
[0111] Step S503: Construct a set mapping model based on the pole vector, and use the set mapping model to obtain the data mapping value corresponding to the second dataset to be processed.
[0112] In the embodiments of this application, the process can be referred to the relevant description in step S303 of FIG3, and will not be repeated here.
[0113] Step S504: Obtain the set size parameter from the data mapping value corresponding to the second dataset to be processed, and construct the second business prediction dataset based on the set size parameter; the second business prediction dataset is used for business prediction.
[0114] In the embodiments of this application, the process can be referred to the relevant description in step S304 of FIG3, and will not be repeated here.
[0115] Step S505: Obtain the business forecast error value from the second business forecast dataset, perform business forecasting based on the business forecast error value, and obtain the second business forecast result.
[0116] In this embodiment of the application, the second business prediction dataset is used to represent the error value of business prediction for a business scenario. Specifically, the computer device can perform business prediction for a business scenario to obtain a first business prediction result; obtain the business prediction error value from the second business prediction dataset, and determine the second business prediction result for the business scenario by summing the first business prediction result and the business prediction error value.
[0117] For example, this business scenario is a data center temperature control scenario, in which the data center temperature is expected to be controlled below a preset value, which can be denoted as T.DC (t)≤T set At this point, the second business forecast result is used to represent the temperature relative to the outside of the data center. The computer equipment can obtain the data center temperature control scenario, the data center temperature value and temperature change at the first moment, and obtain the delay parameter k. The delay parameter is a parameter used to convert the data center temperature value at one moment to the next moment. The sum of the second business forecast result and the temperature change is determined as the temperature parameter. Using the delay parameter, the data center temperature value at the first moment and the temperature parameter are weighted to obtain the data center temperature value at the second moment; the first moment is less than the second moment. This process can be seen in formula (10): T DC (t)=kT DC (t-1)+(1-k)[T out (t-1)+ΔT DC (t-1)] (10)
[0118] As shown in formula (10), the subscript DC represents the data center, and the subscript out represents the data outside the data center. k represents the delay parameter, t-1 represents the first time step, and t represents the second time step. ΔT DC (t-1) is used to represent the temperature change, which refers to the temperature change in a business scenario from the first moment to the second moment, such as the temperature change caused by cooling / heating of air conditioners, servers, etc., from the first moment to the second moment; T out (t-1) represents the second business forecast result, and (t-1) represents the external temperature of the data. At this point, the external temperature of the data consists of the first business forecast result and the business forecast error value, as shown in formula (11): T out (t)=T out,fore (t)+T out,err (t) (11)
[0119] As shown in formula (11), T out,fore (t) represents the first business forecast result at time t, where T is the first business forecast result. out,err (t) represents the service prediction error value at time t, which is the service prediction error value T. out,err (t)∈U P In this case, a statistically guaranteed vertex link (SGVL) was constructed, so that the final business prediction result can meet the constraints based on statistical feasibility. In this business scenario, the constraints based on statistical feasibility can be seen in formula (12): Pr[Pr(T DC (t)≤T set)≥1-ε]≥1-δ (12)
[0120] As shown in formula (12), Pr is used to represent probability; T set This is used to represent the preset range of values for which the data center temperature is expected to be controlled. Optionally, the constraint parameters (i.e., ε and δ) used in the statistical feasibility constraint can be considered as the data splitting parameters used in this application, so that when optimizing the robust optimization problem based on the second business forecast dataset obtained in this application, the optimization result can satisfy the constraint based on statistical feasibility, thereby improving the accuracy of solving the robust optimization problem.
[0121] Optionally, when obtaining latency parameters, the time slot length Δt between the first and second moments can be obtained, as well as the heat capacity C of the data center. DC With thermal resistance R DC Based on the time slot length, heat capacity, and thermal resistance, the temperature influence data is determined. Specifically, the time slot length, heat capacity, and thermal resistance data are integrated to obtain the temperature influence data. The temperature influence data is then subjected to exponential processing to obtain the delay parameter. Optionally, this delay parameter can be found in formula (13):
[0122] As shown in formula (13), the temperature effect data is as follows: For example, in the building temperature management scenario, the business prediction can refer to the prediction of the building temperature. The first business prediction result is the initial building predicted temperature, and the second business prediction result is the building adjustment predicted temperature.
[0123] In all the formulas involved in this application, the superscript T is used to indicate the transpose of a matrix, and the superscript "-1" is used to indicate the inverse matrix.
[0124] This application can be considered a statistically feasible robust scheduling approach. For data center temperature scheduling, it uses the SGVL algorithm proposed in this application to construct a second business prediction dataset. A rolling scheduling method is adopted, with sub-scheduling periods B = 1, 2, 3, 4, 5, 6, 7, 8, and the total scheduling cost is used as the scheduling objective. This total scheduling cost can be seen in Formula 14.
[0125] As shown in formula (14), the subscript grid represents the external power grid, el represents the electrolyzer, fc represents the fuel cell, RE represents new energy, and AC represents the air conditioning system; ρ represents the cost of producing or consuming a unit of power, i.e., the cost per unit of power consumption; and Pw represents the amount of power produced or consumed. In other words, the total dispatch cost... It is the sum of costs incurred by each statistical object. Statistical objects refer to objects that produce or consume power, such as grid, el, etc. as shown in the above formula (14). Other statistical objects can be added to the total cost of rescheduling as needed, or some statistical objects can be deleted.
[0126] For example, the uncertain outdoor temperature is obtained from meteorological datasets (such as the Jena Climate Dataset), and the renewable energy output data comes from renewable energy datasets (such as the Elia Group). The scheduling cycle is 24 hours, and the time interval is 15 minutes (T = 96, used to represent the number of scheduling times). The constraint parameter in the statistical feasibility constraint is ε = δ = 0.05. The experiment was conducted on a personal computer equipped with an i7-12700H 2.30GHz processor and 16GB of RAM, using the MATLAB platform and the GUROBI 10.0.1 solver. The experiment used traditional box-type and elliptical uncertain sets as benchmarks, while this application generates a polyhedral uncertain set. The experimental results are shown in Figures 6 and 7.
[0127] Figure 6 is a schematic diagram of the total scheduling cost provided in an embodiment of this application. As shown in Figure 6, the total cost using polyhedral and elliptical uncertainty sets... Lower than the total cost of using box-shaped uncertainty sets This result demonstrates that introducing statistical feasibility into the framework can reduce the conservatism of the scheduling results. Furthermore, the results for the polyhedral uncertainty set are not significantly better than those for the elliptical uncertainty set. This indicates that, compared to the elliptical uncertainty set, the polyhedral uncertainty set constructed using the SGVL proposed in this application does not worsen the robust optimization results. Moreover, the total scheduling cost increases as the sub-scheduling period B changes from 1 to 8. This trend suggests that conservatism increases with the length of the prediction period. This increase is caused by the fact that prediction errors over longer periods are larger than those over shorter periods. These larger errors lead to a larger uncertainty set, thus increasing conservatism.
[0128] Figure 7 illustrates the average solution time of two uncertainty sets provided in this application and the percentage difference between them. As shown in Figure 7, the solution time using the elliptic uncertainty set (the broken line containing △) is always higher than the solution time using the polyhedral uncertainty set (the broken line containing □). This fact indicates that the polyhedral uncertainty set constructed using the SGVL algorithm can reduce the complexity of the problem to be solved in each time slot. Furthermore, the solution time increases with the increase of B. This is because an increase in B leads to an increase in the dimensionality of the robust optimization problem. The bars in Figure 7 show the percentage difference in solution time under different uncertainty sets, and it can be seen that using the polyhedral uncertainty set can reduce the solution time by 7% to 14%.
[0129] In summary, it can be seen that this application has significant performance improvements compared to box-shaped and elliptical uncertain sets.
[0130] Further, please refer to Figure 8, which is a schematic diagram of a dataset construction device provided in an embodiment of this application. This dataset construction device can be a computer program (including program code, etc.) running on a computer device; for example, the dataset construction device can be an application software. The device can be used to execute the corresponding steps in the method provided in the embodiment of this application. As shown in Figure 8, the dataset construction device 800 can be used in the computer device in the embodiment corresponding to Figure 3. Specifically, the device may include: a data splitting module 11, an extremum processing module 12, a data mapping module 13, and a data construction module 14.
[0131] The data splitting module 11 is used to split the first business forecast dataset into a first dataset to be processed and a second dataset to be processed, and to construct a data uncertainty set based on the first dataset to be processed.
[0132] Pole processing module 12 is used to obtain data poles from the data uncertainty set and form a pole vector from the data poles;
[0133] Data mapping module 13 is used to construct a set mapping model based on the pole vector and use the set mapping model to obtain the data mapping value corresponding to the second dataset to be processed;
[0134] The data construction module 14 is used to obtain the set size parameter from the data mapping value corresponding to the second dataset to be processed, and construct the second business prediction dataset based on the set size parameter; the second business prediction dataset is used for business prediction.
[0135] Specifically, when splitting the first business prediction dataset into a first dataset to be processed and a second dataset to be processed, the data splitting module 11 can be used for:
[0136] Obtain the data splitting parameters, perform logarithmic processing on the data splitting parameters, and obtain the business quantity threshold;
[0137] Based on the business quantity threshold and the number of first business data included in the first business prediction dataset, the first business prediction dataset is split into a first dataset to be processed and a second dataset to be processed; the number of first business data included in the second dataset to be processed is greater than or equal to the business quantity threshold.
[0138] Specifically, when constructing a data uncertainty set based on the first dataset to be processed, the data splitting module 11 can be used for:
[0139] Based on the data distribution of the first dataset to be processed, an initial uncertain set is constructed;
[0140] The center of the initial uncertainty set is rotated to the origin to obtain the data uncertainty set; the origin refers to the origin corresponding to the feature axis used to describe the data distribution of the first business prediction dataset.
[0141] Specifically, when constructing the initial uncertain set based on the data distribution of the first dataset to be processed, the data splitting module 11 can be used for:
[0142] Obtain the covariance matrix and data mean of the first dataset to be processed, and combine the covariance matrix and data mean with the elliptic function to obtain the initial uncertain set;
[0143] When rotating the center of the initial uncertain set to the origin to obtain the data uncertain set, the data splitting module 11 can be used for:
[0144] Eigenvalue decomposition is performed on the covariance matrix to obtain the eigenvalue matrix and the data orthogonal matrix;
[0145] By using the eigenvalue matrix and the data orthogonal matrix, the center of the initial uncertainty set is rotated to the origin, thus obtaining the data uncertainty set.
[0146] Specifically, when obtaining data poles from the uncertain data set and assembling the data poles into a pole vector, the pole processing module 12 can be used for:
[0147] Obtain N eigenvalues from the eigenvalue matrix, and perform pole coordinate transformation on each of the N eigenvalues to obtain N positive data poles and N negative data poles; N is a positive integer; the N positive data poles and N negative data poles correspond to N characteristic axes, and each characteristic axis includes one positive data pole and one negative data pole;
[0148] Combine N positive data poles with N negative data poles to form 2 N There are N pole matrices; each pole matrix contains N data poles, and the N data poles in each pole matrix belong to different characteristic axes;
[0149] 2 N Performing vector transformations on the pole matrices, we obtain 2 N A pole vector.
[0150] Specifically, when constructing a set mapping model based on pole vectors, the data mapping module 13 can be used for:
[0151] Obtain the data orthogonal matrix and the data average value, and based on the rotation method of the initial uncertain set, form a mapping parameter from the data orthogonal matrix and the data average value;
[0152] A set mapping model is constructed based on pole vectors and mapping parameters.
[0153] Specifically, when obtaining the set size parameter from the data mapping value corresponding to the second dataset to be processed, the data construction module 14 can be used for:
[0154] Sort the data mapping values corresponding to the second dataset to be processed to obtain a sequence of data mapping values;
[0155] Based on the amount of first business data included in the second dataset to be processed and the data splitting parameters, a limit determination model is constructed, and the limit determination model is parsed to obtain the data location;
[0156] The data mapping value located at the data position in the data mapping value sequence is determined as the set size parameter.
[0157] Specifically, when constructing the second business forecast dataset based on the set size parameter, the data construction module 14 can be used for:
[0158] Obtain the pole vector and mapping parameters from the set mapping model;
[0159] Based on the set size parameter, the product of the pole vector and the mapping parameter is constrained to obtain the second business prediction dataset.
[0160] The second business prediction dataset is used to represent the error value of business prediction for business scenarios.
[0161] The device 800 also includes:
[0162] Business forecasting module 15 is used to perform business forecasting for business scenarios and obtain the first business forecasting result.
[0163] The business forecasting module 15 is also used to obtain the business forecasting error value from the second business forecasting dataset, and to determine the sum of the first business forecasting result and the business forecasting error value as the second business forecasting result for the business scenario.
[0164] The business scenario is a data center temperature control scenario; the second business prediction result is used to represent the temperature relative to the outside of the data center.
[0165] The device 800 also includes:
[0166] The business prediction module 15 is also used to obtain data center temperature control scenarios, the data center temperature value and temperature change at the first moment, and to obtain delay parameters.
[0167] The business forecasting module 15 is also used to determine the temperature parameter by summing the second business forecasting result with the temperature change.
[0168] The business prediction module 15 is also used to use delay parameters to weight the data center temperature value at the first moment and the temperature parameter to obtain the data center temperature value at the second moment; the temperature at the first moment is less than that at the second moment.
[0169] In acquiring delay parameters, the business prediction module 15 is also used for:
[0170] Obtain the time slot length between the first and second moments, and obtain the thermal capacity and thermal resistance of the data center;
[0171] Based on the time slot length, heat capacity, and thermal resistance, determine the temperature effect data;
[0172] The temperature effect data is processed exponentially to obtain the delay parameter.
[0173] This application provides a dataset construction apparatus. This apparatus can split a first business prediction dataset into a first dataset to be processed and a second dataset to be processed, and construct a data uncertainty set based on the first dataset to be processed. It then obtains data poles from the data uncertainty set and forms a pole vector. Based on the pole vector, it constructs a set mapping model and uses this model to obtain the data mapping value corresponding to the second dataset to be processed. Finally, it obtains a set size parameter from the data mapping value corresponding to the second dataset to be processed and constructs a second business prediction dataset based on the set size parameter. This second business prediction dataset is used for business prediction. Through this method, after constructing the data uncertainty set, it is not directly used for solving robust optimization problems. Instead, after splitting the initial first business prediction dataset into two independent datasets, an uncertainty set is constructed using a single dataset (i.e., the first dataset to be processed). Then, a transformation is performed based on the data poles in the uncertainty set, converting the robust optimization problem solution from a second-order cone constraint to a linear constraint, thereby reducing the complexity and resource consumption of solving the robust optimization problem. Meanwhile, by using the set mapping model, the final determined second business prediction dataset can conform to the distribution of the first business prediction dataset, which means that the robust optimization results meet the feasibility constraints, thereby improving the accuracy of dataset construction.
[0174] Referring to Figure 9, which is a schematic diagram of the structure of a computer device provided in an embodiment of this application, the computer device in this embodiment may include one or more processors 901, a memory 902, and an input / output interface 903. The processor 901, memory 902, and input / output interface 903 are connected via a bus 904. The memory 902 stores a computer program, which includes program instructions. The input / output interface 903 receives and outputs data, such as for data interaction between the computer device and a business device. The processor 901 executes the program instructions stored in the memory 902.
[0175] The processor 901 can perform the following operations:
[0176] The first business forecast dataset is split into a first dataset to be processed and a second dataset to be processed, and a data uncertainty set is constructed based on the first dataset to be processed.
[0177] Obtain data poles from the uncertain data set and form a pole vector from the data poles;
[0178] A set mapping model is constructed based on the pole vector, and the set mapping model is used to obtain the data mapping value corresponding to the second dataset to be processed;
[0179] Obtain the set size parameter from the data mapping value corresponding to the second dataset to be processed, and construct the second business prediction dataset based on the set size parameter; the second business prediction dataset is used for business prediction.
[0180] In some feasible implementations, the processor 901 may be a central processing unit (CPU), but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.
[0181] The memory 902 may include read-only memory and random access memory, and provides instructions and data to the processor 901 and the input / output interface 903. A portion of the memory 902 may also include non-volatile random access memory. For example, the memory 902 may also store device type information.
[0182] In practice, the computer device can execute the implementation methods provided by each step in Figure 3 through its built-in functional modules. For details, please refer to the implementation methods provided by each step in Figure 3, which will not be repeated here.
[0183] This application embodiment provides a computer device including a processor, an input / output interface, and a memory. The processor retrieves a computer program from the memory and executes the steps of the method shown in Figure 3 to perform a dataset construction operation. This application embodiment implements the following: splitting a first business prediction dataset into a first unprocessed dataset and a second unprocessed dataset; constructing a data uncertainty set based on the first unprocessed dataset; obtaining data poles from the data uncertainty set and forming a pole vector; constructing a set mapping model based on the pole vector; using the set mapping model to obtain the data mapping value corresponding to the second unprocessed dataset; obtaining a set size parameter from the data mapping value corresponding to the second unprocessed dataset; and constructing a second business prediction dataset based on the set size parameter. The second business prediction dataset is used for business prediction. By employing the above method, after constructing the uncertain data set, it is not directly used to solve the robust optimization problem. Instead, after splitting the initial first business prediction dataset into two independent datasets, an uncertain set is constructed using one dataset (i.e., the first dataset to be processed). Then, a transformation is performed based on the data poles within this uncertain set, converting the robust optimization problem from a second-order cone constraint to a linear constraint. This reduces the complexity and resource consumption of solving the robust optimization problem. Simultaneously, through the set mapping model, the final determined second business prediction dataset can be made to conform to the distribution of the first business prediction dataset, ensuring that the robust optimization results meet feasibility constraints and thus improving the accuracy of dataset construction.
[0184] This application also provides a computer-readable storage medium storing a computer program adapted to be loaded by a processor and executed by the dataset construction method provided in each step of FIG3. Specific implementations of each step in FIG3 are described below. Furthermore, the beneficial effects of using the same method are also not described in detail here. For technical details not disclosed in the embodiments of the computer-readable storage medium involved in this application, please refer to the description of the method embodiments of this application. As an example, the computer program can be deployed to execute on a single computer device, or on multiple computer devices located in one location, or on multiple computer devices distributed across multiple locations and interconnected via a communication network.
[0185] The computer-readable storage medium can be the dataset construction apparatus provided in any of the foregoing embodiments or the internal storage unit of the computer device, such as the hard disk or memory of the computer device. The computer-readable storage medium can also be an external storage device of the computer device, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., provided on the computer device. Furthermore, the computer-readable storage medium can include both internal storage units and external storage devices of the computer device. The computer-readable storage medium is used to store the computer program and other programs and data required by the computer device. The computer-readable storage medium can also be used to temporarily store data that has been output or will be output.
[0186] This application also provides a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. The processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in the various optional modes shown in Figure 3. This achieves the following: after constructing the uncertain data set, it does not directly use the uncertain data set for solving the robust optimization problem. Instead, after splitting the initial first business prediction dataset into two independent datasets, it constructs an uncertain set using one dataset (i.e., the first dataset to be processed), and then performs a transformation based on the data poles in the uncertain set. This transforms the robust optimization problem solution from a second-order cone constraint to a linear constraint, thereby reducing the complexity and resource consumption of solving the robust optimization problem. Simultaneously, through the set mapping model, the final determined second business prediction dataset can conform to the distribution of the first business prediction dataset, that is, the robust optimization result meets the feasibility constraints, thereby improving the accuracy of dataset construction.
[0187] The terms "first," "second," etc., in the specification, claims, and drawings of this application are used to distinguish different objects, not to describe a specific order. Furthermore, the term "comprising," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, apparatus, product, or device that includes a series of steps or units is not limited to the listed steps or modules, but may optionally include steps or modules not listed, or may optionally include other step units inherent to these processes, methods, apparatuses, products, or devices.
[0188] In this application embodiment, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.
[0189] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described in terms of functionality. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this application.
[0190] The methods and related apparatuses provided in this application are described with reference to the method flowcharts and / or structural diagrams provided in this application. Specifically, each block of the method flowcharts and / or structural diagrams, as well as combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data set building apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data set building apparatus, create means for implementing the functions specified in one or more blocks of the flowcharts and / or one or more blocks of the structural diagrams. These computer program instructions can also be stored in a computer-readable storage medium capable of directing a computer or other programmable data set building apparatus to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means that implement the functions specified in one or more blocks of the flowcharts and / or one or more blocks of the structural diagrams. These computer program instructions may also be loaded onto a computer or other programmable data-building device to cause a series of operational steps to be performed on the computer or other programmable device to produce a computer-implemented process, such that the instructions, which execute on the computer or other programmable device, provide steps for implementing the functions specified in one or more flowcharts and / or one or more structural diagrams in blocks.
[0191] The steps in the method of this application embodiment can be adjusted, combined, or deleted according to actual needs.
[0192] The modules in the device of this application embodiment can be merged, divided, and deleted according to actual needs.
[0193] The above-disclosed embodiments are merely preferred embodiments of this application and should not be construed as limiting the scope of this application. Therefore, any equivalent variations made in accordance with the claims of this application shall still fall within the scope of this application.
Claims
1. A data set construction method characterized by comprising: The method includes: The first business forecast dataset is split into a first dataset to be processed and a second dataset to be processed, and a data uncertainty set is constructed based on the first dataset to be processed. Data poles are obtained from the data uncertainty set and the data poles are arranged into a pole vector; A set mapping model is constructed based on the extreme point vector, and the set mapping model is used to obtain the data mapping value corresponding to the second dataset to be processed; Obtain the set size parameter from the data mapping value corresponding to the second dataset to be processed, and construct a second business prediction dataset based on the set size parameter; the second business prediction dataset is used for business prediction.
2. The method of claim 1, wherein, The step of splitting the first business prediction dataset into a first dataset to be processed and a second dataset to be processed includes: Obtain the data splitting parameters, perform logarithmic processing on the data splitting parameters, and obtain the business quantity threshold; Based on the business quantity threshold and the number of first business data included in the first business prediction dataset, the first business prediction dataset is split into a first dataset to be processed and a second dataset to be processed; the number of first business data included in the second dataset to be processed is greater than or equal to the business quantity threshold.
3. The method of claim 1, wherein, The construction of the data uncertainty set based on the first dataset to be processed includes: Based on the data distribution of the first dataset to be processed, an initial uncertainty set is constructed; The center of the initial uncertainty set is rotated to the origin to obtain the data uncertainty set; the origin refers to the origin corresponding to the feature axis used to describe the data distribution of the first business prediction dataset.
4. The method of claim 3, wherein, The construction of an initial uncertain set based on the data distribution of the first dataset to be processed includes: Obtain the covariance matrix and data mean of the first dataset to be processed, and combine the covariance matrix and the data mean with an elliptic function to obtain the initial uncertain set; The step of rotating the center of the initial uncertain set to the origin to obtain the data uncertain set includes: The covariance matrix is decomposed into eigenvalues to obtain the eigenvalue matrix and the data orthogonal matrix. Using the eigenvalue matrix and the data orthogonal matrix, the center of the initial uncertain set is rotated to the origin to obtain the data uncertain set.
5. The method of claim 4, wherein, The step of obtaining data poles from the data uncertainty set and forming a pole vector from the data poles includes: N eigenvalues are obtained from the eigenvalue matrix, and pole coordinate transformation is performed on the N eigenvalues respectively to obtain N positive data poles and N negative data poles; N is a positive integer; the N positive data poles and N negative data poles correspond to N feature axes, and each feature axis includes one positive data pole and one negative data pole; The N positive data poles and the N negative data poles are combined to form 2 N pole matrices; each pole matrix includes N data poles, and the N data poles in each pole matrix belong to different characteristic axes; vector transforming the 2 N pole matrices to obtain 2 N pole vectors.
6. The method of claim 4, wherein, The construction of the set mapping model based on the pole vector includes: Obtain the data orthogonal matrix and the data average value, and based on the rotation method of the position of the initial uncertain set, form a mapping parameter from the data orthogonal matrix and the data average value; A set mapping model is constructed based on the pole vector and the mapping parameters.
7. The method of claim 1, wherein, The step of obtaining the set size parameter from the data mapping value corresponding to the second dataset to be processed includes: Sort the data mapping values corresponding to the second dataset to be processed to obtain a sequence of data mapping values; Based on the quantity of first business data included in the second dataset to be processed and the data splitting parameters, a limit determination model is constructed, and the limit determination model is parsed to obtain the data location; The data mapping value located at the data position in the data mapping value sequence is determined as the set size parameter.
8. The method of claim 1, wherein, The construction of the second business prediction dataset based on the set size parameter includes: Obtain the pole vector and mapping parameters from the set mapping model; Based on the set size parameter, the product of the pole vector and the mapping parameter is constrained to obtain the second business prediction dataset.
9. The method of claim 1, wherein, The second business prediction dataset is used to represent the error value of business prediction for a business scenario; The method further includes: For the aforementioned business scenario, a business prediction is performed to obtain a first business prediction result; Obtain the business prediction error value from the second business prediction dataset, and determine the sum of the first business prediction result and the business prediction error value as the second business prediction result for the business scenario.
10. The method of claim 9, wherein, The business scenario is a data center temperature control scenario; the second business prediction result is used to represent the temperature relative to the outside of the data center. The method further includes: The data center temperature control scenario is obtained, including the data center temperature value and temperature change at the first moment, and the delay parameter is obtained. The sum of the second business forecast result and the temperature change is determined as the temperature parameter; Using the delay parameter, the data center temperature value at the first moment is weighted and processed with the temperature parameter to obtain the data center temperature value at the second moment; the first moment is less than the second moment.
11. The method of claim 10, wherein, The acquisition of delay parameters includes: Obtain the time slot length between the first time point and the second time point, and obtain the thermal capacity and thermal resistance of the data center; Based on the time slot length, the heat capacity, and the thermal resistance, the temperature influence data is determined; The temperature effect data is subjected to exponential processing to obtain the delay parameter.
12. A data set construction apparatus characterized by comprising: The device includes: The data splitting module is used to split the first business forecast dataset into a first dataset to be processed and a second dataset to be processed, and to construct a data uncertainty set based on the first dataset to be processed. The pole processing module is used to obtain data poles from the data uncertainty set and form a pole vector from the data poles; The data mapping module is used to construct a set mapping model based on the pole vector and use the set mapping model to obtain the data mapping value corresponding to the second dataset to be processed. The data construction module is used to obtain the set size parameter from the data mapping value corresponding to the second dataset to be processed, and construct a second business prediction dataset based on the set size parameter; the second business prediction dataset is used for business prediction.
13. A computer device, comprising: Includes processor, memory, and input / output interfaces; The processor is connected to the memory and the input / output interface respectively, wherein the input / output interface is used to receive data and output data, the memory is used to store computer programs, and the processor is used to call the computer programs so that the computer device executes the method according to any one of claims 1-11.
14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program adapted to be loaded and executed by a processor to cause a computer device having the processor to perform the method of any one of claims 1-11.
15. A computer program product comprising computer programs / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the method described in any one of claims 1-11.