Procurement forecasting method, apparatus, device, and computer-readable storage medium

By combining multiple regression models and time-series prediction models, the data resource procurement decision-making of the public opinion analysis system is optimized, solving the problems of high bandwidth and accuracy in commercial procurement of autonomous crawling methods, and achieving precise procurement of data resources and ensuring user experience.

CN116051157BActive Publication Date: 2026-04-07CHINA MOBILE (SUZHOU) SOFTWARE TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-27
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

In existing technologies, public opinion analysis systems suffer from problems such as high bandwidth and IP requirements for autonomous web crawling methods and unavoidable information omissions in data acquisition. In commercial procurement, accurately predicting data volume and procurement time to avoid affecting user experience is a challenge.

Method used

By acquiring order data and configuration attribute information for the current ordering cycle, and combining multiple regression models and time series prediction models, the data resource call volume for the next ordering cycle is predicted. Based on the existing call volume and the purchased volume, the predicted purchase volume is determined, and a weighted fusion and condition triggering mechanism is used to optimize the procurement decision.

Benefits of technology

It improved the accuracy of data resource procurement, avoided resource waste and user experience impact caused by procurement too early or too late, and achieved precision and cost-effectiveness in procurement timing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116051157B_ABST
    Figure CN116051157B_ABST
Patent Text Reader

Abstract

This application provides a procurement forecasting method, apparatus, device, and computer-readable storage medium. The method includes: determining a first estimated call volume of data resources in the next ordering cycle based on the ordering data, ordering attribute information, and configuration attribute information of each ordering object in the current ordering cycle; determining a second estimated call volume of data resources in the next ordering cycle based on the historical call volume corresponding to each historical ordering cycle within a preset historical time period; determining a target estimated volume based on the first and second estimated call volumes; and determining a predicted procurement volume of data resources based on the target estimated volume, the historical call volume corresponding to each historical ordering cycle, and the already procured amount of data resources. This improves the accuracy of the target estimated volume, thereby improving the accuracy of the predicted procurement volume. Cyclic forecasting according to the ordering cycle avoids waste caused by premature procurement and disruption to normal use by ordering users due to late procurement, thus improving the accuracy of procurement timing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of big data processing technology, and includes, but is not limited to, a procurement forecasting method, apparatus, equipment, and computer-readable storage medium. Background Technology

[0002] A public opinion analysis system should possess the capabilities of webpage content crawling and semantic analysis, enabling real-time monitoring and in-depth analysis of relevant public opinion on the internet. This provides analysts with a comprehensive understanding of public opinion dynamics and a basis for making correct public opinion guidance. The data sources for such systems primarily utilize open-source online data, generally including websites, public accounts, microblogs, forums, blogs, social networking sites, etc. Data acquisition is a technical challenge, and the acquired data is also a key indicator for evaluating the quality of a public opinion analysis product.

[0003] There are two main methods for acquiring data in related technologies: one is to use self-crawling to acquire data. Due to the massive amount of data collected by public opinion analysis systems, self-crawling requires high bandwidth and a large number of Internet Protocol (IP) connections, necessitating the purchase of proxy IPs to circumvent network blocking, and technically, it cannot prevent information omissions. The other method is to acquire data through commercial procurement. Commercial procurement can reduce technical investment costs, ensure data accuracy and availability, improve information coverage, and overcome the shortcomings of crawling technology. However, determining the appropriate amount of data to purchase and the timing of the purchase without affecting the user experience of the public opinion product remains a significant challenge. Summary of the Invention

[0004] In view of this, this application provides a procurement forecasting method, apparatus, equipment, and computer-readable storage medium, which at least solves the problem of low accuracy in predicting procurement quantities when purchasing data resources.

[0005] The technical solution of this application embodiment is implemented as follows:

[0006] At least one embodiment of this application provides a procurement forecasting method, the method comprising:

[0007] Obtain the order data, order attribute information, and configuration attribute information of each order object within the current order period, and obtain the historical call volume of data resources for each historical order period within a preset historical time period;

[0008] Based on the order data, order attribute information and configuration attribute information of each ordering object, determine the first estimated call volume of data resources in the next ordering cycle, and based on the historical call volume corresponding to each historical ordering cycle, determine the second estimated call volume of data resources in the next ordering cycle.

[0009] Based on the first estimated call volume and the second estimated call volume, determine the target estimated amount of data resources in the next ordering cycle;

[0010] Obtain the purchased quantity of data resources, and determine the predicted purchase quantity of data resources based on the target estimated quantity, the historical call volume corresponding to each historical ordering period, and the purchased quantity.

[0011] Furthermore, according to at least one embodiment of this application, determining the first estimated usage volume of data resources in the next ordering cycle based on the ordering data, ordering attribute information, and configuration attribute information of each ordering object includes:

[0012] Obtain a pre-trained multivariate regression model;

[0013] Based on the order data of each ordering object, the order attribute information, and the configuration attribute information, the parameter values ​​of each input parameter in the multiple regression model are determined;

[0014] The parameter values ​​of each input parameter are input into the trained multivariate regression model to obtain the first estimated amount of data resources to be used in the next ordering cycle.

[0015] Furthermore, according to at least one embodiment of this application, determining the parameter values ​​of each input parameter in the multiple regression model based on the order data of each ordering object, the order attribute information, and the configuration attribute information includes:

[0016] Based on the order data, order attribute information and configuration attribute information of each ordering object, the ordering users corresponding to each ordering object are classified to obtain multiple classification results;

[0017] Based on the number of subscribers and the total number of subscribers included in each category, determine the percentage of subscribers in each category;

[0018] Based on the order attribute information and the configuration attribute information, determine the average usage rate of each order attribute included in the order attribute information;

[0019] Based on the attribute values ​​of each configuration attribute included in the configuration attribute information, determine the set of attribute values ​​corresponding to each configuration attribute, and obtain the number of attribute values ​​in each set of attribute values;

[0020] The percentage of users subscribing to each category, the average usage rate of each subscription attribute, and the number of attribute values ​​in each attribute value set are determined as the parameter values ​​of each input parameter in the multiple regression model.

[0021] Furthermore, according to at least one embodiment of this application, the step of classifying the ordering users corresponding to each ordering object based on the ordering data, ordering attribute information, and configuration attribute information of each ordering object to obtain multiple classification results includes:

[0022] The usage saturation of each ordering object is determined based on the ordering attribute information and the configuration attribute information of each ordering object;

[0023] Based on the usage saturation of each subscription object, the subscribers corresponding to each subscription object are classified, resulting in multiple classification results.

[0024] Furthermore, according to at least one embodiment of this application, the step of classifying the ordering users corresponding to each ordering object based on the ordering data, ordering attribute information, and configuration attribute information of each ordering object to obtain multiple classification results includes:

[0025] Determine whether the order attribute information of each ordering object includes additional order attributes;

[0026] The ordering users corresponding to the ordering objects that include additional ordering attributes are classified as the first category result;

[0027] Subscribers whose subscriptions do not include additional subscription attributes are categorized as the second category result.

[0028] Furthermore, according to at least one embodiment of this application, the order attribute information includes an order attribute and the attribute value of the order attribute, and the configuration attribute information includes a configuration attribute and the attribute value of the configuration attribute;

[0029] Based on the order attribute information and the configuration attribute information, determine the average usage rate of each order attribute included in the order attribute information, including:

[0030] Based on the attribute values ​​of each ordering attribute and each configuration attribute in each ordering object, determine the usage rate of each ordering attribute included in each ordering object.

[0031] The average usage rate of each subscription attribute is determined by averaging the usage rates of all subscription objects.

[0032] Furthermore, according to at least one embodiment of this application, determining the second estimated call volume of data resources in the next ordering period based on the historical call volume corresponding to each historical ordering period includes:

[0033] Obtain a pre-trained time series prediction model;

[0034] The historical call volume corresponding to each historical ordering period is input into the trained time series prediction model to obtain the second estimated call volume of data resources in the next ordering period.

[0035] Furthermore, according to at least one embodiment of this application, determining the predicted purchase quantity of data resources based on the target estimated quantity, the historical call volume corresponding to each historical ordering period, and the purchased quantity includes:

[0036] Based on the target estimated amount and the historical call amount corresponding to each historical ordering period, the predicted call amount of data resources is determined;

[0037] When the procurement conditions are met based on the predicted call volume and the already procured volume, the predicted procurement volume is determined based on the data source.

[0038] When it is determined that the procurement conditions have not been met based on the predicted call volume and the already procured volume, the predicted procurement volume is set as the first value.

[0039] Furthermore, according to at least one embodiment of this application, determining the predicted purchase quantity based on the data source includes:

[0040] When the data source is a first commercial data source, the predicted purchase quantity is determined as a second value, so as to purchase data resources with the predicted purchase quantity as the second value from the first commercial data source;

[0041] When the data source is a second commercial data source, determine whether the total number of subscribers has reached a preset user threshold.

[0042] When the total number of subscribers reaches a preset user threshold, the difference between the total amount of data resources in the second commercial data source and the amount already purchased is determined as the predicted purchase amount, so as to purchase all remaining data resources in the second commercial data source.

[0043] When the total number of subscribers does not reach the preset user threshold, the predicted purchase volume is determined as a third value, so that data resources with the predicted purchase volume of the third value can be purchased from the second business data source.

[0044] At least one embodiment of this application provides a procurement forecasting apparatus, the apparatus comprising:

[0045] The first acquisition module is used to acquire the order data, order attribute information and configuration attribute information of each order object in the current order period, and to acquire the historical call volume of data resources in each historical order period within a preset historical time period.

[0046] The first determining module is used to determine the first estimated call volume of data resources in the next ordering cycle based on the ordering data, ordering attribute information and configuration attribute information of each ordering object, and to determine the second estimated call volume of data resources in the next ordering cycle based on the historical call volume corresponding to each historical ordering cycle.

[0047] The second determining module is used to determine the target estimated amount of data resources in the next ordering cycle based on the first estimated call volume and the second estimated call volume.

[0048] The second acquisition module is used to acquire the purchased quantity of data resources;

[0049] The third determining module is used to determine the predicted purchase quantity of data resources based on the target estimated quantity, the historical call quantity corresponding to each historical ordering period, and the purchased quantity.

[0050] At least one embodiment of this application provides a procurement forecasting device, comprising:

[0051] Processor; and

[0052] Memory for storing computer programs that can run on the processor;

[0053] The computer program, when executed by a processor, implements the steps of the above-described procurement forecasting method.

[0054] At least one embodiment of this application provides a computer-readable storage medium storing computer-executable instructions configured to perform the steps of the above-described procurement forecasting method.

[0055] This application provides a procurement forecasting method, apparatus, device, and computer-readable storage medium. The method includes: acquiring order data, order attribute information, and configuration attribute information of each ordering object in the current ordering period, and acquiring the historical call volume of data resources in each historical ordering period within a preset historical time period; determining a first estimated call volume of data resources in the next ordering period based on the order data, order attribute information, and configuration attribute information of each ordering object, and determining a second estimated call volume of data resources in the next ordering period based on the historical call volume corresponding to each historical ordering period; determining a target estimated volume of data resources in the next ordering period based on the first estimated call volume and the second estimated call volume; acquiring the already procured quantity of data resources, and determining the predicted procurement quantity of data resources based on the target estimated volume, the historical call volume corresponding to each historical ordering period, and the already procured quantity. Thus, by analyzing the ordering and usage data of the current ordering period, the first estimated usage volume of data resources within a single period is predicted. By analyzing the usage data from multiple historical ordering periods, the second estimated usage volume of data resources within consecutive ordering periods is predicted. The first and second estimated usage volumes are then combined to obtain the target estimated usage volume of data resources for the next ordering period. This improves the accuracy of the target estimated usage volume for data resources in the next ordering period. Consequently, when determining the predicted purchase volume of data resources based on the target estimated usage volume, historical usage volumes, and already purchased quantities, the accuracy of the predicted purchase volume is improved. Furthermore, cyclically predicting the predicted purchase volume for the next ordering period according to the ordering period avoids waste caused by purchasing data resources too early and also avoids disruption to normal use by ordering data resources too late, thereby improving the accuracy of the purchase timing. Attached Figure Description

[0056] In the accompanying drawings (which are not necessarily drawn to scale), similar reference numerals may describe similar parts in different views. The drawings illustrate, by way of example and not limitation, the various embodiments discussed herein.

[0057] Figure 1 A schematic diagram illustrating an implementation process of the procurement forecasting method provided in this application embodiment;

[0058] Figure 2 A schematic diagram illustrating an implementation process for determining the first and second estimated call volumes of data resources in the procurement forecasting method provided in this application embodiment;

[0059] Figure 3 A schematic diagram illustrating an implementation process for determining the predicted procurement quantity of data resources in the procurement forecasting method provided in this application embodiment;

[0060] Figure 4 A schematic diagram of the public opinion data resource procurement prediction method provided in the embodiments of this application;

[0061] Figure 5 This is a schematic diagram of the composition structure of the procurement forecasting device provided in the embodiments of this application;

[0062] Figure 6 This is a schematic diagram of the composition structure of the procurement forecasting equipment provided in the embodiments of this application. Detailed Implementation

[0063] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0064] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0065] In the following description, the terms "first, second, third" are used merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first, second, third" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.

[0066] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0067] To address the problems existing in related technologies, this application provides a procurement forecasting method, which is applied to a procurement forecasting device. The method provided in this application can be implemented by a computer program, which, when executed, completes each step of the procurement forecasting method provided in this application. In some embodiments, the computer program can be executed by a processor in the procurement forecasting device. Figure 1 This is a schematic diagram illustrating an implementation process of the procurement forecasting method provided in an embodiment of this application, such as... Figure 1 As shown, this procurement forecasting method includes the following steps:

[0068] Step S101: Obtain the order data, order attribute information and configuration attribute information of each order object in the current order period, and obtain the historical call volume of data resources in each historical order period within the preset historical time period.

[0069] The method provided in this application embodiment can be executed by a procurement forecasting device, which can be a user equipment (UE), mobile device, terminal, laptop, tablet, desktop computer, or other device capable of procurement forecasting.

[0070] The procurement forecasting method provided in this application is mainly applied to data resource management systems, such as public opinion analysis systems. The following explanation uses a public opinion analysis system as an example to illustrate the procurement forecasting method provided in this application. A public opinion analysis system should possess webpage content crawling and semantic analysis capabilities, enabling real-time monitoring and in-depth analysis of relevant public opinion on the internet. This provides a basis for public opinion analysts to comprehensively grasp public opinion dynamics and make correct public opinion guidance, and provides public opinion information to users of public opinion products. Due to the massive data collection volume of public opinion analysis systems, the self-crawling method has high bandwidth and IP quantity requirements, necessitating the purchase of proxy IPs to cope with network blocking. Furthermore, from a technical perspective, it is difficult to avoid information omissions. To improve information coverage, this application uses commercially procured data to address the shortcomings of crawling technology. Commercial procurement can reduce technical investment costs and ensure data accuracy and availability. Through the procurement forecasting method provided in this application, it is possible to accurately predict when and how much data resources need to be procured, so as not to affect users' normal use of public opinion products.

[0071] In this embodiment of the application, the data source for the data resources procured by the public opinion analysis system can be open-source data from the Internet, generally including news websites, government websites, portal websites, public accounts, microblogs, various forums, blogs, social networking sites, etc.

[0072] Here, "subscription object" refers to the public opinion analysis products sold by the public opinion analysis system; "subscription data" refers to the sales data of the public opinion products; and "subscription period" refers to the usage period of the public opinion products. The subscription attribute information of the subscription object includes subscription attributes and their attribute values. Subscription attribute information refers to information about the public opinion products sold; subscription attributes refer to the indicators included in the public opinion products purchased by users; and the attribute values ​​of subscription attributes refer to the content of the indicators. The configuration attribute information of the subscription object includes configuration attributes and their attribute values. Configuration attribute information refers to the actual usage information of the public opinion products purchased by users; configuration attributes refer to the indicators actually used by users; and the attribute values ​​of configuration attributes refer to the content of the indicators actually used.

[0073] Subscription attributes include configuration attributes and the attribute value of the subscription attribute being greater than or equal to the value of the configuration attribute. For example, a user subscribes to a package to enjoy public opinion monitoring services, with subscription attributes such as the package including 5 topics and 100 keywords. In actual use, the user may configure 3 topics and 30 keywords. Under the highest usage scenario, the user can only configure a maximum of 5 topics and 100 keywords. If this is still insufficient, the user needs to subscribe to other subscription attributes and public opinion monitoring products with higher subscription attribute values ​​(i.e., upgrade the package).

[0074] After a user completes their order for the desired items within the current ordering period, the procurement forecasting device retrieves the ordering data, ordering attribute information, and configuration attribute information for each item within that period. It also retrieves the historical data resource usage volume for each historical ordering period within a preset historical timeframe. The ordering period can be in months, and the preset historical timeframe can be the predicted duration, such as n ordering periods. The device retrieves the historical data resource usage volume for each month within the current ordering period and the previous n-1 ordering periods. Here, the historical usage volume refers to the amount of data resources actually used by all ordering users. Let's assume that the historical usage volume for each of the i ordering periods can be denoted as y. t-i , i = 1, 2, ..., n, where i = 1 refers to the current ordering period.

[0075] Step S102: Based on the order data, order attribute information and configuration attribute information of each ordering object, determine the first estimated call volume of data resources in the next ordering cycle, and based on the historical call volume corresponding to each historical ordering cycle, determine the second estimated call volume of data resources in the next ordering cycle.

[0076] The first estimated call volume is determined based on the order data, order attribute information, and configuration attribute information of each ordering object within the current ordering period. That is, within a single time period, based on information about all ordering objects ordered by all subscribers within the current ordering period, the amount of data resources expected to be called up in the next ordering period is estimated. The second estimated call volume is predicted in time series based on historical call volumes.

[0077] Step S103: Based on the first estimated call volume and the second estimated call volume, determine the target estimated amount of data resources for the next ordering cycle.

[0078] Here, the first estimated call volume is the data resource call volume estimated based on the ordering and usage data of the current subscription period, and the second estimated call volume is the data resource call volume estimated based on historical call volumes over multiple historical subscription periods. The first and second estimated call volumes are then combined to obtain the target estimated volume of data resources to be called in the next subscription period.

[0079] In practical applications, assuming the current order period is the (t-1)th order period, and the predicted span is n order periods, let y denote the first estimated call volume for the t-th order period. 1,t The estimated call volume for the second ordering cycle in the t-th cycle is denoted as y. 2,t A weighted fusion method can be used to fuse the data resources to obtain the target estimated quantity y′ of the data resources to be called in the t-th ordering cycle. t =c1y 1,t +c2y 2,t Where c1 and c2 are the weights of the first estimated call volume and the second estimated call volume, respectively, and c1+c2=1.

[0080] Step S104: Obtain the purchased quantity of data resources.

[0081] Each time data resources are procured, the public opinion analysis system generates a procurement log or stores the procurement information in a data table. When predicting the procurement volume for the next ordering cycle, the procurement volume of data resources can be obtained by querying the procurement log or the procurement data table.

[0082] Step S105: Based on the target estimated quantity, the historical call quantity and the purchased quantity corresponding to each historical ordering period, determine the predicted purchase quantity of data resources.

[0083] The procurement of data resources for the public opinion analysis system is predicted and initiated in advance to avoid affecting data access for ordering users. Here, the procured quantity is denoted as G, and the estimated target quantity is y′. t Historical call volume y corresponding to each historical ordering period i (where i = t-1, t-2, ..., tn) and the already purchased quantity G, determine the predicted purchase quantity of data resources.

[0084] In one implementation, historical call volume can be accumulated according to the ordering cycle after the last purchase. When the accumulated historical call volume reaches a certain proportion λ of the purchased volume G, the next purchase action is triggered, that is: when y′ t +y t-1 +…+y t-n When the target is estimated, the predicted purchase quantity of data resources is determined based on the target estimated quantity, the historical call quantity and the purchased quantity corresponding to each historical ordering period, and the purchase behavior is triggered, such as triggering the sending of a purchase request to the data source server, which is a server that can provide data resources.

[0085] When the cumulative amount of historical calls does not reach a certain proportion λ of the purchased amount G, it indicates that the data resources purchased last time are sufficient for the next call. At this time, the predicted purchase amount of data resources can be set to 0, and the procurement prediction step can continue to be repeated in the next ordering cycle.

[0086] The procurement forecasting method provided in this application includes: obtaining order data, order attribute information, and configuration attribute information of each ordering object in the current ordering period, and obtaining the historical call volume of data resources in each historical ordering period within a preset historical time period; determining a first estimated call volume of data resources in the next ordering period based on the order data, order attribute information, and configuration attribute information of each ordering object, and determining a second estimated call volume of data resources in the next ordering period based on the historical call volume corresponding to each historical ordering period; determining a target estimated volume of data resources in the next ordering period based on the first estimated call volume and the second estimated call volume; obtaining the purchased quantity of data resources, and determining the predicted purchase quantity of data resources based on the target estimated volume, the historical call volume corresponding to each historical ordering period, and the purchased quantity. By analyzing the ordering and usage data of the current ordering period, the first estimated usage volume of data resources within a single period is predicted. Then, by analyzing the usage data from multiple historical ordering periods, the second estimated usage volume of data resources within consecutive ordering periods is predicted. Finally, by combining the first and second estimated usage volumes, the target estimated usage volume of data resources for the next ordering period is obtained. This improves the accuracy of the target estimated usage volume for the next ordering period. Consequently, when determining the predicted purchase volume of data resources based on the target estimated usage volume, historical usage volumes, and already purchased quantities, the accuracy of the predicted purchase volume is improved. Furthermore, cyclically predicting the predicted purchase volume for the next ordering period according to the ordering period avoids waste caused by purchasing data resources too early and also avoids disruption to normal use by ordering data resources too late, thereby improving the accuracy of the purchase timing.

[0087] In some embodiments, the above Figure 1 In step S102 of the illustrated embodiment, "determining the first estimated data resource usage in the next ordering cycle based on the ordering data, ordering attribute information, and configuration attribute information of each ordering object" can be achieved through... Figure 2 Steps S1021 to S1023 shown are used to achieve the following:

[0088] Step S1021: Obtain the pre-trained multivariate regression model.

[0089] Using the ordering period as a time window, ordering data, ordering attribute information, and configuration attribute information of each ordering object in multiple historical ordering periods are obtained as training samples to construct an initial multiple regression model. The initial multiple regression model is then fitted and trained based on the sample data to obtain a trained multiple regression model.

[0090] For example, Table 1 is a table of order attribute information for the ordering object:

[0091] Table 1 Order Attribute Information of Ordering Objects

[0092]

[0093] Table 2 shows the usage information of the subscription object ordered by user A during this historical subscription period, i.e., the configuration attribute information of the subscription object:

[0094] Table 2 Configuration attribute information of the subscription object ordered by user A

[0095]

[0096]

[0097] Based on the attributes in Tables 1 and 2, an initial multiple regression model was constructed, and the parameter values ​​of each parameter were calculated based on the attribute values ​​in Tables 1 and 2. The parameter data table of the multiple regression model is shown in Table 3.

[0098] Table 3. Parameter Data Table for Multiple Regression Model

[0099]

[0100] In Table 3, when calculating the percentage of users, users can be categorized based on their usage saturation of the subscribed objects. The ratio of the number of subscribed objects in each category to the total number of subscribed objects is then calculated; this ratio represents the percentage of users in that category. The average usage rate can be calculated by first calculating the usage rate of each subscribed object, and then averaging the usage rates of all subscribed objects. The number of deduplicated configuration attributes can be determined by grouping the configuration attributes of all subscribed objects into a set; the number of elements in this set represents the number of deduplicated configuration attributes. Based on the attribute values ​​of the subscribed and configuration attributes of each subscribed object, and the subscription data, the parameter values ​​x for each parameter in each historical subscription period in Table 3 are calculated. The historical call volume y for each historical subscription period is obtained. Based on the historical call volume y and the parameter values ​​x for each parameter in each historical subscription period, the initial multiple regression model is fitted and trained to obtain a trained multiple regression model.

[0101] For example, selecting a month as the basic time window, extract historical order data, order attribute information, configuration attribute information, and the monthly data resource call volume (i.e., historical call volume) for each month within two consecutive years. Calculate the parameter values ​​based on the extraction results, then fit a multiple regression model, denoted as y = b0 + b1x1 + b2x2 + ... + b j x j +ε, where y is the number of historical calls, j is the number of parameters, and b k (k = 1, 2, ..., j) are the model coefficients (i.e., the fitting coefficients), x k(k = 1, 2, ..., j) are model parameters (also called independent variables), b0 is a constant term, representing the part that is not explained by the independent variables and exists for a long time (non-random), i.e., information residue, and ε is a random error term (also called white noise), which is the error between the predicted value and the actual value after removing the constant term in the interpretation space of the independent variables.

[0102] Step S1022: Determine the parameter values ​​of each input parameter in the multiple regression model based on the order data, order attribute information, and configuration attribute information of each ordering object.

[0103] Based on the order data, order attribute information and configuration attribute information of each ordering object within the current ordering period obtained in step S101, calculate the parameter values ​​of each parameter within the current ordering period.

[0104] Step S1023: Input the parameter values ​​of each input parameter into the trained multivariate regression model to obtain the first estimated call volume of data resources in the next ordering cycle.

[0105] Input the parameter values ​​(x1, x2, ..., xj) of each parameter within the current ordering period into the trained multiple regression model y = b0 + b1x1 + b2x2 + ... + b j x j +ε, to obtain the first estimated data resource usage y in the next ordering cycle. 1,t .

[0106] The method provided in this application predicts the first estimated call volume of data resources within a single period based on the ordering and usage of the ordering objects in the current ordering period.

[0107] In some embodiments, step S1022, "determining the parameter values ​​of each input parameter in the multiple regression model based on the order data, order attribute information, and configuration attribute information of each ordering object," can be implemented as follows:

[0108] Step S10221: Based on the order data, order attribute information and configuration attribute information of each ordering object, classify the ordering users corresponding to each ordering object to obtain multiple classification results.

[0109] In this embodiment, subscribers can be categorized based on the saturation of each subscription object. The higher the saturation of a subscriber's usage, the greater their demand for data resources in the next subscription cycle. One implementation of step S10221, which involves categorizing subscribers, is as follows: determining the usage saturation of each subscription object based on its subscription attribute information and configuration attribute information; and then categorizing the subscribers corresponding to each subscription object based on their usage saturation, resulting in multiple categorization results.

[0110] For example, User A subscribes to a basic subscription with a maximum of 5 topics, meaning the "topic" attribute has a value of 5. However, 3 topics are actually used, meaning the "topic" attribute has a value of 3. The usage rate of the "topic" subscription attribute is 0.6 (3 / 5). The usage rate of all subscription attributes is calculated, and the usage saturation of the subscription is determined based on the usage rates of all subscription attributes. Taking the average usage rate of all subscription attributes as an example, according to the data in Tables 1 and 2 above, the usage saturation of User A's subscription is calculated as (3 / 5 + 30 / 100 + 15 / 20 + 4 / 5 + 2 / 5 + 4 / 5 + 200 / 200 + 7 / 10 + 5 / 5) / 9 = 6.35 / 9 = 0.71. Therefore, the usage saturation of User A's subscription is 0.71. Calculate the usage saturation of all users' subscriptions in the current period in the same way, and classify all subscribers according to the size of the usage saturation. For example, subscribers with a usage saturation greater than 0.7 are classified as Class I users, subscribers with a usage saturation greater than 0.3 and less than or equal to 0.7 are classified as Class II users, and subscribers with a usage saturation less than or equal to 0.3 are classified as Class III users.

[0111] Another implementation of step S10221 for classification can be as follows: User needs are classified based on an improved Euclidean distance user classification method. The differences in user needs are represented by the user's subscription to a service package and the actual use of services within the package. Classification of the public opinion product user group is achieved using k-means unsupervised clustering based on the improved Euclidean distance. The classification of the public opinion product user group is predicted using the improved Euclidean distance, and the distance formula is shown in equation (1):

[0112]

[0113] Where r i , These represent the user's subscribed public opinion monitoring product package service items and the user's actual usage of the service items, respectively. α i This indicates the correlation between the weight of each sub-item and the package price, which is normalized to a sum of 1.

[0114] In another implementation, subscribers can be classified according to the subscription attributes of each subscription object. Step S10221 can be implemented as follows: determine whether the subscription attribute information of each subscription object includes additional subscription attributes; classify subscribers corresponding to subscription objects that include additional subscription attributes as the first classification result; classify subscribers corresponding to subscription objects that do not include additional subscription attributes as the second classification result.

[0115] As in the example above, if user A has not subscribed to any additional subscription attributes, then user A is classified into the second category. By judging the subscription attributes of all subscribed users, we can obtain the first category result and the second category result.

[0116] The first and second classification methods mentioned above share the same principle: classifying users based on usage rate. They differ only in their calculation methods, without affecting the classification results. The third classification method differs from the first and second, as it classifies users based on different indicators. In practical applications, multiple different classification indicators can coexist, as shown in numbers 1-3 and 4 in Table 3. Parameters from different classification indicators can simultaneously serve as parameters for a multiple regression model.

[0117] Step S10222: Determine the percentage of each category of subscribers based on the number of subscribers and the total number of subscribers included in each category.

[0118] In the first implementation method mentioned above, if the total number of subscribers is 1000, there are 300 users of type I, 400 users of type II, and 300 users of type III. The calculated percentage of users of type I is 0.3 (300 / 1000), the percentage of users of type II is 0.4 (400 / 1000), and the percentage of users of type III is 0.3 (300 / 1000).

[0119] In the second implementation method mentioned above, the total number of subscribers remains unchanged at 1000. If there are 490 subscribers who subscribed to additional attributes and 510 subscribers who did not subscribe to additional attributes, then the percentage of subscribers who subscribed to additional attributes is 0.49 (=490 / 1000), and the percentage of subscribers who did not subscribe to additional attributes is 0.51 (=510 / 1000).

[0120] Step S10223: Determine the average usage rate of each order attribute included in the order attribute information based on the order attribute information and configuration attribute information.

[0121] The order attribute information includes order attributes and their attribute values, while the configuration attribute information includes configuration attributes and their attribute values. Determining the average usage rate of each order attribute can be achieved as follows: based on the attribute values ​​of each order attribute and each configuration attribute in each order object, determine the usage rate of each order attribute included in each order object; the average usage rate of each order attribute included in all order objects is then determined as the average usage rate of each order attribute. Alternatively, first calculate the usage rate of each order attribute for each order object, then calculate the average usage rate of the same order attribute across all order objects to obtain the average usage rate of that order attribute.

[0122] Step S10224: Based on the attribute values ​​of each configuration attribute included in the configuration attribute information, determine the set of attribute values ​​corresponding to each configuration attribute, and obtain the number of attribute values ​​in each set of attribute values.

[0123] Here, the configuration attributes of all subscribed objects can be grouped into a set. The number of this set is the number of attribute values, which is also the number of deduplicated configuration attributes. For the same configuration attribute, the data resources called are the same. Calculating the number of deduplicated configuration attributes here can avoid counting the same data resources subscribed to by different users repeatedly. The deduplication operation can improve the accuracy of the determined amount of data resources called.

[0124] Step S10225: The percentage of users ordering in each category, the average usage rate of each ordering attribute, and the number of attribute values ​​in each attribute value set are determined as the parameter values ​​of each input parameter in the multiple regression model.

[0125] The data calculated in steps S10221 to S10225 are used as the parameter values ​​of each input parameter and input into the multivariate regression model trained in step S1021 to obtain the first estimated call volume of data resources in the next ordering cycle.

[0126] In some embodiments, the above Figure 1 In step S102 of the illustrated embodiment, "determining the second estimated call volume of data resources in the next ordering period based on the historical call volume corresponding to each historical ordering period" can be achieved through... Figure 2 Steps S1024 and S1025 shown are used to achieve the following:

[0127] Step S1024: Obtain the pre-trained time series prediction model.

[0128] The historical call volume for each historical ordering period is obtained, and the historical call volume for each historical ordering period is used for fitting training to obtain a trained time series prediction model. The longer the historical ordering period, the larger the sample data volume, and the more parameters the trained time series prediction model has, resulting in more accurate predictions of the call volume for the next ordering period. In this embodiment, the trained time series prediction model can be denoted as y = w1y. t-1 +w2y t-2 +…+w p y t-p where w1+w2+…+w p =1, where p is the number of historical ordering periods. Generally, the closer the period is to the current ordering period, the greater its impact on the next ordering period. Therefore, the corresponding weight in the trained time series prediction model is larger, i.e., w1 > w2 > ... > w p .

[0129] Step S1025: Input the historical call volume corresponding to each historical ordering period into the trained time series prediction model to obtain the second estimated call volume of data resources in the next ordering period.

[0130] Input the historical call volume y for each historical ordering period into the trained time series prediction model y = w1y t-1 +w2y t-2 +…+w p y t-p The second estimated call volume y is obtained. 2,t .

[0131] The prediction span period can be a preset value n, where n is a positive integer not greater than p. The larger n is, the longer the prediction span period. The second estimated call quantity y is... 2,t The more accurate the prediction, the greater the computational load, and the more memory is needed to store historical subscription cycles. In practical applications, an appropriate prediction cycle span can be chosen based on the specific circumstances.

[0132] The method provided in this application predicts a second estimated call volume of data resources within a consecutive ordering period based on the call data from multiple historical ordering periods. By fusing the first and second estimated call volumes to obtain the target estimated volume of data resources in the next ordering period, the accuracy of the target estimated volume of data resources in the next ordering period can be improved. Therefore, when determining the predicted purchase volume of data resources based on the target estimated volume, each historical call volume, and the already purchased volume, the accuracy of the predicted purchase volume can be improved.

[0133] In some embodiments, the above Figure 1 Step S105 of the illustrated embodiment, "Determining the predicted purchase quantity of data resources based on the target estimated quantity, the historical call quantity and the purchased quantity corresponding to each historical ordering period," can be achieved through... Figure 3 Steps S1051 to S1054 shown are used to achieve the following:

[0134] Step S1051: Based on the target estimated quantity and the historical call quantity corresponding to each historical ordering period, determine the predicted call quantity of data resources.

[0135] Step S1052: Determine whether the procurement conditions have been met based on the predicted call volume and the purchased volume.

[0136] When the procurement conditions are met based on the predicted call volume and the purchased volume, proceed to step S1053; when the procurement conditions are not met based on the predicted call volume and the purchased volume, proceed to step S1054.

[0137] In one implementation, historical call volume can be accumulated according to the ordering cycle after the last purchase. When the accumulated historical call volume reaches a certain proportion λ of the purchased volume G, the next purchase action is triggered, that is: when y′ t +y t-1 +…+y t-n When the target is estimated, the predicted purchase quantity of data resources is determined based on the target estimated quantity, the historical call quantity and the purchased quantity corresponding to each historical ordering period, and the purchase behavior is triggered, such as triggering the sending of a purchase request to the data source server, which is a server that can provide data resources.

[0138] When the cumulative amount of historical calls does not reach a certain proportion λ of the purchased amount G, it indicates that the data resources purchased last time are sufficient for the next call. At this time, the predicted purchase amount of data resources can be set to 0, and the procurement prediction step can continue to be repeated in the next ordering cycle.

[0139] Step S1053: Determine the predicted purchase quantity based on the data source.

[0140] Different data sources may have different procurement methods. Here, the procurement method can be determined based on information such as procurement costs and procurement contracts. For example, if the procurement method for the first commercial data source is to purchase a fixed amount of data resources each time, then when the data source is the first commercial data source, the predicted procurement quantity is determined as the second value, and data resources with the predicted procurement quantity of the second value are procured from the first commercial data source. As another example, if the procurement method for the second commercial data source is to purchase a fixed amount of data resources each time when the number of subscribers is small, and then purchase all data resources at once when the number of subscribers reaches a preset user threshold, then when the data source is the second commercial data source, it is determined whether the total number of subscribers has reached the preset user threshold; when the total number of subscribers has reached the preset user threshold, the difference between the total amount of data resources in the second commercial data source and the already procured quantity is determined as the predicted procurement quantity, and all remaining data resources in the second commercial data source are procured; when the total number of subscribers has not reached the preset user threshold, the predicted procurement quantity is determined as the third value, and data resources with the predicted procurement quantity of the third value are procured from the second commercial data source.

[0141] Step S1054: Determine the predicted purchase quantity as the first value.

[0142] When it is determined that the procurement conditions have not been met based on the predicted call volume and the purchased volume, the preset procurement volume is set to the first value. This first value is less than the second value and less than the third value. In practical applications, when the predicted data resources are sufficient for the ordering user to call in the next ordering cycle, procurement can be skipped, that is, the first value is set to 0.

[0143] The method provided in this application determines whether the procurement conditions have been met based on the predicted call volume and the already procured volume. This avoids waste caused by premature data resource procurement and also prevents disruption to normal use by subscribing users due to delayed procurement, thereby improving the accuracy of procurement timing. Furthermore, determining the predicted procurement volume based on the data source during procurement enables diversified data collection that differentiates between data sources.

[0144] The following will describe an exemplary application of the embodiments of this application in a real-world application scenario.

[0145] A professional public opinion analysis system (hereinafter referred to as a public opinion analysis system) should have the ability to crawl web page content and perform semantic analysis, and conduct real-time monitoring and in-depth analysis of relevant public opinion on the Internet, so as to provide a basis for public opinion analysts to fully grasp the dynamics of public opinion and make correct public opinion guidance.

[0146] The data sources for public opinion analysis systems mainly refer to open-source online data, generally including news websites, government websites, portal websites, public accounts, microblogs, various forums, blogs, social networking sites, etc. Data acquisition is a technical challenge and a key indicator for evaluating the quality of public opinion analysis products, and is usually achieved through self-developed web crawling.

[0147] Because public opinion analysis systems collect massive amounts of data, self-developed web crawling methods require high bandwidth and a large number of IP addresses, necessitating the purchase of proxy IPs to circumvent network blocking. Furthermore, the possibility of data omissions is technically unavoidable. To improve information coverage, commercially purchased data is used to address the shortcomings of web crawling technology. Commercial procurement can reduce technical investment costs and ensure data accuracy and availability; however, determining the appropriate amount of data to purchase and the timing of such purchases to avoid impacting the user experience of online products remains a significant challenge.

[0148] Regarding resource procurement forecasting methods in the field of big data, the main related technologies include the following:

[0149] 1) A machine learning-based procurement quantity forecasting method is used to obtain daily sales reports of various commodities, construct a linear model between the procurement forecast value and the historical average sales data, train the linear model with the historical average sales data as sample data, optimize the linear parameters, determine the linear model using the optimized linear parameters, and obtain the procurement forecast value by inputting the sales data of the day.

[0150] 2) Based on the Industrial Intelligent Cloud System (INDICS), and using the SMART IoT gateway (SMART Internet of Things), real-time data acquisition and monitoring are carried out, and industrial big data analysis technology is applied to achieve procurement demand forecasting.

[0151] 3) A method for predicting the demand for material procurement based on recurrent neural networks (RNNs) is proposed. This method involves constructing and training a material procurement prediction model based on an RNN network, considering the temporal relationship between historical procurement volumes, and increasing or decreasing the level according to the time series, so that the pattern can be repeated continuously over time.

[0152] In addition, there are other material procurement forecasting methods based on Long Short-Term Memory (LSTM) networks and Gate Recurrent Unit (GRU) networks. However, the existing methods mentioned above either only consider historical procurement trends without correlating them with sales data at the time of forecasting; or, those integrated with public opinion analysis products or systems either focus on monitoring the public opinion analysis system or on improving the accuracy of public opinion analysis through algorithm optimization, lacking a data resource procurement forecasting method based on public opinion products.

[0153] Current data-driven procurement forecasting methods used in fields such as financial operations and industrial technology have the following main shortcomings:

[0154] 1) The metrics collected in the sales process generally include sales revenue, input costs, and average sales volume, but they ignore an important factor in product sales: the target user, and fail to reflect the differences in user needs.

[0155] 2) In terms of the time dimension, some solutions only analyze the relationship between sales and purchases within the current forecast period, using time variables to optimize the parameters of the forecasting model, while others only consider the time-series relationship between historical purchase volumes.

[0156] 3) The forecast results only output the predicted purchase volume, lacking a deep integration with business scenarios.

[0157] The data resource procurement method based on public opinion package content provided in this application fully analyzes the differences in user needs, as well as usage behavior characteristics such as package ordering, configuration of titles and keywords in the package, and procurement data resource patterns. It combines the relationship between sales and procurement in the current forecast period and the trend of procurement volume in continuous time periods to achieve prediction and evaluation of the scale of public opinion product data resource procurement and differentiated decision-making in terms of method.

[0158] This application provides a data resource procurement method based on public opinion product packages. The method includes: 1) classifying user groups using an unsupervised clustering method based on improved Euclidean distance to reflect differentiated user needs; 2) establishing a multiple regression model based on collected indicators within a single time period to fit a procurement volume prediction formula; 3) predicting procurement volume using a weighted average time series prediction method within a continuous time period, and adjusting the procurement volume prediction by integrating the prediction results from the multiple regression model in step 2); and 4) making procurement method decisions based on data resource sales patterns. This method achieves comprehensive understanding of user needs, accurate prediction of procurement volume, and differentiated determination of procurement decisions, thereby saving economic costs.

[0159] The data supply and demand forecasting for public opinion products aims to analyze the relationship between data resource procurement and package subscriptions, user differences, and product promotion scale, providing a reference for subsequent data resource procurement. The data sources procured in this application's embodiment include:

[0160] 1) It includes WeChat official accounts and commercial microblogs as supplementary commercial data sources for procurement, which can effectively prevent omissions and improve public opinion coverage and user satisfaction.

[0161] 2) The amount of data resources purchased is strongly correlated with the user base. Once the resources for a public account reach a certain limit, they can be purchased in one go to serve all users. Commercial microblogs calculate fees based on the amount of data returned by calling the Application Programming Interface (API).

[0162] The public opinion monitoring product offers packages divided into basic and supplementary packages. When the configuration options in the basic package do not meet user needs, users can choose supplementary packages at various price points. Configuration options directly related to data resources include: the number of key individuals in communication programs, the number of key individuals on Weibo, etc. The public opinion monitoring product is priced according to the packages subscribed to by users. Each package allows for the configuration of thematic topics, and each topic can be configured with content keywords, behavioral keywords, protagonist keywords, and regional keywords. Neither the topics nor the keywords can exceed the upper limit. Through the combination of various keywords, and by calling API data interfaces, the product returns procurement data matching the keywords, meeting the service needs of the public opinion monitoring product.

[0163] Table 4. Contents and Examples of Public Opinion Product Packages

[0164]

[0165] After purchasing commercial data resources, users subscribe to packages to enjoy public opinion monitoring services. For example, a user might subscribe to a package that allows them to configure 5 topics and 100 keywords. However, in actual use, users only configure 3 topics, with only 30 keywords per topic, resulting in low package saturation. As the user base grows and individual users adjust their topic and keyword configuration needs, the data volume from the previous purchase is nearing its limit. When the amount of data returned by the API is limited, the public opinion information obtained by users is incomplete, impacting user satisfaction and necessitating a next purchase.

[0166] Figure 4 This is a schematic diagram of the public opinion data resource procurement prediction method provided in the embodiments of this application, such as... Figure 4 As shown, the method includes the following steps:

[0167] Step S401: Obtain user's product order and usage data.

[0168] Step S402, User needs classification.

[0169] The data resource procurement forecast based on the public opinion product package has three main parts: unsupervised user classification module, procurement volume forecast module (single period and continuous period), and procurement decision module.

[0170] In this embodiment, user needs are classified based on an improved Euclidean distance-based user classification method. The differences in user needs are represented by the user's subscription to a service package and the actual usage of services within the package. Classification of the public opinion product user group is achieved using k-means unsupervised clustering based on the improved Euclidean distance. The classification of the public opinion product user group uses the improved Euclidean distance to predict service items within the public opinion product service package, denoted by r, as detailed in Table 5.

[0171] Table 5 Public Opinion Product Package Services

[0172] Serial Number Indicator Name Field identifier 1 Maximum number of topics r1 2 Keyword limit r2 3 TOP keyword recommendations r3 4 Number of key personnel in the communication program r4 5 Number of key people on Weibo r5 6 Sub-account r6 7 Number of SMS push messages r7 8 Push Hot Event Analysis r8 9 Number of times analysis reports are output r9 10 Use fuel pack or not? r10

[0173] The improved Euclidean distance is used to represent the user's package usage tendency. The distance formula is shown in equation (2) below:

[0174]

[0175] Where r i , These represent the user's subscribed public opinion monitoring product package service items and the user's actual usage of the service items, respectively. α i This indicates the correlation between the weight of each sub-item and the package price, which is normalized to a sum of 1.

[0176] Step S403: Establish an indicator system.

[0177] Here, establishing the indicator system includes constructing an initial multiple regression model and an initial time series prediction model. Based on historical ordering data and ordered products and usage data, the initial multiple regression model and the initial time series prediction model are trained respectively to obtain the trained multiple regression model and the trained time series prediction model.

[0178] Step S404, Multiple regression model prediction (single period).

[0179] Multiple regression prediction within a single period. An indicator system is established using time as the baseline dimension, including sales distribution, package usage, keyword configuration, and API return data volume. Assuming users in step 1 are divided into three categories, labeled I, II, and III, user profiles are created for each category based on their subscription and usage characteristics.

[0180] Table 6. Modeling Indicators for the Multiple Regression Model

[0181] Serial Number Indicator Name Field identifier 1 Percentage of Category I users x1 2 Percentage of Category II users x2 3 Percentage of Category III users x3 4 Percentage of users who ordered fuel packages x4 5 Average usage rate of special topics within the package x5 6 Keyword usage rate within the package x6 7 Weibo usage rate by key users x7 8 Key users of the communication program x8 9 Number of topics after deduplication x9 10 Number of duplicate keywords x10 11 The number of data records returned by the API call. y

[0182] Using a month as the basic time window, historical sales data samples from two consecutive years were extracted. A multiple regression model was established, extracting indicators such as user distribution, package subscriptions, usage rate, and keyword count from the monthly granular data samples, and fitting the multiple regression model. Let the dependent variable be the amount of data returned by the API call, y, and the independent variable be the indicator x. i The effective data samples are extracted monthly, and the relationship is shown in the following formula (3):

[0183] y = b0 + b1x1 + b2x2 + ... + b i x i +ε (3);

[0184] Where b0 is a constant term, b i (i = 1, 2, ...) are the model parameters. ε is the random error term, which follows a standard normal distribution, or white noise. The constant term represents the portion of information that is not explained by the independent variables and persists over a long period (non-random), i.e., residual information. The random error is the error between the predicted value and the actual value after removing the constant term within the interpretation space of the independent variables.

[0185] Step S405: Extract the monthly purchase quantity Yt.

[0186] Step S406, Time series model prediction (continuous period).

[0187] Time series forecasting under continuous trends. A time series analysis and forecasting model is established using the weighted moving average method, with weights w. i Since time is a factor, the most recent data is the best predictor of future conditions, and therefore should have a larger weight. This is expressed as equation (4):

[0188] y t =w1y t-1 +w2y t-2 +…+w n y t-n (4);

[0189] Where n is the predicted span, w1+w2+…+w n =1.

[0190] Step S407: Integrate the prediction results.

[0191] A weighted method is used to combine the prediction results from single-period and continuous periods. The weighted fusion of multiple regression and time series prediction models can predict the amount of data y returned by the API call after the user configures keywords in the next month. Assuming it's month M, the predicted value of the multiple regression model in step S404 is y. m0 In step S406, the predicted time series value is y. m1 The final predicted value is shown in equation (5):

[0192] y m =c0y m0 +c1y m1 (5);

[0193] In the formula, c0 + c1 = 1.

[0194] Step S408: Should a purchase be initiated this month?

[0195] Determine whether the procurement conditions are met. The procurement of public opinion products is predicted and initiated in advance. Assuming that the amount of data procured last time is G, when the amount returned by the API call on a monthly basis after procurement reaches a certain threshold λ of G, the next procurement behavior will be triggered. That is, it is determined whether the procurement conditions are met according to equation (6).

[0196] yt+y t-1 +y t-2 +…+y t-n >λG (6);

[0197] If the procurement conditions are not met in this cycle, the procurement forecasting process will be repeated in the next cycle (month).

[0198] Step S409, Decision on procurement method.

[0199] The procurement methods differ for different data resources. As the user base grows, commercial Weibo data is procured using a tiered system, predicting keyword returns for the next period and initiating procurement requests in advance. Due to different data resource sales rules, WeChat official account resources can be converted from tiered procurement to a one-time purchase once the user base reaches a certain scale.

[0200] This application's embodiments can be applied to apparel procurement. Using historical sales data as sample data for data regression, the weights are continuously adjusted along the time vector to ensure that the procurement forecasts generated through machine learning remain up-to-date. Other technologies perform data cleaning on procurement plans and approval data from the past three years, removing outliers. Considering the numerous influencing factors on procurement demand, a neural network is introduced to fit the prediction function.

[0201] The solution provided in this application uses an unsupervised clustering algorithm based on improved Euclidean distance to classify users; it integrates a regression model within a single period with a time series analysis model under continuous time trends to predict the amount of data to be purchased monthly; and it fully considers the contract rules of data service providers and the actual user growth trend of public opinion products to make decisions on the procurement methods of different data sources and minimize economic costs.

[0202] This scheme has the following characteristics:

[0203] 1) Based on the product packages ordered by users and their usage of services within those packages, the differences in user needs can be shown, thus enabling user group classification.

[0204] 2) Collect user classification indicators, package subscriptions and usage, keyword API return data, etc. to make single-cycle predictions for the procurement of public opinion product data resources; and integrate the prediction results of API return data volume within a continuous time interval to adjust the final procurement forecast value.

[0205] 3) In addition to predicting the scale of procurement, the procurement method of data resources is guided by comprehensive consideration of factors such as user scale, data sales method, and cumulative procurement volume. It is divided into two procurement methods: one-time and tiered.

[0206] Based on the foregoing embodiments, this application provides a procurement forecasting device. The various modules and units included in the device can be implemented by a processor in a computer device; of course, they can also be implemented by specific logic circuits. In the implementation process, the processor can be a central processing unit (CPU), a microprocessor unit (MPU), a digital signal processor (DSP), or a field programmable gate array (FPGA), etc.

[0207] This application embodiment further provides a procurement forecasting device. Figure 5 This is a schematic diagram of the composition of the procurement forecasting device provided in the embodiments of this application, as shown below. Figure 5 As shown, the procurement forecasting device 500 includes:

[0208] The first acquisition module 501 is used to acquire the order data, order attribute information and configuration attribute information of each order object in the current order period, and to acquire the historical call volume of data resources in each historical order period within a preset historical time period.

[0209] The first determining module 502 is used to determine the first estimated call volume of data resources in the next ordering cycle based on the ordering data, ordering attribute information and configuration attribute information of each ordering object, and to determine the second estimated call volume of data resources in the next ordering cycle based on the historical call volume corresponding to each historical ordering cycle.

[0210] The second determining module 503 is used to determine the target estimated amount of data resources in the next ordering cycle based on the first estimated call volume and the second estimated call volume.

[0211] The second acquisition module 504 is used to acquire the purchased quantity of data resources;

[0212] The third determining module 505 is used to determine the predicted purchase quantity of data resources based on the target estimated quantity, the historical call quantity corresponding to each historical ordering period, and the purchased quantity.

[0213] In some embodiments, the second determining module 503 is further configured to: obtain a pre-trained multiple regression model; determine the parameter values ​​of each input parameter in the multiple regression model based on the order data of each ordering object, the order attribute information and the configuration attribute information; input the parameter values ​​of each input parameter into the trained multiple regression model to obtain the first estimated call volume of data resources in the next ordering cycle.

[0214] In some embodiments, the second determining module 503 is further configured to: classify the subscribers corresponding to each subscriber based on the subscription data, subscription attribute information, and configuration attribute information of each subscriber, and obtain multiple classification results; determine the proportion of each category of subscribers based on the number of subscribers included in each classification result and the total number of subscribers; determine the average usage rate of each subscription attribute included in the subscription attribute information based on the subscription attribute information and the configuration attribute information; determine the set of attribute values ​​corresponding to each configuration attribute based on the attribute values ​​of each configuration attribute included in the configuration attribute information, and obtain the number of attribute values ​​in each set of attribute values; and determine the proportion of each category of subscribers, the average usage rate of each subscription attribute, and the number of attribute values ​​in each set of attribute values ​​as the parameter values ​​of each input parameter in the multiple regression model.

[0215] In some embodiments, the second determining module 503 is further configured to: determine the usage saturation of each ordering object based on the ordering attribute information and the configuration attribute information of each ordering object; and classify the ordering users corresponding to each ordering object based on the usage saturation of each ordering object to obtain multiple classification results.

[0216] In some embodiments, the second determining module 503 is further configured to: determine whether the order attribute information of each ordering object includes additional order attributes; classify the ordering users corresponding to the ordering objects that include additional order attributes as a first classification result; and classify the ordering users corresponding to the ordering objects that do not include additional order attributes as a second classification result.

[0217] In some embodiments, the ordering attribute information includes ordering attributes and attribute values ​​of the ordering attributes, and the configuration attribute information includes configuration attributes and attribute values ​​of the configuration attributes; the second determining module 503 is further configured to: determine the usage rate of each ordering attribute included in each ordering object based on the attribute values ​​of each ordering attribute and the attribute values ​​of each configuration attribute in each ordering object; and determine the average usage rate of each ordering attribute included in all ordering objects as the average usage rate of each ordering attribute.

[0218] In some embodiments, the second determining module 503 is further configured to: obtain a pre-trained time series prediction model; input the historical call volume corresponding to each historical ordering period into the trained time series prediction model to obtain the second estimated call volume of data resources in the next ordering period.

[0219] In some embodiments, the third determining module 505 is further configured to: determine the predicted call volume of data resources based on the target estimated volume and the historical call volume corresponding to each historical ordering period; when it is determined that the procurement conditions have been met based on the predicted call volume and the purchased volume, determine the predicted purchase volume based on the data source; when it is determined that the procurement conditions have not been met based on the predicted call volume and the purchased volume, determine the predicted purchase volume as a first value.

[0220] In some embodiments, the third determining module 505 is further configured to: when the data source is a first commercial data source, determine the predicted purchase quantity as a second value, so as to purchase data resources with the predicted purchase quantity of the second value in the first commercial data source; when the data source is a second commercial data source, determine whether the total number of ordering users has reached a preset user number threshold; when the total number of ordering users has reached the preset user number threshold, determine the difference between the total amount of data resources in the second commercial data source and the purchased amount as the predicted purchase quantity, so as to purchase all remaining data resources in the second commercial data source; when the total number of ordering users has not reached the preset user number threshold, determine the predicted purchase quantity as a third value, so as to purchase data resources with the predicted purchase quantity of the third value in the second commercial data source.

[0221] It should be noted that the descriptions of the above-described procurement forecasting device embodiments are similar to the descriptions of the methods described above, and have the same beneficial effects as the method embodiments. For technical details not disclosed in the procurement forecasting device embodiments of this application, those skilled in the art should refer to the descriptions of the method embodiments of this application for understanding.

[0222] It should be noted that, in the embodiments of this application, if the above-described procurement forecasting method is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiments of this application, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), magnetic disks, or optical disks. Thus, the embodiments of this application are not limited to any specific hardware and software combination.

[0223] Accordingly, embodiments of this application provide a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps in the procurement forecasting method provided in the above embodiments.

[0224] This application provides a procurement forecasting device. Figure 6 This is a schematic diagram of the composition structure of the procurement forecasting equipment provided in the embodiments of this application. Figure 6The exemplary structure of the procurement forecasting device 600 shown can be used to foresee other exemplary structures of the procurement forecasting device 600. Therefore, the structure described herein should not be regarded as a limitation. For example, some components described below may be omitted, or components not described below may be added to suit the specific needs of certain applications.

[0225] Figure 6 The procurement forecasting device 600 shown includes: a processor 601, at least one communication bus 602, a user interface 603, at least one external communication interface 604, and a memory 605. The communication bus 602 is configured to enable communication between these components. The user interface 603 may include a display screen, and the external communication interface 604 may include standard wired and wireless interfaces. The processor 601 is configured to execute a program of a procurement forecasting method stored in the memory to implement the steps of the procurement forecasting method provided in the above embodiments.

[0226] The descriptions of the above embodiments of the procurement forecasting equipment and storage media are similar to those of the above method embodiments, and have similar beneficial effects. For technical details not disclosed in the embodiments of the procurement forecasting equipment and storage media of this application, please refer to the descriptions of the method embodiments of this application for understanding.

[0227] It should be understood that the phrase "one embodiment" or "an embodiment" throughout the specification means that a specific feature, structure, or characteristic related to the embodiment is included in at least one embodiment of this application. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. It should be understood that in the various embodiments of this application, the sequence numbers of the above-described processes do not imply a sequential order of execution; the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application. The sequence numbers of the above-described embodiments are merely descriptive and do not represent the superiority or inferiority of the embodiments.

[0228] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0229] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods, such as: multiple units or components can be combined, or integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed can be through some interfaces, and the indirect coupling or communication connection between devices or units can be electrical, mechanical, or other forms.

[0230] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units. They may be located in one place or distributed across multiple network units. Some or all of the units may be selected to achieve the purpose of this embodiment according to actual needs.

[0231] In addition, each functional unit in the various embodiments of this application can be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit; the integrated unit can be implemented in hardware or in the form of hardware plus software functional units.

[0232] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media that can store program code, such as mobile storage devices, ROMs, magnetic disks, or optical disks.

[0233] Alternatively, if the integrated units described above are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, or the parts that contribute to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause an AC to execute all or part of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, ROMs, magnetic disks, or optical disks.

[0234] The above description is merely an embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A procurement forecasting method, characterized in that, The method includes: Obtain the order data, order attribute information, and configuration attribute information of each ordering object within the current ordering period, and obtain the historical call volume of data resources for each historical ordering period within a preset historical time period; the configuration attribute information is used to characterize the actual usage information of the order data; Based on the order data, order attribute information and configuration attribute information of each ordering object, determine the first estimated call volume of data resources in the next ordering cycle, and based on the historical call volume corresponding to each historical ordering cycle, determine the second estimated call volume of data resources in the next ordering cycle. Based on the first estimated call volume and the second estimated call volume, determine the target estimated amount of data resources in the next ordering cycle; Obtain the purchased amount of data resources, and determine the predicted usage amount of data resources based on the target estimated amount and the historical usage amount corresponding to each historical ordering period; When the procurement conditions are met based on the predicted call volume and the already procured volume, the predicted procurement volume is determined based on the data source; the data source is used to represent the supply source of the data resources. When it is determined that the procurement conditions have not been met based on the predicted call volume and the already procured volume, the predicted procurement volume is set as the first value.

2. The method according to claim 1, characterized in that, The step of determining the first estimated data resource usage volume for the next ordering cycle based on the ordering data, ordering attribute information, and configuration attribute information of each ordering object includes: Obtain a pre-trained multivariate regression model; Based on the order data of each ordering object, the order attribute information, and the configuration attribute information, the parameter values ​​of each input parameter in the multiple regression model are determined; The parameter values ​​of each input parameter are input into the trained multivariate regression model to obtain the first estimated amount of data resources to be used in the next ordering cycle.

3. The method according to claim 2, characterized in that, The step of determining the parameter values ​​of each input parameter in the multiple regression model based on the order data of each ordering object, the order attribute information, and the configuration attribute information includes: Based on the order data, order attribute information and configuration attribute information of each ordering object, the ordering users corresponding to each ordering object are classified to obtain multiple classification results; Based on the number of subscribers and the total number of subscribers included in each category, determine the percentage of subscribers in each category; Based on the order attribute information and the configuration attribute information, determine the average usage rate of each order attribute included in the order attribute information; Based on the attribute values ​​of each configuration attribute included in the configuration attribute information, determine the set of attribute values ​​corresponding to each configuration attribute, and obtain the number of attribute values ​​in each set of attribute values; The percentage of users subscribing to each category, the average usage rate of each subscription attribute, and the number of attribute values ​​in each attribute value set are determined as the parameter values ​​of each input parameter in the multiple regression model.

4. The method according to claim 3, characterized in that, The step involves classifying the subscribers corresponding to each subscriber based on their order data, order attribute information, and configuration attribute information, resulting in multiple classification results, including: The usage saturation of each ordering object is determined based on the ordering attribute information and the configuration attribute information of each ordering object; Based on the usage saturation of each subscription object, the subscribers corresponding to each subscription object are classified, resulting in multiple classification results.

5. The method according to claim 3, characterized in that, The step involves classifying the subscribers corresponding to each subscriber based on their order data, order attribute information, and configuration attribute information, resulting in multiple classification results, including: Determine whether the order attribute information of each ordering object includes additional order attributes; The ordering users corresponding to the ordering objects that include additional ordering attributes are classified as the first category result; Subscribers whose subscriptions do not include additional subscription attributes are categorized as the second category result.

6. The method according to claim 3, characterized in that, The order attribute information includes order attributes and the attribute values ​​of the order attributes; the configuration attribute information includes configuration attributes and the attribute values ​​of the configuration attributes. Based on the order attribute information and the configuration attribute information, determine the average usage rate of each order attribute included in the order attribute information, including: Based on the attribute values ​​of each ordering attribute and each configuration attribute in each ordering object, determine the usage rate of each ordering attribute included in each ordering object. The average usage rate of each subscription attribute is determined by averaging the usage rates of all subscription objects.

7. The method according to claim 1, characterized in that, The step of determining the second estimated data resource usage volume for the next subscription period based on the historical usage volume corresponding to each historical subscription period includes: Obtain a pre-trained time series prediction model; The historical call volume corresponding to each historical ordering period is input into the trained time series prediction model to obtain the second estimated call volume of data resources in the next ordering period.

8. The method according to claim 1, characterized in that, The step of determining the predicted purchase quantity based on the data source includes: When the data source is a first commercial data source, the predicted purchase quantity is determined as a second value, so as to purchase data resources with the predicted purchase quantity as the second value from the first commercial data source; When the data source is a second commercial data source, determine whether the total number of subscribers has reached a preset user threshold. When the total number of subscribers reaches a preset user threshold, the difference between the total amount of data resources in the second commercial data source and the amount already purchased is determined as the predicted purchase amount, so as to purchase all remaining data resources in the second commercial data source. When the total number of subscribers does not reach the preset user threshold, the predicted purchase volume is determined as a third value, so that data resources with the predicted purchase volume of the third value can be purchased from the second business data source.

9. A procurement forecasting device, characterized in that, The device includes: The first acquisition module is used to acquire the order data, order attribute information and configuration attribute information of each ordering object in the current ordering period, and to acquire the historical call volume of data resources in each historical ordering period within a preset historical time period; the configuration attribute information is used to characterize the actual usage information of the order data; The first determining module is used to determine the first estimated call volume of data resources in the next ordering cycle based on the ordering data, ordering attribute information and configuration attribute information of each ordering object, and to determine the second estimated call volume of data resources in the next ordering cycle based on the historical call volume corresponding to each historical ordering cycle. The second determining module is used to determine the target estimated amount of data resources in the next ordering cycle based on the first estimated call volume and the second estimated call volume. The second acquisition module is used to acquire the purchased quantity of data resources; The third determining module is used to determine the predicted call volume of data resources based on the target estimated volume and the historical call volume corresponding to each historical ordering cycle; when it is determined that the procurement conditions have been met based on the predicted call volume and the purchased volume, the predicted purchase volume is determined based on the data source; the data source is used to represent the supply source of the data resources; when it is determined that the procurement conditions have not been met based on the predicted call volume and the purchased volume, the predicted purchase volume is determined as a first value.

10. A procurement forecasting device, characterized in that, The device includes: Processor; and Memory for storing computer programs that can run on the processor; When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 8.

11. A computer-readable storage medium, characterized in that, The device stores computer-executable instructions configured to perform the steps of the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Linear regression electric quantity prediction method and system based on power utilization characteristic clustering

    CN110348604A

  • Product purchase quantity analysis method

    CN110659946A