Data processing apparatus and method based on data value
By using data processing devices and methods based on data value, high-value data is screened and calculated, and business models are trained for prediction. This solves the problems of high computational cost and significant impact of noisy data in existing technologies, and achieves more efficient business decision-making and greater accuracy.
Patent Information
- Application Number
- CN202210532391.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-11
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2042-05-11
AI Technical Summary
Existing technologies struggle to effectively manage and utilize data, resulting in high computational costs, significant impact from noisy data, and an inability to conduct targeted data management and business decisions.
By using data value-based data processing devices and methods, and utilizing high-value data acquisition and prediction units, high-value data is screened and calculated based on a utility function associated with business revenue, and a business model is trained for prediction.
It reduces the computational burden, improves the accuracy and efficiency of business forecasting, and enables more effective business decision-making.
Smart Images

Figure CN114926204B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to mass data processing, more particularly to a data processing device based on data value and a method thereof BACKGROUND
[0002] At present, various industries are undergoing digital transformation, and digital platforms use data to make intelligent business decisions, but big data has also led to high computing costs and lack of data management.
[0003] Due to the lack of effective management of data, full data is usually used, which is limited by computing resources and affected by noise data.
[0004] For example, an online medical platform wants to establish a prediction model for the demand of each doctor's consultation to optimize resource allocation. If full data is used for training, the computing cost is high, and there are noise data such as brushing in the online platform, that is, some doctors encourage patients to make false evaluations or brush articles and gifts in the platform in order to obtain higher profits.
[0005] However, the existing method is difficult to clarify the value of data points in practice, and cannot manage data specifically, so the use efficiency of data is not high. SUMMARY
[0006] According to an aspect of the present application, a data processing device based on data value is provided, comprising: a high-value data acquisition unit configured to acquire at least a part of high-value data from original data; and an estimation unit configured to make business prediction based on the acquired high-value data; wherein the high-value data acquisition unit calculates the value of each original data according to a utility function associated with business revenue, and acquires the at least a part of high-value data based on the calculation result.
[0007] Optionally, the estimation unit comprises: a training unit configured to train a business model using the acquired high-value data; and a prediction unit configured to make business prediction using the business model trained by the training unit.
[0008] Optionally, the high-value data acquisition unit calculates the data Sharpe value of each original data as the value of each original data according to the utility function associated with the business revenue.
[0009] Optionally, the utility function is associated with the prediction error of the business model trained using at least a part of the original data.
[0010] Optionally, the high-value data acquisition unit randomly sorts the set of all original data in each iteration of the utility function calculation, and keeps the utility function value of any original data for the set consisting of the previous element unchanged in the case that the difference between the utility function value of the set consisting of the any original data and the previous element and the utility function value of the set of all original data is less than a preset threshold.
[0011] Optionally, the high-value data acquisition unit calculates the data Shapley value of each original data by: training a full-set business model based on the set of all original data, and recording the full-set prediction result of each original data under the full-set business model; in each iteration of the utility function calculation, obtaining the prediction result of the original data according to the full-set prediction result of at least one original data close to the original data.
[0012] Optionally, the original data includes at least one period of historical data occurring before the current time, and the estimation unit performs business prediction for the next period of prediction data corresponding to the current time based on the obtained high-value data.
[0013] Optionally, the high-value data acquisition unit adjusts the utility function associated with business revenue according to at least one of business population, business logic, external environment, and time change.
[0014] Optionally, the data processing apparatus further comprises a value display unit configured to construct a graphical display interface based on the obtained high-value data to explain the business model to the user.
[0015] Optionally, the business model relates to an online medical platform, and the business prediction is used to predict the number of diagnoses and treatments that each doctor on the online doctor medical platform will undertake in the next period based on the attribute information of the doctors.
[0016] According to another aspect of the exemplary embodiments of the present application, a data processing method based on data value is provided, comprising: obtaining at least one part of high-value data from original data; and performing business prediction based on the obtained high-value data; wherein the value of each original data is calculated according to a utility function associated with business revenue, and the at least one part of high-value data is obtained based on the calculation result.
[0017] Optionally, in any of the above methods, the step of performing business prediction based on the obtained high-value data comprises: training a business model using the obtained high-value data; and performing business prediction using the trained business model.
[0018] Optionally, in any of the methods above, a data Shapley value of each raw data is calculated as a value of each raw data according to a utility function associated with a business benefit.
[0019] Optionally, in any of the methods above, the utility function is associated with a prediction error of a business model trained using at least part of the raw data.
[0020] Optionally, in any of the methods above, in each iteration of calculating the utility function, a set of all raw data is randomly sorted, and in a case where a difference between a utility function value of a set composed of any raw data and a previous element and a utility function value of the set of all raw data is less than a preset threshold, the utility function value of the set composed of the previous element is kept unchanged.
[0021] Optionally, in any of the methods above, the data Shapley value of each raw data is calculated by training a full-set business model based on the set of all raw data and recording a full-set prediction result of each raw data under the full-set business model, and in each iteration of calculating the utility function, a prediction result of a raw data is obtained according to a full-set prediction result of at least one raw data close to the raw data.
[0022] Optionally, in any of the methods above, the raw data includes at least one period of historical data occurring before a current time, and a next period of business prediction is performed for prediction data corresponding to the current time based on the obtained high-value data.
[0023] Optionally, in any of the methods above, the utility function associated with the business benefit is adjusted according to at least one of a business population, a business logic, an external environment, and a time change.
[0024] Optionally, in any of the methods above, the method further includes constructing a graphical display interface based on the obtained high-value data to explain the business model to a user.
[0025] Optionally, in any of the methods above, the business model relates to an online medical platform, and the business prediction is used to predict a number of consultations to be taken by each doctor on the online medical platform in a next period based on attribute information of the each doctor.
[0026] According to another aspect of the exemplary embodiments of the present application, there is provided a data processing apparatus based on data value, comprising a storage component and a processor, wherein the storage component stores a set of computer executable instructions, which when executed by the processor, performs the steps of: obtaining at least a portion of high value data from raw data; and making business prediction based on the obtained high value data; wherein the value of each raw data is calculated according to a utility function associated with business revenue, and the at least a portion of high value data is obtained based on the calculation result.
[0027] According to another aspect of the exemplary embodiments of the present application, there is provided a computer readable medium for data processing based on data value, wherein the computer readable medium has recorded thereon a computer program for performing the steps of: obtaining at least a portion of high value data from raw data; and making business prediction based on the obtained high value data; wherein the value of each raw data is calculated according to a utility function associated with business revenue, and the at least a portion of high value data is obtained based on the calculation result. BRIEF DESCRIPTION OF DRAWINGS
[0028] Figure 1 is a block diagram illustrating a data processing apparatus according to an exemplary embodiment of the present application;
[0029] Figure 2 is a block diagram illustrating a prediction unit according to an exemplary embodiment of the present application;
[0030] Figure 3 is a flow chart illustrating a data processing method according to an exemplary embodiment of the present application;
[0031] Figure 4 illustrates a prediction system for an online medical platform according to an exemplary embodiment of the present application;
[0032] Figure 5 illustrates the correlation between different physician information variables;
[0033] Figure 6 illustrates the relationship between platform extra cost & revenue and prediction error; and
[0034] Figure 7 illustrates the change in data model performance with stepwise deletion of data. DETAILED DESCRIPTION
[0035] The present application will be understood more easily with the following detailed description, given by way of example and with reference to the enclosed drawings, wherein the same reference numbers designate the same structural elements throughout the figures.
[0036] According to the general concept of the present application, intelligent business decisions can be made in a business system based on past data. By filtering out high-value data from a large amount of historical data, and making business predictions based on the filtered high-value data, not only the computational burden of the business system can be reduced, but also the accuracy of business predictions can be improved.
[0037] According to the exemplary embodiments of the present application, business revenue is used as a basis for measuring data value, thereby effectively associating business promotion with data value. This can more effectively filter out data that helps improve the decision-making effect of the business platform, and facilitate the explanation of the decision-making mechanism.
[0038] As an example, Figure 1 A block diagram of a data processing apparatus according to an exemplary embodiment of the present application is shown. Figure 1 The data processing apparatus 10 shown includes a high-value data acquisition unit 100 and a prediction unit 200.
[0039] As shown Figure 1 The high-value data acquisition unit 100 is configured to acquire at least a portion of high-value data from raw data.
[0040] Here, the raw data can be data generated online, data generated and stored in advance, or data received from the outside through an input device or a transmission medium. These data can relate to attribute information of individuals, enterprises or organizations, such as identity, education, occupation, assets, contact information, liabilities, income, profit, tax, etc. Alternatively, these data can also relate to other attribute information of business-related projects, such as transaction amount, transaction parties, subject matter, transaction location, etc. related to a sales contract. It should be noted that the data content mentioned in the exemplary embodiments of the present application can relate to the performance or nature of any object or transaction in business, and is not limited to the limitation or description of individuals, objects, organizations, units, institutions, projects, events, etc.
[0041] The raw data can be structured or unstructured data of different sources, such as text data or numerical data, etc. These data can come from within the entity for which business predictions are desired, such as banks, enterprises, schools, etc. that desire to obtain prediction results; these data can also come from outside the above-mentioned entities, such as data providers, the Internet (e.g. social networking sites), mobile operators, APP operators, express companies, credit agencies, etc. Alternatively, the above-mentioned internal data and external data can be used in combination to form raw data carrying more information.
[0042] The above raw data can be input to the high-value data acquisition unit 100 through an input device, or automatically generated by the high-value data acquisition unit 100 according to existing data, or can be obtained by the high-value data acquisition unit 100 from a network (for example, a storage medium (for example, a data warehouse) on the network), in addition, an intermediate data exchange device such as a server can help the high-value data acquisition unit 100 to obtain corresponding data from external data sources. Here, the data obtained can be converted into an easy-to-handle format by a data conversion module such as a text analysis module in the high-value data acquisition unit 100. It should be noted that the high-value data acquisition unit 100 can be configured as various modules composed of software, hardware and / or firmware, and some or all of these modules can be integrated or cooperated together to complete a certain function.
[0043] According to an exemplary embodiment of the present application, the high-value data acquisition unit 100 can calculate the value of each raw data according to the utility function associated with business revenue, and acquire at least part of the high-value data based on the calculation result. Here, the high-value data acquisition unit 100 takes business revenue as a measurement element when screening high-value data, realizes the unification of data quality and business operation, and makes it possible to improve the prediction effect more effectively while reducing the computational burden.
[0044] In addition, Figure 1 The data processing device 10 shown can also include a value display unit (not shown) for constructing a graphical display interface based on the acquired high-value data to explain the business model to the user. Here, the value display unit can draw a graphical display interface such as a curve based on the acquired high-value data, alone or in combination with other related data, to help the user view the role of the high-value data in the model and better understand the business logic.
[0045] As Figure 1 As shown, the estimation unit 200 is configured to make business predictions based on the acquired high-value data. Here, the estimation unit 200 can make business predictions based on high-value data in various ways, for example, generating corresponding rules based on the statistical characteristics presented by the high-value data, and then using these rules to make predictions about future data.
[0046] According to an exemplary embodiment of the present application, the raw data can include at least one period of historical data occurring before the current time, and accordingly, the estimation unit 200 makes business predictions for the next period based on the prediction data corresponding to the current time based on the acquired high-value data.
[0047] In addition, Figure 1The shown data processing apparatus 10 can further comprise a feedback unit (not shown) for collecting real business results corresponding to the above-mentioned predicted data, thereby forming new historical data samples, and further obtaining high-value data based on the updated historical data, and so on, to form a data closed loop of the business operation system.
[0048] In the present application, the estimation unit 200 can make intelligent decisions based on high-quality data, for example, train a model (e.g., an artificial intelligence (AI) model such as machine learning) using high-quality data, and then use the trained model to make business predictions.
[0049] Figure 2 is a block diagram illustrating an estimation unit according to an exemplary embodiment of the present application. As an example, the estimation unit 200 comprises a training unit 210 and a prediction unit 220.
[0050] Here, the training unit 210 is configured to train a business model using the obtained high-value data. As an example, the training unit can use the obtained high-value data to train a machine learning model. Here, machine learning is a natural product of the development of artificial intelligence research to a certain stage, which aims to improve the performance of the system itself through computational means using experience. In a computer system, "experience" usually exists in the form of "data", and through machine learning algorithms, a "model" can be generated from the data, that is, providing experience data to a machine learning algorithm can generate a model based on these experience data, and when facing new situations, the model will provide the corresponding judgment, i.e., the prediction result. Machine learning can be implemented in the form of "supervised learning", "unsupervised learning" or "semi-supervised learning", it should be noted that the exemplary embodiments of the present application do not make any specific restrictions on the specific machine learning algorithm. In addition, it should be noted that other means such as statistical algorithms can also be combined in the process of training and applying the model.
[0051] The prediction unit 220 can use the business model trained by the training unit 210 to make business predictions. Here, the training of the model by the training unit 210 can be performed offline or online, and the training here not only includes the first training of the model, but also includes the updating of the model. Accordingly, the prediction unit 220 can use the trained business model to make offline or online predictions, and the exemplary embodiments of the present application do not make any restrictions on this.
[0052] Figure 1 and Figure 2The illustrated apparatuses can be configured as software, hardware, firmware, or any combination thereof, to perform particular functions as described. For example, these apparatuses can correspond to dedicated integrated circuits, to purely software code, or to a combination of software and hardware elements or modules. Moreover, one or more functions performed by these apparatuses can be performed by components of a physical entity device (e.g., a processor, a client, or a server, etc.) in unison.
[0053] A data processing method according to an exemplary embodiment of the present application will be described below with reference to Figure 3
[0054] Here, as an example, Figure 3 The illustrated method can be performed by Figure 1 The illustrated apparatus, can be implemented in software by a computer program entirely, or by a combination of software and hardware elements or modules configured specifically, and Figure 3 The illustrated method can be performed by Figure 3 The illustrated apparatus. Figure 1
[0055] As shown in the figure, in step S100, at least a part of high-value data is acquired from raw data by a high-value data acquisition unit 100.
[0056] Here, as an example, the high-value data acquisition unit 100 can collect raw data in a manual, semi-automatic, or fully automatic manner, or process the collected raw data so that the processed data records have a proper format or form. As an example, the high-value data acquisition unit 100 can collect raw data in batches.
[0057] Here, the high-value data acquisition unit 100 can receive raw data records manually input by a user through an input device (e.g., a workstation). In addition, the high-value data acquisition unit 100 can systematically extract raw data records from data sources in a fully automated manner, e.g., by a timer mechanism implemented in software, firmware, hardware, or a combination thereof to systematically request data sources and obtain the requested raw data from the response. The data sources can include one or more databases or other servers. The fully automated manner of acquiring data can be implemented via an internal network and / or an external network, which can include transmitting encrypted data over the Internet. In the case where the servers, databases, networks, etc. are configured to communicate with each other, the data collection can be automatically performed without human intervention, but it should be noted that there can still be some user input operations in this manner. A semi-automated manner is between a manual manner and a fully automated manner. The difference between a semi-automated manner and a fully automated manner is that a trigger mechanism activated by a user replaces, for example, a timer mechanism. In this case, the request to extract data is generated only upon receiving a specific user input. Each time raw data is acquired, the captured data is preferably stored in a non-volatile memory. As an example, a data warehouse can be utilized to store the raw data collected during acquisition as well as processed data.
[0058] The acquired raw data records described above can be sourced from the same or different data sources, that is, each data record can also be a splicing result of different data records. For example, in addition to acquiring the information data record filled in by a customer when applying for a bank debit card (which includes income, education, position, asset situation, etc. attribute information fields), the high-value data acquisition unit 100 can also acquire other data records of the customer at the bank, such as loan records, daily transaction data, etc. as an example, which can be spliced into a complete data record. In addition, the high-value data acquisition unit 100 can also acquire data sourced from other private sources or public sources, such as data sourced from data providers, data sourced from the Internet (e.g., social networking sites), data sourced from mobile operators, data sourced from APP operators, data sourced from express companies, data sourced from credit agencies, etc.
[0059] Optionally, the high-value data acquisition unit 100 can store and / or process the collected data, e.g., store, classify, and other offline operations, with the aid of a hardware cluster such as a Hadoop cluster, a Spark cluster, etc. In addition, the high-value data acquisition unit 100 can also perform online stream processing on the collected data.
[0060] As an example, the high-value data acquisition unit 100 can include a data conversion module such as a text analysis module, and accordingly, in step S100, the high-value data acquisition unit 100 can convert unstructured data such as text into structured data that is easier to use for further processing or reference later. Text-based data can include emails, documents, web pages, graphics, electronic spreadsheets, call center logs, transaction reports, and the like.
[0061] As an example but not limitation, the example embodiments of the present application can use business data related to at least one of the following: image recognition (e.g., OCR, face recognition (security), object recognition (traffic signs), picture classification), speech recognition (e.g., fusion of natural language processing, including voice assistants, etc.), natural language processing (e.g., review of text (such as contracts, legal documents, customer service records), spam content identification, text classification (sentiment, intent, topic)), automatic control (e.g., energy industry (mines, wind turbine generators), energy saving (air conditioning systems)), intelligent question answering (e.g., chatbots, intelligent customer service), operational decision-making (e.g., financial technology, including marketing and customer acquisition, anti-fraud, anti-money laundering, underwriting and credit scoring, etc.), healthcare (e.g., disease screening and prevention, personalized health management, assisted diagnosis), municipal (e.g., social governance and regulatory law enforcement, resource environment and facility management, industrial development and economic analysis, public service and livelihood security, smart city, etc.), recommendation business (e.g., advertising, consulting, music, video, financial products (financial management, insurance), etc.).
[0062] In step S100, after obtaining the original data, the high-value data acquisition unit 100 acquires at least a part of the high-value data from the original data, where the high-value data acquisition unit 100 can calculate the value of each original data according to a utility function associated with business revenue, and acquire the at least a part of the high-value data based on the calculation result.
[0063] In the example embodiments of the present application, the basis for the high-value data acquisition unit 100 to calculate the value of the data is a utility function related to business revenue. Here, the high-value data acquisition unit 100 can construct a calculation method of business revenue based on the period of business prediction, analyze each component part therein, extract the part related to data contribution therefrom, and construct a utility function for calculating the value of the data based thereon. As an example, the high-value data acquisition unit 100 can calculate the value of the original data based on data Sharpe value, specifically, the high-value data acquisition unit 100 can calculate the data Sharpe value of each original data as the value of each original data according to a utility function associated with business revenue.
[0064] Next, in step S200, a business prediction is made by a prediction unit 200 based on the high-value data acquired by the high-value data acquisition unit 100. Here, the prediction unit 200 can make the prediction only relying on the filtered high-value data. As an example, a business model can be trained by a training unit 210 in the prediction unit 200 using the acquired high-value data. Here, the high-value data used can include only the high-value data acquired in the current period, can include all the high-value data acquired historically, or can be any or several periods of high-value data filtered. As an example, the training unit 210 can use any applicable machine learning algorithm to train the business model. Then, a prediction unit 220 in the prediction unit 200 makes a business prediction using the business model trained by the training unit 210, where the prediction unit 220 can input all the data to be predicted into the business model respectively, so as to obtain the prediction result corresponding to each data to be predicted. Further, a corresponding business decision can be made based on the prediction result, such as business resource allocation, etc.
[0065] Hereinafter, exemplary embodiments of the present application are described taking an online medical platform as an example, however, it should be understood that the exemplary embodiments of the present application are not limited to the online medical platform, but can be applied to any similar business prediction system, such as airport passenger flow distribution prediction, music trend prediction, demand prediction and warehouse planning scheme, Sina Weibo interaction prediction, money fund inflow and outflow prediction, movie box office prediction, agricultural product price prediction analysis, Qinghai-Tibet Plateau lake area prediction based on multi-source data, microblog propagation scale and depth prediction, abalone age prediction, student performance ranking prediction, online car travel flow prediction, red wine quality score, search engine search volume and stock price fluctuation, rural resident income growth prediction, real estate sales influencing factor analysis, stock price trend prediction, national comprehensive transportation volume prediction, earthquake prediction, etc.
[0066] According to the exemplary embodiments of the present application, the business model can be related to the online medical platform, and accordingly, the business prediction is used to predict the number of consultations to be taken by each doctor on the online medical platform in the next period based on attribute information of the doctors. Figure 4 A prediction system for an online medical platform according to an exemplary embodiment of the present application is shown.
[0067] In the rapid development of the digital economy today, the combination of traditional medical resources and digital platforms has given birth to a new medical model - online medical platforms. Compared with traditional hospitals, online platforms eliminate the spatial distance between patients and medical resources, providing patients with more options for seeking medical treatment in the era of the COVID-19 pandemic. Moreover, platforms can make full use of platform data to integrate and utilize resources, greatly improving the efficiency of medical resource allocation. For example, online medical platforms accumulate a large amount of data during operation, and use data to make personalized recommendations and demand matching to improve platform efficiency, thereby obtaining more data, thus forming a " virtuous cycle" of data-operation.
[0068] The typical service mode of such platforms is as follows: patients combine their own disease symptoms and search for the doctors they need on the platform, and then choose the doctors they like for consultation based on the doctor list returned by the platform and the static data of the doctors. For example, for users searching for "hypertension" disease recommended doctors, based on doctor homepage information and data such as "doctor gender", "age", "heart gift", "thank you letter", etc., patients choose the appropriate doctor for consultation.
[0069] Online medical platforms integrate data with traditional medical industries, injecting new momentum into the traditional industry. The governance of platform data is also increasingly concerned. On the one hand, big data can depict medical behavior on the platform more clearly, and the platform can achieve more accurate marketing at a finer granularity based on this. However, the platform needs to balance data computing costs and effectiveness. On the other hand, there is a phenomenon of "brushing data" (such as malicious brushing and inducing false evaluations) in the platform, which distorts the data. Therefore, it is very important for the platform to evaluate the value of data and filter out high-value data. First, it can reduce the cost of data calculation without reducing the effect of the algorithm. Second, it can filter out false data and build a better platform ecosystem.
[0070] Currently, evaluating the value of data in existing technologies is still a challenging problem. The determination of data value can also promote data integration and provide a basis for data governance. Current research on data value is mainly from a top-down economic perspective, and the main measurement methods are cost method, market method and income method. These three methods are limited by data regulation, imperfect data market, and the difficulty of splitting the value of data generated, resulting in a large deviation in the measurement of data value.
[0071] According to an example embodiment of the present application, in an online medical platform, the platform employs doctors to complete online consultation services, and it is very important for the platform to properly configure the doctor resources (sign a contract with a doctor who has more demand for longer online time, and vice versa), so it is necessary to predict the demand that the doctor will face in the future using data. For example, using machine learning, using the static data accumulated by the doctor in a period of time: doctor age, gender, hospital level, gifts, thank-you letters, patient check-in, waiting time, patient voting, etc. to predict the total demand (single quantity) of the doctor in a period of time, and make decisions accordingly (doctor resource allocation). But the amount of data used for prediction is large and there is noise data - there are behaviors such as doctor vote rigging and comment rigging, and simply using the full amount of data will cause losses in both precision and computing cost; Moreover, the black box model used by the platform is complex, and there is a lack of effective explanation of the black box model. In addition, as an example, evaluating the value of the data used in the machine learning process can guide the platform's data management, reduce the platform's computing cost, improve the learning progress, and also explain the underlying logic of the black box model. Among them, the explanation of the model: high-value data points reflect data that significantly affect the model's prediction results, and reflect the data distribution under normal business patterns, such as high consultation volume corresponding to high thank-you letter quantity, which can show which data distribution helps the model to get good results.
[0072] In Figure 4 The data value evaluation framework updated online in the online medical platform is shown. First, the time is discretized into 1, 2, …, T periods, and the framework evaluates and updates the data accumulated by the platform in each period to assist the platform operation. The flowchart shows the complete high-value data update and evaluation framework for periods T-1 and T.
[0073] Here, the prediction + operation process of the platform is described: in order to reasonably allocate doctor resources, the platform uses the static data of the doctor accumulated in the last period to predict the demand faced by the doctor, and the data used for prediction is as follows:
[0074]
[0075] The platform can make decisions on resource allocation based on the above prediction results, and then accumulate the data.
[0076] First, the operation of the platform and the prediction behavior of machine learning need to be combined to construct the corresponding utility function.
[0077] For example, for a medical platform, assume that time is discrete and infinite , the average time for a doctor to serve a patient is , here it is assumed that , the platform's revenue from a service is divided into The platform decides the cooperation agreement with the doctor in period t according to the historical data in period t-1 The platform decides the cooperation agreement with the doctor in period t according to the historical data in period t-1 The platform decides the cooperation agreement with the doctor in period t according to the historical data in period t-1 The platform decides the cooperation agreement with the doctor in period t according to the historical data in period t-1 The platform decides the cooperation agreement with the doctor in period t according to the historical data in period t-1 The platform decides the cooperation agreement with the doctor in period t according to the historical data in period t-1 The platform decides the cooperation agreement with the doctor in period t according to the historical data in period t-1 The platform decides the cooperation agreement with the doctor in period t according to the historical data in period t-1 The platform decides the cooperation agreement with the doctor in period t according to the historical data in period t-1 The platform decides the cooperation agreement with the doctor in period t according to the historical data in period t-1 The platform decides the cooperation agreement with the doctor in period t according to the historical data in period t-1 The platform decides the cooperation agreement with the doctor in period t according to the historical data in period t-1
[0078] The platform decides the cooperation agreement with the doctor in period t according to the historical data in period t-1 The platform decides the cooperation agreement with the doctor in period t according to the historical data in period t-1 The platform decides the cooperation agreement with the doctor in period t according to the historical data in period t-1
[0079] The platform decides the cooperation agreement with the doctor in period t according to the historical data in period t-1 The platform decides the cooperation agreement with the doctor in period t according to the historical data in period t-1
[0080] The platform decides the cooperation agreement with the doctor in period t according to the historical data in period t-1 The platform decides the cooperation agreement with the doctor in period t according to the historical data in period t-1 The platform decides the cooperation agreement with the doctor in period t according to the historical data in period t-1
[0081] The platform decides the cooperation agreement with the doctor in period t according to the historical data in period t-1 The platform decides the cooperation agreement with the doctor in period t according to the historical data in period t-1 The platform decides the cooperation agreement with the doctor in period t according to the historical data in period t-1
[0082] The platform decides the cooperation agreement with the doctor in period t according to the historical data in period t-1 The platform decides the cooperation agreement with the doctor in period t according to the historical data in period t-1 The platform decides the cooperation agreement with the doctor in period t according to the historical data in period t-1
[0083] The platform decides the cooperation agreement with the doctor in period t according to the historical data in period t-1 The platform decides the cooperation agreement with the doctor in period t according to the historical data in period t-1 The platform decides the cooperation agreement with the doctor in period t according to the historical data in period t-1
[0084] The platform decides the cooperation agreement with the doctor in period t according to the historical data in period t-1 The platform decides the cooperation agreement with the doctor in period t according to the historical data in period t-1 The platform decides the cooperation agreement with the doctor in period t according to the historical data in period t-1
[0085] The platform decides the cooperation agreement with the doctor in period t according to the historical data in period t-1 The platform decides the cooperation agreement with the doctor in period t according to the historical data in period t-1 The platform decides the cooperation agreement with the doctor in period t according to the historical data in period t-1 The platform decides the cooperation agreement with the doctor in period t according to the historical data in period t-1
[0086] The platform decides the cooperation agreement with the doctor in period t according to the historical data in period t-1 The platform decides the cooperation agreement with the doctor in period t according to the historical data in period t-1
[0087] That is, according to the exemplary embodiments of the present application, the utility function is associated with the prediction error of the business model trained using at least part of the original data. Through the utility function, the accuracy of the platform data exhibited in the machine learning learning is linked to the platform value.
[0088] In addition, optionally, according to the exemplary embodiments of the present application, the high-value data acquisition unit can construct different utility functions in a similar manner, for example, the utility function associated with business revenue can be adjusted according to at least one of the business population, business logic, external environment, and time variation.
[0089] Next, the data value can be calculated based on the above-mentioned utility function. According to the exemplary embodiments of the present application, the data Shapley value of each original data can be calculated as the value of each original data according to the utility function associated with the business revenue.
[0090] Specifically, Shapley Value is proposed by Lloyd Shapley to solve the problem of fair distribution of cooperative income. Shapley Value is derived from the theory in game theory, which is determined according to the "marginal contribution" of the individual participants, and in cooperative game, the marginal contribution can be regarded as the influence on the cooperative income after a participant joins, and by traversing all combinations of participants, the contribution of each participant can be calculated, and the specific formula is as follows:
[0091]
[0092] wherein K is the universal set, , , is the utility function that reflects the income of the cooperative game. Then represents the contribution of element under the utility function definition, that is, the Shapley Value of element .
[0093] According to the exemplary embodiments of the present application, the can be regarded as each data point or feature in the model, and the loss function or accuracy (such as the utility function) that measures the model effect. Then under this setting, is the contribution of data point or feature to the loss function or accuracy of the entire model, and this value evaluation method retains the completeness, fairness and additivity of shapley value, and can effectively extract the value of data points from the machine learning scene.
[0094] As can be seen from the formulas above, the computational complexity of Shapley Values is exponential, and the "utility" of the model decreases with each calculation. This means retraining the model using a subset of data or a subset of features U, which places a huge computational burden on models that learn from large amounts of data.
[0095] Therefore, as an optional approach, the high-value data acquisition unit can randomly sort the set of all original data in each iteration of the utility function calculation, and keep the utility function value of any original data relative to the set of the preceding elements unchanged if the difference between the utility function value of the set consisting of any original data and its preceding elements and the utility function value of the set of all original data is less than a preset threshold.
[0096] In other words, the complexity of the algorithm is controlled by simplifying the arrangement of data or features. For example, the expression for the Shapley Value can be written in terms of the arrangement of data points or features as follows:
[0097]
[0098] middle It is a sorting method for all elements in a complete set. In sorting, in the elements The set consisting of the preceding elements. In other words, that is... ,in First, the set of elements is randomly arranged. Second, the elements are scanned from the first to the last, and the marginal contribution of an element to the utility of the set of its preceding elements is defined as 0 when the difference between the utility of a given element and the utility of the entire set is less than a certain threshold. Finally, the Shapley Value of each element is updated using the concept of the average value. The pseudocode is as follows:
[0099]
[0100] This method treats the marginal utility of data nodes as zero once a threshold is met, and the order of the element set is arbitrary at each step. Furthermore, this method utilizes a property of Shapley Values, namely, under certain assumptions, as the amount of data or features increases, the utility of the dataset converges to the utility of the entire dataset or feature set. In other words, when the amount of data is small, newly added data contributes significantly to the model; however, when the amount of data is large, the contribution of newly added data to the model decreases and approaches zero.
[0101] In fact, by regulating the iteration number and threshold of the above algorithm, we can obtain the approximate Shapley Value in a shorter time. Compared with the previous calculation efficiency of Shapley Value method, the calculation efficiency has been improved to a certain extent.
[0102] As another example, considering that most machine learning models perform continuously in the feature space, adjacent data points tend to have the same prediction results. Based on this property, the high-value data acquisition unit can calculate the data Shapley value of each original data by the following processing: training a full-set business model based on the set of all original data, and recording the full-set prediction result of each original data under the full-set business model, in each iteration of calculating the utility function, the prediction result of the original data is obtained according to the full-set prediction result of at least one original data close to the original data. In this way, retraining the model can be avoided.
[0103] As an example, training a model in a complete training set can record the prediction of each data point. Further, when the prediction result in the prediction set needs to be calculated in each iteration, the average of K nearest neighbor training set prediction values is used for estimation. That is, for any data point in the test set , , let its K nearest neighbor set in the training set be , and the prediction model be , then the prediction value of has:
[0104]
[0105] Finally, the utility function is calculated according to all the approximate prediction results in the prediction set, and the algorithm can greatly reduce the calculation time without retraining the model.
[0106] It should be understood that the above two optimization processing methods can be used alone or in combination.
[0107] Suppose that an online medical platform has about 200,000 doctor data and can obtain doctor data in 2018, 2019 and 2020. In order to reduce the complexity of calculation, 5000 pieces of doctor information and their corresponding consultation order information are extracted for empirical research, and a prediction model for the demand of doctors in the next year is constructed using static information of doctors and platform behavior information. According to this, the value of the doctor data is studied.
[0108] Based on the characteristics of the data, the doctor data selected in this paper can be divided into the following two categories: one is the doctor registration information, which is the information provided by the doctor when registering on the Internet medical platform, including the gender, title, and hospital level of the doctor, reflecting the static provision of the doctor's inquiry ability. The second is the doctor platform behavior information, which is based on the doctor's inquiry behavior on the platform. It includes "total single volume", "article number", "post-diagnosis patient report number", "patient voting", "thank you letter", "heart gift", "general waiting time", and "comprehensive recommendation heat".
[0109] The actual meaning of the data is introduced in detail in Table 1.
[0110]
[0111] As an example, 5000 data are extracted, and 4000 are randomly selected as training samples, and the rest are used as test sets. As shown in Figure 5 For quantitative variables, the distribution and correlation are shown by scatter plot matrix. The diagonal line is the distribution of single data, the upper triangular part is the correlation coefficient of two data, and the lower triangular part is the scatter plot distribution of two data. From the distribution, except for the comprehensive recommendation heat, other data have a long-tail distribution. Overall, these quantitative data are platform behavior index data, and they all show a significant positive correlation with each other. The article number is concentrated near 0, and the positive correlation with other variables is not significant.
[0112] For qualitative variables "gender", "hospital level", and "general waiting time", they reflect the composition of the platform doctors and the service comfort. From Table 2, the distribution of male and female ratio in the hospital level and waiting time is relatively stable; doctors are concentrated in the highest level of three-A hospitals, and the waiting time for doctor inquiry is mostly "none", which shows that the service efficiency of doctors in the platform is high.
[0113] Table Distribution statistics of doctor information qualitative variables
[0114]
[0115] According to the example embodiment of the present application, XGBoost can be used to predict the total single volume of doctors in 19-20 years based on the doctor behavior information and registration information (Table 1 excluding total single volume) in early 19. According to the preferred mode described above, the value of different data sets or data points is estimated to deal with different data problems in the platform. The following sections will introduce several parts.
[0116] A. The impact of data set size on value
[0117] Under the background of big data, the larger the data volume means the higher the accuracy of machine learning, so the accumulation of data is crucial to the value accumulation of the platform. In online platforms, the cold start problem is that the data volume of platform participants is small, and the platform has difficulty in understanding them, resulting in large errors in operation strategy. When the data volume is large enough, the platform's precise matching is relatively easy, but because online platforms need to respond to user needs quickly, the data volume is limited by computing power. Therefore, how to balance the value of different data sets and the cost of computing is important to the platform.
[0118] Using the actual data of online platforms, different sizes of data sets are taken to calculate their Shapley Value to measure the value of different sizes of data sets, and compared with the classic information entropy (using Kozachenko-Leonenko estimation). For the platform, the quality of the data set is directly related to the prediction accuracy MAE. It can be seen that as the data set increases, the prediction error gradually decreases, and the Shapley value gradually increases - that is, the value of data increases, and Shapley can better describe the value increase of the data set. Entropy is a function of random variables, which is directly determined by its probability distribution, so the entropy of different data sets should be the same. However, in actual calculation, numerical methods (Kozachenko-Leonenko) are needed to estimate it, and the increase in data volume makes its distribution closer to the true value, so it shows a gradual downward trend of convergence. Therefore, Shapley Value is more direct in capturing the value of data volume.
[0119] B. Relationship between prediction accuracy and actual utility of the platform
[0120] In the previous section, Shapley value successfully captured the value increase brought by the increase in data volume, which in turn improved the prediction accuracy. In this section, we selected several different training sets to compare their prediction error (MAE) and platform total revenue (Total Revenue) ).
[0121] Figure 6 Show the relationship between the additional cost and benefit of the platform and the prediction error. As shown in Figure 6 , the left graph shows that as the prediction error increases, the cost loss brought by the platform's operation is divided into two parts: one part is the additional hiring cost paid by the insufficient demand, and the other part is the processing cost of the platform with excess demand; The right graph shows the relationship between the total revenue of the platform in a single period and the prediction error. Overall, the smaller the prediction error, the smaller the additional cost and the larger the revenue.
[0122] C. Explanation of the platform demand prediction model
[0123] Based on the first two parts, Shapley Value successfully links the platform's revenue with the value of data. In addition, it can also be observed how each data point in the model affects the model, that is, how it affects the platform's revenue, thereby interpreting the model and assisting in the formulation of platform operation strategies.
[0124] According to the defined utility function, the Shapley value of each data point in the training set can be calculated. As shown in Table 3, SV_Quantile represents the Shapley value from low to high, and is divided into 5 groups according to every 20% quantile. 0%-20% represents the lowest group of Shapley value, and 80%-100% represents the highest group. The rest of the variables in the table represent the average of the corresponding values in the group.
[0125] First, observe the extreme two groups of data. Compared with other groups of data, the group of data with the highest "value" (80%-100%) has the characteristics of high gift, thank-you letter, and high total order quantity, that is, it meets the high correlation between the variables shown in the Figure 3 However, in the data with the lowest "value" (20%-40%), the gift, thank-you letter, and other variables are lower, but the order quantity is higher, and the overall distribution of the data is inconsistent.
[0126] Table Shapley value and data distribution
[0127]
[0128] Overall, the five groups of data, 20-60% of the two groups of data are relatively close, reflecting the basic situation of a large amount of data, and the data of 60-100% more reflects the situation of high order quantity data. Combined with the long-tail distribution of data, correctly predicting the data points with high total order quantity, the contribution of the utility function is relatively large, so its data value is high. In actual business, that is, the value of helping the platform to correctly select the data points with high service demand from doctors is relatively high.
[0129] In the above model calculation, the data with high "value" can successfully reflect the positive correlation between the quantitative variables such as thank-you letter and gift and the order quantity in the data distribution. The capture of this relationship will help the platform correctly establish the relationship between the doctor data and its service ability; while the low-value data is contrary to the overall data distribution, which will interfere with the platform's designated strategy - reduce the accuracy of the prediction model, and the platform needs to remove the interference of this part of data to ensure the maximum revenue of the platform strategy under the overall sample.
[0130] D. Noise data detection
[0131] Malicious brushing and review brushing in online platforms is a common problem. The honor accumulated by doctors in the platform, such as the number of articles and heart gifts, is an important factor in attracting patients. Some doctors will deliberately brush up these data. This will cause distortion of demand prediction and difficulty in platform operation. The Shapley value can identify the contribution of data in the model. Users with malicious brushing behavior will have a negative impact on the model, resulting in a lower Shapley value. Based on this, the platform can appropriately delete some low-value data points to improve the quality of platform usage data.
[0132] In the experiments of this part, the effect of deleting low-value data on the model is verified. The data points in the training set are arranged according to the Shapley value, and are deleted from low to high. Each time a data point is deleted, the model is retrained, and the change in model performance (MAE, absolute error between predicted and actual) is observed. The data points are deleted in random order for comparison to reflect the correction effect of Shapley value on model prediction. The experimental results are shown in Figure 7 , where the vertical axis is MAE and the horizontal axis is the proportion of deleted data points.
[0133] Figure 7 The change in model performance with step-by-step data deletion. Overall, in both orders of deleting data, the model performance according to the data value deletion of the present application is better than random deletion, indicating that the Shapley idea plays a role in improving the model. Observing the performance of the algorithm according to the present application alone, in the range of deleting 50% of the data, this deletion method will reduce the prediction error. Generally, deleting data points will make the model worse. In the figure, it can be seen that deleting low-value data points using the present application actually makes the loss function further decrease.
[0134] This shows that "low-value" data points affect the overall performance of the model, and this part of low-value data is most likely to correspond to some participants who fake data in the platform. Therefore, it is meaningful for the platform to screen low-value data, which provides high-quality data learning samples for further data value calculation based on doctor demand to achieve future precise doctor recommendation and doctor-patient resource matching.
[0135] In the present application, for the resource allocation scenario of online medical platforms, the value of each data point used by the platform can be effectively measured, and a feasible solution is proposed to improve the operation effect of the platform from data to model prediction to platform decision-making.
[0136] The above refers to Figures 1 to 7A data processing apparatus and a method thereof according to an exemplary embodiment of the present application are described. It should be understood that the above-described method can be implemented by a program recorded on a computer-readable medium, for example, a computer medium according to an exemplary embodiment of the present application can be provided, in which a computer program for performing the following method steps is recorded: obtaining at least a part of high-value data from raw data; and making a business prediction based on the obtained high-value data; wherein the value of each raw data is calculated according to a utility function associated with business revenue, and the at least a part of high-value data is obtained based on the calculation result.
[0137] The computer program in the above-described computer-readable medium can run in an environment deployed in a computer device such as a client, a host, a proxy apparatus, a server, etc. It should be noted that the computer program can also be used to perform additional steps other than the above-described steps or perform more specific processing when performing the above-described steps, the content of which has been described with reference to the above-described exemplary embodiments of the present application and will not be repeated here. Figures 1 to 7 The above-described exemplary embodiments of the present application are described, and here will not be repeated.
[0138] It should be noted that the data processing apparatus according to an exemplary embodiment of the present application can completely rely on the running of the computer program to implement the corresponding functions, i.e., each apparatus corresponds to the steps in the functional architecture of the computer program, so that the whole system is called by a special software package (e.g., lib library) to implement the corresponding functions.
[0139] On the other hand, Figures 1 to 2 Each apparatus shown can also be implemented by hardware, software, firmware, middleware, microcode or any combination thereof. When implemented by software, firmware, middleware or microcode, the program code or code segments for performing the corresponding operations can be stored in a computer-readable medium such as a storage medium, so that the processor can perform the corresponding operations by reading and running the corresponding program code or code segments.
[0140] For example, the exemplary embodiments of the present application can also be implemented as a computing apparatus including a storage component and a processor, the storage component storing a set of computer-executable instructions, when the set of computer-executable instructions is executed by the processor, performing the data processing method according to an exemplary embodiment of the present application.
[0141] Specifically, the computing apparatus can be deployed in a server or a client, or on a node apparatus in a distributed network environment. In addition, the computing apparatus can be a PC computer, a tablet apparatus, a personal digital assistant, a smart phone, a web application or other apparatus capable of executing the above-described set of instructions.
[0142] Here, the computing device need not be a single computing device, but can be a collection of devices or circuits that individually or jointly execute the instructions (or sets of instructions) as described above. The computing device can also be part of a larger system or a system manager, or a portable electronic device that can be configured to interface with a local or remote (e.g., via wireless transmission) interface.
[0143] In the computing device, the processor can include a central processing unit (CPU), a graphics processing unit (GPU), a programmable logic device, a dedicated processor system, a microcontroller, or a microprocessor. By way of example and not limitation, the processor can also include an analog processor, a digital processor, a microprocessor, a multi-core processor, a processor array, a network processor, etc.
[0144] Some of the operations described above can be implemented in software, some can be implemented in hardware, and some can be implemented in a combination of software and hardware. In some embodiments, the operations described above can be performed by a computing device.
[0145] The processor can execute instructions or code stored in one of the storage components, which also can store data. The instructions and data could alternatively be transmitted or received over a network via the network interface device, using any one of a number of well-known transfer protocols.
[0146] The storage components can be integral to the processor, such as RAM or flash memory arranged within an integrated circuit microprocessor, etc. Alternatively, the storage components can include separate devices, such as external disk drives, memory arrays, or other storage devices that can be used by any database system. The storage components and the processor can be operatively coupled, or can communicate with each other, for example, via I / O ports, network connections, etc., so that the processor can read files stored in the storage components.
[0147] In addition, the computing device can include a video display (such as a liquid crystal display) and a user interface interface (such as a keyboard, a mouse, a touch input device, etc.). All of the components of the computing device can be connected via a bus or a network, for example.
[0148] The operations involved in the data processing method according to the exemplary embodiments of the present application can be described as various interconnected or coupled functional blocks or functional diagrams. However, these functional blocks or functional diagrams can be equally implemented as a single logical device or operate in non-exact boundaries.
[0149] For example, as described above, the data processing apparatus for data value-based data processing according to an exemplary embodiment of the present application can include a storage component and a processor, wherein the storage component stores therein a set of computer executable instructions which, when executed by the processor, perform the steps of: obtaining at least a portion of high value data from raw data; and making a business prediction based on the obtained high value data; wherein the value of each raw data is calculated according to a utility function associated with a business revenue, and the at least a portion of high value data is obtained based on the calculation result.
[0150] The above describes various exemplary embodiments of the present application, it should be understood that the above description is only exemplary and is not exhaustive, and the present application is not limited to the disclosed exemplary embodiments. Many modifications and changes are obvious to those skilled in the art without departing from the scope and spirit of the present application. Therefore, the scope of protection of the present application should be subject to the scope of claims.
Claims
1. An image data processing apparatus based on data value, comprising a storage component and a processor, wherein, The storage component stores a set of computer-executable instructions. When the set of computer-executable instructions is executed by the processor, the following steps are performed: The unstructured raw image data is converted into more easily usable structured raw data. The data Shapley value of each raw data is calculated based on the utility function associated with business benefits as the value of each raw data, and at least a portion of high-value data is obtained based on the calculation results. Business forecasting is based on the acquired high-value data, enabling the allocation of business resources based on the forecast results. as well as Based on the acquired high-value data, either alone or in combination with relevant data, a graphical interface is created to illustrate the business model, helping users understand the role of high-value data within the model and comprehend the business logic. In each iteration of the utility function calculation, the set of all original data is randomly sorted. If the difference between the utility function value of the set consisting of any original data and its preceding elements and the utility function value of the set of all original data is less than a preset threshold, the utility function value of that original data with respect to the set consisting of its preceding elements remains unchanged. The data Shapley value of each original data is calculated through the following process: a full-set business model is trained based on the set of all original data, and the full-set prediction result of each original data under the full-set business model is recorded. In each iteration of the utility function calculation, the prediction result of the original data is obtained based on the full-set prediction result of at least one original data that is close to the original data. The business model involves an online medical platform, and the business forecast is used to predict the number of consultations that each doctor will undertake in the next period based on the attribute information of each doctor on the online doctor medical platform. This study uses doctors' static information and platform behavior information to predict the demand for doctors in the following year, and studies the value of doctor data accordingly. Doctors' static information includes doctors' gender, professional title, and hospital level, reflecting their ability to provide consultations. Platform behavior information is based on statistics of doctors' consultation behavior on the platform, including total number of orders, number of articles, number of post-consultation patient check-ins, patient votes, thank-you letters, gifts, average waiting time, and overall recommendation popularity.
2. The image data processing apparatus as described in claim 1, wherein, The business forecasting based on acquired high-value data includes: Use the acquired high-value data to train business models; and Use the trained business model to make business predictions.
3. The image data processing apparatus as described in claim 1, wherein, The utility function is associated with the prediction error of a business model trained using at least a portion of the original data.
4. The image data processing apparatus as described in claim 1, wherein, The raw data includes at least one period of historical data that occurred before the current moment, and the business forecasting based on the acquired high-value data includes making business forecasts for the next period based on the acquired high-value data for the forecast data corresponding to the current moment.
5. The image data processing apparatus as claimed in claim 1, wherein, The utility function associated with business revenue is adjusted based on at least one of the following: the target audience, business logic, external environment, and time changes.
6. An image data processing method based on data value, comprising: The unstructured raw image data is converted into more easily usable structured raw data. The data Shapley value of each raw data is calculated based on the utility function associated with business benefits as the value of each raw data, and at least a portion of high-value data is obtained based on the calculation results. Business forecasting is based on the acquired high-value data, enabling the allocation of business resources based on the forecast results. as well as Based on the acquired high-value data, either alone or in combination with relevant data, a graphical interface is created to illustrate the business model, helping users understand the role of high-value data within the model and comprehend the business logic. In each iteration of the utility function calculation, the set of all original data is randomly sorted. If the difference between the utility function value of the set consisting of any original data and its preceding elements and the utility function value of the set of all original data is less than a preset threshold, the utility function value of that original data with respect to the set consisting of its preceding elements remains unchanged. The data Shapley value of each original data is calculated through the following process: a full-set business model is trained based on the set of all original data, and the full-set prediction result of each original data under the full-set business model is recorded. In each iteration of the utility function calculation, the prediction result of the original data is obtained based on the full-set prediction result of at least one original data that is close to the original data. The business model involves an online medical platform, and the business forecast is used to predict the number of consultations that each doctor will undertake in the next period based on the attribute information of each doctor on the online doctor medical platform. This study uses doctors' static information and platform behavior information to predict the demand for doctors in the following year, and studies the value of doctor data accordingly. Doctor's static information includes doctors' gender, professional title, and hospital level, reflecting the doctors' ability to provide consultations. Platform behavior information is based on statistics of doctors' consultation behavior on the platform, including total number of orders, number of articles, number of post-consultation patient check-ins, patient votes, thank-you letters, gifts, average waiting time, and overall recommendation popularity.
Citation Information
Patent Citations
Method of energy big data acquisition key value extraction
CN106383837A
User data processing method and device, electronic equipment and storage medium
CN111080338A