Downstream data measuring and calculating method, equipment, medium and product
By acquiring and structuring the verification information, and using a linear regression model optimized by a genetic algorithm to predict downstream data, the accuracy and efficiency issues when upstream data is unknown are solved, achieving accurate downstream data evaluation and process simplification.
Patent Information
- Application Number
- CN202510674874.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2026-01-13
AI Technical Summary
When upstream data is unknown, the accuracy and efficiency of downstream data measurement in combined upstream and downstream system businesses are low, which increases the number of information exchanges and the business time span, and reduces service efficiency.
By acquiring the target business's measurement voucher information, performing structured preprocessing, and inputting it into a pre-trained downstream data measurement model, the model uses a linear regression model optimized by a genetic algorithm to predict downstream data, adapting to changes in business regulations and requirements in different scenarios.
It improved the accuracy of downstream data measurement, reduced the number of information exchanges between upstream and downstream units, simplified business processes, and improved execution efficiency.
Smart Images

Figure CN121329293A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of artificial intelligence and financial technology, and in particular to a downstream data measurement method, device, medium and product. Background Technology
[0002] As services become more sophisticated, many business scenarios have emerged that require data support from multiple upstream and downstream systems. Examples include loan amount calculation for combined mortgage loans and warehouse management and logistics scheduling calculations. In these scenarios, due to limitations such as system silos, privacy protection, and data production latency, downstream systems often need to independently provide downstream data to users when they cannot accurately obtain upstream data from upstream systems.
[0003] In related technologies, downstream system service personnel often need to calculate downstream data based on existing prior knowledge and personal experience, during the period when upstream data is still unknown.
[0004] This implementation method generally has low accuracy, and there is a high probability that downstream systems will need to recalculate their data after accurately obtaining accurate upstream data. This increases the number and amount of information exchange between upstream and downstream systems, lengthens the business time span, and significantly reduces the service efficiency when upstream and downstream systems jointly provide business services. Summary of the Invention
[0005] This invention provides a downstream data measurement method, device, medium, and product to solve the problems of low accuracy and low efficiency in downstream data measurement for combined upstream and downstream businesses when upstream data is unknown.
[0006] According to one aspect of the present invention, a downstream data measurement method is provided, comprising:
[0007] Obtain the calculation voucher information for the target business to be combined with upstream and downstream businesses; the calculation voucher information includes: multiple individual status descriptions and multiple related status descriptions;
[0008] The calculated voucher information is preprocessed in a structured manner to obtain structured voucher information;
[0009] Structured voucher information is input into a pre-trained downstream data calculation model to predict the downstream data of the target business when it conducts upstream and downstream combined business, given that the upstream data is unknown. The downstream data calculation model is a linear regression model optimized based on a genetic algorithm.
[0010] According to another aspect of the present invention, a downstream data measurement apparatus is provided, comprising:
[0011] The voucher acquisition module is used to acquire the calculation voucher information of the target business for which upstream and downstream combined business is to be carried out; wherein, the calculation voucher information includes: multiple individual status description information and multiple related status description information;
[0012] The structured module is used to perform structured preprocessing on the calculated voucher information to obtain structured voucher information;
[0013] The downstream calculation module is used to input structured voucher information into a pre-trained downstream data calculation model to predict the downstream data of the target business when it conducts upstream and downstream combined business, given that the upstream data is unknown. The downstream data calculation model is a linear regression model optimized based on a genetic algorithm.
[0014] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:
[0015] At least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores a computer program executable by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the downstream data measurement method according to any embodiment of the present invention.
[0016] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions, the computer instructions being configured to cause a processor to execute and implement the downstream data measurement method described in any embodiment of the present invention.
[0017] According to another aspect of the present invention, a computer program product is also provided, including a computer program that, when executed by a processor, implements the steps of the method as described in any embodiment of the present invention.
[0018] The technical solution of this invention involves acquiring the calculation voucher information of the target business undergoing upstream and downstream combined business; performing structured preprocessing on the calculation voucher information to obtain structured voucher information; and inputting the structured voucher information into a pre-trained downstream data calculation model to predict the downstream data of the target business when conducting upstream and downstream combined business, given that the upstream data is unknown. This achieves unified governance of diverse calculation elements in a structured manner, optimizes and upgrades the model through a genetic algorithm to adapt to differences in business regulations and changing needs under different scenarios, improves prediction accuracy and reduces errors, provides accurate downstream data evaluation support for upstream units, reduces the number of information exchanges between upstream and downstream units, simplifies the processing procedures for upstream and downstream combined business, and improves processing efficiency.
[0019] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0020] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0021] Figure 1 This is a flowchart of a downstream data measurement method provided in Embodiment 1 of the present invention;
[0022] Figure 2 This is a flowchart of another downstream data calculation method provided in Embodiment 2 of the present invention;
[0023] Figure 3 This is a schematic diagram of the structure of a downstream data measurement device provided in Embodiment 3 of the present invention;
[0024] Figure 4 This is a schematic diagram of the structure of an electronic device that implements the downstream data measurement method of this invention. Detailed Implementation
[0025] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0026] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0027] Example 1
[0028] Figure 1 This is a flowchart of a downstream data measurement method provided in Embodiment 1 of the present invention. This embodiment is applicable to situations where upstream data is lacking, and downstream data for combined upstream and downstream business operations is measured. This method can be executed by a downstream data measurement device, which can be implemented in hardware and / or software and is generally configured in an electronic device. For example... Figure 1 As shown, the method includes:
[0029] S110. Obtain the calculation voucher information of the target business for the upstream and downstream combined business to be carried out.
[0030] The calculation voucher information includes: multiple individual status descriptions and multiple related status descriptions.
[0031] In this embodiment of the invention, the upstream and downstream combined business can be specifically understood as: combining and coordinating different links or different business types in the business process to form a complete business. The upstream business is at the front end of the business, responsible for generating or providing the basic data, conditions, products, or services required by the downstream business; the downstream business is at the back end of the business, conducting further analysis and decision-making based on the basic data, conditions, products, or services generated by the upstream business, processing them, or directly providing them to the end customer. The target business can be specifically understood as: in the upstream and downstream combined business, when the upstream business data or the execution status of the upstream business is unknown, using features related to the upstream business to predict a specific business objective of the downstream business. The calculation voucher information can be specifically understood as: all relevant data used to evaluate and predict downstream data in the combined business. The individual status description information can be specifically understood as: data directly related to the upstream and downstream combined business, used to depict its basic status. The associated status description information can be specifically understood as: external or environmental factors related to the upstream and downstream combined business but not inherent to the upstream and downstream combined business itself.
[0032] In a specific example, when the upstream and downstream business involves a combined mortgage loan for home purchase, the upstream business is the application and approval of a housing provident fund loan. This involves the customer submitting a loan application to the housing provident fund management center, which calculates and determines the customer's provident fund loan amount based on information such as the customer's contribution base, contribution period, and credit status. The downstream business is the application and approval of a commercial loan. After the provident fund loan amount is determined, the customer submits a commercial loan application to a commercial bank. The bank determines the customer's commercial loan amount based on information such as income, credit score, and debt situation, combined with the provident fund loan amount. The target business is predicting the commercial loan amount that a target customer might obtain when the provident fund loan amount is unknown.
[0033] Target customers can be specifically understood as those planning to buy a house and needing to apply for both a housing provident fund loan and a commercial loan. Correspondingly, the assessment documentation information can be understood as: data and information reflecting the customer's economic situation, repayment ability, and creditworthiness, used to assess the target customer's commercial loan amount. This includes multiple individual and family situation descriptions. Individual situation descriptions may include: customer name, type of identification document, identification number, personal income certificate (monthly income or annual salary, etc.), credit results (credit history, credit rating, and credit score, etc.), anti-fraud inquiry results, physical health status, occupation (employer and years of service, etc.), personal debt (number and amount of outstanding loans and credit card debt, etc.), asset certificates (deposits and investments, etc.), income-to-debt ratio, etc. Family situation descriptions may include: number of family members, family income, family debt, and family assets (real estate and vehicles, etc.).
[0034] Understandably, when the upstream and downstream combined business is a combined loan for home purchase, before the commercial loan amount is calculated, the customer needs to check documents such as the "Face Information Collection Authorization Letter", "Personal Credit Information Business Authorization Letter", "Anti-Fraud Inquiry Business Authorization Letter", "Personal Information Authorization Letter" and "Family Information Authorization Letter". Only after the customer agrees to the authorization can the calculation voucher information of the target customer to be purchased with a combined loan be obtained.
[0035] In a specific example, when the upstream and downstream combined business involves warehouse management and logistics transportation scheduling, the upstream business is warehouse inventory management. Warehouse managers are responsible for monitoring and managing inventory levels, recording the receipt, issuance, and quantity of goods, ensuring the accuracy and real-time nature of inventory data. The downstream business is logistics transportation scheduling. Logistics dispatchers, based on warehouse inventory data and order information, develop transportation plans, arrange vehicles and routes, and ensure timely delivery of goods to customers. The target business is predicting logistics transportation demand. This involves forecasting future logistics transportation demand using historical sales data, order trends, and seasonal factors, even when warehouse inventory data is unknown. Accordingly, the calculation voucher information can be understood as: the data and information reflecting logistics transportation demand required to assess the target business, which may include: inventory records, order details, and transportation vehicle information. Individual situation description information may include: warehouse storage capacity, quantity and type of inventory items, etc. Related situation description information may include: order urgency, customer geographical distribution, etc.
[0036] S120. Perform structured preprocessing on the calculated voucher information to obtain structured voucher information.
[0037] It is understandable that the sources of calculation voucher information are diverse, and they may exist in various forms such as text, tables, and images. The data content may also contain missing, incomplete, or inconsistent data.
[0038] Specifically, the acquired calculation voucher information is converted into text (e.g., by converting the content in an image into text using optical character recognition technology). Data cleaning is then performed to remove duplicate, missing, or incomplete data. The data is then converted into a format that the model can process. For example, the text data is encoded (when the upstream and downstream combined business is a combined loan home purchase scenario, the occupational attribute in the calculation voucher information is set to 1 if the customer is a self-employed individual, and 0 otherwise; similarly, when the upstream and downstream combined business is a warehouse management and logistics transportation scheduling calculation scenario, the status of the transportation vehicle in the calculation voucher information is set to 1 if the vehicle is available, and 0 otherwise). Based on an open-source machine learning library, the text information is converted into a tabular data frame (a two-dimensional tabular data structure) to obtain structured voucher information.
[0039] S130. Input the structured voucher information into the pre-trained downstream data calculation model to predict the downstream data of the target business when it performs upstream and downstream combined business, given that the upstream data is unknown.
[0040] The downstream data measurement model is a linear regression model optimized using a genetic algorithm.
[0041] In this embodiment of the invention, the pre-trained downstream data calculation model can be specifically understood as: a linear regression algorithm model trained using historical data. A genetic algorithm is used to optimize the parameters of the linear regression model during the pre-training stage. By simulating the process of natural selection, the parameters are continuously adjusted to improve the accuracy of model predictions. Correspondingly, this can be a commercial loan amount calculation model or a logistics transportation scheduling calculation model for different scenarios.
[0042] Understandably, the selection of calculation voucher information can be determined based on business needs and relevant rules.
[0043] In a combined loan scenario for home purchase, the amount of the housing provident fund loan in the combined loan will affect the amount of the commercial loan. Local housing provident fund loan regulations directly affect the housing provident fund loan amount, and indirectly affect the calculation of the commercial loan amount. Therefore, the specific calculation voucher information required can be set in conjunction with local housing provident fund loan regulations. For example, suppose the housing provident fund loan regulations in a certain area state that the loan amount is related to the customer's housing provident fund contribution base and contribution period. Then, when calculating the commercial loan amount in a combined loan, the bank needs to collect information such as the customer's housing provident fund contribution base and contribution period (such as salary details and housing provident fund contribution certificates) and structure it. This information is then input into a pre-trained commercial loan amount calculation model. The model will calculate the customer's commercial loan amount based on the rules determined during pre-training, such as a certain multiple of the contribution base or a coefficient corresponding to the contribution period as a correction coefficient affecting the commercial loan amount. Alternatively, the model can calculate the maximum loan amount from the housing provident fund based on rules determined during pre-training, such as a certain multiple of the contribution base or a coefficient corresponding to the contribution period. Then, it determines the upper limit of the commercial loan amount by subtracting the housing provident fund loan amount from the total house price.
[0044] In warehouse management and logistics transportation scheduling calculations, the status of transport vehicles directly impacts the formulation of logistics transportation plans. Therefore, the specific calculation data required can be set in conjunction with local regulations for transport vehicles. For example, suppose a local logistics transportation regulation states that the allocation of transportation tasks is related to the availability and load capacity of vehicles. Then, when performing logistics transportation scheduling calculations, it is necessary to collect and structure information such as vehicle availability and load capacity. This information is then input into a pre-trained logistics transportation scheduling calculation model. The model will calculate the transportation task volume of each vehicle based on rules determined during pre-training, such as the priority of vehicles in availability or the transportation coefficient corresponding to load capacity. Alternatively, the model will calculate the transportation task volume that each vehicle can handle based on rules determined during pre-training, such as the priority of vehicles in availability or the transportation coefficient corresponding to load capacity. Then, the remaining transportation tasks are allocated based on the total order volume minus the already allocated transportation task volume.
[0045] Optionally, based on the above embodiments, after the prediction is completed, if the actual downstream data (such as the actual commercial loan amount in the combined loan) of the current upstream and downstream business is obtained, the structured certificate information, the predicted downstream data and the actual downstream data can be input into the pre-trained model for incremental training, and the genetic algorithm can be used for parameter optimization training to improve the prediction accuracy of the model.
[0046] The technical solution of this invention involves acquiring the calculation voucher information of the target business for which upstream and downstream combined business is to be conducted; performing structured preprocessing on the calculation voucher information to obtain structured voucher information; and inputting the structured voucher information into a pre-trained downstream data calculation model to predict the downstream data of the target business when conducting upstream and downstream combined business, given that the upstream data is unknown. This achieves unified governance of diverse calculation elements in a structured manner, optimizes and upgrades the model through a genetic algorithm to adapt to differences in industry regulations and changes in business needs across different regions, improves prediction accuracy and reduces errors, provides precise downstream data evaluation support, reduces the number of information interactions between upstream and downstream business units, simplifies the execution process of upstream and downstream combined business, and improves execution efficiency.
[0047] Optionally, based on the above embodiments, obtaining the calculation voucher information of the target business for which upstream and downstream combined business is to be carried out may include:
[0048] Based on the scenario types of upstream and downstream combined businesses, obtain the target knowledge retrieval graph that matches the scenario type;
[0049] Among them, different knowledge retrieval graphs are constructed based on the various types of calculation voucher information used in different scenario types; the target knowledge retrieval graph includes: multiple knowledge nodes, and directed edges for connecting multiple knowledge nodes. Each knowledge node is assigned individual status description information or associated status description information. The directed edges are used to define what permission data and data source are used to retrieve the ending knowledge node starting from the starting knowledge node; at least one target directed edge has a judgment condition, which is used to point from the starting knowledge node to the ending knowledge node when the judgment condition is met.
[0050] Obtain target individual status description information that matches the target business, and locate the target individual status description information in the target knowledge graph;
[0051] Starting from the located position in the knowledge graph, complete calculation voucher information for the target business is obtained level by level in the target knowledge graph; and the calculation voucher information is preprocessed in a structured manner to obtain structured voucher information, which may include:
[0052] The calculation voucher information is structured into tabular voucher information, and the tabular voucher information is divided into at least one structured voucher information through a table splitting operation.
[0053] In this embodiment of the invention, a knowledge graph can be specifically understood as a structured knowledge base used to represent relationships between entities, supporting entity queries and inter-entity reasoning. In a knowledge graph, nodes represent entities, and edges represent relationships between entities. Entities are defined as individual status descriptions or associated status descriptions. A knowledge retrieval graph can be specifically understood as a subset of a knowledge graph constructed based on calculation voucher information for a corresponding scenario. It is used to determine how to retrieve and associate information in a specific scenario, that is, to specify how to retrieve information from different data sources and how to access information according to set permissions and conditions. Directed edges in a knowledge retrieval graph represent retrieval paths from one knowledge node to another, defining retrieval permissions and data sources. For example, in a combined loan scenario, one node could be a customer's income information, and another node could be a customer's credit score. The directed edge defines how to retrieve credit score information from income information. Judgment conditions are set on the directed edges; only when the conditions are met can the path point from the starting node to the ending node be retrieved. For example, in a combined loan scenario, only when a customer's credit score reaches a certain standard can the loan amount node be retrieved from the credit score node.
[0054] Specifically, based on the business scenario, an appropriate knowledge retrieval graph is selected. For example, in a combined loan scenario, a knowledge retrieval graph containing loan-related knowledge and data sources is constructed; in a logistics and transportation scenario, a knowledge retrieval graph containing logistics and transportation-related knowledge and data sources is constructed. Target individual status description information matching the target business is obtained, and the graph position of this information within the target knowledge graph is located. For example, in a combined loan scenario, the position of the customer's income information node is determined; in a logistics and transportation scenario, the position of the cargo weight node for a specific transportation task is determined. Starting from the located node, the graph is searched level by level along the directed edges to obtain complete calculation voucher information. For example, in a combined loan scenario, starting from the customer's income information node, credit scores and debt information, as well as other relevant information, are retrieved; in a logistics and transportation scenario, starting from the cargo weight node, calculation voucher information such as transportation mode and vehicle type can be retrieved and determined.
[0055] The system reads the acquired calculation voucher information and converts it into tabular voucher information based on an open-source machine learning library. This tabular voucher information is then further segmented into at least one structured voucher information, such as splitting a large tabular voucher information into multiple smaller tabular vouchers, each containing a specific type of information. For example, a large tabular voucher information containing all calculation voucher information can be divided into multiple smaller tabular vouchers, each corresponding to a voucher information type (such as income, debt, or family status). Using a knowledge retrieval graph to complete downstream data calculations ensures data accuracy and compliance, improves process transparency and traceability, supports complex business logic, and adapts to different scenario needs. The segmented structured voucher information can be combined with features as needed, facilitating analysis and processing using data processing software and downstream data calculation models, thereby improving the automation and efficiency of data processing and prediction.
[0056] Example 2
[0057] Figure 2 This is a flowchart of another downstream data calculation method provided in Embodiment 2 of the present invention. This embodiment is a refinement of the downstream data calculation method in the above embodiment. Specifically, before inputting the structured voucher information into the pre-trained downstream data calculation model to obtain the predicted downstream data, it may further include: generating multiple training samples and multiple test samples based on the historical structured voucher information and historical downstream data of multiple historical businesses that have previously conducted upstream and downstream combined business; constructing a hyperplane regression model based on the amount of information contained in the structured voucher information; performing multiple iterations to optimize the model parameters of the hyperplane regression model based on each training sample and a preset genetic algorithm with the goal of minimizing the structural risk function, obtaining a candidate calculation model; and determining the candidate calculation model as the downstream data calculation model when it is determined that the prediction accuracy of the candidate calculation model for each test sample meets the prediction performance requirements.
[0058] Correspondingly, such as Figure 2 As shown, the method includes:
[0059] S210. Obtain the calculation voucher information of the target business for the upstream and downstream combined business to be carried out.
[0060] The calculation voucher information includes: multiple individual status descriptions and multiple related status descriptions.
[0061] S220. Perform structured preprocessing on the calculated voucher information to obtain structured voucher information.
[0062] S230. Based on the historical structured voucher information of multiple historical businesses that have previously conducted upstream and downstream combined business, as well as historical downstream data, generate multiple training samples and multiple test samples.
[0063] Specifically, the process involves acquiring historical structured voucher information from multiple historical transactions involving combined upstream and downstream businesses (such as combined mortgage loans for home purchases), as well as actual historical downstream data. The collected data is then divided into training and test sets according to a pre-defined ratio (e.g., 70% training set and 30% test set; the specific ratio can be adjusted based on actual conditions, such as data volume and model complexity). The training samples are used to train the downstream data calculation model; the test samples are used to evaluate the performance of the trained model and verify its predictive accuracy.
[0064] Optionally, based on the above embodiments, the target historical structured voucher information and target historical downstream data are selected by comparing local historical industry regulations with current industry regulations and based on similarity. First, the types of voucher information can be compared for consistency (e.g., comparing historical and current regulations regarding the calculation of voucher information data (e.g., whether the housing provident fund contribution base is based on a certain percentage of monthly average income)). Then, the severity of the regulations is compared; for example, in a combined loan home purchase scenario, the requirements for historical and current credit scores or credit records are compared, or the requirements for checking down payment ratios are compared; in a logistics scheduling scenario, the requirements for vehicle age are compared, etc. Weights are assigned according to the importance of the above factors. For each historical data sample, a similarity score is calculated based on its matching with current regulations. When the similarity score is greater than a preset score threshold, it is used as the target historical structured voucher information and target historical downstream data.
[0065] S240. Based on the amount of information contained in the structured voucher information, construct a hyperplane regression model.
[0066] In this embodiment of the invention, the hyperplane regression model can be specifically understood as a regression model for multidimensional data. The hyperplane regression model finds a hyperplane that minimizes the sum of distances from data points to this hyperplane, thereby achieving a good fit to the data. The dimensionality of the hyperplane regression model is related to the number of information types contained in the structured credential information. For example, in a combined loan home purchase scenario, if the structured credential information in each training sample includes four features—customer income, debt, family size, and credit score—then the constructed hyperplane regression model is four-dimensional.
[0067] Specifically, based on the amount of information contained in the structured voucher information, a hyperplane regression model is constructed: y = a1x1 + a2x2 + ... + a i x i +...+a n x n+b, where y is the target variable to be predicted (i.e., downstream data), and x i Let a be the feature of the i-th structured credential information. i Let b be the weight of the i-th structured document information, b be the bias term, and n be the number of pieces of information contained in the structured document information.
[0068] S250. With the goal of minimizing the structural risk function, the model parameters of the hyperplane regression model are iteratively optimized multiple times based on each training sample and a preset genetic algorithm to obtain alternative calculation models.
[0069] In this embodiment of the invention, the structured risk function can be specifically understood as a metric for evaluating model performance that comprehensively considers the model's performance on training data and the model's complexity. It typically includes a loss function and a regularization term. The regularization term can be either L1 regularization (sum of absolute values of parameters) or L2 regularization (sum of squared parameters) to prevent overfitting. If the goal is to make the coefficients of many features in the model close to or equal to zero, thereby achieving feature selection, L1 regularization can be chosen. If the goal is to make the parameters in the model smaller, but not strictly equal to zero, to prevent overfitting while maintaining model stability, L2 regularization can be chosen. Accordingly, the initial value of the regularization coefficient can be set to a small value and then adjusted based on the model's performance on the test set.
[0070] Taking L2 regularization as an example, minimizing the structural risk function can be specifically understood as: Where N is the number of training samples, L(w,x) i ,y i ) is the loss function for a single sample (measuring the difference between the model's predicted downstream data and the historical downstream data of the training samples, such as mean squared error or mean absolute error), w is the parameter vector to be solved, containing the weights corresponding to each structured voucher information feature in the model, x i It is the feature vector of the i-th training sample (containing all structured voucher information features of the i-th sample), y i It is the historical downstream data of the i-th training sample. This is the regularization term, used to prevent overfitting of the model. λ is the regularization coefficient, used to control the strength of regularization. This represents the square of the L2 norm of the parameter vector w, which is the sum of the squares of all elements in the parameter vector.
[0071] Specifically, model parameters (such as the weights of each structured credential information feature corresponding to each training sample and the model bias for each training sample) are randomly generated. These parameters are then applied to the model. After measuring the downstream data for each training sample, its structural hazard function value is calculated. The smaller the structural hazard function value, the higher the fitness of the individual (the model parameters of the training sample). Fitness refers to an individual's ability to survive and reproduce during natural selection, and is used to measure an individual's relative advantage.
[0072] Individuals are selected for reproduction based on fitness. Individuals with high fitness (low structural risk function value) have a higher probability of being selected. For example, a tournament selection method can be used, where k (2 or 3) individuals are randomly selected from the population, their fitness is compared, and the individual with the highest fitness (corresponding to the smallest structural risk function value) is selected. This process is repeated until a predetermined number of parent individuals are selected. Crossover and mutation operations are performed on the selected individuals to generate new individuals. Downstream data for each training sample is calculated based on the new individuals (model parameters), and their structural risk function value is calculated. The selection, crossover, and mutation operations are repeated until a stopping condition is met (such as reaching the maximum number of iterations or the change in the minimum structural risk function value in adjacent iterations being less than a predetermined threshold, indicating that the change in the minimum structural risk function value is not significant, reaching the convergence condition, and achieving the goal of minimizing the structural risk function), resulting in candidate measurement models.
[0073] S260. When it is determined that the prediction accuracy of the alternative calculation model meets the prediction performance requirements for each test sample, the alternative calculation model is determined as the downstream data calculation model.
[0074] Specifically, when the prediction accuracy of the candidate calculation model on all test samples meets the prediction performance requirements (for example, according to the loss function selected in the structural risk function, the error between the downstream data calculated for each training sample and the corresponding historical downstream data is calculated, and the prediction performance requirements are met when the error of all training samples is within the preset threshold range), the model is selected as the downstream data calculation model.
[0075] S270. Input the structured voucher information into the pre-trained downstream data calculation model to predict the downstream data of the target business when it is conducting upstream and downstream combined business, given that the upstream data is unknown.
[0076] The downstream data measurement model is a linear regression model optimized using a genetic algorithm.
[0077] The technical solution of this invention obtains the calculation voucher information of the target business to be combined with upstream and downstream business, performs structured preprocessing to obtain structured voucher information, and realizes the unified governance of diverse calculation elements in a structured manner; based on the historical structured voucher information of multiple historical businesses that have previously combined upstream and downstream business and historical downstream data, multiple training samples and multiple test samples are generated; based on the amount of information contained in the structured voucher information, a hyperplane regression model is constructed; with the goal of minimizing the structural risk function, the model parameters of the hyperplane regression model are iteratively optimized multiple times based on each training sample and a preset genetic algorithm to obtain a candidate calculation model; when it is determined that the prediction accuracy of the candidate calculation model for each test sample meets the prediction performance requirements, the candidate calculation model is determined as the downstream data calculation model; the structured voucher information is input into the pre-trained downstream data calculation model to predict the downstream data of the target business when combining upstream and downstream business in the case of unknown upstream data. By pre-training and optimizing the model using genetic algorithms, it can adapt to differences in industry regulations and changes in business needs in different regions, improve prediction accuracy and reduce errors, provide accurate commercial loan limit assessment support, reduce the number of information exchanges between upstream and downstream units, simplify the execution process of upstream and downstream combined business, and improve execution efficiency.
[0078] Optionally, based on the above embodiments, with the objective of minimizing the structural risk function, the model parameters of the hyperplane regression model are iteratively optimized multiple times based on each training sample and a preset genetic algorithm to obtain alternative calculation models, which may include:
[0079] Construct an initial population that matches the hyperplane regression model, and use the initial population as the current population. The population contains multiple individuals, and each individual corresponds to a set of model parameters.
[0080] Each training sample is input into the hyperplane regression model to obtain each pre-trained model. Based on the output of each pre-trained model, the fitness of each individual in the population and the structural risk function value matched with the current population are calculated.
[0081] Based on the fitness of each individual in the population, at least one parent individual is selected from the current population, and crossover and mutation operations are performed on the parent individuals to obtain multiple offspring individuals.
[0082] The individuals in each offspring population are used to replace some individuals in the current population to obtain a new current population.
[0083] Return the initial population that performs the construction of the hyperplane regression model, and use the initial population as the current population for the operation until the goal of minimizing the structural risk function is achieved.
[0084] Specifically, an initial population matching the hyperplane regression model is constructed. This population consists of multiple individuals, each representing a set of model parameters. The number of model parameters depends on the number of features in the hyperplane regression model. For example, if the model has 5 features (historical structured credentials contain 5 types of information), then each individual is a vector containing 5 weight values and 1 bias term. This initial population will serve as the current population to begin the optimization process.
[0085] Each training sample is input into a pre-trained model. Then, based on the model's output (i.e., the measured downstream data), the fitness of each individual is calculated (the absolute value of the difference or relative error between the downstream data of the training sample and the corresponding historical downstream data; a smaller result indicates higher fitness, meaning the model's prediction is closer to the actual value). The structural risk function value matching the current population is calculated. Based on the individual's fitness, at least one parent individual is selected from the current population. Individuals with higher fitness have a greater probability of being selected; for example, a tournament selection method can be used to select parent individuals.
[0086] The selected parent individuals undergo crossover and mutation operations to generate multiple offspring individuals. Crossover simulates gene recombination in biological heredity, generating new offspring by exchanging some parameters between two parent individuals. If only one parent individual is selected, it can undergo crossover with itself or with other individuals in the population. Mutation increases population diversity by randomly altering certain parameters of the individuals. The generated offspring individuals randomly replace some individuals in the current population, forming a new current population. The process of initializing the population, calculating fitness and structural risk, selecting parents, crossover and mutation, and updating the population is repeated until the structural risk function is minimized. By initializing the population and iteratively optimizing, combined with fitness assessment and structural risk control, the structural risk function integrates loss and regularization terms, reducing prediction error and preventing model overfitting. Through crossover and mutation, better model parameters are gradually found, improving the model's generalization performance and the accuracy of downstream data calculations.
[0087] Optionally, based on the above embodiments, each training sample is input into a hyperplane regression model to obtain a pre-trained model. Then, based on the output of each pre-trained model, the fitness of each individual in the population and the structural risk function value matching the current population are calculated. This may include:
[0088] Each training sample is input into the hyperplane regression model to obtain each pre-trained model. The absolute value of the difference between the output of each pre-trained model and the corresponding historical downstream data is calculated as the fitness of each individual in the population.
[0089] The loss function is used to calculate the loss value of each output result of the pre-trained model and the corresponding historical downstream data. The loss of all training samples is accumulated and then divided by the number of training samples to obtain the average value of the loss function.
[0090] Calculate the corresponding regularization term based on the model parameters of the current model;
[0091] The average value of the loss function is added to the regularization term to obtain the structural risk function value that matches the current population.
[0092] Specifically, each training sample is input into the pre-trained model to obtain predicted downstream data. The absolute value of the difference between the predicted downstream data of each training sample and the corresponding historical downstream data is calculated as the fitness of each individual in the population. The smaller the absolute value of the difference, the closer the model's prediction is to the actual value, and the higher the fitness.
[0093] The loss function (such as mean squared error or absolute error function) is used to calculate the loss value of the predicted downstream data and the corresponding historical downstream data for each training sample. The loss values of all training samples are summed and divided by the total number of training samples to obtain the average value of the loss function. This average value reflects the overall prediction error of the model on all training samples.
[0094] Based on the parameters of the current model (i.e., the weight vector in the linear regression model), the corresponding regularization term is calculated. The average of the loss function is added to the regularization term to obtain the structural risk function value. By calculating the fitness of each individual in the population and the structural risk function value matching the current population, a basis is provided for subsequent genetic algorithm operations (such as parent selection, crossover, and mutation), ultimately optimizing the model parameters, improving the model's generalization performance, and enhancing the accuracy of downstream data calculations.
[0095] Optionally, based on the above embodiments, selecting at least one parent population individual from the current population according to the fitness of each individual in the population may include:
[0096] The ratio of the fitness of each individual in the population to the sum of the fitness of all individuals in the population is used as the target probability.
[0097] Based on the calculated target probability for each individual in the population, at least one parent individual is selected from the current population.
[0098] Specifically, for each individual, its fitness is divided by the sum of the fitness of all individuals to obtain the relative probability that the individual will be selected as the next generation, i.e., the target probability. Assuming there are N data points, the probability of the i-th data point being selected is: Where, f(x) iLet be the fitness of the i-th data point. Based on the calculated target probability for each individual in the population, the cumulative selection probability of each individual is calculated, i.e., the selection probability of each individual is accumulated sequentially starting from the first individual. A random number between 0 and 1 is generated, and the individual whose cumulative selection probability falls within the range is selected as the parent population individual chosen in the current population. The above operation of calculating the cumulative selection probability and selecting individuals is repeated until a set number of parent population individuals are selected. By simulating the natural selection process, individuals with higher fitness are selected for reproduction, gradually optimizing the performance of the population, thereby optimizing the model parameters, improving the model's generalization performance, and enhancing the accuracy of the model in measuring downstream data.
[0099] Optionally, based on the above embodiments, when it is determined that the prediction accuracy of the candidate calculation model meets the prediction performance requirements for each test sample, the candidate calculation model is determined as the downstream data calculation model, which may include:
[0100] The root mean square error is calculated based on the output of the pre-trained model for each test sample and the corresponding historical downstream data.
[0101] When the root mean square error is less than the preset error threshold, the candidate calculation model is determined to meet the prediction performance requirements for each test sample, and the candidate calculation model is determined as the downstream data calculation model.
[0102] Specifically, for each test sample, the difference between the model's predicted value and the corresponding historical downstream data is calculated, and its square is determined. The sum of the squared errors of all test samples is then averaged to obtain the mean square error (MSE). The square root of the MSE is then taken to obtain the root mean square error (RMSE). If the RMSSE is less than a preset error threshold, the model's prediction results on the test samples are considered sufficiently accurate, meeting the prediction performance requirements, and the candidate calculation model is selected as the downstream data calculation model. By calculating the RMSSE and setting an error threshold, the prediction accuracy of the candidate calculation model can be quantitatively evaluated. When the RMSSE is less than the preset threshold, it indicates that the model's prediction results on the test samples are relatively close to the actual historical downstream data, the prediction performance meets the requirements, and it can be used as a downstream data calculation model, improving the accuracy and reliability of the model in calculating the customer's downstream data.
[0103] Optionally, based on the above embodiments, when setting the root mean square error (RMSE) threshold, local industry regulations (such as regulations on housing provident fund loans in combined mortgage scenarios, or regulations on warehouse quantity safety in logistics scheduling scenarios) can be considered. Since different regions have different requirements for upstream data settings (such as loan amounts, interest rates, and terms in combined mortgage scenarios, or warehouse quantity in logistics scheduling scenarios), these regulations directly affect downstream data settings (such as customers' repayment ability and loan demand in combined mortgage scenarios, or transportation capacity demand in logistics scheduling scenarios). Therefore, the setting of the RMSE threshold needs to be adjusted according to local industry regulations to ensure that the model's prediction results conform to the local reality. For example, if the upper limit of housing provident fund loan amounts in a certain region is low in a combined mortgage scenario, customers may need higher commercial loan amounts. In this case, the RMSE threshold should be appropriately relaxed to accommodate a larger range of loan amount predictions.
[0104] To facilitate understanding, the specific application scenarios applicable to the above-described embodiments are described below. Taking a combined loan for home purchase as an example, increasingly more customers are applying for combined loans (i.e., a housing provident fund loan plus a commercial loan) when purchasing a home. The housing provident fund loan is processed by the local housing provident fund center and enjoys a lower provident fund interest rate, while the commercial loan is processed by various commercial banks and applies a higher interest rate. Generally, customers tend to apply for a housing provident fund loan first, using the available provident fund amount, and then supplementing the shortfall with a commercial loan. When applying for a combined loan, the provident fund management center typically uses a specialized model to review the customer's eligibility and loan amount. The customer then submits the required documents to the bank, which determines the commercial loan amount accordingly. However, when the customer submits the documents first, the bank must first review the customer's eligibility and loan amount for a commercial loan. Because banks lack the dedicated calculation model provided by the provident fund management center, they cannot accurately determine the provident fund loan amount and must rely on employee experience for estimation. This results in inaccurate estimations of the commercial loan amount, increasing the workload for subsequent adjustments and reducing efficiency. To address the aforementioned issues, this invention proposes a downstream data calculation method (in this scenario, a commercial loan limit calculation method for a combined loan home purchase), which specifically includes:
[0105] 1. Collect and calculate voucher information
[0106] Collect and summarize the assessment voucher information in the form of a dataset file. The assessment voucher information may include: customer name, document type, document number, customer income, customer debt, customer family situation (such as a family with two or three children), housing situation, and health status, etc.
[0107] 2. Load the dataset and preprocess it.
[0108] It reads all the data from the data file using an open-source machine learning library and forms a data frame using a tabular data structure.
[0109] 3. Dataset partitioning
[0110] Based on a publicly available open-source machine learning library, the collected data is divided into feature variables (equivalent to the structured voucher information mentioned above) and target variables through table partitioning. For example, customer income, debt, and family circumstances are considered feature variables, while the loan amount calculation is the target variable. Then, according to a preset ratio (e.g., a preset percentage of the total sample to the training or test set), the collected data is divided into a training set and a test set. The training set is used to train the loan amount calculation model (downstream data calculation model), and the test set is used to evaluate the performance of the loan amount calculation model.
[0111] The calculation voucher information included in the dataset should be set in accordance with the calculation elements and standards of the local housing provident fund center. Table 1 shows one possible dataset partitioning method.
[0112] The attributes in the table are used as feature variables that affect the loan calculation results, and the output loan calculation results are used as the target variable.
[0113] 4. Training the model
[0114] First, a linear regression model is created. Linear regression primarily aims to minimize the error between the predicted loan limit and the actual loan limit by fitting a straight line or plane to the model, and is used to predict the linear relationship between the dependent variable and one or more independent variables. Since calculating the loan amount requires multiple independent variable features such as customer credit history and family size, the linear regression needs to be extended to a hyperplane linear regression. The model can use customer scores, customer ratings, anti-fraud query results, whether the family has multiple children, family size, individual monthly income, and family annual income as inputs, and the commercial loan amount for a combined mortgage as the output.
[0115] Table 1
[0116] parameter property Customer ratings Credit report results Customer rating Credit report results Anti-fraud query results Anti-fraud query dummy variables Customer's monthly income Personal monthly income Family size Number of children in a family Income-to-debt gap The difference between wealth and liabilities Income-to-debt ratio Wealth to Liabilities Ratio Personal loan status Number of outstanding loans Customer's physical health status Customer's physical health status Is the individual self-employed (1 if yes, 0 otherwise)? Profession
[0117] Then, based on the historical structured documentation and loan amounts of multiple customers who had previously taken out combined mortgages for home purchases, multiple training samples and multiple test samples were generated. A hyperplane regression model was constructed based on the amount of information contained in the structured documentation, and the root mean square error (RMSE) was used as an evaluation metric to indicate the magnitude of the deviation between the model's predicted values and the actual values. To minimize structural risk, the essence is to find the extreme value of an objective function, which can be specifically set as... Where N is the number of training samples, L(w,x) i ,yi ) is the loss function (such as mean squared error) for a single sample, w is the parameter vector to be solved, which contains the weights corresponding to each feature in the model, and x i It is the feature vector of the i-th sample (containing all structured voucher information features of the i-th sample), y i It is the historical commercial loan amount of the i-th sample. λ is the regularization term, used to prevent overfitting of the model, and λ is the regularization coefficient, used to control the strength of regularization. This represents the square of the L2 norm of the parameter vector w, which is the sum of the squares of all elements in the parameter vector. With the goal of minimizing the structural risk function, the model parameters of the hyperplane regression model are iteratively optimized multiple times based on each training sample and a pre-set genetic algorithm to obtain candidate calculation models. When the prediction accuracy of the candidate calculation model for each test sample meets the prediction performance requirements (e.g., the root mean square error is less than a pre-set error threshold, where the error threshold needs to be set in conjunction with local housing provident fund loan calculation elements and standards), the candidate calculation model is determined as the commercial loan amount calculation model.
[0118] In the process of iteratively optimizing the model parameters of the hyperplane regression model based on the genetic algorithm, before crossover and mutation, it is necessary to select data from the parent individuals (e.g., selecting model parameters based on the data with the smallest difference between the calculated amount and the actual amount), and use these data in the next evolution. Through continuous crossover and mutation operations, the optimal solution is obtained, thereby optimizing the model parameters. The specific data selection method can be as follows: Assuming there are N data points, the probability of the i-th data point being selected is: Where, f(x) i The fitness of the i-th data point is calculated (by the absolute value of the difference between the output of the pre-trained model for the i-th data point and the corresponding historical loan amount). Based on the calculated probability of each data point being selected, the model parameters for the selected data point are determined.
[0119] The structured certificate information to be predicted is input into a pre-trained commercial loan amount calculation model to predict the commercial loan amount for a target customer when taking out a combined loan to purchase a home, assuming the provident fund loan amount is unknown. After the prediction is completed, if the actual commercial loan amount in this combined loan is obtained, the structured certificate information, the predicted commercial loan amount, and the actual commercial loan amount can be input into the pre-trained model, and a genetic algorithm can be used for parameter optimization training to improve the model's prediction accuracy.
[0120] The downstream data calculation method proposed in this invention addresses the issue of inconsistent calculation elements and standards across different housing provident fund centers. It utilizes a publicly available open-source machine learning library to structure calculation voucher information into a tabular format. Table segmentation is then used to achieve unified data governance. A pre-trained linear regression model optimized using a genetic algorithm is then used to calculate the loanable amount for commercial loans, providing this information to the housing provident fund center for reference and increasing its acceptance of banks. Simultaneously, a genetic algorithm is employed to continuously evolve the model, enabling it to adapt to differences in housing provident fund loan regulations across different regions and market changes, improving prediction accuracy and reducing errors. This streamlines the information exchange process between the two institutions and improves the efficiency of combined loan processing.
[0121] Example 3
[0122] Figure 3 This is a schematic diagram of a downstream data measurement device provided in Embodiment 3 of the present invention. Figure 3 As shown, the device includes: a voucher acquisition module 310, a structured module 320, and a downstream calculation module 330, wherein:
[0123] The voucher acquisition module 310 is used to acquire the calculation voucher information of the target business for which upstream and downstream combined business is to be carried out; wherein, the calculation voucher information includes: multiple individual status description information and multiple related status description information;
[0124] The structured module 320 is used to perform structured preprocessing on the calculated voucher information to obtain structured voucher information;
[0125] The downstream calculation module 330 is used to input structured voucher information into a pre-trained downstream data calculation model to predict the downstream data of the target business when it is conducting upstream and downstream combined business in the case of unknown upstream data; wherein, the downstream data calculation model is a linear regression model optimized based on a genetic algorithm.
[0126] The technical solution of this invention involves acquiring the calculation voucher information of the target business undergoing upstream and downstream combined business; performing structured preprocessing on the calculation voucher information to obtain structured voucher information; and inputting the structured voucher information into a pre-trained downstream data calculation model to predict the downstream data of the target business when conducting upstream and downstream combined business, given that the upstream data is unknown. This achieves unified governance of diverse calculation elements in a structured manner, optimizes and upgrades the model through a genetic algorithm to adapt to differences in industry regulations and changes in business needs across different regions, improves prediction accuracy and reduces errors, provides precise downstream data evaluation support, reduces the number of information exchanges between upstream and downstream units, simplifies the execution process of upstream and downstream combined business, and improves execution efficiency.
[0127] Furthermore, based on the above embodiments, the downstream data measurement device may further include: a sample generation module, a model construction module, an iterative optimization module, and a model determination module, wherein:
[0128] The sample generation module is used to generate multiple training samples and multiple test samples based on historical structured voucher information and historical downstream data from multiple historical businesses that have previously conducted upstream and downstream combined business, before the structured voucher information is input into the pre-trained downstream data calculation model to obtain the predicted downstream data.
[0129] The model building module is used to construct a hyperplane regression model based on the amount of information contained in the structured voucher information.
[0130] The iterative optimization module is used to perform multiple iterations of the model parameters of the hyperplane regression model based on each training sample and a preset genetic algorithm, with the goal of minimizing the structural risk function, to obtain alternative calculation models.
[0131] The model determination module is used to determine the candidate calculation model as the downstream data calculation model when the prediction accuracy of the candidate calculation model meets the prediction performance requirements for each test sample.
[0132] Based on the above embodiments, the iterative optimization module is specifically used for:
[0133] Construct an initial population that matches the hyperplane regression model, and use the initial population as the current population. The population contains multiple individuals, and each individual corresponds to a set of model parameters.
[0134] Each training sample is input into the hyperplane regression model to obtain each pre-trained model. Based on the output of each pre-trained model, the fitness of each individual in the population and the structural risk function value matched with the current population are calculated.
[0135] Based on the fitness of each individual in the population, at least one parent individual is selected from the current population, and crossover and mutation operations are performed on the parent individuals to obtain multiple offspring individuals.
[0136] The individuals in each offspring population are used to replace some individuals in the current population to obtain a new current population.
[0137] Return the initial population that performs the construction of the hyperplane regression model, and use the initial population as the current population for the operation until the goal of minimizing the structural risk function is achieved.
[0138] Based on the above embodiments, the iterative optimization module is further used for:
[0139] Each training sample is input into the hyperplane regression model to obtain each pre-trained model. The absolute value of the difference between the output of each pre-trained model and the corresponding historical downstream data is calculated as the fitness of each individual in the population.
[0140] The loss function is used to calculate the loss value of each output result of the pre-trained model and the corresponding historical downstream data. The loss of all training samples is accumulated and then divided by the number of training samples to obtain the average value of the loss function.
[0141] Calculate the corresponding regularization term based on the model parameters of the current model;
[0142] The average value of the loss function is added to the regularization term to obtain the structural risk function value that matches the current population.
[0143] Based on the above embodiments, the iterative optimization module is further used for:
[0144] The ratio of the fitness of each individual in the population to the sum of the fitness of all individuals in the population is used as the target probability.
[0145] Based on the calculated target probability for each individual in the population, at least one parent individual is selected from the current population.
[0146] Based on the above embodiments, the model determination module is specifically used for:
[0147] The root mean square error is calculated based on the output of the pre-trained model for each test sample and the corresponding historical downstream data.
[0148] When the root mean square error is less than the preset error threshold, the candidate calculation model is determined to meet the prediction performance requirements for each test sample, and the candidate calculation model is determined as the downstream data calculation model.
[0149] Based on the above embodiments, the credential acquisition module 310 is specifically used for:
[0150] Based on the scenario types of upstream and downstream combined businesses, obtain the target knowledge retrieval graph that matches the scenario type;
[0151] Among them, different knowledge retrieval graphs are constructed based on the various types of calculation voucher information used in different scenario types; the target knowledge retrieval graph includes: multiple knowledge nodes, and directed edges for connecting multiple knowledge nodes. Each knowledge node is assigned individual status description information or associated status description information. The directed edges are used to define what permission data and data source are used to retrieve the ending knowledge node starting from the starting knowledge node; at least one target directed edge has a judgment condition, which is used to point from the starting knowledge node to the ending knowledge node when the judgment condition is met.
[0152] Obtain target individual status description information that matches the target business, and locate the target individual status description information in the target knowledge graph;
[0153] Starting from the location in the target knowledge graph, complete calculation voucher information for the target business is obtained step by step in the target knowledge graph.
[0154] Based on the above embodiments, the structured module 320 is specifically used for:
[0155] The calculation voucher information is structured into tabular voucher information, and the tabular voucher information is divided into at least one structured voucher information through a table splitting operation.
[0156] The downstream data calculation device provided in the embodiments of the present invention can execute the downstream data calculation method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the method execution.
[0157] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0158] Example 4
[0159] Figure 4 A schematic diagram of an electronic device 10 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0160] like Figure 4As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0161] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0162] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as downstream data measurement methods, i.e.:
[0163] Obtain the calculation voucher information for the target business to be combined with upstream and downstream businesses; the calculation voucher information includes: multiple individual status descriptions and multiple related status descriptions;
[0164] The calculated voucher information is preprocessed in a structured manner to obtain structured voucher information;
[0165] Structured voucher information is input into a pre-trained downstream data calculation model to predict the downstream data of the target business when it conducts upstream and downstream combined business, given that the upstream data is unknown. The downstream data calculation model is a linear regression model optimized based on a genetic algorithm.
[0166] In some embodiments, the downstream data measurement method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the downstream data measurement method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the downstream data measurement method by any other suitable means (e.g., by means of firmware).
[0167] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0168] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0169] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0170] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0171] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0172] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0173] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0174] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A method for downstream data measurement, characterized in that, include: Obtain the calculation voucher information for the target business to be combined with upstream and downstream businesses; the calculation voucher information includes: multiple individual status descriptions and multiple related status descriptions; The calculated voucher information is preprocessed in a structured manner to obtain structured voucher information; Structured voucher information is input into a pre-trained downstream data calculation model to predict the downstream data of the target business when it conducts upstream and downstream combined business, given that the upstream data is unknown. The downstream data calculation model is a linear regression model optimized based on a genetic algorithm.
2. The method according to claim 1, characterized in that, Before inputting the structured voucher information into the pre-trained downstream data measurement model to obtain the predicted downstream data, the following steps are also included: Based on the historical structured voucher information of multiple historical transactions involving upstream and downstream integration, as well as historical downstream data, multiple training samples and multiple test samples are generated. Based on the amount of information contained in the structured voucher information, a hyperplane regression model is constructed; With the goal of minimizing the structural risk function, the model parameters of the hyperplane regression model are iteratively optimized multiple times based on each training sample and a preset genetic algorithm to obtain alternative calculation models; Once it is determined that the candidate calculation model meets the prediction performance requirements for each test sample in terms of prediction accuracy, the candidate calculation model will be selected as the downstream data calculation model.
3. The method according to claim 2, characterized in that, With the objective of minimizing the structural risk function, the model parameters of the hyperplane regression model are iteratively optimized multiple times based on each training sample and a pre-defined genetic algorithm to obtain candidate calculation models, including: Construct an initial population that matches the hyperplane regression model, and use the initial population as the current population. The population contains multiple individuals, and each individual corresponds to a set of model parameters. Each training sample is input into the hyperplane regression model to obtain each pre-trained model. Based on the output of each pre-trained model, the fitness of each individual in the population and the structural risk function value matched with the current population are calculated. Based on the fitness of each individual in the population, at least one parent individual is selected from the current population, and crossover and mutation operations are performed on the parent individuals to obtain multiple offspring individuals. The individuals in each offspring population are used to replace some individuals in the current population to obtain a new current population. Return the initial population that performs the construction of the hyperplane regression model, and use the initial population as the current population for the operation until the goal of minimizing the structural risk function is achieved.
4. The method according to claim 3, characterized in that, Each training sample is input into a hyperplane regression model to obtain a pre-trained model. Based on the output of each pre-trained model, the fitness of each individual in the population and the structural risk function value matching the current population are calculated, including: Each training sample is input into the hyperplane regression model to obtain each pre-trained model. The absolute value of the difference between the output of each pre-trained model and the corresponding historical downstream data is calculated as the fitness of each individual in the population. The loss function is used to calculate the loss value of each output result of the pre-trained model and the corresponding historical downstream data. The loss of all training samples is accumulated and then divided by the number of training samples to obtain the average value of the loss function. Calculate the corresponding regularization term based on the model parameters of the current model; The average value of the loss function is added to the regularization term to obtain the structural risk function value that matches the current population.
5. The method according to claim 3, characterized in that, Based on the fitness of each individual in the population, at least one parent individual is selected from the current population, including: The ratio of the fitness of each individual in the population to the sum of the fitness of all individuals in the population is used as the target probability. Based on the calculated target probability for each individual in the population, at least one parent individual is selected from the current population.
6. The method according to claim 2, characterized in that, Once it is determined that the prediction accuracy of the candidate calculation models meets the prediction performance requirements for each test sample, the candidate calculation models are selected as the downstream data calculation models, including: The root mean square error is calculated based on the output of the pre-trained model for each test sample and the corresponding historical downstream data. When the root mean square error is less than the preset error threshold, the candidate calculation model is determined to meet the prediction performance requirements for each test sample, and the candidate calculation model is determined as the downstream data calculation model.
7. The method according to any one of claims 1-6, characterized in that, Obtain the calculation voucher information for the target business to be combined with upstream and downstream businesses, including: Based on the scenario types of upstream and downstream combined businesses, obtain the target knowledge retrieval graph that matches the scenario type; Among them, different knowledge retrieval graphs are constructed based on the various types of calculation voucher information used in different scenario types; the target knowledge retrieval graph includes: multiple knowledge nodes, and directed edges for connecting multiple knowledge nodes. Each knowledge node is assigned individual status description information or associated status description information. The directed edges are used to define what permission data and data source are used to retrieve the ending knowledge node starting from the starting knowledge node; at least one target directed edge has a judgment condition, which is used to point from the starting knowledge node to the ending knowledge node when the judgment condition is met. Obtain target individual status description information that matches the target business, and locate the target individual status description information in the target knowledge graph; Starting from the located position in the knowledge graph, complete calculation voucher information for the target business is obtained step by step within the target knowledge graph; and The calculated voucher information is preprocessed in a structured manner to obtain structured voucher information, including: The calculation voucher information is structured into tabular voucher information, and the tabular voucher information is divided into at least one structured voucher information through a table splitting operation.
8. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the downstream data measurement method according to any one of claims 1-7.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the downstream data measurement method according to any one of claims 1-7.
10. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the downstream data measurement method according to any one of claims 1-7.