Business model construction method and device, medium and program product
By training and aggregating cyclical models between the main business system and sub-business systems, the challenges of traditional business systems in terms of data privacy, compliance, differences in computing resources, and response speed are solved, enabling cross-system joint modeling and the construction of business models for personalized needs.
Patent Information
- Application Number
- CN202511070058.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-31
- Publication Date
- 2025-11-14
AI Technical Summary
Traditional business systems face challenges in building intelligent business models, including data privacy and compliance requirements, data heterogeneity across sub-business systems, differences in computing resources, and high requirements for response speed. This makes it difficult to achieve cross-system joint modeling and personalized needs.
The main business system distributes the models to be trained to the sub-business systems. Each sub-business system trains its model for the current round and uploads the optimized model back to the main business system for aggregation. The training is repeated until the training round ends, forming a business model suitable for each sub-business system.
While protecting data privacy and compliance, we can achieve joint modeling across sub-business systems to ensure global collaborative optimization, while also taking into account the personalized model requirements of each sub-business system, thereby improving the adaptability and response speed of the model.
Smart Images

Figure CN120952080A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of business model construction technology, and in particular to a business model construction method, device, medium and program product. Background Technology
[0002] With the development of science and technology, various business systems are evolving towards data-driven, intelligent, and personalized approaches. Traditional business systems often rely on rule engines for processing or static model predictions, making it difficult to adapt to dynamic data changes, and significant differences in performance exist between different sub-business systems. Currently, business systems face numerous challenges in building intelligent business models: 1. Data privacy and compliance requirements. Even if the head office can access data from various sub-business systems, traditional centralized model training methods still require data transfer and centralized storage, facing compliance and privacy risks. 2. Heterogeneity and personalized needs of data across sub-business systems. Sub-business systems in different regions have significantly different environments, data, and trends. Data patterns and trends in specific regions cannot be effectively learned, and a unified global model cannot adapt to the needs of all sub-business systems. 3. Differences in computing resources and the challenges of implementing federated learning. There is a significant gap in computing resources between the head office and sub-business systems, with limited resources in sub-business systems, leading to unbalanced computational loads during training and hindering efficient training. 4. High requirements for response speed. A large amount of new data is generated daily, and traditional static models struggle to cope with rapid data changes and cannot afford the time and cost of retraining. Summary of the Invention
[0003] This application provides a method, device, medium, and program product for constructing a business model, which enables joint modeling across sub-business systems and the overall business system while protecting privacy and compliance, ensuring global collaborative optimization, and taking into account the personalized model requirements of each sub-business system.
[0004] According to one aspect of this application, a method for constructing a business model is provided, the method comprising:
[0005] The system distributes the training model to each sub-business system through the main business system, and receives the training model through each sub-business system.
[0006] For each sub-business system, the model is trained in this round based on the model to be trained, and the optimized model for this round of training is obtained.
[0007] If the training round is not completed, the optimized model is uploaded to the main business system through each sub-business system. The main business system then aggregates the models and distributes the aggregated models as new models to be trained to each sub-business system. Training and optimization are then performed on each sub-business system separately until the training round is completed, and the resulting business model is used by the sub-business systems for business processing.
[0008] According to one aspect of this application, a business model construction apparatus is provided, the apparatus comprising:
[0009] The module for determining the model to be trained is used to distribute the model to be trained to each sub-business system through the main business system, and to receive the model to be trained through each sub-business system.
[0010] The model training module is used to train the model in this round for each sub-business system based on the model to be trained, so as to obtain the optimized model for this round of training.
[0011] The upload aggregation module is used to upload the optimized model to the main business system through each sub-business system if the training round is not completed. The main business system then aggregates the models and distributes the aggregated models as new models to be trained to each sub-business system for training and optimization based on each sub-business system until the training round is completed, and the resulting business model is used by the sub-business systems for business processing.
[0012] According to another aspect of this application, an electronic device is provided, the electronic device comprising:
[0013] At least one processor; and
[0014] A memory that is communicatively connected to at least one processor; wherein,
[0015] The memory stores a computer program that can be executed by at least one processor, which enables the at least one processor to execute the method for constructing a business model according to any embodiment of the present application.
[0016] According to another aspect of this application, a computer-readable storage medium is provided, which stores computer instructions for causing a processor to execute a method for constructing a business model according to any embodiment of this application.
[0017] According to another aspect of this application, a computer program product is provided, which includes a computer program that, when executed by a processor, implements a method for constructing a business model according to any embodiment of this application.
[0018] The technical solution of this application embodiment involves distributing training models to each sub-business system through the overall business system, and receiving the training models through each sub-business system. For each sub-business system, a training round is conducted based on the training model to obtain an optimized model for that round. If the training round is not yet complete, the optimized models are uploaded to the overall business system through each sub-business system. The overall business system then aggregates the models and distributes the aggregated models as new training models to each sub-business system for individual training and optimization until the training round is completed, resulting in a business model used by the sub-business systems for business processing. This solution enables joint modeling across sub-business systems and the overall business system while protecting data privacy and compliance, ensuring global collaborative optimization while also considering the personalized model needs of each sub-business system.
[0019] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this application, nor is it intended to limit the scope of this application. Other features of this application will become readily apparent from the following description. Attached Figure Description
[0020] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0021] Figure 1 A flowchart illustrating a method for constructing a business model as provided in an embodiment of this application;
[0022] Figure 2 A flowchart illustrating a method for constructing a business model, as provided in another embodiment of this application;
[0023] Figure 3 A flowchart illustrating a method for constructing a business model, as provided in another embodiment of this application;
[0024] Figure 4 A schematic diagram of a business model construction apparatus provided in an embodiment of this application;
[0025] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0026] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.
[0027] It should be noted that the terms "first," "second," "third," "fourth," "actual," "preset," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0028] Figure 1 This is a flowchart illustrating a method for constructing a business model according to an embodiment of this application. This embodiment is applicable to training business models within a business system. The method can be executed by a business model construction device, which can be implemented in hardware and / or software and can be configured in an electronic device. Figure 1 As shown, the method includes:
[0029] S110. The training model is distributed to each sub-business system through the main business system, and the training model is received by each sub-business system.
[0030] The overall business system can be a higher-level system that manages all sub-business systems. One overall business system can manage multiple sub-business systems. The business functions processed in the overall business system and sub-business systems are not limited and can include classification tasks, prediction tasks, recognition tasks, etc. The model to be trained is a model that has been processed by the overall business system and still needs to be further trained and optimized by each sub-business system.
[0031] For example, during initialization, the main business system can first initialize and set up the model to be trained, and then distribute the model to each sub-business system. The sub-business systems then train and optimize the model. In subsequent processes, the main business system can further process the optimized model reported by each sub-business system after one round of training and optimization, and then distribute it as a new model to be trained to each sub-business system for further processing. Each sub-business system receives the model to be trained from the main business system.
[0032] S120. For each sub-business system, the model is trained in this round based on the model to be trained, and the optimized model for this round of training is obtained.
[0033] For each sub-business system, the process of training the model in this round is executed separately in each sub-business system until the termination condition set in that sub-business system is met, resulting in the optimized model for this round of training. During this round of model training, the training samples used for model training are the business data from that sub-business system. The type of business data is not limited and is adaptively selected based on specific circumstances.
[0034] In this embodiment, when training the model in each sub-business system, each sub-business system can train based on local business data, enabling the business model to be applicable to local business needs and fulfilling the personalized requirements of each sub-business system. Furthermore, since model training is performed locally in each sub-business system, training data does not need to be shared externally, preventing data leakage and protecting data privacy and compliance.
[0035] S130. If the training round is not completed, the optimized model is uploaded to the main business system through each sub-business system. The main business system aggregates the models and distributes the aggregated models as new models to be trained to each sub-business system for training and optimization based on each sub-business system until the training round is completed, and the resulting business model is used by the sub-business systems for business processing.
[0036] For example, in each sub-business system, if the training round is not yet complete after the current model training, further training is required. In this case, each sub-business system uploads its optimized model to the main business system. The main business systems then aggregate the models, learning the global knowledge of the optimized models uploaded by each sub-business system, thus achieving global system optimization. The number of training rounds can be set according to specific circumstances or determined based on specific optimization conditions.
[0037] After the main business system aggregates the optimized models reported by each sub-business system, the aggregated model is used as the new model to be trained. The system then returns to step S110 and distributes the models to each sub-business system for continued model training and optimization, incrementing the training round by one. This process repeats until the total number of training rounds reaches the preset limit. At this point, the training round ends, and the corresponding business model for each sub-business system is obtained for its business processing.
[0038] The technical solution of this application embodiment distributes training models to each sub-business system through the overall business system, and each sub-business system receives the training models. For each sub-business system, it performs a training round based on the training models to obtain an optimized model for that round. If the training round is not completed, each sub-business system uploads the optimized model to the overall business system, which then aggregates the models and distributes the aggregated models as new training models to each sub-business system for training and optimization on each sub-business system until the training round is completed, resulting in a business model used by the sub-business systems for business processing. This solution enables cross-sub-business system and collaborative business system joint modeling while protecting data privacy and compliance, ensuring global collaborative optimization while also considering the personalized model needs of each sub-business system.
[0039] As a non-limiting implementation, before the main business system first distributes the model to be trained to each sub-business system, the method further includes:
[0040] The training samples are input into the base model and multiple iterations are performed to obtain the loss function value for each iteration.
[0041] Based on the loss function values and weights corresponding to each round, the sensitivity score of the model parameters is determined;
[0042] Select a first preset number of model parameters based on the sensitivity score, and determine the model to be trained based on the preset sparsity and the first preset number of model parameters.
[0043] Before distributing the model to be trained to each sub-business system after the initial distribution from the main business system, the method further includes:
[0044] Receive optimization models uploaded by each sub-business system and aggregate the optimization models;
[0045] The aggregated model is used as the new model to be trained.
[0046] For example, before the main business system distributes the model to be trained to each sub-business system for the first time, the main business system needs to initialize the model. Specifically, training samples are input into the base model, and multiple iterations are performed to obtain the corresponding loss function value in each iteration. Based on the loss function value and weights corresponding to each iteration, a sensitivity score for the model parameters is determined, reflecting the degree of influence of model parameter optimization on the decrease of the loss function. Generally, ignoring the weights, the sensitivity score of the model parameters is positively correlated with the degree of decrease in the loss function; that is, the greater the decrease in the loss function, the higher the sensitivity score of the model parameters. Weights can also be set for the loss function values corresponding to each iteration to reflect the importance of the degree of decrease in the loss function in different iterations to the overall model training. Based on the loss function value and weights corresponding to each iteration, the sensitivity score of the model parameters is determined. The sensitivity score reflects the importance of the model parameters in model training. Generally, the importance of the model parameters in model training is positively correlated with the sensitivity score. That is, the higher the sensitivity score, the greater the importance of the model parameters in model training, and the more important it is to adjust these model parameters during model training optimization. To achieve sparsity in the business model, reduce computational resource consumption, and improve training efficiency, a first preset number of model parameters can be selected based on sensitivity scores. These first preset number of parameters are retained for model construction, while other model parameters are set to zero and not involved. Generally, the first preset number of model parameters with the highest sensitivity scores are selected for model construction. The model to be trained is determined based on the preset sparsity and the first preset number of model parameters. In the model network structure, sparsity typically refers to the proportion of non-zero parameters or connections in the neural network, used to measure the simplification of the model structure.
[0047] If the main business system distributes the training models to each sub-business system after the initial distribution, meaning after each sub-business system has trained and optimized its models and returned them to the main business system, the main business system needs to determine a new training model based on the optimized models reported by each sub-business system. The main business system receives the optimized models uploaded by each sub-business system, aggregates them, and uses the aggregated model as the new training model. Aggregating optimized models enables global training and learning, combining the model parameter characteristics of optimized models from various sub-business systems, and leveraging global knowledge to improve performance.
[0048] In this embodiment of the application, the process of the main business system aggregating the optimization models reported by each sub-business system specifically involves using... As a loss function, the personalized federated learning problem can be expressed as: W: The set of parameters for the optimized models reported by all sub-business systems, i.e., W = {w1, w2, wN}. N: The number of sub-business systems participating in training. Xi: The local dataset of the i-th sub-business system. E: Expected loss, representing the average loss calculated across the data distribution of the sub-business systems. L: The loss value calculated using the model on sample xy. This formula indicates that the goal of personalized federated learning is to minimize the average loss of the optimized models across all sub-business systems on their respective data. In other words, it aims to minimize the loss function value calculated across the entire business system by finding the optimal model parameters w. i The model parameters in the optimization models reported by different sub-business systems may differ. The N in the above formula refers to the calculation performed on the same model parameters included in the optimization model. However, after adopting dynamic sparse training, the optimization formula is redefined as follows: Where M i M is a binary mask matrix. M is the set of all row-wise mask matrices, M = {M1, M2, MN}. Mi: A binary mask matrix (elements are 0 or 1) used to select the parameters to be retained in the optimization model of the i-th sub-system. Hadamard product: Under dynamic sparse training, each sub-system maintains a local sparse mask Mi, and the objective function becomes minimizing the average loss of all sub-systems using the sparse model on their respective data.
[0049] Figure 2 This is a flowchart illustrating a method for constructing a business model, as provided in another embodiment of this application. This embodiment is an optimization based on the above embodiment; solutions not described in detail in this embodiment are found in the above embodiment. Figure 2 As shown, the method in this embodiment of the application specifically includes the following steps:
[0050] S210. The training model is distributed to each sub-business system through the main business system, and the training model is received by each sub-business system.
[0051] S220. For each sub-business system, perform adaptive dynamic sparse training on the model to be trained to determine the sparse structure of the model to be trained.
[0052] Adaptive dynamic sparse training is a technique that dynamically adjusts the sparsity of parameters during deep learning model training. It achieves a balance between computational efficiency and model performance by optimizing the neural network connection structure in real time. After each sub-system receives the model to be trained, it can continue to perform adaptive dynamic sparse training on the model, thereby further refining the sparse structure of the model. By determining the sparse structure of the model through adaptive sparse training, it is possible to adaptively determine a sparse structure suitable for the data within the sub-system based on the data characteristics within that sub-system, optimizing the structure of the model to be trained and facilitating rapid optimization while improving training speed.
[0053] In this embodiment of the application, adaptive dynamic sparse training is performed on the model to be trained to determine the sparse structure of the model to be trained, including:
[0054] Based on the pruning rate, the current sparsity of the sparse structure, and the model parameters of the model to be trained, the model gradient index is determined, and the second preset number of model parameters with the largest model gradient index are retained.
[0055] Based on the descent gradient of the loss function value obtained from multiple rounds of training, a third preset number of model parameters is determined; where the third preset number is the difference between the first preset number and the second preset number.
[0056] Based on the current sparsity and a first preset number of model parameters, the sparse structure of the model to be trained is determined.
[0057] The pruning rate refers to the proportion of weights (or neurons) removed from the total parameters of a neural network, usually expressed as a percentage. For example, a 30% pruning rate means that 30% of the weights in the model are permanently set to zero or deleted. The pruning rate directly affects the sparsity of the model; a higher pruning rate can significantly reduce computation, but may lead to a decrease in accuracy. Therefore, a trade-off must be made between compression efficiency and performance.
[0058] Specifically, the model gradient metric is determined based on the pruning rate, the current sparsity of the sparse structure, and the model parameters of the model to be trained. In sparse network architectures of deep learning, the model gradient serves as the core basis for parameter adjustment during training, and its dynamic characteristics directly reflect the network's learning efficiency, generalization ability, and optimization path. Calculating the model gradient metric reflects the training optimization efficiency based on the current model parameters. Specifically, the pruning rate can be... The pruning rate gradually decreases with each cycle. The model gradient index K = (1-f) decay )×S grad ×|w new |,S grad Given the current sparsity, w newThese are the model parameters for the model to be trained. A second preset number of model parameters with the largest gradient indices can be retained; these are the parameters that have the greatest impact on the model training and optimization speed, and are used for subsequent model optimization. To ensure the invariance of the overall number of parameters in the sparse structure, other model parameters can be determined to bring the total number of selected model parameters to a first preset number. Specifically, based on the descent gradient of the loss function value obtained from multiple training rounds, a third preset number of model parameters are determined. This means that based on the descent gradient of the loss function value, other model parameters besides the second preset number are selected, resulting in the third preset number of model parameters, which, together with the second preset number of model parameters, form the first preset number of model parameters for model optimization. Based on the current sparsity and the first preset number of model parameters, the sparse structure of the model to be trained is determined.
[0059] In this embodiment of the application, the sub-business system performs this round of model training based on the model to be trained to obtain the optimized model for this round of training, including:
[0060] Real-time monitoring of local data, calculation of the first data distribution within a first time period and the second data distribution within a second time period; wherein the first time period is earlier than the second time period;
[0061] The difference between the first data distribution and the second data distribution is determined. If the difference is greater than a preset difference threshold, the current sparsity is dynamically adjusted.
[0062] In this embodiment, to ensure the model to be trained meets the business requirements under the current local data distribution and has superior processing capabilities for incremental local data, local data can be used in real-time, and the data distribution characteristics of the local data can be used to participate in model training and optimization. Specifically, local data is monitored in real-time, and a first data distribution within a first time period and a second data distribution within a second time period are calculated. The first time period is earlier than the second time period, and the second time period is generally the time period from the current time to a past moment before the current time. The first time period is generally a time period before a past moment. The difference data between the first time period and the second time period is determined. The difference data reflects whether the data distribution has drifted and whether the model needs to be adjusted to adapt to the second data distribution.
[0063] If the difference in data exceeds a preset difference threshold, the current sparsity is dynamically adjusted. If the difference in data is less than or equal to the preset difference threshold, no adjustment to the current sparsity is needed; training and optimization can proceed directly based on the loss function. Specifically, the first data distribution can be used... This means that the second data distribution can be represented as... The difference between the first and second data distributions can be represented as follows: If the difference data DL > ε, it is determined that the second data distribution has drifted, triggering dynamic adjustment to adjust the current sparsity to suit the processing of the current incremental data. ε is a preset difference data threshold.
[0064] In this embodiment of the application, the dynamic adjustment of the current sparsity includes:
[0065] Based on the difference data and the preset difference data threshold, a reference expansion ratio is determined;
[0066] The minimum value between the reference expansion ratio and the upper limit of the expansion ratio shall be used as the target expansion ratio;
[0067] The updated current sparsity is determined based on the target expansion ratio and the current sparsity.
[0068] In this embodiment, after real-time monitoring shows that local data meets the conditions for dynamic adjustment, the current sparsity is dynamically adjusted. Specifically, this involves determining a reference expansion ratio based on the difference data and a preset difference data threshold, typically the ratio of the difference data to the preset difference data threshold. An upper limit for the expansion ratio can be set as a constraint, and the minimum value between the reference ratio and the upper limit is taken as the target expansion ratio. Specifically, the target expansion ratio... γ max To extend the upper limit of the ratio, The target expansion ratio is used as a reference. The updated current sparsity is determined based on the target expansion ratio and the current sparsity. Specifically, S... grad =1-(1-S) γ S grad Let S be the current sparsity after the update, and S be the current sparsity before the update. When γ = 1, it remains unchanged; when γ > 1 (exceeding the threshold), it increases exponentially, and the sparsity grows faster. By increasing the sparsity, the model will retain those connection branches that have a greater impact on the task during training.
[0069] In this embodiment of the application, the process of dynamically adjusting the current sparsity based on the real-time monitoring results of local data can be performed before S220. That is, the current sparsity is dynamically adjusted based on local data first, and then the training model obtained based on the dynamic adjustment continues to be adaptively dynamically sparsely trained to determine the sparse structure of the training model in order to optimize the network structure of the training model.
[0070] S230. Based on the sparse structure of the model to be trained, the training is optimized to obtain the optimized model for this round of training.
[0071] S240. If the training round is not completed, the optimized model is uploaded to the main business system through each sub-business system. The main business system aggregates the models and distributes the aggregated models as new models to be trained to each sub-business system for training and optimization based on each sub-business system until the training round is completed, and the resulting business model is used by the sub-business systems for business processing.
[0072] In this embodiment, the main business system distributes the training model to each sub-business system, and each sub-business system receives the training model. For each sub-business system, adaptive dynamic sparse training is performed on the training model to determine its sparse structure. Based on the sparse structure of the training model, training optimization is performed to obtain the optimized model for this round of training. If the training round is not yet complete, each sub-business system uploads the optimized model to the main business system, which then aggregates the models and distributes the aggregated models as new training models to each sub-business system for separate training optimization. This process continues until the training round ends, resulting in a business model used by the sub-business systems for business processing. This approach allows for further optimization of the sparse network during model training and updates in each sub-business system. The model automatically deletes unimportant parameters, and during training, the parameters of each layer are redistributed according to their importance, adapting to each training round without increasing computational pressure. Therefore, the optimal solution for the model can be explored by repeatedly deleting and restoring parameters during the update process.
[0073] Figure 3 This is a flowchart illustrating a method for constructing a business model, as provided in another embodiment of this application. This embodiment is an optimization based on the above embodiments; solutions not described in detail in this embodiment are found in the above embodiments. Figure 3 As shown, the method in this embodiment of the application specifically includes the following steps:
[0074] S310. The training model is distributed to each sub-business system through the main business system, and the training model is received by each sub-business system.
[0075] S320. For each sub-business system, the previous round of optimization model is used as the teacher model, and the model to be trained is used as the student model.
[0076] In this embodiment, given that the sparse network structure of the training model in each sub-business system is already determined, training can be performed based on this determined model. To retain historical knowledge during model training and avoid catastrophic forgetting, the optimized model obtained from the previous training round can be frozen during the training process. This optimized model is used as the teacher model, and the current training model is used as the student model, with knowledge distillation constraints applied. The distillation loss can be designed as soft label distillation; specifically, the distillation loss can be determined based on the output probability distributions of the teacher and student models, as well as the feature layer alignment.
[0077] S330. Determine a first loss function based on the output of the teacher model and the output of the student model, and determine a second loss function based on the output of the feature extraction layer in the teacher model and the output of the feature extraction layer in the student model.
[0078] Specifically, the first loss function is determined based on the outputs of the teacher model and the student model. The outputs of the teacher and student models can refer to their probability distributions. The probability distribution of the model output refers to the model's assessment of the likelihood of different outcomes (such as classification labels, generated content, etc.) that the input data might produce. Each outcome corresponds to a probability value (between 0 and 1), and the sum of the probabilities of all outcomes is 1, reflecting the model's confidence level in different options. For example, in an image classification task, the model might output "cat (0.8), dog (0.15), bird (0.05)," indicating that it believes the input image is most likely a cat. The probability distribution is a quantitative expression of the model's uncertainty and directly affects the final decision (such as choosing the outcome with the highest probability). Specifically, the specific form of the first loss function can be... Here, KL represents the divergence, which measures the difference between two probability distributions. || represents the probability distribution output by the teacher model relative to the probability distribution output by the student model.
[0079] The second loss function can be determined based on the outputs of the feature extraction layers in the teacher model and the student model. Specifically, the second loss function... in, This represents the output of the feature extraction layer of the teacher model during the l-th round of model training. This represents the output of the feature extraction layer of the student model during the l-th round of model training.
[0080] S340. Based on the first loss function, the second loss function, and the time decay weighted loss function, determine the total loss function, and train the model to be trained to obtain the optimized model for this round of training.
[0081] For example, the total loss function can be determined based on the first loss function, the second loss function, and the time-decayed weighted loss function. The time-decayed weighted loss function is a loss function design method that dynamically adjusts sample weights during model training. Its core idea is to continuously reduce the loss contribution of historical samples through an exponential decay function, making the model more inclined to focus on the prediction error of recent samples during training, that is, to focus more on the prediction error corresponding to local data generated in the more recent time period. Specifically, the first loss function and the second loss function can be weighted and summed to obtain the output loss function, specifically, Where α represents the weights corresponding to the first loss function, and β represents the weights corresponding to the second loss function. The total loss function can be: ∑δ T-i L(x i y i Let be the time-decay weighted loss function. T is the current time, and i is the time when the training sample was generated. δ is the coefficient.
[0082] For example, the student model is trained and optimized based on the total loss function, thus preserving historical knowledge during training while enabling the model to adapt to learning the latest data.
[0083] S350. If the training round is not completed, the optimized model is uploaded to the main business system through each sub-business system. The main business system aggregates the models and distributes the aggregated models as new models to be trained to each sub-business system for training and optimization based on each sub-business system until the training round is completed, and the resulting business model is used by the sub-business systems for business processing.
[0084] The scheme in this embodiment uses the previous optimized model as the teacher model and the model to be trained as the student model. A first loss function is determined based on the outputs of the teacher model and the student model; a second loss function is determined based on the outputs of the feature extraction layers in the teacher model and the student model; a total loss function is determined based on the first loss function, the second loss function, and the time-decay weighted loss function; and the model to be trained is then trained to obtain the optimized model for this round of training. Knowledge distillation constraint, by using the original model as the teacher model and the new model as the student model, transfers knowledge from the old model using knowledge distillation technology under new data, effectively preventing catastrophic forgetting. This mechanism can enhance the model's continuous learning ability, ensuring that the model maintains its performance on old data while optimizing on new data.
[0085] The solutions described in the above embodiments can be applied to the construction of collection models for both bank headquarters and branches. Background:
[0086] Suppose that a large portion of the customer base of a branch of a certain bank consists of manufacturing workers.
[0087] Over the past year (historical data), the region's economy has been stable, with ample manufacturing orders and secure worker incomes. Historical data distribution shows:
[0088] The customer's average monthly income is stable at X yuan.
[0089] The debt-to-income ratio (debt / monthly income) is mainly concentrated in the Y range.
[0090] Among customers who are 1-30 days overdue, a relatively high percentage (e.g., Z%) are able to repay within 7 days of receiving a collection reminder (high response rate).
[0091] Characteristics: A large proportion of customers have stable income, long working years, and manageable debt.
[0092] Based on this historical data, the collection model (or the global / personalized model of federated learning) trained locally by Branch A may learn the following strategy: for most customers in the early stages of delinquency, it is effective and cost-efficient to prioritize gentle SMS / phone reminders.
[0093] Events that cause data distribution drift:
[0094] Suddenly, a large manufacturing plant in the city where Branch A is located laid off a large number of employees or even went bankrupt due to external factors (such as trade wars in major export markets, technological upgrades and obsolescence, and bankruptcy).
[0095] This incident directly affected a large number of bank clients (as well as clients in its upstream and downstream industries) who worked at the factory.
[0096] Data distribution after drift (current distribution):
[0097] Feature distribution changes:
[0098] Income characteristics: Among the newly generated overdue customers, the average monthly income has decreased significantly (from X yuan to X-Δ yuan), and there are even a large number of "unemployed" status markers.
[0099] Debt-to-income ratio: Due to a sharp drop in income or even to zero, the customer's debt-to-income ratio has soared, far exceeding the historical distribution range Y, and has entered a high-risk area.
[0100] Job stability characteristics: The proportion of current employees is decreasing, while the proportion of those "unemployed" is surging. The distribution of years of service may also shift towards lower values (newly unemployed).
[0101] Residential area characteristics: A sudden increase in customer delinquency rates may be observed in specific industrial areas / factory dormitory areas.
[0102] Label distribution changes:
[0103] Collection response rate: Even with the same mild collection reminder text messages as in the past, the customer's repayment response rate drops significantly (from historical Z% to Z-Δ%). This is because the customer genuinely cannot afford to repay, rather than simply forgetting. The mild reminders become ineffective.
[0104] For these newly unemployed customers whose incomes have plummeted, it may be necessary to intervene earlier with human customer service to understand their difficulties and negotiate repayment plans (such as installment payments or reductions), or more forceful collection methods may be required. The model needs to adapt to this change.
[0105] Figure 4 This is a schematic diagram of a business model construction apparatus provided in an embodiment of this application. This apparatus can execute the business model construction method provided in any embodiment of this application, and has corresponding functional modules and beneficial effects for executing the method. For example... Figure 4 As shown, the device includes:
[0106] The model to be trained module 410 is used to distribute the model to be trained to each sub-business system through the main business system, and to receive the model to be trained through each sub-business system.
[0107] The model training module 420 is used to train the model in this round for each sub-business system based on the model to be trained, so as to obtain the optimized model for this round of training.
[0108] The upload aggregation module 430 is used to upload the optimized model to the main business system through each sub-business system if the training round is not completed. The main business system then aggregates the models and distributes the aggregated models as new models to be trained to each sub-business system for training and optimization based on each sub-business system until the training round is completed, and the resulting business model is used by the sub-business systems for business processing.
[0109] In this embodiment of the application, before the main business system first distributes the model to be trained to each sub-business system, the apparatus further includes a model initialization module, used for:
[0110] The training samples are input into the base model and multiple iterations are performed to obtain the loss function value for each iteration.
[0111] Based on the loss function values and weights corresponding to each round, the sensitivity score of the model parameters is determined;
[0112] Select a first preset number of model parameters based on the sensitivity score, and determine the model to be trained based on the preset sparsity and the first preset number of model parameters.
[0113] The device further includes a model aggregation module, used for:
[0114] Before the main business system distributes the models to be trained to each sub-business system after the first transmission, it receives the optimized models uploaded by each sub-business system and aggregates the optimized models.
[0115] The aggregated model is used as the new model to be trained.
[0116] In this embodiment, the model training module 420 performs this round of model training based on the model to be trained through the sub-business system to obtain the optimized model for this round of training, including:
[0117] Adaptive dynamic sparse training is performed on the model to be trained to determine the sparse structure of the model to be trained.
[0118] The optimized model for this round of training is obtained by training and optimizing based on the sparse structure of the model to be trained.
[0119] In this embodiment, the model training module 420 performs adaptive dynamic sparse training on the model to be trained to determine the sparse structure of the model to be trained, including:
[0120] Based on the pruning rate, the current sparsity of the sparse structure, and the model parameters of the model to be trained, the model gradient index is determined, and the second preset number of model parameters with the largest model gradient index are retained.
[0121] Based on the descent gradient of the loss function value obtained from multiple rounds of training, a third preset number of model parameters is determined; where the third preset number is the difference between the first preset number and the second preset number.
[0122] Based on the current sparsity and a first preset number of model parameters, the sparse structure of the model to be trained is determined.
[0123] In this embodiment, the model training module 420 performs the current round of model training based on the model to be trained through the sub-business system, including:
[0124] Real-time monitoring of local data, calculation of the first data distribution within a first time period and the second data distribution within a second time period; wherein the first time period is earlier than the second time period;
[0125] The difference between the first data distribution and the second data distribution is determined. If the difference is greater than a preset difference threshold, the current sparsity is dynamically adjusted.
[0126] In this embodiment, the model training module 420 dynamically adjusts the current sparsity, including:
[0127] Based on the difference data and the preset difference data threshold, a reference expansion ratio is determined;
[0128] The minimum value between the reference expansion ratio and the upper limit of the expansion ratio shall be used as the target expansion ratio;
[0129] The updated current sparsity is determined based on the target expansion ratio and the current sparsity.
[0130] In this embodiment, the model training module 420 performs this round of model training based on the model to be trained through the sub-business system to obtain the optimized model for this round of training, including:
[0131] The previous optimized model is used as the teacher model, and the model to be trained is used as the student model.
[0132] A first loss function is determined based on the output of the teacher model and the output of the student model, and a second loss function is determined based on the output of the feature extraction layer in the teacher model and the output of the feature extraction layer in the student model.
[0133] The total loss function is determined based on the first loss function, the second loss function, and the time decay weighted loss function. The model to be trained is then trained to obtain the optimized model for this round of training.
[0134] The business model construction apparatus provided in this application embodiment can execute the business model construction method provided in any embodiment of this application, and has the corresponding functional modules and beneficial effects of the execution method.
[0135] Figure 5 A schematic diagram of an electronic device 10, which can be used to implement embodiments of this application, is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the application described and / or claimed herein.
[0136] like Figure 5As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0137] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0138] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as methods for constructing business models.
[0139] In some embodiments, the business model construction method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the business model construction method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the business model construction method by any other suitable means (e.g., by means of firmware).
[0140] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0141] Computer programs used to implement the methods of this application may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other means of building a programmable business model, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The computer programs may execute entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0142] In the context of this application, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0143] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0144] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0145] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0146] This invention also provides a computer program product, including a computer program that, when executed by a processor, implements the method for constructing a business model as provided in any embodiment of this application.
[0147] In implementing the computer program product, computer program code for performing the operations of this invention can be written in one or more programming languages or a combination thereof. Programming languages include object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0148] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this application can be executed in parallel, sequentially, or in different orders, as long as the desired information of the technical solution of this application can be achieved, and this is not limited herein.
[0149] The specific embodiments described above do not constitute a limitation on the scope of protection of this application. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A method for constructing a business model, characterized in that, The method includes: The system distributes the training model to each sub-business system through the main business system, and receives the training model through each sub-business system. For each sub-business system, the model is trained in this round based on the model to be trained, and the optimized model for this round of training is obtained. If the training round is not completed, the optimized model is uploaded to the main business system through each sub-business system. The main business system then aggregates the models and distributes the aggregated models as new models to be trained to each sub-business system. Training and optimization are then performed on each sub-business system separately until the training round is completed, and the resulting business model is used by the sub-business systems for business processing.
2. The method according to claim 1, characterized in that, Before the model to be trained is first distributed to each sub-business system by the main business system, the method further includes: The training samples are input into the base model and multiple iterations are performed to obtain the loss function value for each iteration. Based on the loss function values and weights corresponding to each round, the sensitivity score of the model parameters is determined; Select a first preset number of model parameters based on the sensitivity score, and determine the model to be trained based on the preset sparsity and the first preset number of model parameters. Before distributing the model to be trained to each sub-business system after the initial distribution from the main business system, the method further includes: Receive optimization models uploaded by each sub-business system and aggregate the optimization models; The aggregated model is used as the new model to be trained.
3. The method according to claim 1, characterized in that, The sub-business system performs this round of model training based on the model to be trained, and obtains the optimized model for this round of training, including: Adaptive dynamic sparse training is performed on the model to be trained to determine the sparse structure of the model to be trained. The optimized model for this round of training is obtained by training and optimizing based on the sparse structure of the model to be trained.
4. The method according to claim 3, characterized in that, Adaptive dynamic sparse training is performed on the model to be trained to determine the sparse structure of the model, including: Based on the pruning rate, the current sparsity of the sparse structure, and the model parameters of the model to be trained, the model gradient index is determined, and the second preset number of model parameters with the largest model gradient index are retained. Based on the descent gradient of the loss function value obtained from multiple rounds of training, a third preset number of model parameters is determined; where the third preset number is the difference between the first preset number and the second preset number. Based on the current sparsity and a first preset number of model parameters, the sparse structure of the model to be trained is determined.
5. The method according to claim 1, characterized in that, The current round of model training is conducted through the sub-business system based on the model to be trained, including: Real-time monitoring of local data, calculation of the first data distribution within a first time period and the second data distribution within a second time period; wherein the first time period is earlier than the second time period; The difference between the first data distribution and the second data distribution is determined. If the difference is greater than a preset difference threshold, the current sparsity is dynamically adjusted.
6. The method according to claim 5, characterized in that, Dynamically adjust the current sparsity, including: Based on the difference data and the preset difference data threshold, a reference expansion ratio is determined; The minimum value between the reference expansion ratio and the upper limit of the expansion ratio shall be used as the target expansion ratio; The updated current sparsity is determined based on the target expansion ratio and the current sparsity.
7. The method according to any one of claims 1-6, characterized in that, The sub-business system performs this round of model training based on the model to be trained, and obtains the optimized model for this round of training, including: The previous optimized model is used as the teacher model, and the model to be trained is used as the student model. A first loss function is determined based on the output of the teacher model and the output of the student model, and a second loss function is determined based on the output of the feature extraction layer in the teacher model and the output of the feature extraction layer in the student model. The total loss function is determined based on the first loss function, the second loss function, and the time decay weighted loss function. The model to be trained is then trained to obtain the optimized model for this round of training.
8. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the method for constructing the business model according to any one of claims 1-7.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that, when executed by a processor, implement the method for constructing the business model according to any one of claims 1-7.
10. A computer program product, characterized in that, It includes a computer program that, when executed by a processor, implements the method for constructing the business model as described in any one of claims 1-7.