A QoS-based optimization method for service provider combination of technology consulting platform
Through the research on QoS of service providers and the multi-objective mathematical optimization model combined with the Q-ACA algorithm, the problem of unclear description of service provider capabilities in the science and technology consulting platform is solved, and the optimization of service provider combination solutions is achieved, taking into account the interests of multiple parties, improving matching efficiency and effect.
Patent Information
- Application Number
- CN202210535353.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-17
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2042-05-17
AI Technical Summary
In the matching of service providers, there are problems such as unclear, incomplete and unstandard descriptions of service provider capabilities, and it is difficult to take into account the interests of the platform, employers and service providers, resulting in difficulty in selecting service provider combinations.
By studying the service quality QoS of service providers, a multi-objective mathematical optimization model is constructed, combined with the Q-ACA algorithm, the service provider combination plan is optimized, and the platform profit, employer satisfaction, service provider satisfaction and coordination are considered to achieve service provider capability assessment and combination optimization.
It realizes the precise portrayal and optimization of service provider combination solutions, promotes the precise matching of service provider combinations, takes into account the interests of multiple participating entities, and improves the efficiency and effectiveness of matching.
Smart Images

Figure CN115222088B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for optimizing service providers, and in particular to a method for optimizing a combination of service providers of a technology consulting platform based on QoS. Background Art
[0002] In the complex task-service provider matching model of the technology consulting platform, after the complex tasks submitted by the employer are broken down into multiple task packages, the platform selects a suitable service provider combination plan based on the different service provider combination plans formed by the candidate service providers, and subcontracts the complex tasks to the service providers in the plan to complete them, thereby achieving accurate matching of complex tasks and service providers.
[0003] Since the complex tasks of the science and technology consulting platform have attributes such as large task scale, large transaction amount, and high degree of personalization, they have high requirements for matching service providers. Therefore, in addition to considering traditional influencing factors such as time and cost, we should also consider the impact of factors such as the service provider's QoS and the degree of coordination between service providers on the matching of complex tasks and service providers. Summary of the Invention
[0004] To address the shortcomings of existing technologies, this paper provides a QoS-based method for optimizing the combination of technology consulting platform service providers. This method addresses issues such as unclear, incomplete, and non-standard descriptions of technology consulting platform service providers' capabilities. It studies the QoS of technology consulting platform service providers, identifies their service capabilities, and constructs a multi-objective mathematical optimization model that balances the interests of the platform, employers, and service providers, thereby obtaining a suitable service provider combination solution. This effectively solves the problem of balancing the interests of multiple participating parties when matching service providers.
[0005] To achieve the above objectives, the present invention proposes a QoS-based method for optimizing the combination of service providers for a technology consulting platform, comprising:
[0006] Evaluate the service capabilities of service providers and obtain the service providers with the best service capabilities;
[0007] Taking into account the interests of employers, service providers and the platform, a multi-objective mathematical optimization model is constructed with the goals of maximizing the profits of the technology consulting platform, maximizing employer satisfaction, maximizing service provider satisfaction and maximizing the synergy between service providers;
[0008] The multi-objective mathematical optimization model is solved based on the predefined Q-ACA algorithm to obtain the optimal solution for the service provider combination plan.
[0009] Preferably, the evaluation of the service capabilities of the service providers and obtaining the service provider with the best service capabilities also includes: determining the service provider with the best service capabilities by maximizing the service provider satisfaction and the degree of synergy between service providers through calculation of the service provider satisfaction and the degree of synergy between service providers.
[0010] Furthermore, the satisfaction of the service provider is calculated by the following formula:
[0011] The satisfaction of the service providers is affected by the expected workload range of each service provider's task package and the actual workload of the task package. The calculation formula is:
[0012]
[0013] when When the satisfaction of the jth candidate service provider of the i-th task package is when When, satisfaction when When, satisfaction when When, satisfaction when When the workload of the task package exceeds the maximum workload expected by the service provider, the satisfaction gradually decreases.
[0014] Furthermore, the degree of collaboration between the service providers is calculated using the following formula:
[0015]
[0016] Where: t ij is the time required for the jth candidate service provider of the i-th task package to complete the task package i independently, t i'j' is the time required for the j'th candidate service provider of the i'th task package to complete the task package i' independently, t ij-i'j' The time required for the jth candidate service provider of the i-th task package and the j'th candidate service provider of the i'-th task package to collaboratively complete task package i and task package i'.
[0017] Preferably, the multi-objective mathematical optimization model is determined by the following formula:
[0018]
[0019]
[0020]
[0021]
[0022]
[0023]
[0024] Tan ij ×x ij ≥Tan min
[0025] Rel ij ×x ij ≥Rel min
[0026] Res ij ×x ij ≥Res min
[0027] Ass ij ×x ij ≥Ass min
[0028] Emp ij ×x ij ≥Emp min
[0029]
[0030]
[0031] x ij ∈{0,1},i=1,2,...,n; j=1,2,...,m
[0032] In the formula, tv is the transaction amount signed between the employer and the platform; c ij The unit workload quotation of the jth candidate service provider for the i-th task package; wl i is the rated workload of the i-th task package; n is the number of task packages; is the historical employer satisfaction of the jth candidate service provider for the i-th task package; is the satisfaction of the jth candidate service provider for the i-th task package; col ij-i'j' Cap is the degree of collaboration between the jth candidate service provider of the i-th task package and the j'th candidate service provider of the i'th task package; ij is the existing idle working capacity of the jth candidate service provider for the i-th task package; T max The upper limit of the time required to complete a complex task; C max is the highest subcontracting cost expected by the platform; Tan ij is the QoS tangibility dimension score of the jth candidate service provider for the i-th task package; Tan minThe minimum requirement for the platform to score the QoS tangible dimension of the service provider; Rel ij Rel is the QoS reliability dimension score of the jth candidate service provider of the i-th task package; min The minimum requirement for the platform to score the QoS reliability dimension of the service provider; Res ij Res is the QoS responsiveness dimension score of the jth candidate service provider for the i-th task package; min The minimum requirement for the platform's QoS responsiveness score for service providers; Ass ij Ass is the QoS assurance dimension score of the jth candidate service provider of the i-th task package; min The minimum requirement for the platform to evaluate the QoS guarantee dimension of the service provider, Emp ij is the QoS empathy dimension score of the jth candidate service provider for the i-th task package; Emp min The minimum requirement for the platform's QoS empathy score for service providers; is the lower limit of the expected workload of the j-th candidate service provider for the i-th task package; is the initial satisfaction of the jth candidate service provider for the i-th task package; is the upper limit of the expected workload of the jth candidate service provider for the i-th task package; link ij-i'j' If the jth candidate service provider of the i-th task package has a task completion process association with the j'th candidate service provider of the i'th task package, it is 1; otherwise, it is 0; ij It is 1 when the jth candidate service provider of the i-th task package matches the i-th task package, otherwise it is 0.
[0033] Furthermore, the pre-definition of the Q-ACA algorithm includes: combining the Q-learning algorithm and the ant colony algorithm (ACA), using the Q value table in the Q-learning algorithm as the pheromone initial value in the ant colony algorithm (ACA), and performing non-dominated sorting on the solution results of the ant colony algorithm (ACA); and selecting the optimal solution from the Pareto solution set through the GRA and TOPSIS methods.
[0034] Furthermore, the Q-ACA algorithm specifically includes:
[0035] a. Initialize the Q-learning algorithm parameters so that the cumulative reward function value of all states is 0;
[0036] b. Each subtask is considered a state, and the set of candidate service providers for the subtask is considered the action in each state. When selecting an action in each state, the action that maximizes the cumulative reward function value of the state is selected. After the action is executed, the cumulative reward function value of the state is updated based on the state-action reward value.
[0037] c. Determine whether it is the terminal state. If it is the terminal state, perform a new iteration; otherwise, enter the next state and repeat step b;
[0038] d. Determine whether the cumulative reward function value of all states converges. If so, end the iteration; otherwise, perform a new iteration;
[0039] e. Standardize the Q-value tables of each objective function and perform arithmetic averaging to generate a comprehensive Q-value table; the objective functions are maximizing platform profits, maximizing employer satisfaction, and maximizing average service provider satisfaction;
[0040] f. ACA parameter initialization, where the initial pheromone value is the comprehensive Q value table;
[0041] g. Generate the initial solution. Each ant starts from the initial state and ends at the terminal state, visiting an action in each state;
[0042] h. Calculate the objective function based on the actions of each ant and update the pheromone concentration on the actions visited by the ants;
[0043] i. Determine whether the maximum number of iterations is met. If so, proceed to step j; otherwise, proceed to step h;
[0044] j. Perform non-dominated sorting on all solutions and use GRA-TOPSIS to select the optimal solution in the Pareto solution set.
[0045] Furthermore, the ant colony algorithm ACA specifically includes: a comprehensive Q value table Q, a Q value table Q of each objective function i , the number of ants M and related data parameters P are input, and the output is the Pareto solution set S;
[0046] Let the initial number of iterations be 0, according to the probability Select the action in the current state;
[0047] Among them, τ ij Indicates the pheromone accessing action j in state i, and its initial value comes from the comprehensive Q value table Q; η ij Indicates the heuristic function value of accessing action j in state i, whose value comes from the Q value table Q of a randomly selected objective function iThe corresponding value; calculate the objective function value according to the action visited by each ant, and perform non-dominated sorting to obtain the Pareto level F of each ant k ; Calculate the pheromone concentration released by each ant when visiting the action Where G is the total amount of pheromone released by an ant in one iteration; Update pheromones.
[0048] Compared with the closest prior art, the present invention also has the following beneficial effects:
[0049] The present invention relates to a method for optimizing the combination of service providers of a science and technology consulting platform based on QoS. First, in order to address the problems of unclear, incomplete and non-standard descriptions of the capabilities of service providers of the science and technology consulting platform, the service quality QoS of the service providers is studied based on historical transaction data and service provider information, so as to evaluate the ability of service providers to meet the needs of employers and obtain service providers with the best service capabilities. Through the above scheme, a QoS evaluation system for service providers of the science and technology consulting platform is constructed, which can accurately characterize the service characteristic attributes of the service providers, establish a "portrait" of the service providers, and effectively promote the optimization of the combination of service providers.
[0050] Secondly, based on the evaluation of service provider QoS, and taking into account the interests of the platform, employers, and service providers, a multi-objective mathematical optimization model is constructed, considering multiple attributes such as service provider QoS, quotation dimension, time dimension, and service provider collaboration. A new Q-ACA algorithm is designed by combining reinforcement learning and heuristic algorithms to solve the multi-objective mathematical optimization model and thus determine the optimal solution for the service provider combination. This solves the problem of the difficulty of balancing the interests of multiple stakeholders in the optimal combination of service providers on the science and technology consulting platform. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] In order to more clearly illustrate the specific implementation of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for the specific implementation or the description of the prior art.
[0052] Figure 1 This is a flow chart of a QoS-based technology consulting platform service provider combination optimization method provided in a specific embodiment of the present invention;
[0053] Figure 2 This is a schematic diagram of a combination optimization scheme for technology consulting platform service providers provided in a specific embodiment of the present invention;
[0054] Figure 3 is a diagram of a service provider satisfaction function provided in a specific embodiment of the present invention;
[0055] Figure 4 is a flow chart of the Q-ACA algorithm provided in a specific embodiment of the present invention;
[0056] Figure 5 is a four-dimensional spatial distribution diagram of the Pareto solution set in Example 1 of the present invention;
[0057] Figure 6 It is a box plot of the simulation experiment results in Example 1 of the present invention. DETAILED DESCRIPTION
[0058] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0059] like Figure 2 As shown in the figure, in the complex task-service provider matching model of the science and technology consulting platform, the service provider of the complex task matching will directly affect the final task completion. Therefore, the goal of complex task-service provider matching is to assign each task package to the most suitable service provider, so that the matching solution can maximize the benefits of all participating entities while meeting the conditions such as skill requirements and time constraints. Based on this, Figure 1 As shown, the specific embodiment of the present invention proposes a QoS-based technology consulting platform service provider combination optimization method, the specific steps of the method are as follows:
[0060] S1 evaluates the service capabilities of service providers and obtains the service providers with the best service capabilities;
[0061] S2 comprehensively considers the interests of employers, service providers and platforms, and constructs a multi-objective mathematical optimization model with the goals of maximizing the profits of the technology consulting platform, maximizing employer satisfaction, maximizing service provider satisfaction and maximizing the synergy between service providers;
[0062] S3 solves the multi-objective mathematical optimization model based on a predefined Q-ACA algorithm to obtain the optimal solution for the service provider combination plan.
[0063] In step S1, the service capability of the service provider is evaluated by calculating the service provider QoS using DEMATEL-TOPSIS to obtain the QoS of the service provider.
[0064] In a specific embodiment of the present invention, a service provider QoS calculation method combining the DEMATEL and TOPSIS methods is proposed for evaluating the service capability of the service provider:
[0065] ①DEMATEL method determines the weight of QoS evaluation indicators
[0066] Step 1 C={c1,c2,L,ci ,L,c j ,L,c k ,L,c n} is the set of QoS evaluation indicators, and n is the number of QoS evaluation indicators. m} represents the expert team, and m represents the number of experts. Use the language in Table 1-1 to evaluate the mutual impact between the two indicators.
[0067] Table 1-1 Level 5 Language Assessment Set
[0068]
[0069] The opinions of the expert team are averaged to obtain the direct relationship matrix Z, as shown below. In the matrix Z, z ij Indicates index c i For indicator c j Since there is no self-influence relationship in the DEMATEL operation rules, the diagonal elements in the matrix Z are 0.
[0070]
[0071] Step 2: Get the normalized direct influence matrix N by the following formula:
[0072]
[0073]
[0074] Step 3: The comprehensive influence matrix T is obtained by the following formula, where I is the unit matrix.
[0075] T=N(IN) -1
[0076]
[0077] Step 4 calculates the sum of the elements of the comprehensive influence matrix T in the horizontal and vertical directions by the following formula.
[0078]
[0079]
[0080] Where D i Indicates the degree of influence of indicator i on other indicators, R i Indicates the degree to which indicator i is affected by other indicators. Calculate the centrality of each indicator (D i +R i ) and causal degree (D i -R i ), for index i, if (Di -R i ) is a positive number, then the indicator is a cause indicator; otherwise, the indicator is a result indicator. Then, draw a centrality-cause relationship diagram, with the horizontal axis and vertical axis using centrality (D i +R i ) and causal degree (D i -R i )express.
[0081] Step 5: According to the centrality-causality relationship diagram, the weight of each indicator is determined by the following formula.
[0082]
[0083] Finally, the normalized weight of each indicator is obtained by the following formula.
[0084]
[0085] Where, is the normalized weight of the indicator and
[0086] ② QoS priority ranking of service providers based on TOPSIS
[0087] Step 1: Assume that there is a multi-attribute decision-making problem, where there are m service providers P = {p1, p2, ..., p m} and n QoS evaluation indicators C={c1,c2,...,c n For the service providers evaluated by n QoS evaluation indicators, a decision matrix X is constructed. W={w1,w2,...,w n} is the relative weight of QoS evaluation index obtained by DEMATEL method, where
[0088] Step 2: Use the following formula to obtain the normalized decision matrix R.
[0089]
[0090] Among them, I is the benefit-type indicator set, and J is the cost-type indicator set.
[0091] Then the weighted normalized decision matrix R* is obtained by the following formula.
[0092]
[0093] Step 3: The ideal solution includes the positive ideal solution (PIS) and the negative ideal solution (NIS), which are obtained by the following formulas:
[0094]
[0095]
[0096] Step 4: Calculate the service provider p by the following formula i Distance to PIS and NIS.
[0097]
[0098]
[0099] Step 5 defines the closeness coefficient C i Determine the ranking order of all service providers. The closeness coefficient can be obtained by the following formula.
[0100]
[0101] Step 6 arranges the service providers in descending order according to the closeness coefficient values and visualizes the QoS of the service providers in each dimension.
[0102] After calculating and obtaining the QoS of the service provider in each dimension in step S1, the service provider with the best service capability is selected from the service provider based on the QoS of the service provider, and the service provider with the best service capability is used as the candidate service provider. By calculating the service provider satisfaction of the candidate service providers and the degree of coordination between the service providers, the service provider with the best capability that maximizes the service provider satisfaction and the degree of coordination between the service providers is determined.
[0103] 1. Calculation of service provider satisfaction
[0104] The satisfaction of the service providers is affected by the expected workload range of each service provider's task package and the actual workload of the task package. The calculation formula is:
[0105]
[0106] The service provider satisfaction function diagram is as follows: Figure 3 As shown; when When the satisfaction of the jth candidate service provider of the i-th task package is when When, satisfaction when When, satisfaction when When, satisfaction when When the workload of the task package exceeds the maximum workload expected by the service provider, the satisfaction gradually decreases.
[0107] 2. Calculation of service provider collaboration
[0108] The degree of coordination between service providers is mainly reflected by time, so this paper uses the completion time of the task package to evaluate the degree of coordination between service providers. For example, the coordination between the jth candidate service provider of the i-th task package and the j'th candidate service provider of the i'th task package is calculated as follows:
[0109]
[0110] Where, t ij is the time required for the jth candidate service provider of the i-th task package to complete the task package i independently, t i'j' is the time required for the j'th candidate service provider of the i'th task package to complete the task package i' independently, t ij-i'j' The time required for the jth candidate service provider of the i-th task package and the j'th candidate service provider of the i'-th task package to collaboratively complete task package i and task package i'.
[0111] Because complex tasks are composed of multiple task packages, each containing one or more basic relationship structures, the different relationship structures between task packages affect the calculation of their collaborative completion time. When calculating the collaborative completion time between task packages, the main types of relationships are serial, parallel, and interactive coupling. Under different relationship structures, the collaborative completion time of task packages is calculated using the following formula.
[0112]
[0113] Where: ij-i'j' is the interaction coupling coefficient between the jth candidate service provider of the i-th task package and the j'th candidate service provider of the i'th task package, ξ ij-i'j' ∈[-1,1],ξ ij-i'j' It mainly depends on the number of previous cooperation and the effect of cooperation. The more times of previous cooperation or the better the effect of cooperation, the ij-i'j' The smaller the value, the smaller the ij-i'j' The larger the value.
[0114] Therefore, the synergy matrix among technology consulting platform service providers is as follows:
[0115]
[0116] In step S2, this embodiment considers factors such as the subcontracting cost of the task package, the total time to complete the task, and the QoS of the service provider, and establishes a multi-objective mathematical optimization model for the optimal combination of technology consulting platform service providers based on QoS. This model considers the interests of multiple participating entities, so the objective function is as follows:
[0117] ① Maximizing the profits of technology consulting platforms. The platform earns a profit margin when matching employers and service providers, which is the foundation of its sustainable development. The platform's profit is the difference between the transaction amount signed between the employer and the platform and the cost of subcontracting the task package to the service provider.
[0118] Maximize employer satisfaction. As a service-oriented crowdsourcing service, tech consulting requires employers to be satisfied not only with the task completion but also with the service experience. This time, employer satisfaction is predicted based on the service provider's historical employer satisfaction scores.
[0119] ③ Maximize service provider satisfaction. As key participants in technology consulting services, the platform should fully consider their perspectives. Only in this way can the entire platform achieve healthy and sustainable development. Predict service provider satisfaction based on the workload expected by the service provider.
[0120] ④ Maximize the degree of collaboration among service providers. Since complex tasks require the collaboration of multiple service providers, the degree of collaboration among service providers will directly affect the completion process of complex tasks and the final delivery plan.
[0121] 2.1 Model Assumptions:
[0122] Based on the actual modeling situation, this paper makes the following assumptions:
[0123] ① There is no duplication of service providers in the service provider candidate set of each task package, that is, a service provider can only appear in the service provider candidate set of one task package;
[0124] ② The unit workload quotation of each service provider for the task package, the service provider's expected workload range for the task package, the service provider's existing idle work capacity, and the degree of coordination between service providers are known.
[0125] Related symbols and instructions:
[0126] The problem description, model assumptions, and parameters and variables involved in the mathematical model are shown in the following table.
[0127]
[0128]
[0129] The parameter unit for transaction amount and cost is RMB, the parameter unit for unit workload quotation is RMB / man-day, the parameter unit for rated workload is man-day, the parameter unit for existing idle working capacity is man (the existing idle working capacity of the service provider is the relative value of the industry average working capacity. Due to differences in capacity levels, the existing idle working capacity of the service provider may be non-integer), and the parameter unit for time is day.
[0130] 2.2 Model construction:
[0131] The multi-objective mathematical optimization model is determined by the following formula:
[0132]
[0133]
[0134]
[0135]
[0136]
[0137]
[0138] Tan ij ×x ij ≥Tan min
[0139] Rel ij ×x ij ≥Rel min
[0140] Res ij ×x ij ≥Res min
[0141] Ass ij ×x ij ≥Ass min
[0142] Emp ij ×x ij ≥Emp min
[0143]
[0144]
[0145] x ij ∈{0,1},i=1,2,...,n; j=1,2,...,m
[0146] In the formula, tv is the transaction amount signed between the employer and the platform; c ij The unit workload quotation of the jth candidate service provider for the i-th task package; wl i is the rated workload of the i-th task package; n is the number of task packages; is the historical employer satisfaction of the jth candidate service provider for the i-th task package; is the satisfaction of the jth candidate service provider for the i-th task package; colij-i'j' Cap is the degree of collaboration between the jth candidate service provider of the i-th task package and the j'th candidate service provider of the i'th task package; ij is the existing idle working capacity of the jth candidate service provider for the i-th task package; T max The upper limit of the time required to complete a complex task; C max is the highest subcontracting cost expected by the platform; Tan ij is the QoS tangibility dimension score of the jth candidate service provider for the i-th task package; Tan min The minimum requirement for the platform to score the QoS tangible dimension of the service provider; Rel ij Rel is the QoS reliability dimension score of the jth candidate service provider of the i-th task package; min The minimum requirement for the platform to score the QoS reliability dimension of the service provider; Res ij Res is the QoS responsiveness dimension score of the jth candidate service provider for the i-th task package; min The minimum requirement for the platform's QoS responsiveness score for service providers; Ass ij Ass is the QoS assurance dimension score of the jth candidate service provider of the i-th task package; min The minimum requirement for the platform to evaluate the QoS guarantee dimension of the service provider, Emp ij is the QoS empathy dimension score of the jth candidate service provider for the i-th task package; Emp min The minimum requirement for the platform's QoS empathy score for service providers; is the lower limit of the expected workload of the j-th candidate service provider for the i-th task package; is the initial satisfaction of the jth candidate service provider for the i-th task package; is the upper limit of the expected workload of the jth candidate service provider for the i-th task package; link ij-i'j' If the jth candidate service provider of the i-th task package has a task completion process association with the j'th candidate service provider of the i'th task package, it is 1; otherwise, it is 0; ij It is 1 when the jth candidate service provider of the i-th task package matches the i-th task package, otherwise it is 0.
[0147] Since the multi-objective mathematical optimization model constructed in step S2 of the previous section is a typical combinatorial optimization model, and its objective functions influence and restrict each other, step S3 of the embodiment of the present invention solves the multi-objective mathematical optimization model based on the predefined Q-ACA algorithm, and designs a Q-ACA algorithm that combines the Q-learning algorithm and the Ant Colony Algorithm (ACA) to solve it.
[0148] 3.1 Q-ACA Algorithm Principle and Process
[0149] In a specific embodiment of the present invention, ACA is used as a framework to improve problem-solving performance by combining methods such as Q-learning, non-dominated sorting, Grey Relation Analysis (GRA), and TOPSIS, targeting the characteristics of the problem being solved. Due to its powerful and efficient decision-making capabilities, reinforcement learning has been widely used to solve combinatorial optimization problems. The optimal selection of service providers for technology consulting platforms is a typical combinatorial optimization problem, and reinforcement learning methods are effective in solving this problem.
[0150] The Q-learning algorithm is a typical reinforcement learning algorithm with significant advantages in solving discrete optimization problems. It consists of four elements: state s, action a, state-action reward r, and cumulative reward function Q(s, a). During iteration, each potential action a under state s is examined, and the optimal strategy is found using the state-action reward r and cumulative reward function Q(s, a). The iterative equation for the Q-learning algorithm is as follows:
[0151]
[0152] Where, Q(s t ,a t ) is in state s t Next take action a t The cumulative reward value obtained; α is the learning step size, which controls the learning efficiency of the algorithm; r t+1 For state s t Next take action a t The state-action reward obtained; γ is the discount factor, γ∈[0,1], which reflects the relationship between immediate rewards and future rewards; For state s t+1 The maximum cumulative reward value among all actions under .
[0153] The Q-learning algorithm is a simple and efficient model-free reinforcement learning algorithm, which is widely used to solve decision-making problems with a limited number of states and actions. The Q-learning algorithm does not need to model the environment, so it can effectively solve the problem of optimizing the combination of service providers. Since the mathematical model constructed in the previous step is a multi-objective mathematical optimization model, the Q value table in the Q-learning algorithm is used as the initial value of the pheromone in ACA, and the solution results of ACA are non-dominated sorted, and finally the optimal solution is selected from the Pareto solution set through the GRA and TOPSIS methods. Based on this, the main steps of the Q-ACA algorithm predefined by the present invention are as follows, and the flow chart is as follows. Figure 4 shown.
[0154] Step a: Initialize the Q-learning algorithm parameters and set the cumulative reward function value of all states to 0.
[0155] In step b, each subtask is considered a state, and the set of candidate service providers for the subtask is considered the action in each state. When selecting an action in each state, the action that maximizes the cumulative reward function value for that state is chosen. After executing the action, the cumulative reward function value for that state is updated based on the state-action reward value.
[0156] Step c determines whether it is the terminal state. If it is the terminal state, a new iteration is performed; otherwise, the next state is entered and step b is repeated.
[0157] Step d determines whether the cumulative reward function values of all states converge. If so, the iteration ends; otherwise, a new iteration is performed.
[0158] Step e: Standardize the Q-value tables of each objective function and perform arithmetic averaging to generate a comprehensive Q-value table. The objective functions here only consider maximizing platform profits, maximizing employer satisfaction, and maximizing average service provider satisfaction.
[0159] Step f: ACA parameter initialization, where the initial pheromone value is the comprehensive Q value table.
[0160] In step g, the initial solution is generated. Each ant starts from the initial state and ends at the terminal state, visiting an action in each state.
[0161] In step h, the objective function is calculated based on the actions that each ant has taken, and the pheromone concentration on the actions visited by the ants is updated.
[0162] Step i determines whether the maximum number of iterations is met. If so, proceed to step j; otherwise, proceed to step h.
[0163] In step j, all solutions are non-dominated and the optimal solution is selected from the Pareto solution set using GRA-TOPSIS.
[0164] 3.2 About Q-value table generation strategy design:
[0165] Based on the idea of Q-learning algorithm, the present invention designs a Q-value table generation strategy for the problem of optimizing the combination of service providers of technology consulting platform based on QoS. Obtaining the Q-value table in the Q-learning algorithm includes:
[0166] Assume that the Q-learning algorithm consists of four elements: the starting state s, the action a, the state-action reward r, and the cumulative reward function Q(s, a); during iteration, examine each potential action a under state s and select an action a in a t , and get r t+1 and s t+1 ; Through the iterative equation of the Q-learning algorithm, find the optimal strategy based on the state-action reward r and the cumulative reward function Q(s, a); generate the Q value table of each objective function, and take the arithmetic average to generate a comprehensive Q value table.
[0167] The iterative equation of the Q-learning algorithm is as follows:
[0168]
[0169] Where, Q(s t ,a t ) is in the initial state s t Next take action a t The cumulative reward value obtained; α is the learning step size, which controls the learning efficiency of the algorithm; r t+1 For state s t Next take action a t The state-action reward obtained; γ is the discount factor, γ∈[0,1] reflects the relationship between immediate reward and future reward; For state s t+1 The maximum cumulative reward value among all actions under .
[0170] 3.3 Improved ant colony algorithm design:
[0171] Based on the idea of ACA and the Q-value table generated by the Q-learning algorithm, this paper designs an improved ACA algorithm for the problem of optimizing the combination of service providers of the technology consulting platform, which integrates the Q-value table Q and the Q-value table Q of each objective function. i , the number of ants M and related data parameters P are input, and the output is the Pareto solution set S;
[0172] Let the initial number of iterations be 0, according to the probability Select the action in the current state;
[0173] Among them, τ ij Indicates the pheromone accessing action j in state i, and its initial value comes from the comprehensive Q value table Q; η ij Indicates the heuristic function value of accessing action j in state i, whose value comes from the Q value table Q of a randomly selected objective function i The corresponding value; calculate the objective function value according to the action visited by each ant, and perform non-dominated sorting to obtain the Pareto level F of each ant k ; Calculate the pheromone concentration released by each ant when visiting the action Where G is the total amount of pheromone released by an ant in one iteration; Update pheromones.
[0174] 3.4 Design of optimal solution generation strategy
[0175] According to the ideas of GRA and TOPSIS methods, based on the Pareto solution set generated by improved ACA, a GRA-TOPSIS-based strategy is designed to select the optimal solution from the Pareto solution set. The pseudo code is shown in Table 3-1.
[0176] Table 3-1 Pseudocode of optimal solution generation strategy based on GRA-TOPSIS
[0177]
[0178] Example 1:
[0179] Based on the above specific implementation methods, implementation 1 of the present invention adopts the same technical concept of a QoS-based technology consulting platform service provider combination optimization method to propose a service provider combination optimization solution, which is as follows:
[0180] Based on the corresponding set of service provider candidates, by calculating the QoS of the candidate service providers in each category, the candidate service provider information for this complex task is generated, as shown in Table 2-1, where C is the service provider's unit workload price, sa C is the historical employer satisfaction of the service provider, WL is the expected workload range of the service provider, is the initial satisfaction of the service provider, Cap is the idle working capacity of the service provider, Tan is the QoS (tangible dimension) score of the service provider, Rel is the QoS (reliability dimension) score of the service provider, Res is the QoS (responsiveness dimension) score of the service provider, Ass is the QoS (assurance dimension) score of the service provider, and Emp is the QoS (empathy dimension) score of the service provider.
[0181] Table 2-1 Information on candidate service providers
[0182]
[0183]
[0184] ① Model solution
[0185] Based on the above candidate service provider information, the Q-ACA algorithm designed in this paper solves the mathematical model of QoS-based technology consulting platform service provider combination optimization. The model-related data and algorithm parameters are shown in Table 2-2 and Table 2-3. Through case simulation experiments, the Pareto solution set generated is as follows: Figure 5 The service provider portfolio optimization model constructed in this paper is a multi-objective optimization model that considers the interests of the platform, employers, and service providers. Therefore, in different situations, the interests of different participants may be prioritized, meaning that the solution will produce different service provider portfolios as the situation changes. To address this, by adjusting the weights, we generate service provider portfolios for different scenarios, as shown in Table 2-5.
[0186] Table 2-2 Model related data settings
[0187]
[0188] Table 2-3 Algorithm-related parameter settings
[0189]
[0190]
[0191] Table 2-4 Pareto solution set for optimal service provider combination
[0192]
[0193] Note: [1,3,3,3,1,1,1,4,1,1] represents the service provider combination {SP 11 ,SP 23 ,SP 33 ,SP 43 ,SP 51 ,SP 61 ,SP 71 ,SP 84 ,SP 91 ,SP 10,1}
[0194] Table 2-5 Selection of service provider combinations with different orientations
[0195]
[0196] ②Algorithm validity verification
[0197] In order to verify the effectiveness and superiority of the Q-ACA algorithm designed in this paper in solving the combination optimization problem of technology consulting platform service providers, ACA and NSGA-Ⅱ were used to solve the above case respectively, where the algorithm parameters are shown in Table 2-6. The three algorithms were run 10 times respectively, and the simulation results are shown in Table 2-7 and Figure 6 shown.
[0198] Table 2-6 Algorithm-related parameter settings
[0199]
[0200] Table 2-7 Comparative analysis of simulation experiment results
[0201]
[0202] In Table 2-7, z1, z2, z3 and z4 represent the objective functions, namely platform profit, employer satisfaction, service provider average satisfaction and service provider collaboration. Figure 6 As can be seen, Q-ACA's results are superior and more stable than those of ACA and NSGA-II in terms of platform profit, employer satisfaction, and average provider satisfaction. Q-ACA's results are slightly inferior to those of ACA and NSGA-II in terms of provider collaboration, but Q-ACA's results are more stable. Therefore, the Q-ACA used in this example is highly effective in solving the QoS-based service provider portfolio optimization model for technology consulting platforms.
[0203] It should be noted that for a more detailed description of the steps and beneficial effects of the embodiment of this method, please refer to the aforementioned surveying and mapping data visualization platform, which will not be repeated here.
[0204] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and such modifications or substitutions are intended to be within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be subject to the scope of protection of the claims.
Claims
1. A QoS-based method for selecting a combination of service providers for a technology consulting platform, comprising: Evaluate the service capabilities of service providers and obtain the service providers with the best service capabilities; Taking into account the interests of employers, service providers and the platform, a multi-objective mathematical optimization model is constructed with the goals of maximizing the profits of the technology consulting platform, maximizing employer satisfaction, maximizing service provider satisfaction and maximizing the synergy between service providers; Solve the multi-objective mathematical optimization model based on the pre-defined Q-ACA algorithm to obtain the optimal solution for the service provider combination plan; After evaluating the service capabilities of the service providers and obtaining the service provider with the best service capabilities, the method further includes: determining the service provider with the best service capabilities with the maximum service provider satisfaction and the maximum service provider synergy by calculating the service provider satisfaction and the synergy between the service providers; The satisfaction of the service provider is calculated by the following formula: The satisfaction of the service providers is affected by the expected workload range of each service provider's task package and the actual workload of the task package. The calculation formula is: when When the satisfaction of the jth candidate service provider of the i-th task package is when When, satisfaction when When, satisfaction when When, satisfaction when When the workload of the task package exceeds the maximum workload expected by the service provider, the satisfaction gradually decreases. The degree of synergy between the service providers is calculated by the following formula: Where: t ij is the time required for the jth candidate service provider of the i-th task package to complete the task package i independently, t i'j' is the time required for the j'th candidate service provider of the i'th task package to independently complete the task package i', t ij-i'j' The time required for the jth candidate service provider of the i-th task package and the j'th candidate service provider of the i'th task package to collaboratively complete task package i and task package i'; The multi-objective mathematical optimization model is determined by the following formula: Tan ij ×x ij ≥Tan min Rel ij ×x ij ≥Rel min Res ij ×x ij ≥Res min Ass ij ×x ij ≥Is min Emp ij ×x ij ≥Emp min x ij ∈{0,1},i=1,2,...,n;j=1,2,...,m In the formula, tv is the transaction amount signed between the employer and the platform; c ij The unit workload quotation of the jth candidate service provider for the i-th task package; wl i is the rated workload of the i-th task package; n is the number of task packages; is the historical employer satisfaction of the jth candidate service provider for the i-th task package; is the satisfaction of the jth candidate service provider for the i-th task package; col ij-i'j' Cap is the degree of collaboration between the jth candidate service provider of the i-th task package and the j'th candidate service provider of the i'th task package; ij is the existing idle working capacity of the jth candidate service provider for the i-th task package; T max The upper limit of the time to complete a complex task; C max is the highest subcontracting cost expected by the platform; Tan ij is the QoS tangibility dimension score of the jth candidate service provider for the i-th task package; Tan min The minimum requirement for the platform to score the QoS tangible dimension of the service provider; Rel ij Rel is the QoS reliability dimension score of the jth candidate service provider of the i-th task package; min The minimum requirement for the platform to score the QoS reliability dimension of the service provider; Res ij Res is the QoS responsiveness dimension score of the jth candidate service provider for the i-th task package; min The minimum requirement for the platform's QoS responsiveness score for service providers; Ass ij Ass is the QoS assurance dimension score of the jth candidate service provider of the i-th task package; min The minimum requirement for the platform to evaluate the QoS guarantee dimension of the service provider, Emp ij is the QoS empathy dimension score of the jth candidate service provider for the i-th task package; Emp min The minimum requirement for the platform's QoS empathy score for service providers; is the lower limit of the expected workload of the j-th candidate service provider for the i-th task package; is the initial satisfaction of the jth candidate service provider for the i-th task package; is the upper limit of the expected workload of the jth candidate service provider for the i-th task package; link ij-i'j' If the jth candidate service provider of the i-th task package has a task completion process association with the j'th candidate service provider of the i'th task package, it is 1; otherwise, it is 0; ij It is 1 when the jth candidate service provider of the i-th task package matches the i-th task package, otherwise it is 0.
2. The method according to claim 1, wherein The pre-definition of the Q-ACA algorithm includes: combining the Q-learning algorithm with the ant colony algorithm (ACA), using the Q value table in the Q-learning algorithm as the pheromone initial value in the ant colony algorithm (ACA), and performing non-dominated sorting on the solution results of the ant colony algorithm (ACA); and selecting the optimal solution from the Pareto solution set using the GRA and TOPSIS methods.
3. The method according to claim 2, wherein The Q-ACA algorithm specifically includes: a. Initialize the Q-learning algorithm parameters so that the cumulative reward function value of all states is 0; b. Each subtask is considered a state, and the set of candidate service providers for the subtask is considered the action in each state. When selecting an action in each state, the action that maximizes the cumulative reward function value of the state is selected. After the action is executed, the cumulative reward function value of the state is updated based on the state-action reward value. c. Determine whether it is the terminal state. If it is the terminal state, perform a new iteration; otherwise, enter the next state and repeat step b; d. Determine whether the cumulative reward function value of all states converges. If so, end the iteration; otherwise, perform a new iteration; e. Standardize the Q-value tables of each objective function and perform arithmetic averaging to generate a comprehensive Q-value table; the objective functions are maximizing platform profits, maximizing employer satisfaction, and maximizing average service provider satisfaction; f. ACA parameter initialization, where the initial pheromone value is the comprehensive Q value table; g. Generate the initial solution. Each ant starts from the initial state and ends at the terminal state, visiting an action in each state; h. Calculate the objective function based on the actions of each ant and update the pheromone concentration on the actions visited by the ants; i. Determine whether the maximum number of iterations is met. If so, proceed to step j; otherwise, proceed to step h; j. Perform non-dominated sorting on all solutions and use GRA-TOPSIS to select the optimal solution in the Pareto solution set.
4. The method according to claim 2, wherein The ant colony algorithm ACA specifically includes: a comprehensive Q value table Q, a Q value table Q of each objective function i , the number of ants M and related data parameters P are input, and the output is the Pareto solution set S; Let the initial number of iterations be 0, according to the probability Select the action in the current state; Among them, τ ij Indicates the pheromone accessing action j in state i, and its initial value comes from the comprehensive Q value table Q; η ij Indicates the heuristic function value of accessing action j in state i, whose value comes from the Q value table Q of a randomly selected objective function i The corresponding value; calculate the objective function value according to the action visited by each ant, and perform non-dominated sorting to obtain the Pareto level F of each ant k ; Calculate the pheromone concentration released by each ant when visiting the action Where G is the total amount of pheromone released by an ant in one iteration; Update pheromones.