A service quality prediction method combining multiple relationships between users and services
By building a multi-layer relationship diagram between users and services and performing matrix decomposition, the challenge of accurately predicting the quality of Web services in complex network environments is solved, and the accuracy and user experience of QoS prediction are improved.
Patent Information
- Application Number
- CN202211203346.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-29
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2042-09-29
AI Technical Summary
In a complex and dynamic network environment, accurately predicting the quality of service (QoS) attribute values of users using Web services under different circumstances is a challenging problem. The prior art is difficult to effectively consider the multi-layer relationship between users and services, resulting in insufficient accuracy of QoS value prediction.
By representing the relationship between users and users, the relationship between services, and the relationship between users and services in the form of a graph, a global graph is constructed and edges with edge weights smaller than the threshold value are cropped to form a subgraph, fuse the response time QoS matrix, and perform matrix decomposition to predict the QoS value.
This method improves the accuracy of QoS prediction of service quality by comprehensively considering the direct and potential relationship between users and services, and enhances the user's experience in calling services.
Smart Images

Figure CN115660148B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of communication technology, and in particular relates to a method for predicting service quality by combining multiple relationships between users and services. Background Art
[0002] In today's era, information networks have become a huge driving force for social development. Since the Internet entered the commercialization stage in the 1990s, the Internet has experienced a period of rapid development, and more and more Web applications have been built and deployed on the Internet. Internet Web services have entered the commercialization stage. At this time, the Internet is developing at a rapid pace, and more and more scenarios will be applied to Web applications. The number of published Web services has grown rapidly in the past few years. The popularity of Web services and service-oriented architecture (SOA) allows the construction of different service-oriented applications to meet the increasingly complex business needs of various users. However, as mobile Internet technology and distributed computing technology have also developed very rapidly in recent years, the complexity of the network environment has also increased. So, how to accurately predict the quality of service (QoS) attribute values of various Web services used by each user in different situations in a complex and dynamic network environment is a challenging problem. QoS is generally used to reflect the evaluation criteria of Web services, measured by indicators such as service scores, response time and throughput from the user's perspective. In order to solve the differences in user needs and satisfy users' better Web service quality in different application scenarios, the network needs to allocate and schedule relevant data according to the personalized needs of target users, and provide different service quality evaluation results for different data streams. The service quality of the target user's demand service is predicted based on the historical data stream, so that users can get a better calling service experience in different scenarios. Intuitively, the accuracy of the prediction algorithm has been significantly improved, but there are still defects. For example, the data stream application is single, so that the user's needs are not personalized. Therefore, for the problem of QoS value prediction, improvements are made on the existing matrix decomposition technology: according to the existing data set, the data characteristics are considered in many aspects, and the direct relationship and potential relationship between the user and the service are comprehensively considered in the form of a graph, and the QoS feature requirements of the user who has not called the service or the user wants to call a service at the next moment are predicted. Based on this method, the accuracy of QoS value prediction is improved, thereby improving the user experience. Summary of the invention
[0003] In view of this, an object of the present invention is to provide a service quality prediction method combining multiple relationships between users and services.
[0004] To achieve the above object, the present invention provides the following technical solution: a method for predicting service quality by combining multiple relationships between users and services, comprising the following steps:
[0005] Step 1: Based on the historical data stream, the relationship between users is represented by the user graph G UU To represent the relationship between services, the service graph G SS To represent the relationship between users and services, the user-service graph G US To express.
[0006] Step 2: Based on the user graph G UU 、Service Graph G SS and user service graph G US To construct the global graph G, and cut off the edges in G whose weights are less than the threshold θ1, forming K subgraphs G1, G2, …G K .
[0007] Step 3: Use the response time QoS matrices corresponding to the K subgraphs to weight the response time QoS matrix A between the original user and the service, and then merge it into a new response time QoS matrix B.
[0008] Step 4: Perform matrix decomposition on the fused new matrix B to obtain the minimum loss function.
[0009] Step 5: Evaluate the performance of the multi-relation prediction method using mean absolute error and root mean square error.
[0010] The advantages and beneficial effects of the present invention are as follows:
[0011] The innovation of the present invention is mainly the coordination of steps 1, 2, and 3. Step 1 displays the user data flow, service data flow, and user and service data flow in the form of a graph, avoiding the problem that the traditional method is too one-sided. In the process of graph establishment, not only the direct relationship between the user and the service is considered, but also the potential relationship between the user and the service is fully considered. Step 3 merges the processed subgraph with the original graph to make the data features more distinct. Comprehensive global considerations avoid single considerations in user and service classification. Classification in the form of a graph solves the above problem, and then matrix decomposition technology predicts the QoS value. The matrix decomposition technology (Multi Relationship Matrix Factorization, MRMF) based on the joint user and service multi-relationship has more obvious advantages to a certain extent. In short, this method combines the advantages of graphs in graph theory and matrix decomposition and gradient descent in convex optimization, avoids the noise of prediction, and can improve the accuracy of service quality QoS prediction through this method, while improving the user experience of calling services. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be described in detail below in conjunction with the accompanying drawings, wherein:
[0013] Figure 1 A relationship diagram simulating a user calling a service for a specific implementation mode of the present invention;
[0014] Figure 2 QoS value matrix A of a specific implementation mode of the present invention;
[0015] Figure 3 Matrix reduction decomposition for a specific implementation of the present invention;
[0016] Figure 4 A model of a service quality prediction mechanism combining multiple relationships between users and services in a specific embodiment of the present invention;
[0017] Figure 5 This is a comparison diagram of the Loss function of the MF algorithm and the MRMF algorithm in a specific embodiment of the present invention;
[0018] Figure 6 The MAE comparison between the MF algorithm and the MRMF algorithm in the specific implementation mode of the present invention is shown;
[0019] Figure 7 This is a comparison of RMSE between the MF algorithm and the MRMF algorithm according to a specific implementation manner of the present invention. DETAILED DESCRIPTION
[0020] The following describes the embodiments of the present invention by specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner, and the following embodiments and features in the embodiments can be combined with each other without conflict.
[0021] Among them, the drawings are only used for illustrative explanations, and they only represent schematic diagrams rather than actual pictures, and should not be understood as limitations on the present invention. In order to better illustrate the embodiments of the present invention, some parts of the drawings may be omitted, enlarged or reduced, and do not represent the size of actual products. For those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.
[0022] The same or similar numbers in the drawings of the embodiments of the present invention correspond to the same or similar parts; in the description of the present invention, it should be understood that if the terms "upper", "lower", "left", "right", "front", "rear", etc. indicate the orientation or position relationship, they are based on the orientation or position relationship shown in the drawings, which is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operate in a specific orientation. Therefore, the terms describing the position relationship in the drawings are only used for illustrative purposes and cannot be understood as limiting the present invention. For ordinary technicians in this field, the specific meanings of the above terms can be understood according to specific circumstances.
[0023] A model diagram of a service quality prediction method combining multiple relationships between users and services is shown in Figure 1 As shown, the following steps are included:
[0024] Step 1: Based on the historical data stream, the relationship between users is represented by graph G. UU To represent the relationship between services, use graph G SS To represent the relationship between users and services, we use graph G US To express;
[0025] Step 2: Based on the user graph G UU 、Service Graph G SS and user service graph G US To construct the global graph G, and cut off the edges in G whose weights are less than the threshold θ1 to reduce the noise between the relations, thus forming K subgraphs G1, G2, …G K ;
[0026] Step 3: In order to comprehensively consider the direct relationship and potential relationship between each node, the response time QoS matrix corresponding to the K subgraphs is used to weight the response time QoS matrix A between the original user and the service, so as to merge it into a new response time QoS matrix B;
[0027] Step 4: Perform matrix factorization (MF) on the fused new matrix B to obtain the minimum loss function;
[0028] Step 5: Use the mean absolute error (MAE) and root mean square error (RMSE) to evaluate the performance of the multi-relationship service quality prediction method.
[0029] Step 1: Based on the historical data stream, the relationship between users is represented by a graph G. UU To represent the relationship between services, use graph G SSTo represent the relationship between users and services, we use graph G US To indicate, specifically including:
[0030] The relationship between users: including the longitude and latitude distance relationship of users, the relationship between the autonomous domains of the users' environments, and whether the result of comprehensive linear superposition can form a relationship diagram.
[0031] The user graph is graph G UU , the user node set is regarded as V U , user u a With user u b The edge connection between them is measured by weight, expressed as E UU ( Belong to E UU ), then G UU = {V U ,E UU}. The similarity relationship of users is affected by the autonomous system (AS) of the network environment. The formula for measuring the similarity relationship of users in the AS domain is shown in (1):
[0032]
[0033] W ASU (u a ,u b ) represents the similarity under the autonomous domain of the network environment, Represents user u a The network environment autonomous domain where it is located, Represents user u b The network environment autonomous domain where the user is located. From formula (1), we can see that if user u a With user u b In the same network environment autonomous domain, the weight for measuring similarity is 1, otherwise it is 0.
[0034] The latitude and longitude of the user's location are used to calculate the similarity relationship of this part. Since the user is on the earth, the actual distance between two users is calculated using formula (2):
[0035] S(u a ,u b )=R·arccos[cosβ1 cosβ2cos(α1-α2)+sinβ1 sinβ2] (2)
[0036] In the above formula, R is the radius of the earth, and user u is a The latitude angle is β1, and the longitude angle is α1; user u b The latitude angle is β2 and the longitude angle is α2.
[0037] In order to normalize the distance between users and make the distance weight reasonable, the normalization formula is shown in (3):
[0038]
[0039] W DIS (u a ,u b ) represents user u a With user u b The distance similarity weight is, when the actual distance S(u a ,u b ) is smaller, then the similarity weight W of two users is DIS (u a ,u b ) is larger, and its interval is [0,1].
[0040] Taking into account the factors that affect the similarity relationship between users, the data streams of the user's network environment autonomous domain and the user's geographical location are fully utilized, and the weights of the two (including the similarity weight of the user under the network environment autonomous domain and the similarity weight under the normalized distance between users) are linearly weighted to obtain formula (4):
[0041]
[0042] Among them, E(u a ,u b ) represents user u a With user u b λ1 is the weight ratio of the similarity weight under the autonomous domain of the network environment where the user is located, and λ2 is the weight ratio of the similarity weight under the geographical location where the user is located. The relationship between the two is: λ1+λ2=1,λ1>0,λ2>0. Among them, χ1 is the threshold for controlling the edge between user nodes. If the weighted sum is greater than the threshold, there is an edge between the two nodes, otherwise there is no edge between the two user nodes.
[0043] The relationship between services: includes the similarity relationship between word vectors obtained by the service interface crawler, the relationship between the autonomous domains of the service environment, and whether the result of comprehensive linear superposition can form a relationship graph.
[0044] The service graph is graph G SS , the service node set is regarded as V S , Services a With Services b The edge connection between them is measured by weight, expressed as Then G SS = {V S ,E SSServices in the same AS domain system have similar characteristics, so the similarity of services is affected by the AS domain system, as shown in formula (5):
[0045]
[0046] Where W ASS (s a ,s b ) indicates service a With Services b Similar weights in the autonomous domain system, if service s a With Services b In the same autonomous domain system, the weight is 1, otherwise it is 0. Indicates services a The network environment autonomous domain where it is located, Indicates services b The autonomous domain of the network environment.
[0047] Using crawler technology, we visit each service and crawl the content of each service webpage. First, we find the keywords of the content of the two service webpages. We take out several keywords from each service webpage and merge them into a set. We calculate the word frequency of each service webpage for the words in this set, generate the feature vector word of each webpage, and calculate the similarity of the service webpages based on the cosine similarity, as shown in formula (6):
[0048]
[0049] where x i and i The services are a and Services b The feature vector of the web page keywords is obtained, and the first z keywords are taken out to form the feature vector word, and the similarity weight of the service web page is obtained. The larger the value, the more similar it is.
[0050] Taking into account the factors affecting the service relationship, the study of the data set found that the network environment autonomous domain system where the service is located and the content of the service webpage itself play a key role in the similarity of the service. The linear superposition and summation of the two are shown in formula (7):
[0051]
[0052] The above formula comprehensively considers the relationship between the two, where λ3 and λ4 are the weight ratios of the service relationship between the network environment autonomous domain system where the service is located and the service webpage content. χ2 is the threshold between services. If the weighted sum is greater than the threshold, there is an edge between the two service nodes, otherwise there is no edge between the two nodes. SS (s a ,s b) indicates service a and Services b The similarity weight of .
[0053] The relationship between users and services: includes the response time when users call services and the throughput when users request services.
[0054] The user service diagram is shown in Figure G US , including three elements: user nodes, service nodes, and the edge of user-service interaction, which is G US = {V U ,V S ,E US}. The interaction information between users and services includes response time and throughput. This interaction information can intuitively present the dynamic response of users calling services, and is generally reflected in the form of QoS values. The smaller the value of response time (RT), the shorter the waiting time required for the service called by the user, which means that the user has a higher preference for this service. Throughput (TP) refers to the number of requests processed by the system per unit time. The larger the value of TP, the more requests the service responds to users and the shorter the response time. For non-concurrent application systems, throughput and response time are strictly inversely proportional. Since the present invention requires normalized data, it can be seen from the above that the larger the edge weight of the user graph, the closer the connection between users, the larger the edge weight of the service graph, the closer the connection between services, and the larger the TP value, the closer the connection between users and services. Therefore, the present invention uses the TP value to create u a With s b The side information between tp ab Represented as services b Processing user u in unit time a The number of requests and throughput are processed as shown in formula (8):
[0055]
[0056] where tp ab is the normalized value, so the throughput is normalized as shown in formulas (9)-(13):
[0057]
[0058] Among them, tp max ,tp min are the maximum TP value and the minimum TP value when all users call all services, tp ab 'For u a Calls bThe normalized value of the throughput. The larger each element is, the more user requests the service responds to, which means the closer the relationship between the user and the service is.
[0059] Step 2: Based on the user graph G UU 、Service Graph G SS and user service graph G US To construct the global graph G, and cut off the edges in G whose weights are less than the threshold θ1 to reduce the noise between the relations, thus forming K subgraphs G1, G2, …G K ; Specifically include:
[0060] Global graph G: User graph, service graph, user and service graph are constructed into a complete weighted undirected graph G, then G = {G UU ,G SS ,G US}. It can also be expressed as G = {V U ,V S ,E UU ,E SS ,E US}, including the user node set V U , the service node set is V S , the edge set of the relationship between users is E UU , the edge set of the relationship between services is E SS , the edge set of the relationship between users and services is E US .
[0061] Edge weight: The normalized edge weight range is [0,1]. The larger the weight, the closer the node connection.
[0062] Noise: edges with weights less than the threshold θ1, including edges between users, edges between services, and edges between users and services.
[0063] Subgraph: To reduce the noise between relations, the edges of the global graph G are cut to form subgraphs G1, G2, …G K , each subgraph contains two factors: nodes and edges.
[0064] Step 3: In order to comprehensively consider the direct relationship and potential relationship between each node, the response time QoS matrix corresponding to the K subgraphs is used to weight the response time QoS matrix A between the original user and the service, so as to merge into a new response time QoS matrix B, which specifically includes:
[0065] Direct relationship: A direct relationship is when the user directly calls the service.
[0066] Potential relationships: including relationships between multiple closely connected users and the same service. Figure 1In user neighbor group 1, users u1, u2, and u3 are closely related. In historical data, only u2 and u3 have called service s1. Therefore, there is a potential relationship between user u1 and service s1, and there is also a potential connection between users u1 and u3 and service s2.
[0067] Matrix A: The QoS matrix of the initial user call service response time in the historical data stream. The data of user call services can be organized into a User-Service matrix, where each row represents a user and each column represents a service. If a user has called a service, the intersection of the row corresponding to the user and the column corresponding to the service in the matrix represents the response time of the user calling the service. This User-Service matrix is called the QoS matrix A. The "?" indicates that the user has not called the service yet, such as Figure 2 shown.
[0068] Matrix B: The response time QoS matrix corresponding to the K subgraphs is linearly weighted by the response time QoS matrix A between the original user and the service to obtain matrix B, as shown in formula (12).
[0069] The step 4: performing matrix decomposition technology on the fused new matrix B to obtain the minimum loss function, specifically includes:
[0070] Matrix decomposition technology: Since the initial matrix is relatively sparse, in order to solve this problem, matrix decomposition technology is proposed. The main idea is to decompose the original rating matrix M (m*n) into two matrices P (m*k) and Q (k*n). At the same time, only the items with call values in the original matrix are examined to see whether the decomposition results are accurate. The judgment criteria are the mean absolute error and the root mean square error.
[0071] For the original QoS matrix R, we assume that there are three types of implicit features, so we decompose the matrix R (4*5) into the user feature matrix P (4*3) and the service feature matrix Q (3*5). Investigating the QoS value of User1 calling Service1, it can be considered that User1's interest in the three types of implicit features Class1, Class2, and Class3 is P11, P12, and P13 respectively, and the relevance of these three types of implicit features to Service1 is Q11, Q21, Q31, and Q41 respectively. Figure 3 shown. Figure 3 The expression is formula (11). It can be found that the QoS value obtained by user U for the final call of service S is the sum of the interest of U in S under each implicit feature dimension. The interest of user U in service S is represented by the interest of user U in the current implicit feature multiplied by the correlation between service S and the current implicit feature.
[0072] Then, for the matrix M(m*n), it is considered that the matrix has k implicit factors, then it is decomposed into P(m*k) and Q(k*n). At this time, for the position M in the original matrix where the QoS value is called US For example, its corresponding value in the decomposed matrix is M' US As shown in formula (10):
[0073]
[0074] P U,k , Q k,S are two low-order random normal matrices generated, K represents the latent factor
[0075] Then for the entire service quality matrix, the total loss function is shown in formula (11):
[0076] SSE=E U,S 2 =∑ U,S (M U,S -M' U,S ) 2 (11)
[0077] Among them, E U,S Represents the original QoS matrix M U,S and reconstructing M' U,S The error value between .
[0078] Gradient descent method: The core idea is to iterate step by step in the direction of gradient descent. Gradient is a vector, which indicates that a function changes fastest and at the highest rate in the direction of gradient at that point, and the direction of gradient descent refers to the negative gradient direction. Stochastic gradient descent method is mainly used to solve optimization problems in the form of summation. The idea is: since it is cumbersome to calculate the gradient for each item in the summation, randomly select one of the items to calculate the gradient and use it as the total gradient.
[0079] The specific application to the objective function is shown in the following formula (12):
[0080]
[0081] SSE is a multivariate function of P and Q (where P and Q represent two randomly generated low-order matrices.) After U and S are randomly selected, all k need to be enumerated and P U,k and Q k,S Find the partial derivative. In the whole formula, only P U,k Q k,S This term is related to it. Through the chain rule (such as formula (13) (14)), we can know that:
[0082]
[0083]
[0084] In actual operations, in order to update all values in P and Q, U and S corresponding to the points with call values in the QoS matrix are generally selected for iteration in an online learning manner.
[0085] Regularization: Adding a regularization term is a classic method to prevent overfitting. The specific method is to add an L2 regularization term after the loss function. The loss function formula is as follows (15), that is:
[0086]
[0087] Among them, λ is the regularization coefficient, and the entire solution process can still be completed using stochastic gradient descent. U , Q S Regularization terms for the user matrix and service matrix respectively.
[0088] Loss function: It includes the sum of the differences between the predicted value and the true value. In order to prevent the objective function from overfitting, a regularization term is added to the loss function, as shown in formulas (16) and (17).
[0089] B=μ1M U,S +μ2M K U,S (16)
[0090]
[0091] Among them, matrix B is the newly fused matrix after data feature processing, λ' is the regularization coefficient in the new matrix loss function, is the matrix decomposition formula of the new matrix. Where μ1 represents the weight coefficient of the original matrix A, μ2 represents the weight coefficient of the sub-matrix, and M K U,S represents the submatrix, P' U , Q' S Regularization terms of the fused user matrix and service matrix respectively.
[0092] The step 5: evaluating the performance of the improved prediction method by using mean absolute error and root mean square error, specifically includes:
[0093] Mean absolute error: It is the average value of absolute error, which is actually a general form of error mean. It represents the average difference between the predicted result and the actual value. All individual differences are weighted. The specific calculation method of MAE is shown in formula (18):
[0094]
[0095] Where R(u,s) is the initial response time matrix, R'(u,s) is the predicted response time matrix, and L is the total number of elements in the matrix.
[0096] Root mean square error: It is the average size of the measurement error, which is the square root of the average of the squared differences between the predicted value and the actual observed value. The specific calculation method of RMSE is shown in formula (19):
[0097]
[0098] The QoS prediction accuracy is compared with the improved algorithm through matrix decomposition technology, and the efficiency and effectiveness of the improved method of the present invention are evaluated by evaluating the accuracy of QoS prediction as mean absolute error and root mean square error.
[0099] The present invention proposes a service quality prediction method combining multiple relationships between users and services. The overall flow chart of the service quality prediction mechanism combining multiple relationships between users and services is as follows: Figure 4 As shown in the figure, under the same conditions of data set and server, the matrix factorization (MF) algorithm decomposes the original rating matrix M(m*n) into two matrices P(m*k) and Q(k*n), and only examines whether the decomposition results of the items with call values in the original matrix are accurate. The judgment criteria are the mean absolute error and the root mean square error. Compared with the traditional MF algorithm, the method of combining multiple relationships between users and services for matrix decomposition has obvious advantages in the degree of convergence of the loss function, and has advantages in the judgment criteria of mean absolute error and root mean square error. The results of running this scheme on different data set densities show that the accuracy of the prediction is improved, such as Figure 6 , Figure 7 shown.
[0100] Figure 5 This is a comparison chart of the Loss function of the MF algorithm and the MRMF algorithm. From the experimental results shown, as the density of the data set continues to increase, the convergence trend of the loss function of the MF algorithm and the MRMF algorithm becomes faster and faster, which is due to the influence of the implicit factors in the matrix decomposition. In addition, at the initial position of the iteration, the Loss value increases with the increase of the density of the data set. Since the matrix decomposition is caused by the synthesis of randomly generated low-order matrices, it is necessary to increase the iteration step size until the Loss function converges. At this moment, the corresponding two low-order matrices are the optimal matrices. After multiplying the two low-order matrices, a full-rank matrix is obtained, which is the predicted matrix. As can be seen from the above figure, the loss function value in the MRMF algorithm has always been below the loss function value of the MF algorithm. The Loss value of the MRMF algorithm becomes smaller and smaller with the increase of the number of iteration steps, which helps to improve the accuracy of QoS prediction.
[0101] Figure 6 The figure shows the comparison between the mean absolute error of the QoS prediction value obtained by the MF algorithm and the MRMF algorithm. As can be seen from the figure, the results of running under different data set densities are that the MAE of the MRMF algorithm is smaller than the MAE of the MF algorithm. To further illustrate the error reduction of the MRMF algorithm, Figure 7 The accuracy improvement rate of QoS prediction values at different data set densities is shown. When the data set density is 10*10, MRMF reduces the MAE of MF by 96%. Since the data density is relatively small, the effect is more obvious; when the data set density is 30*30, MRMF reduces the MAE of MF by 36%; when the data set density is 50*50, MRMF reduces the MAE of MF by 20%; when the data set density is 100*100, MRMF reduces the MAE of MF by 20%; when the data set density is 200*200, MRMF reduces the MAE of MF by 18%. This shows that MRMF has better performance than MF, which can be attributed to the fact that MRMF uses the graph model to effectively integrate the multiple relationships between users and services, not only considering the direct relationship between users and services, but also further considering the potential relationship between users and services, reducing MAE and improving QoS prediction accuracy.
[0102] Figure 7 The RMSE comparison between MF and MRMF is shown. It can be seen from the figure that the performance of MRMF algorithm is better than that of MF algorithm. Under different data set densities, the RMSE value of MRMF is always lower than the RMSE value of MF. Figure 7 The error reduction rate of the two algorithms in RMSE indicators is shown. When the data set density is 10*10, MRMF reduces the RMSE of MF by 60%; when the data set density is 30*30, MRMF reduces the RMSE of MF by 38%; when the data set density is 50*50, MRMF reduces the RMSE of MF by 22%; when the data set density is 100*100, MRMF reduces the RMSE of MF by 22%; when the data set density is 200*200, MRMF reduces the RMSE of MF by 18%. From the overall results, compared with the MF algorithm, the MRMF algorithm has significantly improved the QoS prediction accuracy, which is mainly attributed to the comprehensive consideration of the multiple relationships between joint users and services, the mining of the potential relationship between users and services, and the reconstruction of the new matrix through in-depth research on the characteristics of the data set, so that the potential relationship can be discovered and applied. The original matrix of QoS values is adjusted and fused to construct a new matrix, which finally reflects its value in the QoS prediction accuracy, and the accuracy is improved to a certain extent.
Claims
1. A method for predicting service quality by combining multiple relationships between users and services, characterized in that: The following steps are involved: Step 1: Based on the historical data stream, the relationship between users is represented by the user graph G UU To represent the relationship between services, the service graph G SS To represent the relationship between users and services, the user-service graph G US To express; Step 2: Based on the user graph G UU 、Service Graph G SS and user service graph G US To construct the global graph G, and cut off the edges in G whose weights are less than the threshold θ1, forming K subgraphs G1, G2, …G K ; Specifically include: Global graph G: User graph, service graph, user and service graph are constructed into a complete weighted undirected graph G, then G = {G UU ,G SS ,G US }, then G={V U ,V S ,E UU ,E SS ,E US }, including the user node set V U , the service node set is V S , the edge set of the relationship between users is E UU , the edge set of the relationship between services is E SS , the edge set of the relationship between users and services is E US ; Edge weight: The normalized edge weight ranges from [0,1]. The larger the weight, the closer the node connection. Noise: edges with weights less than the threshold θ1, including edges between users, edges between services, and edges between users and services; Subgraph: Cut the edges of the global graph G to form subgraphs G1, G2, …G K ,Each subgraph contains two factors: nodes and edges; Step 3: Use the response time QoS matrices corresponding to the K subgraphs to weight the response time QoS matrix A between the original user and the service, so as to merge it into a new response time QoS matrix B; specifically, it includes: Direct relationship: A direct relationship is when a user directly calls a service; Potential relationship: includes the relationship between multiple closely connected users and the same service. If in the historical data, users u1, u2, and u3 are closely related in user neighbor group 1, and services s1 and s2 are closely related in service neighbor group 1, and only u2 and u3 have called service s1, then user u1 has a potential relationship with service s1, and users u1 and u3 also have a potential relationship with service s2. Matrix A: The initial QoS matrix of the response time of user calls to services in the historical data stream. The data of user calls to services are organized into a User-Service matrix. Each row in the matrix represents a user, and each column represents a service. If a user has called a service, the intersection of the row corresponding to the user and the column corresponding to the service in the matrix represents the response time of the user calling the service. This User-Service matrix is called the QoS matrix A. Matrix B: Linearly weight the response time QoS matrix A between the original user and the service by the response time QoS matrix corresponding to the K subgraphs to obtain matrix B; Step 4: Perform matrix decomposition on the fused new matrix B to obtain the minimum loss function.
2. The method for predicting service quality by combining multiple relationships between users and services according to claim 1, characterized in that: The user graph G UU , G UU = {V U ,E UU },V U is the user node set, E UU For user u a With user u b The edge connection weight between them; the formula for measuring the user similarity relationship in the network environment autonomous domain is shown in (1): W ASU (u a ,u b ) represents the similarity weight under the autonomous domain of the network environment, Represents user u a The network environment autonomous domain where it is located, Represents user u b The autonomous domain of the network environment; From formula (1), we can see that if user u a With user u b In the same network environment autonomous domain, the weight for measuring similarity is 1, otherwise it is 0; The similarity relationship is calculated using the longitude and latitude of the user's location, and the actual distance between two users is calculated using formula (2): S(u a ,u b )=R·arccos[cosβ1cosβ2cos(α1-α2)+sinβ1sinβ2] (2) Where R is the radius of the earth, user u a The latitude angle is β1, and the longitude angle is α1; user u b The latitude angle is β2, and the longitude angle is α2; Normalize the distance between users, as shown in formula (3): W dis (u a ,u b ) represents the distance similarity weight between user a and user b. When the actual distance S(u a ,u b ) is smaller, then the similarity weight W of two users is dis (u a ,u b ) is larger, and its interval is [0,1]; The weights of the two are linearly weighted to obtain formula (4): Among them, λ1 is the weight ratio of the similarity weight under the autonomous domain of the network environment where the user is located, λ2 is the weight ratio of the similarity weight under the geographical location where the user is located, and χ1 is the threshold for controlling the edge between user nodes. If the weighted sum is greater than the threshold, there is an edge between the two nodes, otherwise there is no edge between the two user nodes. The service graph G SS , G SS = {V S ,E SS },V S is the set of service nodes, E SS For services a With Services b The edge connection weights between them are similar. Services in the same AS domain system have similar characteristics. Therefore, the similarity of services is affected by the AS domain system, as shown in formula (5): Where W ASS (s a ,s b ) indicates service a With Services b Similar weights in the autonomous domain system, if service s a With Services b In the same autonomous domain system, the weight is 1, otherwise it is 0; The cosine similarity is used to calculate the similarity of service web pages, as shown in formula (6): where x i and i The services are a and Services b The feature vector of the web page keywords is obtained, and the first z keywords are taken out to form the feature vector words, and the similarity weight of the service web page is obtained. The larger the value, the more similar it is; The network environment autonomous domain system where the service is located and the content of the service webpage itself play a key role in the similarity of the service. The two are linearly superimposed and summed, as shown in formula (7): Where λ3, λ4 are the weight ratios of the service relationship between the autonomous domain system of the network environment where the service is located and the service webpage content, and χ2 is the threshold between services. If the weighted sum is greater than the threshold, there is an edge between the two service nodes, otherwise there is no edge between the two nodes. The user service graph G US , including three elements: user nodes, service nodes, and the edge of user-service interaction, which is G US = {V U ,V S ,E US }; Create u using throughput a With s b The side information between tp ab Represented as user u a Calling Services b The throughput at this time is calculated as shown in formula (8): where tp ab is the normalized value, so the throughput is normalized as shown in formula (9): Among them, tp max ,tp min are the maximum and minimum throughputs when all users call all services, tp ab 'For u a Calls b The normalized value of the throughput. The larger each element is, the more user requests the service responds to, which means the closer the relationship between the user and the service is.
3. The method for predicting service quality based on multiple relationships between users and services according to claim 1, characterized in that: The step 4 specifically includes, for the matrix M(m*n), it is considered that the matrix has k implicit factors, then it is decomposed into P(m*k), Q(k*n), at this time, for the position M with the call QoS value in the response time matrix A US For example, its corresponding value in the decomposed matrix is shown in formula (10): P U,k , Q k,S are two low-order random normal matrices generated, K represents the latent factor; For the entire matrix A, the total loss function is shown in formula (11): SSE=E 2 =∑ U,S (M U,S -M' U,S ) 2 (11)。 4. The method for predicting service quality by combining multiple relationships between users and services according to claim 3, characterized in that: In the process of solving the minimum loss function, the gradient descent method is used to iterate step by step, which is specifically applied to the objective function as shown in the following formula (12): SSE is a multivariate function of P and Q. After randomly selecting U and S, we need to enumerate all k and calculate P. U,k and Q k,S Find the partial derivative. Only P is U,k Q k,S This one is related, and by the chain rule: In order to update all the values in P and Q, the U and S corresponding to the points with call values in the QoS matrix are selected for iteration in an online learning manner.
5. The method for predicting service quality by combining multiple relationships between users and services according to claim 4, characterized in that: In order to prevent overfitting, a regularization term is added and the loss function formula is:
6. The method for predicting service quality by combining multiple relationships between users and services according to any one of claims 1 to 5, characterized in that: It also includes evaluating the performance of multi-relationship service quality prediction methods using mean absolute error and root mean square error.
7. The method for predicting service quality by combining multiple relationships between users and services according to claim 6, characterized in that: The mean absolute error is expressed as follows: Where R(u,s) is the initial response time matrix, R'(u,s) is the predicted response time matrix, and L is the total number of elements in the matrix; The root mean square error is expressed as follows: Where R(u,s) is the initial response time matrix, R'(u,s) is the predicted response time matrix, and L is the total number of elements in the matrix.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the service quality prediction method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Web service quality prediction method based on random walk
CN105117325A
A service quality prediction method based on decentralized matrix decomposition
CN109376901A