Recommended software performance detection method based on big data analysis

By constructing an isolated forest model and an ISODATA clustering model through big data analysis, the abnormal scores and confidence levels of preference information are obtained, which solves the problem of poor performance detection of software recommendation systems and realizes accurate detection of recommendation software performance.

CN120950359AInactive Publication Date: 2025-11-14ZHUNJIAN HEBEI TESTING TECH SERVICE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511106396.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-08
Publication Date
2025-11-14
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing software recommendation systems have poor performance in testing, and traditional methods mainly rely on response time thresholds for judgment, which leads to large errors.

Method used

By employing a big data analytics approach, we collect the amount of preference information, the time of receiving requests, and the time of processing requests. We then construct an isolated forest model and an ISODATA clustering model to obtain the anomaly scores and confidence levels of preference information. Combined with anomaly score correction coefficients, we achieve accurate performance testing of recommendation software.

Benefits of technology

It enables accurate detection of the performance of recommendation software, reduces errors, and improves the accuracy and reliability of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120950359A_ABST
    Figure CN120950359A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of big data processing, in particular to a recommendation software performance detection method based on big data analysis, which comprises the following steps: acquiring the data volume of preference information, request receiving time and request processing time; obtaining an isolated forest model of the request receiving time, an isolated forest model of the request processing time and a plurality of class clusters according to the data volume of the collected preference information, the request receiving time and the request processing time; according to the isolated forest model for receiving the request time, the isolated forest model for processing the request time and the plurality of class clusters, obtaining an overall abnormal score of each piece of preference information; and detecting the performance of the recommended software according to the overall abnormal score of each piece of preference information. According to the method, the response time of the personalized content is obtained by analyzing the processing preference information of the recommendation software, so that the performance detection of the recommendation software is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of big data processing technology, and more specifically to a method for performance testing of recommendation software based on big data analysis. Background Technology

[0002] To understand user preferences and interests, user behavior data needs to be collected. By employing data mining and machine learning techniques, recommendation software can build accurate and reliable recommendation models, providing users with personalized recommendations and improving user satisfaction and engagement. However, to monitor various aspects of the recommendation system during operation, such as accuracy, response speed, and system stability, performance monitoring of the recommendation software is necessary. Currently, performance testing for software recommendation systems mainly relies on response time. When the response time exceeds a set threshold, it indicates an anomaly in the recommendation system. However, response time includes the time it takes for the server to receive and process requests and return responses. If any of these components freezes and exceeds the threshold, or if the data processing volume is too large, causing each response time to be longer and the cumulative time to exceed the threshold, errors will occur when using the time threshold to judge the performance of the recommendation software. In other words, traditional performance testing methods for software recommendation systems are ineffective. Summary of the Invention

[0003] This invention provides a method for performance testing of recommendation software based on big data analysis to solve the existing problem that traditional methods for testing the performance of software recommendation systems are ineffective.

[0004] The recommendation software performance testing method based on big data analysis of the present invention adopts the following technical solution:

[0005] Includes the following steps:

[0006] The amount of data collected for preference information, the time of receiving the request, and the time of processing the request, wherein one piece of preference information corresponds to one time of receiving the request and one time of processing the request.

[0007] Based on the receiving request time and processing request time of the collected preference information, an isolated forest model is established for receiving request time and processing request time respectively; clusters are obtained based on the quantity of different types of preference information.

[0008] Based on the isolated forest models of request reception time and request processing time, and combined with the preference information in different clusters, the anomaly scores of preference information in the isolated forest models of request reception time and request processing time are obtained respectively; based on the anomaly scores of preference information in the isolated forest model of request processing time, the confidence level of preference information in the isolated forest model of request processing time is obtained; and combined with the anomaly scores of preference information in the isolated forest model of request reception time and the difference between the anomaly scores of preference information in the two isolated forest models, the anomaly score correction coefficients of preference information in the isolated forest models of request reception time and request processing time are obtained respectively.

[0009] Based on the anomaly score correction coefficient and anomaly score of the preference information in the isolated forest models for both request receiving time and request processing time, the true anomaly scores of the preference information in the isolated forest models for both request receiving time and request processing time are obtained respectively. Based on the true anomaly scores of the preference information in the two isolated forest models, the overall anomaly score of the preference information is obtained. Based on the overall anomaly score of the preference information, the performance of the recommendation software is tested.

[0010] Preferably, the specific method for establishing an isolated forest model based on the receiving request time and processing request time of the collected preference information includes:

[0011] The number of isolated trees in the pre-defined isolated forest model The number of samples in the sample set of each isolated tree Based on the request time of receiving all preference information and the request time of processing, we construct two isolated forest models: one for request time and one for request time. The number of isolated trees in both models is kept constant. And the number of samples in the sample set of each isolated tree is 1. .

[0012] Preferably, the specific method for clustering based on the quantity of different preference information to obtain several clusters includes:

[0013] The data volume of all preference information is input into the ISODATA clustering model, and the initial K value of the ISODATA clustering model is set to 1. ,in The preset clustering parameters are used to cluster all species preference information, resulting in several clusters.

[0014] Preferably, the specific method for obtaining the anomaly scores of preference information in the isolated forest models based on the request receiving time and request processing time, combined with preference information in different clusters, includes:

[0015] In the isolated forest model for receiving request time, the first... First, obtain the preference information in the isolated forest model based on the request receiving time. The cluster containing the preference information is denoted as the i.e., ... species clusters, and the first The preference information in the species clusters is denoted as the target data; the isolation forest model that obtains the request receiving time is the first... Let the isolated trees containing the preference information be denoted as the i-th one. Plant isolated trees;

[0016] For the The first Plant an isolated tree, and the first The first The number of target data in an isolated tree is denoted as ; will the first The first The height of an isolated tree is recorded as ,Will Recorded as the number The first The feature values ​​of isolated trees are used to obtain all the eigenvalues ​​of ... isolated trees. The eigenvalues ​​of the isolated tree, all the eigenvalues ​​of the isolated tree. The mean eigenvalue of the isolated tree is used as the first... Average tree height when planted in isolation;

[0017] Using the Isolation Forest algorithm, based on the first The average tree height of isolated trees in an isolated forest model is used to obtain the request receiving time. Abnormal scores for individual preference information;

[0018] The abnormal scores of preference information in the isolated forest model for receiving request time and the abnormal scores of preference information in the isolated forest model for processing request time are obtained; the method for obtaining the abnormal scores of preference information in the isolated forest model for processing request time is the same as the method for obtaining the abnormal scores of preference information in the isolated forest model for receiving request time.

[0019] Preferably, the specific method for obtaining the confidence level of the preference information in the isolated forest model based on the anomaly score of the preference information in the isolated forest model for processing request time is as follows:

[0020] In the isolated forest model for processing request time, the first... First, the preference information; The preference information in the category cluster is denoted as the target data; then, the anomaly scores of all target data in the isolated forest model for processing request time are obtained; the mean of the anomaly scores of all target data in the isolated forest model for processing request time is compared with the mean of the anomaly scores of the i-th category in the isolated forest model for processing request time. The absolute value of the difference between the abnormal scores of each preference information is used as the value of the first (i) in the isolated forest model at request time. The first value of each preference information is denoted as ;Will In the isolated forest model for processing request time, the first The confidence level of each preference information, among which This represents an exponential function with the natural constant as its base.

[0021] Preferably, the specific method for obtaining the anomaly score correction coefficients of preference information in the isolated forest model for both request reception time and request processing time is as follows:

[0022] In the isolated forest model for receiving request time, the first... The first preference information; obtaining the first request time in the isolated forest model. The preference information, and the processing time in the isolated forest model. The abnormal score of the first preference information; then the first Preference information within the category clusters is denoted as target data; anomaly scores for all target data in the isolated forest model at the time of request reception are obtained; based on the anomaly scores of the isolated forest model at the time of request reception... In the isolation forest model, the anomaly score of the preference information and the request reception time are... The anomaly score of the preference information in the isolated forest model for processing request time, the anomaly scores of all target data in the isolated forest model for receiving request time, combined with the anomaly score of the preference information in the isolated forest model for receiving request time. The confidence level of the preference information in the isolated forest model for processing request time is used to obtain the confidence level of the i-th preference information in the isolated forest model for receiving request time. Anomaly score correction coefficient for preference information;

[0023] The true anomaly score of preference information in the isolated forest model for receiving request time and the true anomaly score of preference information in the isolated forest model for processing request time are obtained. The method for obtaining the true anomaly score of preference information in the isolated forest model for processing request time is the same as the method for obtaining the true anomaly score of preference information in the isolated forest model for receiving request time.

[0024] Preferably, in the isolated forest model for obtaining the time of receiving the request, the first... The abnormal score correction coefficient for each preference information includes the following specific calculation formula:

[0025]

[0026] In the formula, In the isolated forest model, the time of receiving requests represents the first... Anomaly score correction coefficient for preference information; In the isolated forest model, the time of receiving requests represents the first... Abnormal scores for individual preference information; Indicates the quantity of target data; In the isolated forest model, the time of receiving requests represents the first... Anomaly scores for each target data point; In the isolated forest model, the time of receiving requests represents the first... Anomaly scores for individual preference information in an isolated forest model during request processing time; In the isolated forest model, the time of receiving requests represents the first... The confidence level of preference information in an isolated forest model during request processing time; Indicates the activation function; This represents absolute value operations.

[0027] Preferably, the specific method for obtaining the true anomaly scores of preference information in the isolated forest model based on the anomaly score correction coefficient and anomaly score of preference information in the isolated forest model at the request receiving time and request processing time is as follows:

[0028] In the isolated forest model for receiving request time, the first... The first preference information; obtaining the first request time in the isolated forest model. The anomaly score correction coefficient and anomaly score for the first preference information; the first in the isolated forest model for receiving request time. The sum of the anomaly score correction coefficients for each preference information and 1, multiplied by the request reception time in the isolated forest model. The anomaly score of the preference information is used to obtain the th in the isolated forest model for receiving request times. The true outlier score for each preference information;

[0029] The true anomaly score of preference information in the isolated forest model for receiving request time and the true anomaly score of preference information in the isolated forest model for processing request time are obtained. The method for obtaining the true anomaly score of preference information in the isolated forest model for processing request time is the same as the method for obtaining the true anomaly score of preference information in the isolated forest model for receiving request time.

[0030] Preferably, the method for obtaining the overall anomaly score of preference information based on the true anomaly scores of preference information in the two isolated forest models includes:

[0031] For the First, obtain the first preference information. The true outlier scores of the preference information in the isolated forest model at request reception time and the true outlier scores in the isolated forest model at request processing time are then obtained. Among the true outlier scores in the isolation forest model for receiving requests and those for processing requests, the largest outlier score is used as the first... The overall abnormal score of the preference information.

[0032] Preferably, the specific method for detecting the performance of the recommendation software based on the overall anomaly score of the preference information includes:

[0033] First, a threshold for abnormal scores is preset. For the first The preference information, when the first The overall anomaly score of the preference information is greater than In the first Each piece of preference information is recorded as an abnormal response. All abnormal responses are obtained, and the ratio of the number of abnormal responses to the number of all preference information is recorded as the error rate.

[0034] Set another error rate threshold. When the error rate exceeds If so, it indicates that the recommended software has performance issues.

[0035] The beneficial effects of the technical solution of this invention are as follows: This invention collects the amount of preference information data, the time of receiving requests, and the time of processing requests; based on the amount of preference information data collected, the time of receiving requests, and the time of processing requests, it obtains an isolated forest model for receiving requests, an isolated forest model for processing requests, and several clusters. Through the isolated forest model for receiving requests, the isolated forest model for processing requests, and several clusters, it obtains the overall anomaly score of each preference information, so that the overall anomaly score of each preference information can accurately describe the degree of anomaly of the preference information, thereby obtaining anomaly data. Finally, based on the proportion of anomaly data in all preference information, the performance of the recommendation software can be detected. Attached Figure Description

[0036] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0037] Figure 1 This is a flowchart illustrating the steps of the recommendation software performance testing method based on big data analysis according to the present invention.

[0038] Figure 2 This is a flowchart illustrating how the performance of the recommendation software is detected based on the amount of preference information data, the time of receiving requests, and the time of processing requests in this embodiment. Detailed Implementation

[0039] To further illustrate the technical means and effects adopted by the present invention to achieve its intended purpose, the following, in conjunction with the accompanying drawings and preferred embodiments, details the specific implementation, structure, features, and effects of the big data analysis-based recommendation software performance testing method proposed according to the present invention. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.

[0040] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0041] The specific solution of the recommendation software performance testing method based on big data analysis provided by the present invention will be described in detail below with reference to the accompanying drawings.

[0042] Please see Figure 1 The diagram illustrates a flowchart of a recommendation software performance testing method based on big data analysis according to an embodiment of the present invention. The method includes the following steps:

[0043] Step S001: Collect the amount of preference information, the time to receive requests, and the time to process requests.

[0044] It should be noted that this embodiment, as a recommendation software performance testing method based on big data analysis, aims to utilize big data analysis technology to understand user preferences and interests by collecting user behavior data. By applying data mining and machine learning techniques, the recommendation software can establish an accurate and reliable recommendation model, provide users with personalized recommendation content, and improve user satisfaction and user stickiness.

[0045] It needs further explanation that in order to monitor various aspects of the recommendation system, such as accuracy, response speed, and system stability, performance monitoring of the recommendation software is necessary. However, current performance testing of software recommendation systems mainly relies on response time. When the response time exceeds a set time threshold, it indicates an anomaly in the recommendation system. But response time includes the time it takes for the server to receive and process requests and return responses. If any of these components freezes and exceeds the time threshold, or if the data processing volume is too large, causing each response time to be longer and the cumulative time to exceed the time threshold, errors will occur when using the time threshold to judge the performance of the recommendation software. Therefore, this embodiment analyzes the response time of the recommendation software in processing preference information to obtain personalized content, thereby achieving performance testing of the recommendation software. Therefore, the first step is to obtain the response time of the recommendation software in processing preference information.

[0046] Specifically, by inserting timestamps into the program code on the receiving end and the processing end of the server of the recommendation software, the start and end times of the program code on the receiving end and the processing end are recorded respectively. The time for processing each preference information is used as the server's receiving request time and processing request time. The amount of data for each preference information is obtained and recorded once after each processing. The amount of preference information, receiving request time and processing request time recorded for each processing are obtained.

[0047] It should be noted that the response time includes the time it takes for the server to receive and process the request, but does not include the time it takes to return a response, because the time to return a response is only related to hardware factors such as bandwidth. Therefore, the response time in this embodiment only includes the time it takes for the server to receive and process the request.

[0048] Thus, we obtain the amount of preference information recorded in each processing session, the time of receiving the request, and the time of processing the request.

[0049] Step S002: Based on the receiving request time and processing request time of the collected preference information, establish isolated forest models for receiving request time and processing request time respectively; cluster according to the quantity of different types of preference information to obtain several clusters.

[0050] It should be noted that there is a certain relationship between the request time for receiving preference information, the request time for processing preference information, and the amount of preference information data. That is, we can build an isolated forest model based on the request time for receiving preference information and the request time for processing preference information, respectively. Then, based on the two isolated forest models, we can construct the confidence score of the amount of data, the request time for receiving preference information, and the request time for processing preference information in the isolated forest model. Then, we can weight the anomaly scores based on the confidence scores, and finally obtain the number of anomalies in all processing times. The performance of the recommendation software can be judged based on the proportion or number of anomalies.

[0051] It should be further explained that the Isolation Forest model constructs a sample set by randomly selecting several samples from all samples, and then constructs several isolated trees based on the amount of preference information, the time of receiving the request, and the time of processing the request. Samples with greater differences in the sample set are more likely to be decided prematurely during the decision tree construction, meaning that the trees of samples with greater differences in the decision results are taller, thus achieving the purpose of anomaly extraction. However, when using recommendation software to process preference information, the different amounts of data lead to differences in the receiving and processing times, resulting in differences in the decision results of the Isolation Forest. Therefore, this embodiment constructs an Isolation Forest model based on the time of receiving the request and an Isolation Forest model based on the time of processing the request. By using the preference information with the same amount of data in the two Isolation Forest models as a reference, the decision results of the Isolation Forest are adjusted, thereby identifying the requests with anomalies among all requests processed by the recommendation software.

[0052] Specifically, the number of isolated trees in the pre-defined isolated forest model. The number of samples in the sample set of each isolated tree , and The specific values ​​can be set according to the actual situation. This embodiment does not impose strict requirements. In this embodiment, the values ​​are respectively... , The following describes the process: Based on the request time of receiving all preference information and the request time of processing, construct an isolated forest model for the request time of receiving and an isolated forest model for the request time of processing, ensuring that the number of isolated trees in both models is equal. And the number of samples in the sample set of each isolated tree is 1. Since the specific process of constructing an isolated forest model is a well-known existing technology, it will not be described in detail in this embodiment.

[0053] Thus, we have obtained the isolated forest model for receiving request time and the isolated forest model for processing request time.

[0054] It should be noted that since there is more than one type of preference information, and the data volume of different types of preference information may be different, and the time for receiving and processing requests may differ due to bandwidth limitations, it is necessary to analyze preference information with different data volumes separately.

[0055] Specifically, the amount of data containing all preference information is input into the ISODATA clustering model, and the initial K value of the ISODATA clustering model is set to [value missing]. ,in These are the preset clustering parameters. The specific value can be set according to the actual situation. This embodiment does not make a hard requirement. In this embodiment, it is used as... The description is as follows: clustering all the preference information to obtain several clusters; since the specific process of clustering according to the ISODATA clustering model is a well-known existing technology, it will not be described in detail in this embodiment.

[0056] Thus, several clusters were obtained.

[0057] Step S003: Based on the isolated forest models of request reception time and request processing time, and combined with the preference information in different clusters, obtain the anomaly scores of preference information in the isolated forest models of request reception time and request processing time respectively; based on the anomaly scores of preference information in the isolated forest model of request processing time, obtain the confidence level of preference information in the isolated forest model of request processing time; and combined with the anomaly scores of preference information in the isolated forest model of request reception time and the difference between the anomaly scores of preference information in the two isolated forest models, obtain the anomaly score correction coefficients of preference information in the isolated forest models of request reception time and request processing time respectively.

[0058] It should be noted that, in addition to the performance of the recommendation software, the data volume also affects the time for receiving and processing requests. If the recommendation software is running normally, the data volume of requests received and processed should be approximately equal. Therefore, this embodiment constructs an isolated forest model with the time for receiving and processing requests as the dividing features for the preference information of all processing times. Using this model, the anomaly scores of the time for receiving and processing requests for each preference information are obtained. Based on the anomaly scores of other preference information with approximately the same data volume, the anomaly score of the time for receiving requests for each preference information is obtained. Similarly, the anomaly score of the time for processing requests is obtained. Finally, based on the anomaly score results of the isolated forest model for receiving requests and the isolated forest model for processing requests, the anomaly score of each preference information is obtained.

[0059] It should be further explained that since the logic and process of analyzing the isolated forest model for receiving request time is the same as that for processing request time, this embodiment takes the isolated forest model for receiving request time as an example for analysis to obtain the anomaly score of each preference information in the isolated forest model for receiving request time.

[0060] Specifically, in the isolated forest model for receiving request times, the first... First, obtain the preference information in the isolated forest model based on the request receiving time. The cluster containing the preference information is denoted as the i.e., ... species clusters, and the first The preference information in the species clusters is denoted as the target data; the isolation forest model that obtains the request receiving time is the first... Let the isolated trees containing the preference information be denoted as the i-th one. Plant isolated trees, according to each first The number of target data, the amount of preference information, and the number of each isolated tree in the isolated tree The height of an isolated tree and the first The number of isolated trees planted, to obtain the first The formula for calculating the average height of an isolated tree is as follows:

[0061]

[0062] In the formula, Indicates the first Average tree height when planted in isolation; Indicates the first The number of isolated trees planted; Indicates the first The first The height of an isolated tree; Indicates the first The first The number of target data in an isolated tree; This indicates the number of samples in the sample set for each predefined isolated tree.

[0063] It should be noted that, since the sample set of the isolated tree is obtained through random selection, if the sample set of the isolated tree is the first... This particular preference information is salient, meaning its data volume differs significantly from other preference information. This can lead to premature or delayed decision-making, resulting in errors in the decision outcome. Therefore, this embodiment addresses this issue by... Indicates the first The first High confidence level for planting isolated trees The larger the value, the more likely it is to be in the first place. The first The first of the isolated trees The preference information is not prominent, if the first one is not prominent. If the preference information is normal, it is not easy to distinguish it too early; thus, the first preference can be obtained more accurately. The average tree height of an isolated tree, according to the first The average tree height of isolated trees can be used to obtain the first... Abnormal scores for preference information.

[0064] And in obtaining the first After determining the average tree height of isolated trees, one can then proceed according to the first... The average tree height of isolated trees in an isolated forest model is used to obtain the request receiving time. The abnormal scores of each preference information are obtained. Since the specific process of obtaining abnormal scores based on the height of the isolated tree is a well-known existing technology, it will not be described in detail in this embodiment. Similarly, the abnormal scores of all preference information in the isolated forest model for receiving request time and the abnormal scores of all preference information in the isolated forest model for processing request time are obtained.

[0065] It should be further explained that the size of the preference information data directly affects the receiving and processing process of recommendation software. The larger the data volume, the longer the receiving and processing time will be. Therefore, the receiving request time and processing request time of preference information with similar data volume will be more similar. Thus, the anomaly scores in the isolated forest model with similar receiving and processing request times should be approximately the same. Furthermore, the relationship between processing request time and data volume is also directly proportional. Therefore, in the two isolated forest models, the time taken for normal preference information to run is fixed after the isolated forest model makes a decision. Thus, the anomaly scores of the two models will also be similar. Therefore, this can be used as a basis to obtain the anomaly score correction coefficient.

[0066] Specifically, in the isolation forest model for processing request time, the first... First, the preference information; The preference information in the category cluster is denoted as the target data; then, the anomaly scores of all target data in the isolated forest model for processing request time are obtained. Based on the anomaly scores of all target data in the isolated forest model for processing request time, the first... The confidence level of a preference information is calculated using the following formula:

[0067]

[0068] In the formula, In the isolated forest model representing the request processing time, the first... Confidence level of individual preference information; In the isolated forest model representing the request processing time, the first... Abnormal scores for individual preference information; Indicates the quantity of target data; In the isolated forest model representing the request processing time, the first... Anomaly scores for each target data point; This represents the absolute value operation; This represents an exponential function with the natural constant as its base.

[0069] It should be noted that, The smaller the value, the better the processing time in the isolated forest model. The smaller the difference between the preference information and the tree height of its cluster, the faster the processing time in the isolated forest model. The higher the anomaly score of the preference information, the better it can be used as a reference in the isolation forest model for processing request time. A reference for the abnormal scores of preference information.

[0070] It should be further explained that in the isolated forest model that obtains the request processing time, the first... After determining the confidence level of the first preference information, the isolation forest model can be used to determine the processing time of the request. The confidence level of the preference information, combined with the request receiving time in the isolated forest model, is used to determine the confidence level of the preference information. The anomaly score of the preference information is used to obtain the th in the isolated forest model. Anomaly score correction coefficient for preference information.

[0071] Specifically, in the isolated forest model for receiving request times, the first... The first preference information; obtaining the first request time in the isolated forest model. The preference information, and the processing time in the isolated forest model. The abnormal score of the first preference information; then the first Preference information within the category clusters is denoted as target data; anomaly scores for all target data in the isolated forest model at the time of request reception are obtained; based on the anomaly scores of the isolated forest model at the time of request reception... In the isolation forest model, the anomaly score of the preference information and the request reception time are... The anomaly score of the preference information in the isolated forest model for processing request time, the anomaly scores of all target data in the isolated forest model for receiving request time, combined with the anomaly score of the preference information in the isolated forest model for receiving request time. The confidence level of the preference information in the isolated forest model for processing request time is used to obtain the confidence level of the i-th preference information in the isolated forest model for receiving request time. The specific formula for calculating the abnormal score correction coefficient for each preference information is as follows:

[0072]

[0073] In the formula, In the isolated forest model, the time of receiving requests represents the first... Anomaly score correction coefficient for preference information; In the isolated forest model, the time of receiving requests represents the first... Abnormal scores for individual preference information; Indicates the quantity of target data; In the isolated forest model, the time of receiving requests represents the first... Anomaly scores for each target data point; In the isolated forest model, the time of receiving requests represents the first... Anomaly scores for individual preference information in an isolated forest model during request processing time; In the isolated forest model, the time of receiving requests represents the first... The confidence level of preference information in an isolated forest model during request processing time; This represents the activation function, which is used for normalization in this embodiment. This represents absolute value operations.

[0074] It should be noted that, The larger the value, the higher the threshold for request reception time in the isolated forest model. In the isolation forest model of preference information and request processing time, the first... The greater the difference in outlier scores among the preference information, the higher the chance of the 1st preference being rejected in the isolated forest model during request time. The more likely the abnormal score of a preference information is to be caused by an abnormality in the time of request reception; The larger the value, the higher the threshold of the isolated forest model in terms of request reception time. The first preference information and the second If the reception request times for other preference information within the same cluster differ significantly, then the... The time it takes to receive preference information requests may be abnormal.

[0075] Thus, the anomaly score correction coefficients for all preference information in the isolated forest model for receiving request time are obtained, and similarly, the anomaly score correction coefficients for all preference information in the isolated forest model for processing request time are obtained.

[0076] It should be noted that the method for obtaining the anomaly score correction coefficients of all preference information in the isolated forest model for processing request time is the same as the method for obtaining the anomaly score correction coefficients of all preference information in the isolated forest model for receiving request time.

[0077] Step S004: Based on the anomaly score correction coefficient and anomaly score of the preference information in the isolated forest models for receiving request time and processing request time, respectively obtain the true anomaly scores of the preference information in the isolated forest models for receiving request time and processing request time; based on the true anomaly scores of the preference information in the two isolated forest models, obtain the overall anomaly score of the preference information; based on the overall anomaly score of the preference information, test the performance of the recommendation software.

[0078] It should be noted that this embodiment, as a recommendation software performance testing method based on big data analysis, achieves recommendation software performance testing by analyzing the response time of personalized content obtained by the recommendation software in processing preference information. After obtaining the anomaly score correction coefficients for all preference information in the isolated forest model for receiving request time and the anomaly score correction coefficients for processing request time in step S003, the true anomaly score can be obtained by combining the anomaly score correction coefficients with the anomaly score.

[0079] It should be further explained that since the process and logic of obtaining the true anomaly scores of all preference information in the isolated forest model for processing request time are the same as those for obtaining the true anomaly scores of all preference information in the isolated forest model for receiving request time, this embodiment will be described using the example of obtaining the true anomaly scores of all preference information in the isolated forest model for receiving request time.

[0080] Specifically, in the isolated forest model for receiving request times, the first... The first preference information; obtaining the first request time in the isolated forest model. The anomaly score correction coefficient and anomaly score for the first preference information; based on the request time in the isolated forest model... The anomaly score correction coefficient and anomaly score of the preference information are used to obtain the first anomaly score in the isolated forest model at the time of request reception. The specific formula for calculating the true anomaly score of each preference information is as follows:

[0081]

[0082] In the formula, In the isolated forest model, the time of receiving requests represents the first... The true outlier score for each preference information; In the isolated forest model, the time of receiving requests represents the first... Anomaly score correction coefficient for preference information; In the isolated forest model, the time of receiving requests represents the first... Abnormal scores for preference information.

[0083] Obtain the true anomaly scores of all preference information in the isolated forest model for receiving request time, and the true anomaly scores of all preference information in the isolated forest model for processing request time.

[0084] It should be noted that after obtaining the true anomaly scores of all preference information in the isolated forest model for receiving requests and the true anomaly scores of all preference information in the isolated forest model for processing requests, the overall anomaly score of each preference information is obtained based on the true anomaly scores of each preference information in the isolated forest model for receiving requests and the isolated forest model for processing requests.

[0085] Specifically, for the first First, obtain the first preference information. The true outlier scores of the preference information in the isolated forest model at request reception time and the true outlier scores in the isolated forest model at request processing time are then obtained. Among the true outlier scores in the isolation forest model for receiving requests and those for processing requests, the largest outlier score is used as the first... The overall abnormal score of the preference information.

[0086] Furthermore, an anomaly score is obtained for the overall set of all preference information. This overall anomaly score can then be used to assess the performance of the recommendation software.

[0087] Specifically, firstly, a threshold for abnormal scores is preset. , The specific value can be set according to the actual situation. This embodiment does not make a hard requirement. In this embodiment, it is used as... To describe, for the first The preference information, when the first The overall anomaly score of the preference information is greater than In the first Each piece of preference information is recorded as an abnormal response. All abnormal responses are obtained, and the ratio of the number of abnormal responses to the number of all preference information is recorded as the error rate.

[0088] Set another error rate threshold. , The specific value can be set according to the actual situation. This embodiment does not make a hard requirement. In this embodiment, it is used as... Describe the situation when the error rate exceeds [a certain threshold]. If so, it indicates that the recommended software has performance issues.

[0089] This concludes the embodiment.

[0090] In this embodiment, a flowchart is used to detect the performance of the recommendation software by measuring the amount of preference information data, the time of receiving requests, and the time of processing requests, as shown below. Figure 2 As shown.

[0091] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for performance testing of recommendation software based on big data analysis, characterized in that, The method includes the following steps: The amount of data collected for preference information, the time of receiving the request, and the time of processing the request, wherein one piece of preference information corresponds to one time of receiving the request and one time of processing the request. Based on the receiving request time and processing request time of the collected preference information, an isolated forest model is established for receiving request time and processing request time respectively; clusters are obtained based on the quantity of different types of preference information. Based on the isolated forest models of request reception time and request processing time, and combined with the preference information in different clusters, the anomaly scores of preference information in the isolated forest models of request reception time and request processing time are obtained respectively; based on the anomaly scores of preference information in the isolated forest model of request processing time, the confidence level of preference information in the isolated forest model of request processing time is obtained; and combined with the anomaly scores of preference information in the isolated forest model of request reception time and the difference between the anomaly scores of preference information in the two isolated forest models, the anomaly score correction coefficients of preference information in the isolated forest models of request reception time and request processing time are obtained respectively. Based on the anomaly score correction coefficient and anomaly score of the preference information in the isolated forest models for both request receiving time and request processing time, the true anomaly scores of the preference information in the isolated forest models for both request receiving time and request processing time are obtained respectively. Based on the true anomaly scores of the preference information in the two isolated forest models, the overall anomaly score of the preference information is obtained. Based on the overall anomaly score of the preference information, the performance of the recommendation software is tested.

2. The method for performance testing of recommendation software based on big data analysis according to claim 1, characterized in that, The specific method for establishing an isolated forest model based on the receiving request time and processing request time of the collection preference information includes: The number of isolated trees in the pre-defined isolated forest model The number of samples in the sample set of each isolated tree Based on the request time of receiving all preference information and the request time of processing, we construct two isolated forest models: one for receiving request time and the other for processing request time. The number of isolated trees in both models is kept constant. And the number of samples in the sample set of each isolated tree is 1. .

3. The method for performance testing of recommendation software based on big data analysis according to claim 1, characterized in that, The specific method for clustering based on the quantity of different preference information to obtain several clusters includes: The data volume of all preference information is input into the ISODATA clustering model, and the initial K value of the ISODATA clustering model is set to 1. ,in The preset clustering parameters are used to cluster all species preference information, resulting in several clusters.

4. The method for performance testing of recommendation software based on big data analysis according to claim 2, characterized in that, The method for obtaining anomaly scores of preference information in the isolated forest model based on the request receiving time and request processing time, combined with preference information in different clusters, includes the following specific methods: In the isolated forest model for receiving request time, the first... First, obtain the preference information in the isolated forest model based on the request receiving time. The cluster containing the preference information is denoted as the i.e., ... species clusters, and the first The preference information within a category cluster is denoted as the target data; In the isolated forest model for obtaining the time of receiving requests, the first... Let the isolated trees containing the preference information be denoted as the i-th. Plant isolated trees; For the The first Plant an isolated tree, and the first The first The number of target data in an isolated tree is denoted as ; will the first The first The height of an isolated tree is recorded as ,Will Recorded as the number The first The feature values ​​of isolated trees are used to obtain all the eigenvalues ​​of ... isolated trees. The eigenvalues ​​of the isolated tree, all the eigenvalues ​​of the isolated tree. The mean eigenvalue of the isolated tree is used as the first... Average tree height when planted in isolation; Using the Isolation Forest algorithm, based on the first The average tree height of isolated trees in an isolated forest model is used to obtain the request receiving time. Abnormal scores for individual preference information; Obtain the anomaly scores of preference information in the isolated forest model for receiving request time, and the anomaly scores of preference information in the isolated forest model for processing request time. The method for obtaining the anomaly score of preference information in the isolated forest model for processing request time is the same as the method for obtaining the anomaly score of preference information in the isolated forest model for receiving request time.

5. The method for performance testing of recommendation software based on big data analysis according to claim 1, characterized in that, The specific method for obtaining the confidence level of preference information in the isolated forest model based on the anomaly score of preference information in the processing request time includes: In the isolated forest model for processing request time, the first... First, the preference information; The preference information within a category cluster is denoted as the target data; Then, obtain the anomaly scores of all target data in the isolated forest model for processing request time; compare the mean of all target data anomaly scores in the isolated forest model for processing request time with the mean of the anomaly scores of the i-th target data in the isolated forest model for processing request time. The absolute value of the difference between the abnormal scores of each preference information is used as the value of the first (i) in the isolated forest model at request time. The first value of each preference information is denoted as ;Will In the isolated forest model for processing request time, the first The confidence level of each preference information, among which This represents an exponential function with the natural constant as its base.

6. The method for performance testing of recommendation software based on big data analysis according to claim 1, characterized in that, The specific method for obtaining the anomaly score correction coefficients for preference information in the isolated forest model, which includes the request receiving time and request processing time respectively, is as follows: In the isolated forest model for receiving request time, the first... The first preference information; obtaining the first request time in the isolated forest model. The preference information, and the processing time in the isolated forest model. The abnormal score of the first preference information; then the first The preference information within a category cluster is denoted as the target data; Obtain the anomaly scores of all target data in the isolated forest model at the time of receiving the request; based on the anomaly scores of the isolated forest model at the time of receiving the request... In the isolation forest model, the anomaly score of the preference information and the request reception time are... The anomaly score of the preference information in the isolated forest model for processing request time, the anomaly scores of all target data in the isolated forest model for receiving request time, combined with the anomaly score of the preference information in the isolated forest model for receiving request time. The confidence level of the preference information in the isolated forest model for processing request time is used to obtain the confidence level of the i-th preference information in the isolated forest model for receiving request time. Anomaly score correction coefficient for preference information; The true anomaly score of preference information in the isolated forest model for receiving request time and the true anomaly score of preference information in the isolated forest model for processing request time are obtained. The method for obtaining the true anomaly score of preference information in the isolated forest model for processing request time is the same as the method for obtaining the true anomaly score of preference information in the isolated forest model for receiving request time.

7. The method for performance testing of recommendation software based on big data analysis according to claim 6, characterized in that, In the isolated forest model for obtaining the time of receiving the request, the first... The abnormal score correction coefficient for each preference information includes the following specific calculation formula: In the formula, In the isolated forest model, the time of receiving requests represents the first... Anomaly score correction coefficient for preference information; In the isolated forest model, the time of receiving requests represents the first... Abnormal scores for individual preference information; Indicates the quantity of target data; In the isolated forest model, the time of receiving requests represents the first... Anomaly scores for each target data point; In the isolated forest model, the time of receiving requests represents the first... Anomaly scores for individual preference information in an isolated forest model during request processing time; In the isolated forest model, the time of receiving requests represents the first... The confidence level of preference information in an isolated forest model during request processing time; Indicates the activation function; This represents absolute value operations.

8. The method for performance testing of recommendation software based on big data analysis according to claim 1, characterized in that, The specific method for obtaining the true anomaly scores of preference information in the isolated forest model based on the anomaly score correction coefficient and anomaly score of the preference information in the isolated forest model at the request receiving time and request processing time is as follows: In the isolated forest model for receiving request time, the first... The first preference information; obtaining the first request time in the isolated forest model. The anomaly score correction coefficient and anomaly score for the first preference information; the first in the isolated forest model for receiving request time. The sum of the anomaly score correction coefficients for each preference information and 1, multiplied by the request reception time in the isolated forest model. The anomaly score of the preference information is used to obtain the th in the isolated forest model for receiving request times. The true outlier score for each preference information; The true anomaly score of preference information in the isolated forest model for receiving request time and the true anomaly score of preference information in the isolated forest model for processing request time are obtained. The method for obtaining the true anomaly score of preference information in the isolated forest model for processing request time is the same as the method for obtaining the true anomaly score of preference information in the isolated forest model for receiving request time.

9. The method for performance testing of recommendation software based on big data analysis according to claim 1, characterized in that, The specific method for obtaining the overall anomaly score of preference information based on the true anomaly scores of preference information in the two isolated forest models includes: For the First, obtain the first preference information. The true outlier scores of the preference information in the isolated forest model at request reception time and the true outlier scores in the isolated forest model at request processing time are then obtained. Among the true outlier scores in the isolation forest model for receiving requests and those for processing requests, the largest outlier score is used as the first... The overall abnormal score of the preference information.

10. The method for performance testing of recommendation software based on big data analysis according to claim 1, characterized in that, The specific method for detecting the performance of the recommendation software based on the overall anomaly score of preference information includes: First, a threshold for abnormal scores is preset. For the first The preference information, when the first The overall anomaly score of the preference information is greater than In the first Each piece of preference information is recorded as an abnormal response. All abnormal responses are obtained, and the ratio of the number of abnormal responses to the number of all preference information is recorded as the error rate. Set another error rate threshold. When the error rate exceeds If so, it indicates that the recommended software has performance issues.