Method, apparatus, device and storage medium for grading site pages
By analyzing the historical access logs of the site pages for clustering and grading, the problem of inaccurate grading of site pages in the existing technology is solved, and the accuracy and security of monitoring are achieved, and malicious attacks and resource waste are avoided.
Patent Information
- Application Number
- CN202210591689.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-27
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2042-05-27
AI Technical Summary
In the prior art, site page grading methods cannot accurately reflect the importance of the page, resulting in mismatch between monitoring frequency and intensity, and the inability to detect malicious attacks in time, resulting in losses or waste of resources.
By obtaining the site's historical access log, analyzing the access status data for clustering, determining the weight values of each clustered set, and grading based on the weight values, eliminating the data that is not accessed by users to ensure the authenticity of the data.
It realizes accurate grading based on real access records, improves the accuracy of site page monitoring, avoids missed discovery of malicious attacks and waste of resources, and reduces losses.
Smart Images

Figure CN114817818B_ABST
Abstract
Description
Background Art
[0002] With the development of technology and the growth of business, the number of site pages is also increasing continuously. To ensure the security of site pages, monitoring and processing of site pages are currently proposed. When monitoring site pages, monitoring is carried out according to the importance level of the site pages, and different frequencies and intensities of monitoring are carried out for different importance levels, so that the resources such as network and computing consumed by monitoring are controllable. Therefore, how to classify site pages is crucial.
[0003] Currently, there are mainly two ways to classify site pages, namely: using a crawler to build a page tree and determining the importance level according to the layer where the site page is located, and using the PageRank algorithm to measure according to the mutual references between site pages.
[0004] Among them, the site page layer depends on the site organizational structure. Site pages with a higher layer are not necessarily accessed frequently. Similarly, site pages with a lower layer may be accessed frequently and have a higher access popularity. However, the monitoring frequency and monitoring intensity corresponding to site pages with a higher layer are higher, while the monitoring frequency and monitoring intensity corresponding to site pages with a lower layer are lower. At this time, due to the higher access popularity of site pages with a lower layer, site pages with a lower layer are more likely to be maliciously attacked such as being tampered with and having malware installed; also, due to the lower monitoring frequency and monitoring intensity, problems cannot be discovered in time, resulting in huge losses.
[0005] Under the PageRank algorithm, new site pages have few references pointing to them, so the importance level determined based on the number of references pointing to them is low. However, such new site pages often have a high access popularity. Similarly, new site pages are more likely to be maliciously attacked such as being tampered with and having malware installed. However, due to the low importance level, the corresponding monitoring frequency and monitoring intensity are low, resulting in problems not being discovered in time and causing huge losses.
[0006] Therefore, how to classify site pages, ensure the accuracy of site page classification, further ensure the accuracy of site page monitoring, and timely discover whether site pages have been maliciously attacked to avoid losses are the technical problems that need to be solved currently. Summary of the Invention
[0007] This application provides a method, device, equipment and storage medium for classifying site pages to ensure the accuracy of site page classification.
[0008] In a first aspect, an embodiment of this application provides a method for classifying site pages, and the method includes:
[0009] Obtain the historical access log corresponding to the site, where the historical access log is used to record the access record data when accessing at least one site page in the site through each source access address;
[0010] Determine the access status data corresponding to each of at least one site page based on historical access logs;
[0011] Perform clustering processing on at least one site page based on the access status data, and determine the weight value of each clustering set;
[0012] Based on each weight value, respectively determine the importance level of the site pages included in the corresponding clustering set.
[0013] In a second aspect, an embodiment of the present application provides a device for grading site pages, and the device includes:
[0014] An obtaining unit, configured to obtain the historical access log corresponding to the site, and the historical access log is used to record the access record data when at least one site page in the site is respectively accessed through each source access address;
[0015] A first determination unit, configured to determine the access status data corresponding to each of at least one site page based on the historical access log;
[0016] A clustering unit, configured to perform clustering processing on at least one site page based on the access status data, and determine the weight value of each clustering set;
[0017] A second determination unit, configured to respectively determine the importance level of the site pages included in the corresponding clustering set based on each weight value.
[0018] In a possible implementation manner, the device further includes an elimination unit;
[0019] After the obtaining unit obtains the historical access log corresponding to the site, and before the first determination unit determines the access status data corresponding to each of at least one site page based on the historical access log, the elimination unit is configured to:
[0020] Based on the record information in the access record data corresponding to each source access address, eliminate the source access addresses that do not meet the conditions and the corresponding access record data in the historical access log.
[0021] In a possible implementation manner, the elimination unit is specifically configured to:
[0022] For one source access address among each source access address, respectively perform at least one of the following operations:
[0023] If the record information includes the access times of a source access address and the access times reach the times threshold, then eliminate the source access address and the corresponding access record data;
[0024] If the record information contains the access time interval of a source access address and the access time interval is less than the interval threshold, then a source access address and the corresponding access record data are excluded;
[0025] If the record information contains the access success rate of a source access address and the access success rate is lower than the success rate threshold, then a source access address and the corresponding access record data are excluded.
[0026] In a possible implementation, the access status data includes at least one of the number of times a respective site page is visited, the access success rate, and the access time interval.
[0027] In a possible implementation, if the access status data includes the access success rate, then the access success rate corresponding to one of the at least one site pages is determined by the first determination unit in the following manner:
[0028] Determine at least one source access address corresponding to a site page, and the respective sub - access success rates when accessing a site page through at least one access source address;
[0029] Based on at least one sub - access success rate, determine the target quantity corresponding to the same sub - access success rate;
[0030] Based on the target quantity and the first total quantity of sub - access success rates, through the formula Determine the access success rate;
[0031] where p represents the access success rate, n represents the sub - access success rate, m represents the target quantity, and s represents the first total quantity.
[0032] In a possible implementation, if the access status data includes the access time interval, then the access time interval corresponding to one of the at least one site pages is determined by the first determination unit in the following manner:
[0033] Determine each sub - access time interval corresponding to a site page, and the second total quantity of sub - access time intervals;
[0034] Based on the second total quantity and each sub - access time interval, through the formula Determine the access time interval;
[0035] where q represents the access time interval, x i represents each sub - access time interval, represents the average time interval determined based on each sub - access time interval, and r represents the second total quantity.
[0036] In a possible implementation, if the access status data includes the number of visits, the access success rate, and the access time interval corresponding to each of at least one site page, the clustering unit is specifically configured to:
[0037] For each clustering set, respectively determine at least one target site page included therein, and based on the number of visits, the access success rate, and the access time interval corresponding to each of the at least one target site page, respectively determine the corresponding central number of visits, central access success rate, and central access time interval;
[0038] Based on the central number of visits, the central access success rate, and the central access time interval corresponding to each clustering set, through the formula Determine their respective weight values;
[0039] Where, w represents the weight, p′ represents the central access success rate, b represents the empirical value, q′ represents the central access time interval, and o′ represents the central number of visits.
[0040] In a possible implementation, the second determination unit is specifically configured to:
[0041] Sort the corresponding clustering sets according to the magnitudes of the determined weight values, and determine the importance levels of the site pages included in each clustering set based on the sorting result.
[0042] In a third aspect, an embodiment of the present application provides an electronic device, including: a memory and a processor, wherein the memory is used to store computer instructions; the processor is used to execute the computer instructions to implement the steps of the method for grading site pages provided by the embodiment of the present application.
[0043] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, where the computer-readable storage medium stores computer instructions, and when the computer instructions are executed by a processor, the steps of the method for grading site pages provided by the embodiment of the present application are implemented.
[0044] In a fifth aspect, an embodiment of the present application provides a computer program product, which includes computer instructions stored in a computer-readable storage medium; when the processor of an electronic device reads the computer instructions from the computer-readable storage medium, the processor executes the computer instructions, so that the electronic device executes the steps of the method for grading site pages provided by the embodiment of the present application.
[0045] The beneficial effects of the present application are as follows:
[0046] The embodiments of the present application provide a method, apparatus, device, and storage medium for site page grading. First, obtain the historical access logs corresponding to the site, where the historical access logs are used to record the access record data when accessing at least one site page in the site through each source access address. Then, based on the historical access logs, determine the access status data corresponding to each of the at least one site page, and based on the access status data, perform clustering processing on the at least one site page, and determine the weight value of each clustering set. After determining the weight value, based on each weight value, determine the importance level of the site pages included in the corresponding clustering set respectively. By directly analyzing based on the historical access records and grading the site pages according to the real access records, it is closer to the actual situation compared with the method of grading based on the site page hierarchy, and compared with the PageRank algorithm, it solves the grading problems of new site pages with high popularity and original site pages with low popularity, and more accurately reflects the importance level of the site pages. Further, reasonably monitor the site pages according to the importance level of the site pages to ensure the accuracy of site page monitoring, so as to avoid the problem of not being able to detect malicious attacks such as site page tampering and hanging horses in time due to low monitoring frequency and intensity, that is, ensure the security of the site pages; and avoid the problem of resource waste caused by high monitoring frequency and intensity, that is, reduce losses. And the implementation method proposed in the present application does not require actions such as regular crawler crawling, avoiding bringing additional access load.
[0047] Other features and advantages of the present application will be described in the following specification, and, in part, will be obvious from the specification, or will be understood by implementing the present application. The objectives and other advantages of the present application can be realized and obtained by the structures specifically pointed out in the written specification, claims, and drawings. Brief Description of the Drawings
[0048] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0049] Figure 1 It is a schematic diagram of an application scenario provided by the embodiments of the present application;
[0050] Figure 2 It is a flowchart of a method for site page grading provided by the embodiments of the present application;
[0051] Figure 3 It is a flowchart of a specific implementation method for site page grading provided by the embodiments of the present application;
[0052] Figure 4 Structural diagram of a device for site page grading provided by an embodiment of the present application;
[0053] Figure 5 Structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0054] In order to make the objectives, technical solutions and beneficial effects of the present application clearer and more understandable, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all of them. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.
[0055] To facilitate better understanding of the technical solutions of the present application by those skilled in the art, some concepts involved in the present application are introduced below.
[0056] A site refers to a web page that can be accessed through the Internet. A site can consist of one page or multiple pages.
[0057] The source access address is the Internet Protocol Address (IP), which is a logical address assigned to each network and each host on the Internet. In the embodiments of the present application, it can be a logical address assigned to each terminal device.
[0058] The Uniform Resource Locator (URL) is a representation method used by the World Wide Web service program on the Internet to specify the location of information. In the embodiments of the present application, the URL is used to specify the site page. A site page corresponds to a unique URL. It can also be said that the URL is the unique identifier of the site page. Therefore, accessing the site page can also be referred to as accessing the URL, and the two can be used interchangeably in the embodiments of the present application.
[0059] The term "exemplary" used hereinafter means "serving as an example, embodiment or illustration". Any embodiment described as "exemplary" does not have to be construed as superior to or better than other embodiments.
[0060] The terms "first" and "second" in the text are only used for descriptive purposes and cannot be construed as explicitly or implicitly indicating relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more features. In the description of the embodiments of the present application, unless otherwise stated, the meaning of "a plurality" is two or more.
[0061] The design concept of the embodiments of the present application will be briefly introduced below:
[0062] With the development of technology and the growth of business, the number of site pages is also increasing continuously. To ensure the security of site pages, it is currently proposed to monitor the site pages. When monitoring the site pages, they are monitored according to the importance level of the site pages, and different frequencies and intensities of monitoring are carried out for different importance levels, so that the resources such as network and computing consumed by the monitoring are controllable. Therefore, how to classify the site pages is crucial.
[0063] In the related art, there are mainly two ways to classify the site pages, which are respectively:
[0064] Method 1: Use a crawler to construct a page tree and determine the importance level according to the level where the site page is located.
[0065] The level of the site page depends on the site organizational structure. The importance level of the site page with a higher level is higher, and the importance level of the site page with a lower level is lower. At this time, when monitoring the site pages according to the importance level, the monitoring frequency and monitoring intensity corresponding to the site page with a higher level are higher, while the monitoring frequency and monitoring intensity corresponding to the site page with a lower level are lower.
[0066] However, the site page with a higher level is not necessarily accessed frequently. Similarly, the site page with a lower level may be accessed frequently, and the access popularity is higher. Moreover, the site page with a higher access popularity is more likely to be maliciously attacked such as being tampered with or having a Trojan horse implanted. At this time, because the site page with a lower level has a higher access popularity, the site page with a lower level is more likely to be maliciously attacked such as being tampered with or having a Trojan horse implanted. However, the monitoring frequency and monitoring intensity of the site page with a lower level are lower, and it cannot be found in time that the site page has been attacked, that is, the problem cannot be found in time and huge losses will be caused.
[0067] Method 2: Use the PageRank algorithm to measure according to the mutual references between site pages.
[0068] The PageRank algorithm defines a random walk model on a directed graph formed by website pages, describing the behavior of a random walker randomly visiting each node along the directed graph. Moreover, the PageRank algorithm determines the importance level of a website page based on the number of links from the website page to other website pages. At this time, if an original website page is linked by many other website pages, it indicates that the importance level of this original website page is relatively high. However, for a new website page, the number of other website pages linked by this new website page is small, and the importance level of this new website page is relatively low. At this time, when monitoring website pages according to the importance level, the monitoring frequency and intensity corresponding to the original website page are relatively high, while the monitoring frequency and intensity corresponding to the new website page are relatively low.
[0069] However, new website pages often have a high number of visits, that is, a high access popularity. Therefore, new website pages are more likely to be maliciously attacked such as being tampered with and implanted with malware. However, the monitoring frequency and intensity of new website pages are relatively low, and it is impossible to promptly discover that the website page has been attacked, that is, it is impossible to promptly discover problems, resulting in huge losses.
[0070] Therefore, how to classify website pages, ensure the accuracy of website page classification, further ensure the accuracy of website page monitoring, promptly discover whether website pages have been maliciously attacked such as being tampered with and implanted with malware, and avoid losses are technical problems that need to be solved currently.
[0071] In view of this, the embodiments of the present application provide a method, device, electronic device, and storage medium for classifying website pages. In the embodiments of the present application, first, obtain the historical access log corresponding to the website, where the historical access log is used to record the access record data when accessing at least one website page in the website through each source access address. Then, based on the historical access log, determine the access status data corresponding to each of the at least one website page, and based on the access status data, perform clustering processing on the at least one website page, and determine the weight value of each clustering set. After determining the weight value, based on each weight value, determine the importance level of the website pages included in the corresponding clustering set respectively.
[0072] Analyze directly based on historical access records, and classify site pages according to real access records. This is more in line with reality compared to the method of classifying site pages based on the page hierarchy of the site, and compared to the PageRank algorithm, it solves the classification problems of new site pages with high popularity and original site pages with low popularity, and more accurately reflects the importance of site pages. Further, reasonably monitor site pages according to the importance of site pages to ensure the accuracy of site page monitoring, so as to avoid problems such as the inability to detect malicious attacks such as site pages being tampered with and infected with malware in a timely manner due to low monitoring frequency and intensity, that is, to ensure the security of site pages; and avoid problems such as resource waste caused by high monitoring frequency and intensity, that is, to reduce losses. Moreover, the implementation method proposed in this application does not require actions such as regular web crawler crawling, avoiding bringing additional access load to the site.
[0073] The preferred embodiments of the present application will be described below in conjunction with the accompanying drawings of the specification. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present application, and are not used to limit the present application. And without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other.
[0074] Refer to Figure 1 , Figure 1 is a schematic diagram of the application scenario of the embodiment of the present application. In this application scenario, it includes a terminal device 110 and a server 120, and the terminal device 110 and the server 120 can communicate through a communication network.
[0075] In an alternative embodiment, the communication network can be a wired network or a wireless network. Therefore, the terminal device 110 and the server 120 can be directly or indirectly connected through wired or wireless communication methods. For example, the terminal device 110 can be indirectly connected to the server 120 through a wireless access point, or the terminal device 110 can be directly connected to the server 120 through the Internet. The present application does not make any restrictions here.
[0076] In the embodiment of the present application, the terminal device 110 includes, but is not limited to, devices such as mobile phones, tablet computers, laptop computers, desktop computers, e-book readers, intelligent voice interaction devices, intelligent home appliances, vehicle-mounted terminals, etc.; various clients can be installed on the terminal device, and the client can be an application program (such as a browser, a game software, etc.), or a web page, a small program, etc.;
[0077] Server 120 is a background server corresponding to the client installed in the terminal device 110. Server 120 can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.
[0078] It should be noted that the method for grading site pages in the embodiments of the present application can be deployed in an electronic device, and the electronic device can be a server, where the server can be Figure 1 the server 120 shown in Figure 1 or other servers other than the server 120 shown in
[0079] Figure 1 The above is only an example, and actually the number of terminal devices 110 and servers 120 is not limited and is not specifically limited in the embodiments of the present application.
[0080] In the embodiments of the present application, when the number of servers 120 is multiple, the multiple servers 120 can form a blockchain, and the server 120 is a node on the blockchain; for example, the method for grading site pages disclosed in the embodiments of the present application, where the access record data, access status data, etc. involved can be saved on the blockchain.
[0081] Based on the above application scenarios, the method for grading site pages provided in the exemplary embodiments of the present application will be described below in combination with the above-described application scenarios according to the accompanying drawings. It should be noted that the above application scenarios are only shown for the convenience of understanding the spirit and principle of the present application, and the embodiments of the present application are not limited in this regard.
[0082] Refer to Figure 2 , Figure 2 An exemplary method for grading site pages in the embodiments of the present application is provided, and the method includes:
[0083] Step S200, obtaining the historical access logs corresponding to the site, where the historical access logs are used to record the access record data when accessing at least one site page in the site through each source access address.
[0084] Exemplarily, obtain the access record data of each source access address accessing the site page in the past week, or obtain the access record data of each source access address accessing the site page in the past month; among them, one week and one month are examples, and the time can also be set longer or shorter and will not be elaborated here.
[0085] In a possible implementation, the record information included in the access record data includes, but is not limited to:
[0086] The number of accesses corresponding to each source access address, the number of accessed site pages, whether each access is successful, the access success rate, the access time, and the access time interval.
[0087] Exemplarily, in the embodiments of the present application, the access record data can be stored in the form of a table, as specifically shown in Table 1, where Table 1 takes only one source access address as an example:
[0088] Table 1
[0089]
[0090] It should be noted that the data in Table 1 is only for illustration to facilitate understanding of the implementation manner of the present application; and based on the data in Table 1, the number of accessed site pages, the access time interval, and the access success rate of the IP1 address can also be derived.
[0091] Similarly, the data corresponding to all source access addresses can be stored and determined in the above manner, and will not be repeated here.
[0092] Step S201, based on the historical access log, determine the access status data corresponding to each of at least one site page.
[0093] Since the historical access log records data such as the access time of each IP address to the site page and whether the access is successful; therefore, based on the historical access log, the source access address corresponding to each site page, that is, the source access address that has accessed the site page, and the access time of each source access address to the site page, and whether the access is successful at each time point and other data can be determined. As shown in Table 2, where Table 2 takes only one site page as an example:
[0094] Table 2
[0095]
[0096] It should be noted that the data in Table 2 is only for illustration to facilitate understanding of the implementation manner of the present application; and based on the data in Table 2, the number of times the site page URL1 is accessed, the access success rate, and the access time interval and other data can also be derived.
[0097] Next, the specific implementation manners for determining the number of times accessed, the access success rate, and the access time interval in the embodiments of the present application will be described respectively.
[0098] 1. Determine the number of times o the site page is accessed:
[0099] Since each source access address corresponds to an access time each time it accesses the site page, the number of times the site page is accessed, denoted as o, can be directly determined based on the access time.
[0100] II. Determine the access success rate p corresponding to the site page:
[0101] When determining the access success rate corresponding to the site page, first determine at least one source access address corresponding to the site page, and the respective sub-access success rates corresponding to accessing the site page through at least one access source address;
[0102] Since each source access address corresponds to a data indicating whether the access is successful each time it accesses the site page, based on the data indicating whether the access is successful corresponding to one source access address, the corresponding sub-access success rate when this source access address accesses this site page can be determined, that is, the success rate when this source access address accesses this site page; that is to say, a sub-access success rate = the number of successful accesses of a source access address to this site page / the total number of accesses of a source access address to this site page.
[0103] At this time, the sub-access success rates corresponding to each source access address accessing this site page can be determined; for example, taking the data in Table 2 as an example: the success rate of IP1 accessing URL1 is 80%, the success rate of IP2 accessing URL1 is 60%, the success rate of IP3 accessing URL1 is 100%, and the success rate of IP4 accessing URL1 is 100%. At this time, the at least one sub-access success rates are respectively: 80%, 60%, 100%, 100%.
[0104] After determining the at least one sub-access success rate, based on the at least one sub-access success rate, determine the target quantity corresponding to the same sub-access success rate; for example: the target quantity for 80% is 1, the target quantity for 60% is 1, and the target quantity for 100% is 2.
[0105] After determining the target quantity, based on the target quantity and the first total quantity of the sub-access success rates, through the formula Determine the access success rate; where p represents the access success rate, n represents each sub-access success rate, m represents the target quantity corresponding to each access success rate, and s represents the first total quantity.
[0106] Based on the above method, the access success rate corresponding to each site page can be accurately determined.
[0107] III. Determine the access time interval q corresponding to the site page:
[0108] When determining the access time interval corresponding to the site page, first determine the respective sub-access time intervals corresponding to one site page, and the second total quantity of the sub-access time intervals.
[0109] Since each source access address corresponds to an access time each time it accesses the site page, the site page corresponds to multiple accessed times. Then, based on the time order of the multiple accessed times, the time interval between two adjacent accessed times is determined, and the time interval between the two accessed times is used as a sub-accessed time corresponding to the site page, and the total number of sub-accessed time intervals is counted, that is, the second total number in the embodiments of the present application.
[0110] After determining each sub-accessed time interval and the second total number, based on the second total number and each sub-accessed time interval, through the formula the accessed time interval is determined; where q represents the accessed time interval, x i represents each sub-accessed time interval, represents the average time interval determined based on each sub-accessed time interval, and r represents the second total number.
[0111] It should be noted that the accessed time interval q in the embodiments of the present application is the standard deviation of the accessed time interval of the site page, which describes the degree of dispersion of the accessed interval of the site page.
[0112] Based on the above method, the accessed time interval corresponding to each site page can be accurately determined.
[0113] Therefore, in the embodiments of the present application, the access status data includes at least one of the accessed times, accessed success rates, and accessed time intervals corresponding to each site page.
[0114] It should be noted that the access status data corresponding to all site pages can be stored and determined in the above manner, and will not be repeated here.
[0115] Step S202: Based on the access status data, perform clustering processing on at least one site page, and determine the weight value of each clustering set.
[0116] In a possible implementation manner, when the access status data includes the accessed times, accessed success rates, and accessed time intervals corresponding to at least one site page respectively, the site pages are clustered according to the accessed times, accessed success rates, and accessed time intervals, that is, clustered into K categories, and K clustering sets are determined, where K is a positive integer.
[0117] When performing clustering processing, first select K initialized samples as the initial clustering centers:
[0118] a(o,p,q) = a1, a2…a k
[0119] Then, for the site pages in the site, calculate the distances from the site pages to the K clustering centers, and assign the site pages to the clustering set corresponding to the clustering center with the minimum distance, and recalculate the clustering centers for each clustering set;
[0120] In a possible implementation, through the formula where a j represents the clustering center, c i represents the clustering set, and s represents the site page, i.e., the sample data.
[0121] After obtaining the clustering sets, perform the following operations for each clustering set respectively to determine the weight value corresponding to the clustering set:
[0122] First, determine at least one target site page included in the clustering set, and respectively determine the corresponding central access times, central access success rate, and central access time interval based on the access times, access success rate, and access time interval corresponding to each of the at least one target site page;
[0123] Among them, the central access times, central access success rate, and central access time interval can be the data corresponding to the finally determined clustering center, or the mean values determined according to the access times, access success rate, and access time interval corresponding to each of the at least one target site page.
[0124] Then, based on the central access times, central access success rate, and central access time interval corresponding to the clustering set, through the formula determine their respective weight values; where w represents the weight, p′ represents the central access success rate, b represents the empirical value, q′ represents the central access time interval, and o′ represents the central access times.
[0125] Based on the above method, the weight value corresponding to each clustering set can be accurately determined, and the weight value is used to represent the importance degree. Therefore, based on the weight value, the site pages can be classified.
[0126] Step S203, based on each weight value, respectively determine the importance level of the site pages included in the corresponding clustering set.
[0127] In a possible implementation, sort the corresponding clustering sets according to the magnitudes of the determined weight values, and determine the importance level of the site pages included in each clustering set based on the sorting result.
[0128] Exemplarily, the larger the weight value, the higher the importance level of the site pages included in the clustering set corresponding to the weight value.
[0129] In this application, analysis is directly based on historical access records, and site page grading is performed according to real access records, which can more accurately reflect the importance of site pages.
[0130] In the embodiments of this application, considering that in the historical access log, in addition to the real user access data, there is also non-user access data generated by devices such as web crawler scanners. Therefore, in order to further ensure the authenticity of the data and the accuracy of site page grading, in the embodiments of this application, after obtaining the historical access log corresponding to the site page and before determining the access status data corresponding to at least one site page based on the historical access log, based on the record information in the access record data corresponding to each source access address, the source access addresses that do not meet the conditions and the corresponding access record data in the historical access log are excluded;
[0131] In a possible implementation, if the record information includes the access times of a source access address and the access times reach the times threshold, then the source access address and the corresponding access record data are excluded; or determine the access times of each source access address, sort them in descending order of access times, and exclude the source access addresses with the top-ranked access times and the corresponding access record data, where the top-ranked access times can be directly determined based on the ranking or determined according to a ratio.
[0132] In a possible implementation, if the record information contains the access time interval of a source access address and the access time interval is less than the interval threshold, then the source access address and the corresponding access record data are excluded;
[0133] In the embodiments of this application, the access time interval here is the average access time interval, and the average access time interval is determined in the following way:
[0134] According to the set time interval, calculate the average access time interval of each source access address within the set time interval. The access time interval is the difference between the next access time point and the previous access time point, and then take the average of all access time intervals to obtain the average access time interval. Among them, the set time interval can be 15 minutes, 20 minutes, and no setting is made here.
[0135] In a possible implementation, if the record information contains the access success rate of a source access address and the access success rate is lower than the success rate threshold, then the source access address and the corresponding access record data are excluded.
[0136] In this application, the historical access records are cleared, and the non-user real access data such as those from web crawler scanners are excluded. Based on the data after the clearing process, the site pages are graded, which further ensures the authenticity of the data and makes the grading of the site pages more accurate.
[0137] Please refer to Figure 3 , Figure 3 Exemplarily, a specific implementation method flowchart for grading site pages in an embodiment of this application is provided, including the following steps:
[0138] Step S300: Obtain the historical access log corresponding to the site. The historical access log is used to record the access record data when at least one site page in the site is accessed through each source access address.
[0139] Step S301: Based on the record information in the access record data corresponding to each source access address, exclude the source access addresses that do not meet the conditions and the corresponding access record data in the historical access log.
[0140] Step S302: Based on the historical access log after the exclusion process, determine the number of times each site page is accessed, the access success rate, and the access time interval respectively corresponding to at least one site page.
[0141] Step S303: Based on the number of times accessed, the access success rate, and the access time interval, perform clustering processing on at least one site page, and determine the weight value of each clustering set.
[0142] Step S304: Based on each weight value, determine the importance level of the site pages included in the corresponding clustering set respectively.
[0143] In this application, when grading the site pages, the analysis is directly based on the historical access records, and the site pages are graded according to the actual access records. This is closer to the reality compared to the method of grading based on the site page hierarchy, and compared to the PageRank algorithm, it solves the grading problems of new site pages with high popularity and original site pages with low popularity, and more accurately reflects the importance of the site pages. Moreover, the historical access records are cleared, excluding non-user real access data such as those from web crawler scanners, and the site pages are graded based on the data after the clearing process, which further ensures the authenticity of the data and the accuracy of the site page grading. Further, the site pages are reasonably monitored according to their importance levels to ensure the accuracy of the site page monitoring, so as to avoid problems such as the inability to detect malicious attacks such as site page tampering and malware injection in a timely manner due to low monitoring frequency and intensity, that is, to ensure the security of the site pages; and to avoid problems such as resource waste caused by high monitoring frequency and intensity, that is, to reduce losses. And the implementation method proposed in this application does not require actions such as regular crawler crawling, avoiding bringing additional access loads to the site.
[0144] It should be noted that the user information involved in the embodiments of this application is obtained under the permission of the user.
[0145] Based on the same inventive concept as the above method embodiment of this application, an apparatus for grading site pages is also provided in the embodiments of this application. The principle of the apparatus for solving problems is similar to that of the method in the above embodiment. Therefore, the implementation of the apparatus can refer to the implementation of the above method, and the repeated parts will not be elaborated.
[0146] Please refer to Figure 4 , Figure 4 Exemplarily, an apparatus 400 for grading site pages is provided in the embodiments of this application. The apparatus 400 for grading site pages includes:
[0147] An acquisition unit 401, configured to acquire the historical access log corresponding to the site, where the historical access log is used to record the access record data when accessing at least one site page in the site through each source access address;
[0148] A first determination unit 402, configured to determine the access status data corresponding to each of at least one site page based on the historical access log;
[0149] A clustering unit 403, configured to perform clustering processing on at least one site page based on the access status data, and determine the weight value of each clustering set;
[0150] A second determination unit 404, configured to determine the importance level of the site pages included in the corresponding clustering set respectively based on each weight value.
[0151] In a possible implementation, the device further includes an elimination unit 405;
[0152] After the acquisition unit 401 acquires the historical access logs corresponding to the sites, before the first determination unit 402 determines the access status data corresponding to each of at least one site page based on the historical access logs, the elimination unit 405 is configured to:
[0153] Based on the record information in the access record data corresponding to each source access address, eliminate the source access addresses that do not meet the conditions and the corresponding access record data in the historical access logs.
[0154] In a possible implementation, the elimination unit 405 is specifically configured to:
[0155] For one source access address among each source access address, respectively perform at least one of the following operations:
[0156] If the record information includes the access times of a source access address and the access times reach the times threshold, then eliminate a source access address and the corresponding access record data;
[0157] If the record information includes the access time interval of a source access address and the access time interval is less than the interval threshold, then eliminate a source access address and the corresponding access record data;
[0158] If the record information includes the access success rate of a source access address and the access success rate is lower than the success rate threshold, then eliminate a source access address and the corresponding access record data.
[0159] In a possible implementation, the access status data includes at least one of the number of times visited, the access success rate, and the access time interval corresponding to each of at least one site page.
[0160] In a possible implementation, if the access status data includes the access success rate, the access success rate corresponding to one site page among at least one site page is determined by the first determination unit 402 in the following manner:
[0161] Determine at least one source access address corresponding to a site page, and the respective sub-access success rates when accessing a site page through at least one access source address;
[0162] Based on at least one sub-access success rate, determine the target quantity corresponding to the same sub-access success rate;
[0163] Based on the target quantity and the first total quantity of the sub-access success rates, through the formula Determine the access success rate;
[0164] Among them, p represents the success rate of being visited, n represents the success rate of sub - being visited, m represents the target quantity, and s represents the first total quantity.
[0165] In a possible implementation manner, if the access status data includes the time interval of being visited, the time interval of being visited corresponding to one of the at least one site pages is determined by the first determination unit 402 in the following manner:
[0166] Determine the respective sub - time intervals of being visited corresponding to a site page, and the second total quantity of the sub - time intervals of being visited;
[0167] Based on the second total quantity and the respective sub - time intervals of being visited, through the formula Determine the time interval of being visited;
[0168] Among them, q represents the time interval of being visited, x i represents the respective sub - time intervals of being visited, represents the average time interval determined based on the respective sub - time intervals of being visited, and r represents the second total quantity.
[0169] In a possible implementation manner, if the access status data includes the number of times of being visited, the success rate of being visited, and the time interval of being visited corresponding to each of the at least one site pages, the clustering unit 403 is specifically configured to:
[0170] For each clustering set, respectively determine at least one target site page included therein, and based on the number of times of being visited, the success rate of being visited, and the time interval of being visited corresponding to each of the at least one target site pages, respectively determine the central number of times of being visited, the central success rate of being visited, and the central time interval of being visited corresponding thereto;
[0171] Based on the central number of times of being visited, the central success rate of being visited, and the central time interval of being visited corresponding to each clustering set, through the formula Determine the respective weight values;
[0172] Among them, w represents the weight, p′ represents the central success rate of being visited, b represents the empirical value, q′ represents the central time interval of being visited, and o′ represents the central number of times of being visited.
[0173] In a possible implementation manner, the second determination unit 404 is specifically configured to:
[0174] Sort the corresponding clustering sets according to the magnitudes of the determined respective weight values, and determine the importance levels of the site pages included in each clustering set based on the sorting result.
[0175] For the convenience of description, the above parts are divided into respective units (or modules) according to their functions and described separately. Of course, when implementing the present application, the functions of the respective units (or modules) can be implemented in the same or multiple software or hardware.
[0176] Those skilled in the art to which the present application pertains can understand that various aspects of the present application can be implemented as a system, a method, or a program product. Therefore, various aspects of the present application can be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software aspects, which can be collectively referred to herein as "circuitry", "module", or "system".
[0177] After introducing the method and apparatus for hierarchical classification of site pages in the exemplary embodiments of the present application, next, an electronic device for hierarchical classification of site pages according to another exemplary embodiment of the present application will be introduced.
[0178] Based on the same inventive concept as the above method embodiment of the present application, an electronic device is also provided in the embodiment of the present application, and this electronic device can be a server. In this embodiment, the structure of the electronic device can be as Figure 5 shown, including a memory 501, a communication module 503, and one or more processors 502.
[0179] The memory 501 is used to store the computer program executed by the processor 502. The memory 501 may mainly include a program storage area and a data storage area. Among them, the program storage area may store an operating system and programs required to run the instant messaging function, etc.; the data storage area may store various instant messaging information and operation instruction sets, etc.
[0180] The memory 501 can be a volatile memory, such as a random-access memory (RAM); the memory 501 can also be a non-volatile memory, such as a read-only memory, a flash memory, a hard disk drive (HDD), or a solid-state drive (SSD); or the memory 501 is any other medium that can be used to carry or store a desired computer program in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 501 can be a combination of the above memories.
[0181] The processor 502 may include one or more central processing units (CPUs) or be a digital processing unit, etc. The processor 502 is used to implement the above-mentioned method for grading site pages when calling the computer program stored in the memory 501.
[0182] The communication module 503 is used to communicate with terminal devices and other servers.
[0183] In the embodiments of the present application, the specific connection medium between the above-mentioned memory 501, communication module 503 and processor 502 is not limited. In the embodiments of the present application Figure 5 it is described that the memory 501 and the processor 502 are connected through a bus 504. The bus 504 is depicted in thick lines in Figure 5 The connection manners between other components are only schematically illustrated and are not limited thereto. The bus 504 can be divided into an address bus, a data bus, a control bus, etc. For ease of description, Figure 5 it is only depicted by a thick line in
[0184] The memory 501 stores a computer storage medium. The computer storage medium stores computer-executable instructions. The computer-executable instructions are used to implement the method for grading site pages in the embodiments of the present application. The processor 502 is used to execute the above-mentioned method for grading site pages.
[0185] In some possible implementation manners, various aspects of the method for grading site pages provided in the present application can also be implemented in the form of a program product, which includes a computer program. When the program product runs on an electronic device, the computer program is used to cause the electronic device to execute the steps in the method for grading site pages according to various exemplary embodiments of the present application described above in this specification.
[0186] The program product can adopt any combination of one or more readable media. The readable media can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0187] The program product of an embodiment of the present application may be a portable compact disc read-only memory (CD-ROM) and include a computer program, and may be run on a computing device. However, the program product of the present application is not limited thereto. In this document, a readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with a command execution system, device, or component.
[0188] A readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which a readable computer program is carried. Such a propagated data signal may take many forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. A readable signal medium may also be any readable medium other than a readable storage medium that can send, propagate, or transmit a program for use by or in conjunction with a command execution system, device, or component.
[0189] The computer program contained on the readable medium may be transmitted by any appropriate medium, including but not limited to wireless, wired, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0190] It should be noted that although several units or subunits of the device are mentioned in the above detailed description, this division is merely exemplary and not mandatory. In fact, according to the embodiments of the present application, the features and functions of two or more of the above-described units may be embodied in one unit. Conversely, the features and functions of one unit described above may be further divided and embodied by a plurality of units.
[0191] In addition, although the operations of the method of the present application are described in a specific order in the drawings, this does not require or imply that these operations must be performed in that specific order, or that all of the illustrated operations must be performed to achieve the desired result. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution.
[0192] Those skilled in the art should understand that the embodiments of the present application may be provided as a method, system, or computer program product. Therefore, the present application may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable computer programs.
[0193] Although the preferred embodiments of the present application have been described, additional changes and modifications can be made to these embodiments by those skilled in the art once they learn the basic creative concept. Therefore, the appended claims are intended to be interpreted to include the preferred embodiments as well as all changes and modifications that fall within the scope of the present application.
[0194] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these modifications and variations.
Claims
1. A method for grading a site page, characterized in that, The method includes: Obtaining a historical access log corresponding to a site, where the historical access log is used to record access record data when at least one site page in the site is accessed through each source access address; Based on the historical access log, determining access status data corresponding to each of the at least one site page; Based on the access status data, performing clustering processing on the at least one site page and determining a weight value for each clustering set; Based on each of the weight values, respectively determining the importance level of the site pages included in the corresponding clustering set; If the access status data includes the number of times of being accessed, the success rate of being accessed, and the time interval of being accessed corresponding to each of the at least one site page, then determining the weight value for each clustering set includes: For each of the clustering sets, respectively determining at least one target site page included therein, and based on the number of times of being accessed, the success rate of being accessed, and the time interval of being accessed corresponding to each of the at least one target site page, respectively determining the central number of times of being accessed, the central success rate of being accessed, and the central time interval of being accessed corresponding thereto; Based on the respective center visit frequencies, center visit success rates, and center visit time intervals corresponding to each clustering set, through the formula determine their respective weight values; Among them, w represents the weight, p ′ represents the success rate of being visited at the center, b represents the experience value, q ′ represents the time interval between visits to the center, o ′ represents the number of times the center has been visited.
2. The method according to claim 1, characterized in that, After obtaining the historical access log corresponding to the site and before determining the access status data corresponding to each of the at least one site page based on the historical access log, it further includes: Based on the record information in the access record data corresponding to each source access address, removing the source access addresses that do not meet the conditions and the corresponding access record data from the historical access log.
3. The method according to claim 2, characterized in that The removing the source access addresses that do not meet the conditions and the corresponding access record data from the historical access log based on the record information in the access record data corresponding to each source access address includes: For one of the source access addresses in each of the source access addresses, respectively performing at least one of the following operations: If the record information includes the number of access times of the one source access address and the number of access times reaches a threshold number of times, then removing the one source access address and the corresponding access record data; If the record information includes the access time interval of the one source access address and the access time interval is less than a threshold interval, then removing the one source access address and the corresponding access record data; If the record information includes the access success rate of the one source access address and the access success rate is lower than a threshold success rate, then removing the one source access address and the corresponding access record data.
4. The method according to claim 1, wherein The access status data includes at least one of the number of times of being accessed, the success rate of being accessed, and the time interval of being accessed corresponding to each of the at least one site page.
5. The method according to claim 4, wherein If the access status data includes the success rate of being accessed, then the success rate of being accessed corresponding to one of the at least one site page is determined by the following method: Determining at least one source access address corresponding to the one site page and the corresponding sub-success rate of being accessed when accessing the one site page through the at least one access source address; Based on the at least one sub-success rate of being accessed, determining the target quantity corresponding to the same sub-success rate of being accessed; Based on the target quantity and the first total quantity of the sub-interview success rate, through the formula determine the interview success rate; Wherein, p represents the success rate of being visited, n represents the success rate of sub - being visited, m represents the target quantity, and s represents the first total quantity.
6. The method according to claim 4, characterized in that, If the access status data includes the time interval of being visited, the time interval of being visited corresponding to one of the at least one site pages is determined in the following manner: Determine each sub - being - visited time interval corresponding to the one site page, and the second total quantity of the sub - being - visited time intervals; Based on the second total quantity and each sub-access time interval, through the formula determine the access time interval; wherein, q represents the accessed time interval, and x i represents each sub-accessed time interval, and represents the average time interval determined based on each sub-accessed time interval, and r represents the second total quantity.
7. The method according to claim 1, characterized in that, Based on each of the weight values, respectively determining the importance level of the site pages included in the corresponding clustering set includes: Sort the corresponding clustering sets according to the magnitudes of the determined weight values, and determine the importance level of the site pages included in each clustering set based on the sorting result.
8. An apparatus for grading site pages, characterized in that, The apparatus includes: An acquisition unit, configured to acquire the historical access log corresponding to the site, where the historical access log is used to record the access record data when accessing at least one site page in the site through each source access address; A first determination unit, based on the historical access log, determines the access status data corresponding to each of the at least one site pages; A clustering unit, configured to perform clustering processing on the at least one site page based on the access status data, and determine the weight value of each clustering set; A second determination unit, configured to respectively determine the importance level of the site pages included in the corresponding clustering set based on each of the weight values; If the access status data includes the number of times of being visited, the success rate of being visited, and the time interval of being visited corresponding to each of the at least one site pages, the clustering unit is specifically configured to: For each clustering set, respectively determine at least one target site page included therein, and based on the number of times of being visited, the success rate of being visited, and the time interval of being visited corresponding to each of the at least one target site pages, respectively determine the central number of times of being visited, the central success rate of being visited, and the central time interval of being visited corresponding thereto; Based on the respective center visit times, center visit success rates, and center visit time intervals corresponding to each clustering set, through the formula determine their respective weight values; Among them, w represents the weight, and p ′ represents the success rate of being visited at the center, b represents the experience value, and q ′ represents the time interval between visits to the center, and o ′ represents the number of times the center has been visited.
9. An electronic device, characterized in that, The electronic device includes: a memory and a processor, wherein: The memory is used to store a computer program; The processor is configured to execute the computer program to implement the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer - readable storage medium stores a computer program, and when the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 7.
11. A computer program product, characterized in that, It includes computer instructions, and when the computer instructions are executed by the processor, it implements the steps of the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Website risk assessment method and device
CN112039885A