Rule matching method and system for browser, electronic equipment and medium

By generating user portraits and calculating interest correlation and rule matching probability, dynamically adjusting the order of execution of browser rules, the invalid matching problem caused by fixed priority is solved, and the efficiency of rule matching and user experience is improved.

CN120386949AActive Publication Date: 2025-07-29BEIJING FULE TECH CO LTD

Patent Information

Application Number
CN202510456792.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-12
Publication Date
2025-07-29
Estimated Expiration
2045-04-12

AI Technical Summary

Technical Problem

During the rule matching process, existing browsers cannot dynamically adjust the fixed priority, resulting in invalid matching, which reduces the efficiency of rule matching.

Method used

By obtaining user historical access data, the user's image is generated, the user's interest correlation degree of visit frequency and duration of each website is calculated, the number of triggers of the rule set is combined, the rule matching probability is calculated, and the rule matching sequence is generated, and the rule execution order is dynamically adjusted.

Benefits of technology

Improve the efficiency of browser rule matching, avoid invalid rule matching, and provide a more personalized web browsing experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120386949A_ABST
    Figure CN120386949A_ABST
Patent Text Reader

Abstract

The invention discloses a rule matching method and system for a browser, electronic equipment and a medium, and relates to the technical field of data processing. The method comprises the following steps: acquiring historical access data of a user in a browser, and generating a user portrait corresponding to the historical access data; according to the user portrait, determining access frequencies and access durations of the user in the plurality of websites, and calculating interest association degrees corresponding to the access frequencies and the access durations of the websites; for each website, determining a plurality of rule sets triggered in the website and triggering times corresponding to each rule set, and calculating a rule matching probability of the corresponding rule set in combination with the interest association degree and each triggering times; performing priority ranking on the rule matching probability of each rule set to generate a rule matching sequence; and when the user accesses any website, performing rule execution according to the rule matching sequence corresponding to the website. By implementing the technical scheme provided by the invention, the effect of improving the rule matching efficiency of the browser is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing, and specifically relates to a rule matching method, system, electronic device and medium for a browser. Background Art

[0002] With the rapid development of Internet technology, browsers have become the main entry point for users to access network resources. To improve the user experience and achieve personalized services, browsers usually integrate various rule engines to process access requests for different websites. These rules may include security protection rules, advertisement filtering rules, content recommendation rules, etc.

[0003] Currently, existing browsers use a preset rule set to process users' website access. These rule sets are usually matched and executed in a fixed priority order. However, in actual applications, due to differences in the network usage habits of different users, it is often difficult to make dynamic adjustments according to user differences only by using the existing fixed-priority rule matching method, resulting in a large number of invalid rule matches, thus reducing the rule matching efficiency of the browser. Summary of the Invention

[0004] This application provides a rule matching method, system, electronic device and medium for a browser, which has the effect of improving the rule matching efficiency of the browser.

[0005] In a first aspect, this application provides a rule matching method for a browser, including: Obtaining the historical access data of a user in the browser and generating a user portrait corresponding to the historical access data; Determining the access frequency and access duration of the user on multiple websites according to the user portrait, and calculating the interest correlation degree corresponding to the access frequency and access duration of each website; For each website, determining multiple rule sets triggered within the website and the number of trigger times corresponding to each rule set, and combining the interest correlation degree and each trigger time to calculate the rule matching probability of the corresponding rule set; Performing priority sorting on the rule matching probabilities of each rule set to generate a rule matching sequence; When the user accesses any website, performing rule execution according to the rule matching sequence corresponding to the website.

[0006] In a second aspect of this application, there is provided a rule matching system for a browser, the system includes: A user portrait generation module, configured to obtain the historical access data of a user in the browser and generate a user portrait corresponding to the historical access data; An interest correlation calculation module, where the user determines the access frequency and access duration of the user on multiple websites according to the user profile, and calculates the interest correlation corresponding to the access frequency and access duration of each website; A matching probability calculation module, which is used to determine, for each website, multiple rule sets triggered within the website and the number of trigger times corresponding to each rule set, and combine the interest correlation and each trigger time to calculate the rule matching probability of the corresponding rule set; A rule matching module, which is used to prioritize the rule matching probabilities of each rule set to generate a rule matching sequence; when the user accesses any website, the rules are executed according to the rule matching sequence corresponding to the website.

[0007] In a third aspect of the present application, an electronic device is provided, including a memory, a processor, and a program stored on the memory and executable on the processor. When the program is loaded and executed by the processor, it can implement a rule matching method for a browser.

[0008] In a fourth aspect of the present application, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the processor is enabled to implement a rule matching method for a browser.

[0009] In summary, one or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages: By adopting the above technical solutions, a user profile is generated by obtaining the historical access data of the user in the browser, and the interest correlation is calculated based on the access frequency and access duration of the user on multiple websites determined according to the user profile. The rule matching probability is calculated by combining the trigger times of the rule sets within the website and the interest correlation, so as to generate a targeted rule matching sequence, enabling the browser to dynamically adjust the execution order of the rules according to the actual usage habits of the user, avoiding a large number of invalid rule matching problems caused by the fixed priority rule matching method, and thus improving the rule matching efficiency of the browser. Description of the Drawings

[0010] Figure 1 is a schematic flowchart of a rule matching method for a browser provided by an embodiment of the present application; Figure 2 is a schematic structural diagram of a rule matching system for a browser provided by an embodiment of the present application; Figure 3 is a schematic structural diagram of an electronic device provided by an embodiment of the present application.

[0011] Description of reference numerals: 300, electronic device; 301, processor; 302, communication bus; 303, user interface; 304, network interface; 305, memory. DETAILED DESCRIPTION

[0012] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the drawings in the embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this application, not all of the embodiments.

[0013] In the description of the embodiments of this application, words such as "for example" or "for instance" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as "for example" or "for instance" in the embodiments of this application should not be construed as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "for example" or "for instance" is intended to present the relevant concepts in a concrete manner.

[0014] In the description of the embodiments of the present application, the term "multiple" means two or more. For example, multiple systems refer to two or more systems, and multiple screen terminals refer to two or more screen terminals. In addition, the terms "first" and "second" are used for descriptive purposes only and are not to be understood as indicating or implying relative importance or implicitly indicating the indicated technical features. Thus, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. The terms "including", "comprising", "having" and their variations all mean "including but not limited to", unless otherwise specifically emphasized.

[0015] The present application embodiment provides a rule matching method for a browser. In one embodiment, please refer to Figure 1 , Figure 1 This is a flow chart of a rule matching method for a browser provided in an embodiment of the present application. This method can be implemented using a computer program, which can be integrated into an application or run as a standalone tool application. This method can also be implemented using a single-chip microcomputer or run on a rule matching system for a browser based on the von Neumann architecture. Specifically, this method can include the following steps: Step 101: Obtain the user's historical access data in the browser and generate a user profile corresponding to the historical access data.

[0016] Historical access data refers to a series of access records generated when users access websites through browsers. This data includes, but is not limited to, the URLs of websites visited, access timestamps, page dwell time, website categories, webpage content tags, user interactions on the pages (such as scrolling and clicking), and triggered browser rules.

[0017] A user profile is a user characteristic model built based on historical visit data. This model is composed of multiple user preference feature vectors, each of which contains information on dimensions such as website type, visit time, and user behavior. A user profile can be understood as a digital, structured description of a user's online behavior habits. It quantifies and organizes the characteristics of a user's browsing behavior over different time periods to form a comprehensive set of user behavior characteristics. User profiles are primarily used to assess user interest in and visit tendencies for different types of websites, providing basic data for calculating website interest relevance.

[0018] Specifically, the system first obtains the user's historical access data through the browser's history record interface. Because user browsing behavior may vary across time periods, the obtained historical access data is divided into pre-set time periods (such as daily, weekly, or monthly) to generate multiple sub-access records. This time division better captures the changing characteristics of user behavior over time. For each sub-access record, the system extracts three key features: website type, access time, and user operation behavior. Website type features include the website's subject category and content tags, which are used to understand user interests. Access time features include the time distribution and average visit duration, reflecting the user's online habits. User operation features include page scrolling distance, page dwell time, and page click behavior, reflecting the user's level of interest in the content. By analyzing these features, the system generates a corresponding user preference feature vector for each sub-access record. These feature vectors contain information about the user's interests, preferences, and behavioral patterns during that time period. Finally, the user preference feature vectors for each time period are combined to form a complete user profile. This time-series-based user profile construction method not only comprehensively characterizes user interests but also reflects the dynamic changes in user interests. The user profiles generated in this way provide a reliable data foundation for subsequent rule matching optimization, helping the system more accurately predict the types of rules that a user might trigger on different websites, thereby improving rule matching efficiency. For example, if the user profile shows that the user frequently visits shopping websites and has a high level of interaction, the system can prioritize the execution of rules related to e-commerce security, providing more targeted protection.

[0019] Based on the above embodiments, as an alternative embodiment, in step 101: generating a user profile corresponding to the historical access data, this step may further include the following steps: Step 201: Obtain the website access records in the historical access data and divide the website access records into sub-access records corresponding to multiple preset time periods.

[0020] Specifically, first obtain the user's website access records from the data storage module of the browser. Considering that the user's browsing behavior may change over time, it is necessary to divide the obtained access records in terms of time dimension. In specific implementation, the system will set preset time periods (for example, 24 hours can be used as a period, or 7 days can be used as a period), and then, according to the timestamp information in the access records, group all the access records according to these preset time periods to obtain multiple sub-access records. This way of dividing based on time periods can ensure that the time characteristics of user behavior are fully considered in subsequent analysis, which helps to capture the changes in user behavior patterns in different periods.

[0021] Step 202: Respectively extract the corresponding website type features, access time features, and user operation behavior features from each sub-access record. The website type features include the theme category and content tags of the website, the access time features include the time period distribution of access and the average access duration, and the user operation features include the page scrolling distance, page stay duration, and page click behavior.

[0022] Specifically, the system extracts features from each sub-access record. First, extract the website type features. By parsing the website URL and content analysis, determine the theme category of the website (such as news, shopping, social, etc.) and content tags (such as technology, sports, fashion, etc.); then extract the access time features. By analyzing the timestamps of the access records, count the access distribution of the user in each time period (such as morning, afternoon, evening), and calculate the average access duration; finally, extract the user operation features. By analyzing the user interaction data recorded by the browser, obtain the scrolling distance of the user on the page (which can be obtained by monitoring the scroll event), the stay duration (calculated according to the page opening and closing times), and the click behavior (recording the position and frequency of the user's clicks). This multi-dimensional feature extraction method can comprehensively depict the user's browsing habits and interest preferences.

[0023] Step 203: Generate corresponding user preference feature vectors based on the website type features, access time features, and user operation behavior features of each sub-access record; combine the user preference feature vectors into the user's user profile.

[0024] Specifically, first, various extracted features are numerically processed. The website type features are converted into categorical encoding vectors, and the time features and operation features are normalized. Then, these processed features are organized into a user preference feature vector according to a predefined structure. Each sub-access record generates a corresponding feature vector, and these feature vectors contain the complete behavior characteristics of the user within that time period. Finally, the system combines the feature vectors of all time periods, which can be done in a weighted average or serialized manner, to construct a complete user profile. This method of constructing a user profile based on multi-dimensional feature vectors can not only accurately describe the user's interest preferences but also reflect the changing trends of these preferences over time, providing a reliable data basis for subsequent rule matching optimization.

[0025] Step 102: According to the user profile, determine the access frequency and access duration of the user on multiple websites, and calculate the interest correlation degree corresponding to the access frequency and access duration of each website.

[0026] Among them, the access frequency refers to the number of valid accesses of the user to a specific website within the valid access period. Specifically, the access frequency here is not simply the total number of times the user opens the website, but the valid access volume after data cleaning and screening, that is, after excluding access records with a stay duration lower than the time threshold, the actual number of times the user browses and interacts with the content. In the embodiments of the present application, it can be understood that the access frequency is an important indicator to measure the access intensity of the user to a specific website, and it reflects the actual usage degree of the user to the website within a real and effective time period.

[0027] The access duration refers to the average stay time of the user in a single valid access to a specific website, which is obtained by calculating the ratio of the total stay time of the user on the website to the valid access frequency. Specifically, the access duration here is not simply the sum of all the time the user spends on the website, but the average stay time calculated after excluding short visits lower than the time threshold. This time includes the actual browsing, scrolling, and interaction time of the user on the page. In the embodiments of the present application, it can be understood that the access duration is a quantitative indicator to describe the degree of the user's attention to the website content, and it reflects the average depth of participation of the user in each access to the website.

[0028] Specifically, first, according to the time period distribution information in the user profile, determine the user's effective access time period, that is, the time period when the user is frequently active. During these effective access time periods, the system will count the access volume and stay duration of each website. To improve the reliability of the data, the system will set a time threshold to filter out those access records with too short stay duration (which may be due to user misoperation or quick jump), and thus obtain the effective access volume of each website. Based on these effective access volumes, the system calculates the effective access frequency of each website during the effective access time period, and at the same time, according to the cumulative stay duration and effective access frequency of each website, calculates the average stay duration of the user on each website. Next, the system will combine the content characteristics of the website, construct a website content feature vector according to the website category and content tags, and calculate the cosine similarity between this vector and the user preference feature vector to obtain the content matching degree. The system will also normalize the access frequency and access duration to obtain a unified access intensity index. Finally, by performing a weighted calculation on the content matching degree and the access intensity index, the interest association degree of each website is obtained. This calculation method that comprehensively considers access behavior and content matching can more accurately reflect the actual interest degree of the user in different websites, providing a more reliable basis for subsequent rule matching. For example, even if a user visits a certain website a large number of times, but if the stay time is short and the content matching degree is low, the interest association degree of this website will not be very high, which can avoid interference caused by some temporary or functional accesses.

[0029] Based on the above embodiments, as an optional embodiment, in step 102: According to the user profile, determine the access frequency and access duration of the user on multiple websites. This step may further include the following steps: Step 301: According to the time period distribution in the user profile, determine the user's effective access time period; during the effective access time period, count the access volume and stay duration of each website, and filter out the access records with a stay duration lower than the time threshold to obtain the effective access volume of each website.

[0030] Specifically, the system first needs to extract the time period distribution information from the user profile, calculate the proportion of the access record quantity in each time period in all access records, and obtain the corresponding time period proportion. When the proportion of a certain time period exceeds the preset proportion threshold (for example, it can be set to 15%), mark this time period as a candidate time period. These candidate time periods are combined to form the user's effective access time period, which can exclude the interference brought by the time periods with less user activity. Within the determined effective access time period, the system starts to count the original access data of each website, including the access volume and the stay duration. To improve the data quality, the system will set a time threshold (for example, it can be set to 10 seconds). When the stay duration of a certain access is lower than this threshold, it is considered that this access may be a user's misoperation or quick jump, and it will be excluded from the statistical results. Finally, the effective access volume of each website after screening is obtained. This data cleaning method based on the time dimension can effectively improve the accuracy of subsequent analysis.

[0031] Based on the above embodiments, as an optional embodiment, in step 301: determining the user's effective access time period according to the time period distribution in the user profile, this step may further include the following steps: Step 311: Determine multiple time periods in the time period distribution; calculate the proportion of the access record quantity in each time period in all access record quantities to obtain the corresponding time period proportion; screen out the time periods with a time period proportion greater than the proportion threshold as candidate time periods.

[0032] Specifically, when determining the user's effective access period, in order to accurately capture the user's active time characteristics, the system first needs to reasonably divide the 24 hours of a day. In specific implementation, the day is divided into multiple consecutive time periods. For example, it can be divided into 12 or 6 time periods according to every 2 hours or 4 hours. For example, the period from 0:00 to 2:00 is set as the first time period, the period from 2:00 to 4:00 is set as the second time period, and so on. This way of dividing time periods can ensure the appropriateness of the time granularity and will not increase the calculation burden due to overly fine division. For each time period, the system will count the total number of access records within this time period in the historical access data corresponding to the user profile, and divide it by the total sum of access records in all time periods to obtain the access proportion of this time period. For example, if the number of access records of a certain user in the time period from 8:00 to 10:00 accounts for 25% of the total access records, then the period proportion of this time period is 0.25. In order to screen out the truly active time periods of the user, the system sets a proportion threshold (for example, it can be set to 0.1 or 0.15). When the period proportion of a certain time period exceeds this threshold, it is marked as a candidate time period. This time period screening mechanism based on access frequency can effectively identify the user's main activity periods and avoid including the time periods of accidental or low-frequency access by the user in the consideration range. The candidate time periods determined in this way can more accurately reflect the user's Internet surfing habits and provide a reliable time range for subsequent access statistics. For example, if the system finds that a certain user is mainly active in the two time periods from 9:00 to 11:00 and from 14:00 to 16:00, then the access behaviors within these time periods will be considered more representative, and the corresponding access data will obtain higher weights in subsequent analysis. This data processing method considering time characteristics not only improves the quality of the data but also helps the system better understand and predict the user's access patterns.

[0033] Step 302: Based on each effective access volume, determine the effective access frequency of the corresponding website during the effective access period.

[0034] Specifically, the system calculates the effective access frequency of each website according to the filtered effective access volume in combination with the time span of the effective access period. In specific implementation, the system will count the cumulative effective number of times the user accesses each website during the effective access period and perform normalization processing according to the period span to obtain a comparable access frequency index. This calculation method ensures the comparability of data in different time periods and can truly reflect the user's access preferences for each website.

[0035] Step 303: Based on the stay duration and effective access frequency of each website, determine the average stay duration of a single access to the corresponding website; use each effective access frequency and each average stay duration as the access frequency and access duration of the user on each website respectively.

[0036] Specifically, the system first divides the total dwell time of each website during the effective access period (the time after excluding records below the threshold) by the effective access frequency of the website to obtain the average dwell time of the website. This average value can reflect the typical access behavior characteristics of users on the website. Then, the system uses the calculated effective access frequency and average dwell time as the normalized access metrics for users on each website respectively. This normalized metric system provides a reliable data basis for calculating the interest correlation degree of websites in the subsequent steps. For example, if a website has a high effective access frequency and a long average dwell time, it indicates that users show continuous and in-depth interest in the website, which will affect the matching priority of relevant rules.

[0037] Based on the above embodiments, as an alternative embodiment, in step 102: calculating the interest correlation degree corresponding to the access frequency and access duration of each website, this step may further include the following steps: Step 304: For each website, construct a website content feature vector based on the website category and content tags of the website.

[0038] Specifically, the system needs to digitally represent the content features of each website. First, the system will convert the website category (such as news, shopping, social, etc.) into a category code according to a predefined website classification system, and at the same time vectorize the content tags of the website (such as technology, sports, fashion, etc.). In specific implementation, the existing one-hot encoding method can be used to process the category features, and methods such as TF-IDF can be used to process the tag features. For example, for a news website, its category may be encoded as [1, 0, 0], and the content features containing "technology" and "finance" tags may be represented as [0.6, 0.4, 0, 0, 0]. Combine these encoded features according to a predefined structure to form the content feature vector of the website. This structured feature representation method enables the content features of the website to be quantitatively compared and mathematically operated.

[0039] Step 305: Calculate the cosine similarity between the website content feature vector of the website and the user preference feature vector to obtain the content matching degree.

[0040] Specifically, the system calculates the similarity between the website's content feature vector and the user preference feature vector in the user portrait. Since both vectors are high-dimensional vectors that have been standardized, the cosine similarity calculation method is most appropriate. During the specific calculation, the dot product of the two vectors is divided by the product of their moduli. The result range is between [-1, 1]. The closer the value is to 1, the higher the match. For example, if the user preference feature vector shows a high interest in science and technology news, then the cosine similarity calculated between the content feature vector of the science and technology news website and it will also be high. This matching calculation method based on vector similarity can accurately reflect the degree of fit between the website content and the user's interests.

[0041] Step 306: normalize the access frequency and access duration to obtain an access intensity index; perform weighted processing on the content matching degree and the access intensity index to obtain the interest relevance of the corresponding website.

[0042] Specifically, the system first normalizes the previously obtained access frequency and access duration data so that their values fall within the range [0, 1]. This can be achieved using a maximum-minimum normalization method: subtracting the minimum value from the original value and dividing the result by the difference between the maximum and minimum values. The normalized access frequency and access duration are then weighted averaged to produce an access intensity index. The weights can be adjusted based on the specific application scenario; for example, the access frequency weight can be set to 0.6 and the access duration weight to 0.4. The system then weights the normalized access intensity index and content match to produce the final interest relevance score. Specifically, the weight of the access intensity index can be set to α (e.g., 0.7) and the weight of the content match to 1-α (e.g., 0.3). The weighted sum of the two is used as the website's interest relevance score. This calculation method, which comprehensively considers both behavioral and content characteristics, provides a more comprehensive assessment of a user's actual interest in a website. For example, even if a user frequently visits a website, if the content matching degree is low, the final interest relevance will be reduced accordingly, which can avoid misjudging purely functional access behavior as interest tendencies.

[0043] Step 103: For each website, determine the multiple rule sets triggered in the website and the number of triggering times corresponding to each rule set, and calculate the rule matching probability of the corresponding rule set in combination with the interest association and the number of triggering times.

[0044] Among them, a rule set refers to a collection of rules in a browser that have the same functional purpose or processing logic. The rule set includes, but is not limited to, a security rule set for website security protection (such as XSS attack prevention rules, SQL injection prevention rules, etc.), a filtering rule set for ad filtering (such as pop-up ad blocking rules, banner ad filtering rules, etc.), a recommendation rule set for content recommendation (such as personalized content matching rules, related content recommendation rules, etc.), and a display rule set for page optimization (such as page layout adjustment rules, content display optimization rules, etc.). In the embodiments of the present application, it can be understood that the rule set is the basic execution unit when the browser processes website access requests, and each rule set contains attribute information such as trigger conditions, execution actions, and priorities.

[0045] The trigger count refers to the cumulative number of times that the trigger condition of a certain rule set is satisfied and the corresponding action is executed during the user's access to a specific website. When the user performs operations such as browsing, clicking, and inputting on the website, if these operation behaviors or website response contents meet the trigger conditions defined in the rule set (such as detecting suspicious cross-site scripts, discovering ad features, matching recommendation conditions, etc.), the system will record one trigger of the rule set. In the embodiments of the present application, it can be understood that the trigger count is a statistical indicator for measuring the actual call frequency of the rule set on a specific website, and it reflects the degree of association between the rule set and the user's access behavior.

[0046] Specifically, when calculating the rule matching probability, only considering the number of times a rule is triggered may not accurately reflect the actual importance of the rule to the user. Therefore, it is necessary to conduct a comprehensive evaluation in combination with the interest correlation degree of the website. First, all the rule set information triggered in each website will be recorded, including security protection rule sets, advertisement filtering rule sets, content recommendation rule sets, etc., and the number of times each rule set is triggered when the user visits the website will be counted. To evaluate the relative frequency of rule triggering, the system will obtain the total number of visits to the website in the historical access data. By dividing the number of times each rule set is triggered by the total number of visits, the trigger ratio of each rule set is obtained. This trigger ratio reflects the basic trigger probability of the rule set on this website. Next, the system multiplies the interest correlation degree of the website by the trigger ratio of the rule set to obtain the correlation strength of the rule set. This calculation method ensures that even if a rule is triggered frequently, if the user's interest in the website is low, the correlation strength of the corresponding rule set will also decrease. Then, the system will analyze the historical execution records of each rule set and count the ratio of the number of successful executions of the rule to the total number of triggers to obtain the execution success rate of the rule set. This success rate reflects the reliability and effectiveness of the rule set. Finally, the correlation strength of the rule set is multiplied by the execution success rate to obtain the rule matching probability of the rule set. For example, if a certain security protection rule set is frequently triggered on an e-commerce website, and the website has a high interest correlation degree, and the historical execution effect of this rule set is good, then this rule set will obtain a high matching probability. This probability calculation method considering multiple dimensions can not only accurately evaluate the importance of the rule set, but also help the system make more reasonable priority arrangements in subsequent executions.

[0047] Based on the above embodiments, as an optional embodiment, in step 103: Combining the interest correlation degree and the number of times each is triggered to calculate the rule matching probability of the corresponding rule set, this step may further include the following steps: Step 401: Obtain the total number of visits to the website in the historical access data, and calculate the trigger ratio of the number of times each rule set is triggered in the total number of visits.

[0048] Specifically, the system first extracts the total number of visits to a specific website from the user's historical access data. This total number of visits reflects the overall interaction frequency between the user and the website. Then, for each rule set triggered on this website, the system will count the number of times it is triggered and divide the number of times it is triggered by the total number of visits to obtain the trigger ratio of this rule set. For example, if the total number of visits to a certain website is 1000 times, and the security protection rule set is triggered 200 times, then the trigger ratio of this rule set is 0.2. This normalization processing method based on the total number of visits makes the trigger situations of different rule sets comparable and also eliminates the influence caused by differences in user access frequencies.

[0049] Step 402: Take the product of the interest correlation degree of the website and the trigger ratio of each rule set as the correlation strength of the corresponding rule set.

[0050] Specifically, the system multiplies the interest correlation degree of the website by the trigger ratio of the rule set to obtain the correlation strength of the rule set. This calculation method takes into account two important factors: one is the actual interest degree of the user in the website (reflected by the interest correlation degree), and the other is the application frequency of the rule set on this website (reflected by the trigger ratio). For example, if the interest correlation degree of a certain website is 0.8 and the trigger ratio of a certain rule set is 0.3, then the correlation strength of this rule set is 0.24. This comprehensive evaluation method ensures that the importance of the rule set is both related to the user's interest and consistent with its actual usage, avoiding the deviation that may be brought by a single - dimension evaluation.

[0051] Step 403: Obtain the historical execution results of each rule set, and determine the execution success rate of each rule set based on the historical execution results; calculate the product of the rule correlation strength and the execution success rate of each rule set to obtain the rule matching probability of the corresponding rule set.

[0052] Specifically, the system first analyzes the historical execution records of each rule set to obtain the number of successful executions and the number of failed executions of each rule set. A successful execution means that the rule set has completed the expected processing action and produced the expected effect, such as successfully intercepting malicious advertisements or accurately recommending relevant content. The system divides the number of successful executions by the total number of executions to obtain the execution success rate of the rule set. This success rate reflects the reliability and effectiveness of the rule set. Then, the system multiplies the correlation strength of the rule set by the execution success rate to obtain the final rule matching probability. For example, if the correlation strength of a certain rule set is 0.24 and the execution success rate is 0.9, then its rule matching probability is 0.216. This probability calculation method considering the execution effect can, while ensuring the importance of the rule, preferentially select the rule set with better execution effect, thereby improving the overall rule execution efficiency and user experience.

[0053] Step 104: Perform a priority ranking on the rule matching probabilities of each rule set to generate a rule matching sequence; when the user accesses any website, execute the rules according to the rule matching sequence corresponding to the website.

[0054] Among them, the rule matching sequence refers to an ordered execution queue formed by sorting multiple rule sets according to their rule matching probabilities for a specific website. The rule matching sequence contains key information such as the unique identifier, execution priority, and rule matching probability of the rule set. These rule sets may include different types such as security protection rule sets, advertisement filtering rule sets, and content recommendation rule sets. In the embodiments of the present application, it can be understood that the rule matching sequence is a dynamically generated and updated data structure, which reflects the relative importance and execution order of different rule sets on a specific website.

[0055] Specifically, during the actual website access process, due to limited system resources and timing requirements for rule execution, it is necessary to reasonably arrange the execution order of rule sets. The system first sorts the rule sets involved in each website based on the rule matching probabilities calculated above. Specifically, when implementing, the quicksort algorithm is used to sort the rule matching probabilities in descending order. The rule set with a higher rule matching probability will be given a higher execution priority. After sorting, the system will generate a unique rule matching sequence for each website, which contains the identifier of the rule set and the corresponding execution priority information. For example, for an e-commerce website, the generated rule matching sequence may be: Payment security protection rule set (probability 0.85), advertisement filtering rule set (probability 0.72), content recommendation rule set (probability 0.56), etc. When a user accesses the website, the system will automatically load the corresponding rule matching sequence and execute the rule sets in the priority order specified in the sequence. This dynamic priority mechanism based on probability ensures that the most important and most likely to be triggered rule sets can be executed first, thus improving the efficiency of rule execution. At the same time, since the rule matching sequence is dynamically generated based on historical data and can be adaptively adjusted as the user's behavior changes, it ensures the real-time and accuracy of rule execution. For example, if the system detects a change in the user's access pattern to a certain website, the corresponding rule matching probability will be updated, which will in turn lead to the reordering of the rule matching sequence, enabling the rule execution to better adapt to the user's current needs. This dynamically optimized execution mechanism not only improves the performance of the browser but also provides a more personalized web browsing experience for users.

[0056] Based on the above embodiments, as an optional embodiment, in step 104: sorting the rule matching probabilities of each rule set by priority to generate a rule matching sequence, this step may further include the following steps: Step 501: Extract rule priority identifiers from each rule set, and the rule priority identifiers include rule types and rule levels.

[0057] Specifically, the system needs to extract rule priority identification information from the metadata of each rule set, and these identification information are used to preliminarily classify and grade the rule sets. The rule type reflects the functional attributes of the rule set. For example, the rule set can be divided into different types such as security protection, advertisement filtering, content recommendation, etc.; the rule level indicates the importance of the rule set in this type, and can be set to different levels such as high, medium, and low. During the extraction process, the system will parse the configuration information of the rule set, standardize the rule type and rule level information, and store it in a format convenient for subsequent processing. This method of extracting identification based on rule attributes provides a basic classification basis for subsequent rule grouping. For example, a rule set for preventing XSS attacks may be identified as a high-level rule in the security protection category, while an ordinary advertisement filtering rule set may be identified as a medium-level rule in the advertisement filtering category.

[0058] Step 502: Group each rule set based on each rule type and each rule level to obtain multiple rule priority groups, where each rule priority group includes multiple rule sets with the same rule type and the same rule level.

[0059] Specifically, the system performs grouping processing on the rule sets based on the extracted rule priority identification. In specific implementation, first create a grouping matrix, where the rows of the matrix represent different rule types and the columns represent different rule levels. Then, the system traverses all rule sets and assigns the rule sets with the same rule type and rule level to the corresponding rule priority groups. For example, all rule sets identified as security protection and high level will be assigned to the same rule priority group. This multi-dimensional grouping method based on type and level can aggregate rule sets with similar functions and similar importance together, facilitating more detailed priority sorting in the future. Through this grouping process, the system can first determine the general priority range when executing rules, and then perform refined sorting within each group, improving the efficiency of rule scheduling. For example, when the system needs to process security-related rules, it can directly locate to the high-level rule group in the security protection category without having to traverse all rule sets.

[0060] Step 503: Within each rule priority group, sort in descending order according to the rule matching probability of each rule set to obtain the rule sorting result of the corresponding rule priority group.

[0061] Specifically, the system needs to perform a more refined sorting on the rule sets within each rule priority group. In the specific implementation, the system adopts the quicksort algorithm, using the rule matching probability as the key, and sorts the rule sets within each rule priority group in descending order. For example, in the high-level rule group of security protection, if there are three rule sets with rule matching probabilities of 0.85, 0.92, and 0.78 respectively, the sorted order will be: the rule set with a probability of 0.92 is ranked first, followed by the rule set with a probability of 0.85, and finally the rule set with a probability of 0.78. This in-group sorting method based on probability ensures that among rule sets of the same type and level, the rule set with a higher matching probability can be executed first. Through this refined sorting, the system can further optimize the execution order of rules on the basis of ensuring the priority of rule types and levels, and improve the accuracy of rule execution.

[0062] Step 504: Combine the sorted results of each rule in sequence according to the reference priority order of each rule priority group to generate a rule matching sequence.

[0063] Specifically, the system needs to integrate the sorted results of each rule priority group into the final rule matching sequence. First, the system will pre-define the reference priority order of each rule priority group, which is usually determined based on the importance of rule types and the levels of rules. For example, it can be set that the reference priority of the high-level rule group of security protection is the highest, followed by the high-level rule group of advertisement filtering, then the high-level rule group of content recommendation, and then in turn are the medium-level rule groups and low-level rule groups of each type. The system adds the sorted rule sets within each rule priority group to the final rule matching sequence in sequence according to this pre-defined reference priority order. This sequence generation method based on multi-level priorities not only ensures the priority execution of important rule types and high-level rules, but also realizes refined sorting based on matching probability among similar rules. The rule matching sequence generated in this way has a clear hierarchical structure and can better guide the system to execute rules. For example, when a user accesses a certain website, the system will first execute the rule set with the highest matching probability in the high-level rule group of security protection, and then execute other rule sets in sequence, thus ensuring the timely execution of key rules and improving the overall efficiency of rule execution.

[0064] Refer to Figure 2 , a system for a multi-channel based pixel filtering defect measurement method provided by an embodiment of the present application, the system includes: a user portrait generation module, an interest correlation calculation module, a matching probability calculation module, and a rule matching module, wherein: The user portrait generation module is used to obtain the historical access data of the user in the browser and generate a user portrait corresponding to the historical access data; An interest correlation degree calculation module, where the user determines the access frequency and access duration of the user on multiple websites according to the user profile, and calculates the interest correlation degree corresponding to the access frequency and access duration of each website; A matching probability calculation module, which is used to determine, for each website, multiple rule sets triggered within the website and the number of times each rule set is triggered, and calculate the rule matching probability of the corresponding rule set by combining the interest correlation degree and each number of trigger times; A rule matching module, which is used to sort the rule matching probabilities of each rule set by priority to generate a rule matching sequence; when the user accesses any website, the rules are executed according to the rule matching sequence corresponding to the website.

[0065] Based on the above embodiments, the user profile generation module is further used to obtain the website access records in the historical access data, and divide the website access records into sub-access records corresponding to multiple preset time periods; respectively extract the corresponding website type features, access time features, and user operation behavior features from each sub-access record, where the website type features include the theme category and content tags of the website, the access time features include the time period distribution of the access and the average access duration, and the user operation features include the page scrolling distance, the page stay duration, and the page click behavior; generate corresponding user preference feature vectors based on the website type features, access time features, and user operation behavior features of each sub-access record; combine each user preference feature vector into the user profile of the user.

[0066] Based on the above embodiments, the interest correlation degree calculation module is further used to determine the effective access period of the user according to the time period distribution in the user profile; within the effective access period, count the access volume and stay duration of each website, and screen out the access records with a stay duration lower than the time threshold to obtain the effective access volume of each website; based on each effective access volume, determine the effective access frequency of the corresponding website within the effective access period; based on the stay duration and effective access frequency of each website, determine the average stay duration of a single access to the corresponding website; use each effective access frequency and each average stay duration as the access frequency and access duration of the user on each website respectively.

[0067] Based on the above embodiments, the interest correlation degree calculation module is further used to determine multiple time periods in the time period distribution; calculate the proportion of the number of access records in each time period in the total number of all access records to obtain the corresponding period proportion; screen out the time periods with a period proportion greater than the proportion threshold as candidate time periods.

[0068] Based on the above embodiments, the interest correlation degree calculation module is further configured to, for each website, construct a website content feature vector based on the website category and content tags of the website; calculate the cosine similarity between the website content feature vector of the website and the user preference feature vector to obtain a content matching degree; perform normalization processing on the access frequency and access duration to obtain an access intensity index; perform weighted processing on the content matching degree and the access intensity index to obtain the interest correlation degree of the corresponding website.

[0069] Based on the above embodiments, the matching probability calculation module is further configured to obtain the total number of accesses of the website in the historical access data, and calculate the trigger proportion of the trigger times of each rule set in the total number of accesses; use the product of the interest correlation degree of the website and the trigger proportion of each rule set as the association intensity of the corresponding rule set; obtain the historical execution results of each rule set, and determine the execution success rate of each rule set based on the historical execution results; calculate the product of the rule association intensity and the execution success rate of each rule set to obtain the rule matching probability of the corresponding rule set.

[0070] Based on the above embodiments, the rule matching module is further configured to extract rule priority identifiers from each rule set, and the rule priority identifiers include rule types and rule levels; group each rule set based on each rule type and each rule level to obtain a plurality of rule priority groups, and each rule priority group includes a plurality of rule sets with the same rule type and the same rule level; within each rule priority group, sort the rule sets in descending order according to the rule matching probability of each rule set to obtain the rule sorting result of the corresponding rule priority group; combine the rule sorting results in sequence according to the benchmark priority order of each rule priority group to generate a rule matching sequence.

[0071] It should be noted that: when the device provided in the above embodiments implements its functions, only the above-mentioned division of each functional module is used for illustration. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the device and method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process is detailed in the method embodiments, which will not be repeated here.

[0072] This application also discloses an electronic device. Refer to Figure 3 , Figure 3 is a schematic structural diagram of an electronic device disclosed in an embodiment of the present application. The electronic device 300 may include: at least one processor 301, at least one network interface 304, a user interface 303, a memory 305, and at least one communication bus 302.

[0073] Among them, the communication bus 302 is used to realize the connection and communication between these components.

[0074] Among them, the user interface 303 may include a display interface and a camera interface. Optionally, the user interface 303 may further include a standard wired interface and a wireless interface.

[0075] Among them, the network interface 304 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface).

[0076] Among them, the processor 301 may include one or more processing cores. The processor 301 connects various parts within the entire server through various interfaces and lines. By running or executing instructions, programs, code sets, or instruction sets stored in the memory 305, and by calling data stored in the memory 305, the processor 301 performs various functions of the server and processes data. Optionally, the processor 301 may be implemented in at least one of the following hardware forms: digital signal processing (DSP), field-programmable gate array (FPGA), and programmable logic array (PLA). The processor 301 may integrate one or a combination of several of the following: a central processing unit (CPU), a graphics processing unit (GPU), and a modem. Among them, the CPU mainly processes the operating system, user interface graphics, and application programs, etc.; the GPU is responsible for rendering and drawing the content to be displayed on the display screen; the modem is used to process wireless communications. It can be understood that the above-mentioned modem may not be integrated into the processor 301 and may be implemented separately through a single chip.

[0077] Among them, the memory 305 may include a Random Access Memory (RAM), or may also include a Read-Only Memory. Optionally, the memory 305 includes a non-transitory computer-readable storage medium. The memory 305 can be used to store instructions, programs, codes, code sets, or instruction sets. The memory 305 may include a program storage area and a data storage area. Among them, the program storage area can store instructions for implementing an operating system, instructions for at least one function (such as a touch function, a sound playback function, an image playback function, etc.), instructions for implementing the above-mentioned method embodiments, etc.; the data storage area can store data involved in the above-mentioned method embodiments. Optionally, the memory 305 may also be at least one storage device located far from the aforementioned processor 301. Refer to Figure 3 , the memory 305 as a computer storage medium may include an operating system, a network communication module, a user interface module, and an application program for a rule matching method for a browser.

[0078] In Figure 3 In the electronic device 300 shown, the user interface 303 is mainly used to provide an interface for the user to obtain user input data; and the processor 301 can be used to call an application program for a rule matching method for a browser stored in the memory 305. When executed by one or more processors 301, the electronic device 300 executes the method of one or more of the above embodiments. It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be in other sequences or performed simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by this application.

[0079] In the above embodiments, the descriptions of the various embodiments have their own emphases. For parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0080] In several implementation manners provided by this application, it should be understood that the disclosed device can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling, direct coupling, or communication connection to each other can be through some service interfaces. The indirect coupling or communication connection of the device or unit can be in electrical or other forms.

[0081] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0082] In addition, each functional unit in various embodiments of this application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0083] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable memory. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of this application. And the aforementioned memory includes: various media such as USB flash drives, mobile hard disks, magnetic disks, or optical discs that can store program codes.

[0084] The above are only exemplary embodiments of this disclosure and cannot be used to limit the scope of this disclosure. That is, any equivalent changes and modifications made according to the teachings of this disclosure still fall within the scope covered by this disclosure. Those skilled in the art will easily think of other implementation schemes of this disclosure after considering the specification and the practice of the disclosure.

[0085] This application aims to cover any variations, uses, or adaptive changes of this disclosure. These variations, uses, or adaptive changes follow the general principles of this disclosure and include common general knowledge or conventional technical means in this technical field not recorded in this disclosure. The specification and the embodiments are only regarded as exemplary.

Claims

1. A rule matching method for a browser, characterized in that, Including: Obtain the historical access data of the user in the browser and generate a user portrait corresponding to the historical access data; According to the user portrait, determine the access frequency and access duration of the user on multiple websites, and calculate the interest correlation degree corresponding to the access frequency and access duration of each website; For each website, determine multiple rule sets triggered within the website and the corresponding trigger times of each rule set, and combine the interest correlation degree and each trigger time to calculate the rule matching probability of the corresponding rule set; Perform a priority sorting on the rule matching probabilities of each rule set to generate a rule matching sequence; When the user accesses any website, execute the rules according to the rule matching sequence corresponding to the website.

2. The rule matching method for a browser according to claim 1, characterized in that, The generating a user portrait corresponding to the historical access data includes: Obtain the website access records in the historical access data and divide the website access records into sub-access records corresponding to multiple preset time periods; Respectively extract the corresponding website type features, access time features, and user operation behavior features from each sub-access record. The website type features include the theme category and content tags of the website, the access time features include the time period distribution of access and the average access duration, and the user operation features include the page scrolling distance, page stay duration, and page click behavior; Based on the website type features, access time features, and user operation behavior features of each sub-access record, generate a corresponding user preference feature vector; Combine each user preference feature vector into the user portrait of the user.

3. The rule matching method for a browser according to claim 1, wherein The determining the access frequency and access duration of the user on multiple websites according to the user portrait includes: Determine the effective access period of the user according to the time period distribution in the user portrait; Within the effective access period, count the access volume and stay duration of each website, and screen out the access records with a stay duration lower than the time threshold to obtain the effective access volume of each website; Based on each effective access volume, determine the effective access frequency of the corresponding website within the effective access period; Based on the stay duration and effective access frequency of each website, determine the average stay duration of a single access to the corresponding website; Respectively use each effective access frequency and each average stay duration as the access frequency and access duration of the user on each website.

4. The method for rule matching for a browser according to claim 3, characterized in that The determining the effective access period of the user according to the time period distribution in the user portrait includes: Determine multiple time periods in the time period distribution; Calculate the proportion of the number of access records in each time period in all access records to obtain the corresponding period proportion; Screen out the time periods with a period proportion greater than the proportion threshold as candidate time periods.

5. The rule matching method for a browser according to claim 1, characterized in that, The calculating the interest correlation degree corresponding to the access frequency and access duration of each website includes: For each website, construct a website content feature vector based on the website category and content tags of the website; Calculate the cosine similarity between the website content feature vector of the website and the user preference feature vector to obtain the content matching degree; Normalize the access frequency and the access duration to obtain an access intensity index; Perform a weighting process on the content matching degree and the access intensity index to obtain the interest correlation degree of the corresponding website.

6. The rule matching method for a browser according to claim 1, characterized in that, Calculating the rule matching probability of the corresponding rule set by combining the interest correlation degree and each trigger count includes: Obtain the total number of accesses of the website in the historical access data, and calculate the trigger proportion of each trigger count of the rule sets in the total number of accesses; Use the product of the interest correlation degree of the website and the trigger proportion of each rule set as the correlation intensity of the corresponding rule set; Obtain the historical execution results of each rule set, and determine the execution success rate of each rule set based on the historical execution results; Calculate the product of the rule correlation intensity and the execution success rate of each rule set to obtain the rule matching probability of the corresponding rule set.

7. The rule matching method for a browser according to claim 1, characterized in that, Performing a priority sorting on the rule matching probabilities of each rule set to generate a rule matching sequence includes: Extract rule priority identifiers from each rule set, where the rule priority identifiers include rule types and rule levels; Group each rule set based on each rule type and each rule level to obtain a plurality of rule priority groups, where each rule priority group includes a plurality of rule sets with the same rule type and the same rule level; Within each rule priority group, perform a descending order sorting according to the rule matching probabilities of each rule set to obtain the rule sorting result of the corresponding rule priority group; Combine the rule sorting results in sequence according to the benchmark priority order of each rule priority group to generate a rule matching sequence.

8. A rule matching system for a browser, characterized in that The system includes: A user profile generation module, configured to obtain the historical access data of the user in the browser and generate a user profile corresponding to the historical access data; An interest correlation degree calculation module, configured to determine, according to the user profile, the access frequency and access duration of the user on multiple websites, and calculate the interest correlation degree corresponding to the access frequency and access duration of each website; A matching probability calculation module, configured to, for each website, determine a plurality of rule sets triggered in the website and the trigger count corresponding to each rule set, and combine the interest correlation degree and each trigger count to calculate the rule matching probability of the corresponding rule set; A rule matching module, configured to perform a priority sorting on the rule matching probabilities of each rule set to generate a rule matching sequence; when the user accesses any website, execute rules according to the rule matching sequence corresponding to the website.

9. An electronic device, characterized in that, Including a processor, a memory, a user interface, and a network interface, where the memory is used to store instructions, the user interface and the network interface are used to communicate with other devices, and the processor is used to execute the instructions stored in the memory so that the electronic device executes the rule matching method for the browser according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions, and when the instructions are executed, the rule matching method for the browser according to any one of claims 1-7 is executed.

Citation Information

Patent Citations

  • Method, device and system for improving data transmission speed of website

    CN102033883A

  • Method and device for generating user interest label

    CN103870512A

  • Rules based data processing system and method

    CN105308558A

  • Website protection method and device, website protection equipment and readable storage medium

    CN107580005A

  • Integrated system for rule editing, simulation, version control, and business process management

    CN110892375A

Cited By

  • Website ranking method and system based on big data analysis, medium and product

    CN121636847A