Auxiliary lesson preparation system and method based on AI intelligence

Through AI intelligent optimization of web crawler path planning and real-time monitoring, the problem of inaccurate path planning in existing technologies is solved, efficient and accurate acquisition of lesson preparation resources is achieved, and the system's anti-crawling capabilities and dynamic adaptability are enhanced.

CN120632183AInactive Publication Date: 2025-09-12HEBEI BOTU COMMUNICATION TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510751837.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-09-12
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the existing technology, the path planning of web crawlers in the process of assisting lesson preparation is not accurate enough, resulting in low efficiency and accuracy in resource acquisition, and insufficient anti-crawling mechanism capabilities, making it difficult to cope with website anti-crawling strategy upgrades or resource update delays.

Method used

Through an AI-based intelligent method, anti-crawling friendliness is calculated using response delay tolerance, proxy IP compatibility, and verification complexity, the crawling path of the web crawler is planned, and the primary and backup paths are determined through cluster analysis and a dual screening mechanism. The response time and success rate of the crawling path are monitored in real time to optimize the path switching strategy.

Benefits of technology

It improves the accuracy of web crawlers in planning crawling paths for lesson preparation resources, enhances the system's adaptability to dynamic environments, reduces resource waste and task interruptions, and ensures the efficiency and accuracy of resource acquisition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120632183A_ABST
    Figure CN120632183A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, in particular to an auxiliary lesson preparation system and method based on AI intelligence, and the method comprises the steps: obtaining lesson preparation information and website information, and planning a plurality of website crawling paths of web crawlers based on the anti-crawling friendliness and the resource updating frequency of websites; determining a main crawling path and a plurality of standby crawling paths based on the average resource quality density and the path independence degree of the plurality of website crawling paths; determining whether to switch from the main crawling path to the standby crawling path based on whether the average response time of the main crawling path fluctuates or whether the real-time success rate is abnormal; and determining the switched standby crawling path based on the coverage rate of the main crawling path limiting node and the unblocked node of the standby crawling path and the resource association degree of the main crawling path and the standby crawling path. According to the method, the retrieval efficiency of lesson preparation resources is improved by improving the accuracy of crawling path planning of the web crawler.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to an AI-based auxiliary lesson preparation system and method. Background Art

[0002] In the traditional lesson preparation process, resource acquisition relies on manual retrieval and screening, which has problems such as low efficiency, information overload and uneven quality, and obstruction of anti-crawling mechanisms, making it difficult for crawler tools to efficiently obtain data. The current technical solutions for auxiliary lesson preparation have the following main defects: only recommending websites based on simple keyword matching, without fine-grained screening based on dimensions such as teaching stage and resource type; weak anti-crawling avoidance capabilities: the anti-crawling friendliness of the website is not quantitatively evaluated, and crawling parameters such as proxy IP and request interval cannot be designed in a targeted manner; path planning is static: once the crawling path is set, it is difficult to adjust dynamically, and there is a lack of flexibility when facing website anti-crawling strategy upgrades or resource update delays; the backup path switching strategy is single: resource relevance and node coverage are not comprehensively considered, which may lead to a decline in resource quality or an increase in anti-crawling risk after switching.

[0003] For example, Chinese patent application publication number: CN117056612B discloses a method and system for pushing lesson preparation data based on AI assistance. By utilizing the chapter learning behavior data of sample subject chapters and the subject chapter embedding recognition rules, it is possible to accurately identify sample subject chapters related to weak knowledge points in the subject, and cluster the learning behavior data and update the parameters based on these sample subject chapters, thereby generating a target subject weak point prediction network for weak knowledge points in different subjects. At the same time, using the loaded learning behavior data of any target subject chapter, the confidence level of the weak knowledge points of each subject under the chapter can be determined, further guiding the push of lesson preparation data. In other words, the application can quickly predict students' weak knowledge points in different subject chapters, and accordingly push the corresponding number of lesson preparation materials to teachers. Teachers can better understand students' learning needs and provide targeted teaching support and educational resources.

[0004] However, the existing technology has the problem that the crawling path planning of lesson preparation resources using web crawlers is not accurate enough, resulting in low efficiency and low accuracy in obtaining lesson preparation resources. Summary of the Invention

[0005] To this end, the present invention provides an auxiliary lesson preparation method based on AI intelligence to overcome the problem in the prior art that the crawling path planning of lesson preparation resources using web crawlers is not accurate enough, resulting in low efficiency and low accuracy in obtaining lesson preparation resources.

[0006] To achieve the above objectives, the present invention provides an AI-based auxiliary lesson preparation method, comprising:

[0007] Step S1, obtaining lesson preparation information and website information, extracting lesson preparation keywords based on the lesson preparation information, and matching the lesson preparation keywords with the websites to determine the websites to be searched;

[0008] Step S2, calculating the anti-crawling friendliness based on the response delay tolerance, proxy IP compatibility, and verification complexity of the website to be retrieved, and planning several website crawling paths of the web crawler based on the anti-crawling friendliness and the resource update frequency of the website;

[0009] Step S3: determining a primary crawling path and several backup crawling paths based on the average resource quality density and path independence of the several website crawling paths;

[0010] Step S4: Monitor in real time the time interval between the web crawler sending a request to the primary crawling path and receiving data, as well as the percentage of requests in which the crawler successfully obtains valid data per unit time. Determine whether to switch from the primary crawling path to the backup crawling path based on whether the average response time of the primary crawling path fluctuates or whether the real-time success rate is abnormal.

[0011] Step S5: determining the switched backup crawling path based on the coverage of the restricted nodes of the primary crawling path and the unblocked nodes of the backup crawling path and the resource association degree between the primary crawling path and the backup crawling path.

[0012] Furthermore, calculating the anti-crawling friendliness based on the response delay tolerance, proxy IP compatibility, and verification complexity of the website to be retrieved includes weighted summing the response delay tolerance, proxy IP compatibility, and verification complexity to obtain the anti-crawling friendliness.

[0013] Furthermore, planning several website crawling paths of the web crawler based on the anti-crawling friendliness and the resource update frequency of the website includes:

[0014] The website score is calculated by weighting the anti-crawling friendliness and resource update frequency;

[0015] Sort the searched websites based on their scores, and select the searched websites with scores higher than a preset threshold as target websites;

[0016] Performing cluster analysis on the target website to form a number of path groups;

[0017] The websites in each path group are sorted in descending order of scores to form several initial crawling paths.

[0018] Furthermore, determining the primary crawling path and the plurality of backup crawling paths includes:

[0019] Get the average resource quality density and path independence of several website crawling paths;

[0020] The website crawling path with an average resource quality density greater than a preset average resource quality density and a path independence degree greater than a preset path independence degree is used as a backup crawling path;

[0021] The backup crawling path with the maximum value obtained by weighted average sum of the average resource quality density and the path independence degree in the backup crawling path is used as the main crawling path.

[0022] Furthermore, the average resource quality density of several website crawling paths is determined based on the unit path resource quantity and the similarity between the unit path resources and the lesson preparation keywords, and the path independence degree is determined based on the path node overlap rate and the resource source difference.

[0023] Furthermore, determining to switch from the primary crawling path to the backup crawling path includes:

[0024] If the average response time of the main crawling path fluctuates or the real-time success rate is abnormal, it is determined to switch from the main crawling path to the backup crawling path.

[0025] Furthermore, fluctuations in the average response time of the main crawling path include the average response time of the main crawling path being greater than the preset response time, and abnormalities in the real-time success rate of the main crawling path include the real-time success rate being less than the preset success rate for three consecutive detections.

[0026] Furthermore, determining the backup crawling path after switching includes:

[0027] Calculate the coverage of the restricted nodes of the main crawling path and the unblocked nodes of the backup crawling path, as well as the resource correlation degree between the main crawling path and the backup crawling path;

[0028] The weighted sum of coverage and resource association is used to obtain the priority score of the alternative crawling path;

[0029] The backup crawling path with the highest backup crawling path priority score is used as the backup crawling path after switching.

[0030] Furthermore, in step S5, calculating the coverage of the restricted nodes of the primary crawling path and the unblocked nodes of the backup crawling path includes:

[0031] Identify the main path restriction node;

[0032] Identify available nodes on the backup path;

[0033] The ratio of the coverable nodes of the backup crawling path in the main crawling path restriction nodes to the main crawling path restriction nodes is used as the coverage ratio of the main crawling path restriction nodes to the unblocked nodes of the backup crawling path.

[0034] An auxiliary lesson preparation system applied to the AI-based auxiliary lesson preparation method, comprising:

[0035] A data acquisition module is used to obtain lesson preparation information and website information, extract lesson preparation keywords based on the lesson preparation information, and match the lesson preparation keywords with websites to determine the websites to be searched;

[0036] A website crawling path planning module, connected to the data acquisition module, is used to calculate the anti-crawling friendliness based on the response delay tolerance, proxy IP compatibility, and verification complexity of the website to be retrieved, and plan several website crawling paths for the web crawler based on the anti-crawling friendliness and the resource update frequency of the website;

[0037] a crawling path division module, connected to the website crawling path planning module, for determining a main crawling path and a plurality of backup crawling paths based on the average resource quality density and path independence of a plurality of website crawling paths;

[0038] A path switching module, connected to the crawling path division module, is used to monitor in real time the time interval between the web crawler sending a request to the primary crawling path and receiving data, as well as the proportion of requests in which the crawler successfully obtains valid data per unit time, and to determine whether to switch from the primary crawling path to the backup crawling path based on whether the average response time of the primary crawling path fluctuates or whether the real-time success rate is abnormal;

[0039] The backup crawling path determination module is connected to the path switching module and is used to determine the switched backup crawling path based on the coverage of the main crawling path restriction nodes and the unblocked nodes of the backup crawling path and the resource association degree between the main crawling path and the backup crawling path.

[0040] Compared with the existing technology, the beneficial effect of the present invention lies in that the present invention converts the abstract anti-crawling capability into a computable quantitative score through the three core indicators of response delay tolerance (T), proxy IP compatibility (P), and verification complexity (V), avoiding the ambiguity of traditional empirical judgment and directly reflecting the verification cost through numerical differences. The system can give priority to websites with low verification complexity, reducing delays or failures caused by crawlers cracking verification codes. The above method improves the accuracy of crawling path planning for lesson preparation resources using web crawlers, thereby improving the efficiency and accuracy of obtaining lesson preparation resources.

[0041] Furthermore, the present invention achieves a scientific balance between risk control and resource quality through the weighted summation of anti-crawling friendliness and resource update frequency, and divides websites into path groups (such as "courseware resource group" and "test question resource group") through cluster analysis to avoid repeated crawling of websites of the same type and reduce resource waste. Websites in the group are arranged in descending order of score, and high-scoring websites are visited first to ensure rapid acquisition of core resources. When preparing lessons urgently, the system gives priority to crawling websites that are anti-crawling friendly and frequently updated, thereby enhancing the system's adaptability to dynamic environments. The above method improves the accuracy of crawling path planning for lesson preparation resources using web crawlers, thereby improving the efficiency and accuracy of obtaining lesson preparation resources.

[0042] Furthermore, the present invention avoids path failure caused by a single indicator defect through dual screening of average resource quality density (measurement of content relevance and richness) and path independence (measurement of anti-blocking ability). For example, if a path has high resource quality but a node overlap rate of 80% (such as relying on the same website cluster), its path independence is low and it will be automatically excluded from the backup path to prevent no effective alternative when the main path is blocked. The path independence is quantified by the node overlap rate and the resource source difference, ensuring that the backup path and the main path are as independent as possible in terms of infrastructure (website nodes) and content sources (operating entities), reducing the risk of "one loss for all". The backup path must meet the resource quality and independence thresholds at the same time to ensure that it can immediately take over the crawling task when the main path fails, avoiding delays caused by temporary screening paths. By weighted summation of average resource quality density and path independence, the optimal balance point between resource quality and crawling stability is found to avoid the overall efficiency reduction due to excessive pursuit of a certain indicator. The above method improves the accuracy of crawling path planning for lesson preparation resources using web crawlers, thereby improving the efficiency and accuracy of lesson preparation resource acquisition.

[0043] Furthermore, the present invention avoids the path failure problem caused by a single indicator through dual screening of average resource quality density (measurement of content relevance and richness) and path independence (measurement of anti-blocking ability). The path independence is quantified by "path node overlap rate" and "resource source difference", requiring the backup path and the main path to be as independent as possible in terms of infrastructure (website node) and content source (operating entity) to reduce the risk of "one loss, all losses"; through dual threshold screening (such as the case requiring average resource quality density>800 and path independence>1.2), low-quality or high-dependence paths (such as path 2 in the case due to resource quality) are quickly excluded. Those that do not meet the standards are excluded), and focus on high-quality backup paths; find the optimal balance between "content accuracy" and "crawling reliability" by weighted summation of the average resource quality density and the degree of path independence (such as the weight of 0.7:0.3 in the case), and avoid excessive pursuit of a single indicator (such as only focusing on resource quality, which makes the path easy to be blocked, or only focusing on independence, which leads to insufficient content relevance); the main path gives priority to the path with the highest comprehensive score (such as path 1 in the case was selected as the first choice due to its higher weighted score), and the backup path serves as a reliable backup to ensure that when the main path is blocked or fails, it can quickly switch to the independent path to continue crawling, thereby reducing task interruption time.

[0044] Furthermore, the present invention filters out occasional failures (such as instantaneous network jitter) based on the statistical characteristics of historical success rates, and only performs switching in the event of persistent anomalies to ensure the stability of decision-making. At the same time, it monitors the average response time and real-time success rate, covering two core abnormal scenarios: network delay (such as server congestion) and content acquisition failure (such as anti-crawling mechanism triggering). Even if the response time is normal, but multiple consecutive requests fail (such as returning a 403 status code), path switching will still be triggered to ensure that the crawler task is not affected by a single point of failure. When an abnormality is detected in the main path, the system automatically switches to a backup path (such as the high-independence path screened in step S3) to avoid crawling interruptions caused by manual intervention or temporary path selection. Sensitive monitoring of response time fluctuations (such as a threshold of mean + 2 times the standard deviation) can provide early warning of potential risks (such as a precursor to server current limiting). By switching to the backup path in time, the main path can be avoided from being permanently blocked due to excessive requests, thereby extending its available period.

[0045] Furthermore, the present invention accurately evaluates the ability of the backup path to replace the main path by calculating the ratio of the restricted nodes of the main path (such as the blocked website A) to the available nodes of the backup path (such as websites C and F). Switching is only performed when the unblocked nodes of the backup path can effectively cover the restricted nodes of the main path, preventing crawling failures caused by node mismatch. The priority scoring mechanism ensures that the system selects the path with the best comprehensive performance among multiple backup paths (such as the score of path 3 in the case is higher than other backup paths), avoiding suboptimal decisions caused by a single indicator (such as only considering coverage). When the main path fails due to the anti-crawling mechanism (such as IP blocking) or server failure, the system can quickly find a backup path that has both node replacement capability (coverage) and content relevance (resource correlation). By giving priority to backup paths with high resource correlation with the main path, content mismatch problems caused by path switching (such as switching from courseware to irrelevant exercises) are reduced, thereby improving the overall availability of crawled resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1 This is a workflow diagram of the AI-based auxiliary lesson preparation method according to an embodiment of the present invention;

[0047] Figure 2 This is a workflow diagram for planning several website crawling paths of a web crawler in the AI-based auxiliary lesson preparation method according to an embodiment of the present invention;

[0048] Figure 3 This is a workflow diagram for determining a main crawling path and several backup crawling paths in an AI-based auxiliary lesson preparation method according to an embodiment of the present invention;

[0049] Figure 4 The figure is a structural diagram of an auxiliary lesson preparation system applied to the AI ​​intelligent auxiliary lesson preparation method according to an embodiment of the present invention. DETAILED DESCRIPTION

[0050] In order to make the objects and advantages of the present invention more clearly understood, the present invention is further described below in conjunction with embodiments; it should be understood that the specific embodiments described herein are merely used to explain the present invention and are not intended to limit the present invention.

[0051] The preferred embodiments of the present invention are described below with reference to the accompanying drawings. It should be understood by those skilled in the art that these embodiments are only used to explain the technical principles of the present invention and are not intended to limit the scope of protection of the present invention.

[0052] Furthermore, it should be noted that, in the description of the present invention, unless otherwise expressly specified or limited, the terms "mounted," "connected," and "connected" should be understood in a broad sense. For example, they may refer to fixed connections, detachable connections, or integral connections; mechanical connections or electrical connections; direct connections or indirect connections through an intermediate medium; and internal communication between two components. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on specific circumstances.

[0053] See also Figure 1-Figure 3 As shown, Figure 1 This is a workflow diagram of the AI-based auxiliary lesson preparation method according to an embodiment of the present invention; Figure 2 This is a workflow diagram for planning several website crawling paths of a web crawler in the AI-based auxiliary lesson preparation method according to an embodiment of the present invention; Figure 3 This is a workflow diagram for determining a main crawling path and several backup crawling paths in an AI-based auxiliary lesson preparation method according to an embodiment of the present invention.

[0054] The AI-based auxiliary lesson preparation method according to an embodiment of the present invention includes:

[0055] Step S1, obtaining lesson preparation information and website information, extracting lesson preparation keywords based on the lesson preparation information, and matching the lesson preparation keywords with the websites to determine the websites to be searched;

[0056] Step S2, calculating the anti-crawling friendliness based on the response delay tolerance, proxy IP compatibility, and verification complexity of the website to be retrieved, and planning several website crawling paths of the web crawler based on the anti-crawling friendliness and the resource update frequency of the website;

[0057] Step S3: determining a primary crawling path and several backup crawling paths based on the average resource quality density and path independence of the several website crawling paths;

[0058] Step S4: Monitor in real time the time interval between the web crawler sending a request to the primary crawling path and receiving data, as well as the percentage of requests in which the crawler successfully obtains valid data per unit time. Determine whether to switch from the primary crawling path to the backup crawling path based on whether the average response time of the primary crawling path fluctuates or whether the real-time success rate is abnormal.

[0059] Step S5: determining the switched backup crawling path based on the coverage of the restricted nodes of the primary crawling path and the unblocked nodes of the backup crawling path and the resource association degree between the primary crawling path and the backup crawling path.

[0060] The lesson preparation information in the embodiment of the present invention includes but is not limited to "teaching objectives, teaching content and resource type requirements", and the website information includes but is not limited to "basic website attributes, resource classification labels, anti-crawl mechanism parameters and resource update records", and the cluster analysis can be performed according to subject or resource type.

[0061] In the embodiment of the present invention, NLP technology is used to perform word segmentation, part-of-speech tagging and entity recognition on the lesson preparation information, and core keywords are extracted in combination with the education field dictionary. Examples of website information are collected and stored, and keywords are matched with websites, including subject knowledge point matching, resource type matching and educational attribute matching, so as to obtain the website to be searched. The weights appearing in the present invention can be experimented by selecting multiple sets of data weights, and then the appropriate weight size is selected according to the experimental results. For example, different weight schemes are applied to parallel crawler instances to compare the following indicators: the amount of effective resource acquisition per unit time, the anti-crawling blocking rate (such as the proportion of 403 responses) and the verification failure rate. For example, weight scheme A (T = 40%, P = 30%, V = 30%) and scheme B (T = 30%, P = 40%, V = 30%). If the blocking rate of scheme B is reduced by 15% and the resource acquisition amount is equivalent, scheme B is adopted for a long time. However, the above values ​​are not limited to this, and those skilled in the art can also adjust the values ​​according to actual needs.

[0062] Specifically, in step S2, the anti-crawling friendliness is calculated based on the response delay tolerance, proxy IP compatibility and verification complexity of the website to be retrieved, including weighted summation of the response delay tolerance, proxy IP compatibility and verification complexity to obtain the anti-crawling friendliness.

[0063] In the embodiment of the present invention, it is assumed that when a teacher of a certain subject needs to prepare lessons, three websites to be searched are determined by the system: website A (education resource platform), website B (subject question bank website), and website C (teaching video library). The system needs to calculate the anti-crawling friendliness of each website to plan the crawler path; response delay tolerance (T), definition: the acceptable range of response delay of the website to the crawler request (unit: milliseconds), the larger the value, the more "friendly" to the delay; collection method: send 100 requests by simulating the crawler, and calculate the average response time (T_avg). If T_avg≤500ms, the tolerance is 100 points; deduct 10 points for every 50ms exceeding, and add 10 points for every 200ms below; example data: website A: T_avg=300ms T=90 points; website B: T_avg=600ms T=80 points; website C: T_avg=250ms T = 100 points; Proxy IP compatibility (P), definition: the degree of support for proxy IP by the website. The larger the value, the stronger the ability to be compatible with proxy IP (reducing the risk of IP being blocked); Collection method: Use 3 different proxy IPs to test access and count the percentage of successful times (compatibility rate); Sample data: Website A: compatibility rate 80% P = 80 points; Website B: compatibility rate 50% P = 50 points; Website C: compatibility rate 100% P = 100 points; Verification complexity (V), definition: the complexity of the website's anti-crawling verification mechanism (such as verification code, login verification, etc.). The larger the value, the more complex the verification (the stricter the anti-crawling); Scoring rules: No verification: 100 points, simple verification code (such as number / : 60 points; Behavioral verification code (such as slider, image and text selection): 30 points; Login / user permission verification required: 0 points; Example data: Website A: Simple verification code V = 60 points; Website B: Behavioral verification code V = 30 points; Website C: No verification V = 100 points; Anti-crawling friendliness calculation (weighted sum), assuming the system preset weights are: Response delay tolerance (T): 40%, Proxy IP compatibility (P): 30%, Verification complexity (V): 30% () Note: Verification complexity is a "negative indicator" and needs to be inverted before calculation, that is, the actual calculation value is (100-V). Calculation formula: Anti-crawling friendliness = T × 0.4 + P × 0.3 + (100-V) × 0.3. The calculation results show that Website A scored 72, Website B scored 68, and Website C scored 70. The weights corresponding to response delay tolerance, proxy IP compatibility, and verification complexity can be determined through weight verification and iterative methods. Different weighting schemes are applied to parallel crawler instances, and the following indicators are compared: effective resource acquisition per unit time, anti-crawling blocking rate (such as the proportion of 403 responses), and verification failure rate. For example, if weighting scheme A (T = 40%, P = 30%, V = 30%) and scheme B (T = 30%, P = 40%, V = 30%) are compared, if the blocking rate of scheme B is reduced by 15% and the resource acquisition volume is comparable, then scheme B is adopted long-term. However, the above values ​​are not limited to this, and those skilled in the art can also adjust the values ​​according to actual needs.

[0064] The present invention converts the abstract anti-crawling capability into a computable quantitative score through the three core indicators of response delay tolerance (T), proxy IP compatibility (P), and verification complexity (V), avoiding the ambiguity of traditional empirical judgment and directly reflecting the verification cost through numerical differences. The system can give priority to websites with low verification complexity, reducing delays or failures caused by crawlers cracking verification codes. The above method improves the accuracy of crawling path planning for lesson preparation resources using web crawlers, thereby improving the efficiency and accuracy of obtaining lesson preparation resources.

[0065] Specifically, in step S2, the steps of planning several website crawling paths of the web crawler based on the anti-crawling friendliness and the resource update frequency of the website include:

[0066] Step S2201: Calculate the website score by weighted summing the anti-crawling friendliness and resource update frequency;

[0067] Step S2202: sorting the websites to be searched based on their website scores, and selecting the websites to be searched with scores higher than a preset threshold as target websites;

[0068] Step S2203: performing cluster analysis on the target website to form a plurality of path groups;

[0069] Step S2204: Arrange the websites in each path group in descending order of scores to form several initial crawling paths.

[0070] The preset threshold in the embodiment of the present invention can be determined by the following method: statistically analyzing the scores of all websites to be retrieved, calculating the average, median and standard deviation of the scores, understanding the overall distribution of the scores, and determining the threshold range based on the statistical results. For example, if the average score is 70 points and the standard deviation is 10 points, the threshold can be set at one standard deviation below the average (i.e., 60 points) to filter out websites with lower scores. Planning the crawling path of the web crawler includes first calculating the anti-crawling friendliness (F). Website A (education resources) Website A (source platform): 72 points; Website B (subject question bank): 68 points; Website C (teaching video library): 70 points; Resource update frequency (R), definition: the frequency of website content updates, quantified as a percentage based on the "update cycle": real-time updates: 100 points; daily updates: 90 points; weekly updates: 70 points; monthly updates: 50 points; Example data: Website A: weekly updates R = 70 points; Website B: monthly updates R = 50 points; Website C: daily updates R = 90 points; Calculate website score, formula: Website score = α × F +(1-α)×R; Weight setting: Assuming the system presets α=0.6 (anti-crawling friendliness takes precedence over update frequency), the calculation results are: Website A scores 71.2 points, Website B scores 60.8 points, and Website C scores 78 points; 3. Step S2202: Screen target websites, preset threshold: 65 points (filter low-scoring websites to avoid low crawling efficiency or high anti-crawling risk), screening results: Website A (71.2 points) and Website C (78 points) pass the threshold and are included in the target websites; Website B (60.8 points) is excluded and only serves as a backup candidate ; Cluster analysis forms path groups, clustering dimensions: subject area: all are "high school physics", no need to split by subject; resource type: website A: courseware, lesson plan (theoretical teaching resources); website C: experimental video, virtual experiment (practical teaching resources); clustering results: path group 1: theoretical teaching resource group (website A); path group 2: practical teaching resource group (website C); the websites in each path group are arranged in descending order according to the score, forming several initial crawling paths, and arranged in descending order according to the website score (if there are multiple websites in the same group, the one with the higher score will be given priority).

[0071] The present invention achieves a scientific balance between risk control and resource quality through the weighted summation of anti-crawling friendliness and resource update frequency, and divides websites into path groups (such as "courseware resource group" and "test question resource group") through cluster analysis to avoid repeated crawling of websites of the same type and reduce resource waste. Websites in the group are arranged in descending order of score, and high-scoring websites are visited first to ensure rapid acquisition of core resources. When preparing lessons urgently, the system gives priority to crawling websites that are anti-crawling friendly and frequently updated, thereby enhancing the system's adaptability to dynamic environments. The above method improves the accuracy of crawling path planning for lesson preparation resources using web crawlers, thereby improving the efficiency and accuracy of obtaining lesson preparation resources.

[0072] Specifically, in step S3, the steps of determining the main crawling path and the plurality of backup crawling paths include:

[0073] Step S3301: Obtain the average resource quality density and path independence of several website crawling paths;

[0074] Step S3302: Using a website crawling path with an average resource quality density greater than a preset average resource quality density and a path independence degree greater than a preset path independence degree as a backup crawling path;

[0075] Step S3303: The backup crawling path with the maximum value obtained by weighted average sum of average resource quality density and path independence degree in the backup crawling paths is used as the main crawling path.

[0076] In the embodiment of the present invention, the average resource quality density of the several website crawling paths is determined according to the unit path resource amount and the similarity between the unit path resource and the lesson preparation keyword. The path independence degree is determined according to the path node overlap rate and the resource source difference. The preset average resource quality density can be calculated by analyzing the resource data crawled in the past, and the quartiles of the average resource quality density are calculated. The upper quartile (Q3) or the median is used as the threshold. For example, if the average resource quality density distribution in the historical data is [500, 650, 750, 850, 1000], the median is 750, and the upper quartile is 850, then the basic teaching scenario takes the median 750 and the high-order scenario takes the upper quartile 850; the preset path independence degree is the average value of the path independence degrees of several groups of main crawling paths and backup crawling paths, but the above value is not limited to this. Those skilled in the art can also adjust the value according to actual needs.

[0077] The present invention avoids path failure caused by a single indicator defect through dual screening of average resource quality density (measurement of content relevance and richness) and path independence (measurement of anti-blocking ability). For example, if a path has high resource quality but a node overlap rate of 80% (such as relying on the same website cluster), its path independence is low and it will be automatically excluded from the backup path to prevent no effective alternative when the main path is blocked. The path independence is quantified by the node overlap rate and the resource source difference, ensuring that the backup path and the main path are as independent as possible in terms of infrastructure (website nodes) and content sources (operating entities), reducing the risk of "one loss for all". The backup path must meet both resource quality and independence thresholds to ensure that it can immediately take over the crawling task when the main path fails, avoiding delays caused by temporary screening paths. By weighted summation of average resource quality density and path independence, the optimal balance between resource quality and crawling stability is found to avoid a decrease in overall efficiency due to excessive pursuit of a certain indicator. The above method improves the accuracy of crawling path planning for lesson preparation resources using web crawlers, thereby improving the efficiency and accuracy of lesson preparation resource acquisition.

[0078] In the embodiment of the present invention, it is assumed that three initial crawling paths are planned through step S2 (each path contains multiple websites, arranged in the order of crawling): Path 1: Website C (teaching video library, score 78) moves to Website A (education resource platform, score 71.2), resource type: experimental video (website C) moves to courseware, lesson plan (website A); Path 2: Website D (new test question bank, score 71) moves to Website E (courseware library, score 69); resource type: test question analysis (website D) moves to dynamic courseware (website E); Path 3: Website C (teaching video library, score 78) moves to Website F (virtual laboratory, score 75); resource type: experimental video (website C) moves to virtual experiment tool (website F); average resource The formula for calculating source quality density is: average resource quality density = (unit path resource quantity × resource-keyword similarity) / path length (number of websites); unit path resource quantity: the number of valid resources in the path (such as the number of downloadable courseware and videos); resource-keyword similarity: the cosine similarity between resource tags and lesson preparation keywords (such as "Newton's Second Law" and "Experimental Design") is calculated through NLP, with a value of 0-100 points; the resource-keyword similarity of path 1 is 85 points, the path length is 2, and the average resource quality density is 1062.5; the resource-keyword similarity of path 2 is 70 points, the path length is 2, and the average resource quality density is 700; the resource-keyword similarity of path 3 is 90 points, the path length is 2, and the average The resource quality density is 810, and the calculation formula for path independence is: path independence = 1-path node overlap rate + resource source difference, path node overlap rate: the proportion of websites shared with other paths (for example, if path 1 and path 3 both include website C, then the overlap rate = 1 / 2 = 50%); resource source difference: the difference in the operating entities of the websites in the path (for example, website C and website A are different organizations, with high difference; website D and website E are platforms under the same company, with low difference), quantified from 0 to 100 points; the path node overlap rate of path 1 is 50%, the resource source difference is 90 points, and the path independence is 1.4; the path node overlap rate of path 2 is 0, the resource source difference is 60 points, the path independence is 1.6, and the path 3 is 0. The path node overlap rate is 50%, the resource source diversity is 85 points, the path independence is 1.35, the preset average resource quality density threshold is 800 points, the preset path independence threshold is 1.2, and the screening condition is that both the average resource quality density > 800 and the path independence > 1.2 are met. Screening results: Path 1: 1062.5 > 800, 1.4 > 1.2, it is included in the backup path. Path 2: 700 < 800, it is excluded. Path 3: 810 > 800, 1.35 > 1.2, it is included in the backup path. The backup crawling path list includes Path 1 and Path 3. The primary crawling path is determined. Assuming the average resource quality density weight in the backup crawling path is set to 0.7, the value of Path 1 is 744.17, the value of path 3 is 567.405, then path 1 is the main crawling path, and path 3 is the backup crawling path.

[0079] The present invention avoids the path failure problem caused by a single indicator through dual screening of average resource quality density (measurement of content relevance and richness) and path independence (measurement of anti-blocking ability). The path independence is quantified by "path node overlap rate" and "resource source difference", requiring the backup path and the main path to be as independent as possible in terms of infrastructure (website node) and content source (operating entity) to reduce the risk of "one loss, all losses"; through dual threshold screening (such as the case requiring average resource quality density>800 and path independence>1.2), low-quality or high-dependence paths (such as path 2 in the case due to poor resource quality) are quickly excluded. The target is excluded), and the focus is on high-quality backup paths; by weighted summation of the average resource quality density and the degree of path independence (such as the weight of 0.7:0.3 in the case), the optimal balance between "content accuracy" and "crawling reliability" is found to avoid excessive pursuit of a single indicator (such as only focusing on resource quality, which makes the path easy to be blocked, or only focusing on independence, which leads to insufficient content relevance); the main path gives priority to the path with the highest comprehensive score (such as path 1 in the case is selected as the first choice due to its higher weighted score), and the backup path serves as a reliable backup to ensure that when the main path is blocked or fails, it can quickly switch to the independent path to continue crawling, thereby reducing task interruption time.

[0080] Specifically, in step S4, when it is determined to switch from the main crawling path to the backup crawling path, it is determined whether to switch from the main crawling path to the backup crawling path based on whether the average response time of the main crawling path fluctuates or whether the real-time success rate is abnormal;

[0081] When the average response time of the primary crawling path fluctuates or the real-time success rate becomes abnormal, it is determined to switch from the primary crawling path to the backup crawling path;

[0082] When the average response time of the primary crawling path does not fluctuate and the real-time success rate does not show abnormalities, it is determined that there is no need to switch from the primary crawling path to the backup crawling path.

[0083] In an embodiment of the present invention, fluctuations in the average response time of the main crawling path include the average response time of the main crawling path being greater than the preset response time. The preset response time can be calculated by analyzing the historical response time data of the main crawling path, and the average value and standard deviation. For example, if the historical average response time is 2 seconds and the standard deviation is 0.5 seconds, the preset response time can be set to the average value plus 1-2 standard deviations, that is, 3-4 seconds; abnormalities in the real-time success rate of the main crawling path include the real-time success rate being less than the preset success rate for three consecutive times. The preset success rate can be calculated by analyzing the historical success rate data of the main crawling path, and the average value and standard deviation. For example, if the historical average success rate is 90% and the standard deviation is 5%, the preset success rate can be set to the average value minus 1-2 standard deviations, that is, 80%-85%. However, the above values ​​are not limited to this, and those skilled in the art can also adjust the values ​​according to actual needs.

[0084] The present invention filters out occasional failures (such as instantaneous network jitter) based on the statistical characteristics of historical success rates, and only performs switching in the event of persistent anomalies to ensure the stability of decision-making. At the same time, it monitors the average response time and real-time success rate, covering two core abnormal scenarios: network delay (such as server congestion) and content acquisition failure (such as anti-crawling mechanism triggering). Even if the response time is normal, but multiple requests fail in succession (such as returning a 403 status code), path switching will still be triggered to ensure that the crawler task is not affected by a single point of failure. When an abnormality is detected in the main path, the system automatically switches to the backup path (such as the high-independence path screened in step S3) to avoid crawling interruptions caused by manual intervention or temporary path selection. Sensitive monitoring of response time fluctuations (such as a threshold of mean + 2 times the standard deviation) can provide early warning of potential risks (such as a precursor to server current limiting). By switching to the backup path in time, the main path can be avoided from being permanently blocked due to excessive requests, thereby extending its available period.

[0085] Specifically, in step S5, the step of determining the backup crawling path after switching includes:

[0086] Step S5501, calculating the coverage of the restricted nodes of the primary crawling path and the unblocked nodes of the backup crawling path, as well as the resource association degree between the primary crawling path and the backup crawling path;

[0087] Step S5502: weighted sum of coverage and resource association to obtain a priority score for the backup crawling path;

[0088] Step S5503: The backup crawling path with the highest backup crawling path priority score is used as the backup crawling path after switching.

[0089] In an embodiment of the present invention, calculating the coverage rate of the main crawling path restriction nodes and the non-blocked nodes of the backup crawling path includes first identifying the main path restriction nodes and the available nodes of the backup path, and taking the ratio of the coverable nodes of the backup crawling path in the main crawling path restriction nodes to the main crawling path restriction nodes as the coverage rate of the main crawling path restriction nodes and the non-blocked nodes of the backup crawling path; calculating the resource association degree between the main crawling path and the backup crawling path includes calculating the similarity (for example, cosine similarity) of the resource content in the main crawling path and the backup crawling path.

[0090] In the embodiment of the present invention, the main crawling path (path 1): website C (experimental video) moves to website A (courseware, lesson plan), the restricted node: website A is inaccessible due to the IP blocking of the anti-crawling mechanism; the backup crawling path (path 3): website C (experimental video) moves to website F (virtual experiment tool); unblocked nodes: website C and website F are both normally accessible; the coverage rate and resource correlation degree are calculated, assuming that the coverage rate weight γ = 0.6, and the resource association degree weight 1-γ = 0.4; formula: priority score = γ × coverage rate + (1-γ) × resource association degree; the backup crawling path with the highest priority score is used as the backup crawling path after switching.

[0091] The present invention accurately evaluates the ability of the backup path to replace the main path by calculating the ratio of the restricted nodes of the main path (such as the blocked website A) to the available nodes of the backup path (such as websites C and F). Switching is only performed when the unblocked nodes of the backup path can effectively cover the restricted nodes of the main path, preventing crawling failures due to node mismatch. The priority scoring mechanism ensures that the system selects the path with the best comprehensive performance among multiple backup paths (such as the score of path 3 in the case is higher than other backup paths), avoiding suboptimal decisions caused by a single indicator (such as only considering coverage). When the main path fails due to the anti-crawling mechanism (such as IP blocking) or server failure, the system can quickly find a backup path that has both node replacement capability (coverage) and content relevance (resource relevance). By giving priority to backup paths with high resource relevance to the main path, content mismatch problems caused by path switching (such as switching from courseware to irrelevant exercises) are reduced, thereby improving the overall availability of crawled resources.

[0092] See also Figure 4 As shown, Figure 4 The figure is a structural diagram of an auxiliary lesson preparation system applied to the AI ​​intelligent auxiliary lesson preparation method according to an embodiment of the present invention.

[0093] Specifically, an auxiliary lesson preparation system applied to the AI-based auxiliary lesson preparation method includes:

[0094] A data acquisition module is used to obtain lesson preparation information and website information, extract lesson preparation keywords based on the lesson preparation information, and match the lesson preparation keywords with websites to determine the websites to be searched;

[0095] A website crawling path planning module, connected to the data acquisition module, is used to calculate the anti-crawling friendliness based on the response delay tolerance, proxy IP compatibility, and verification complexity of the website to be retrieved, and plan several website crawling paths for the web crawler based on the anti-crawling friendliness and the resource update frequency of the website;

[0096] a crawling path division module, connected to the website crawling path planning module, for determining a main crawling path and a plurality of backup crawling paths based on the average resource quality density and path independence of a plurality of website crawling paths;

[0097] A path switching module, connected to the crawling path division module, is used to monitor in real time the time interval between the web crawler sending a request to the primary crawling path and receiving data, as well as the proportion of requests in which the crawler successfully obtains valid data per unit time, and to determine whether to switch from the primary crawling path to the backup crawling path based on whether the average response time of the primary crawling path fluctuates or whether the real-time success rate is abnormal;

[0098] The backup crawling path determination module is connected to the path switching module and is used to determine the switched backup crawling path based on the coverage of the main crawling path restriction nodes and the unblocked nodes of the backup crawling path and the resource association degree between the main crawling path and the backup crawling path.

[0099] Thus far, the technical solutions of the present invention have been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art may make equivalent changes or substitutions to the relevant technical features, and the technical solutions after such changes or substitutions will fall within the scope of protection of the present invention.

Claims

1. An AI-based auxiliary lesson preparation method, characterized in that: include: Step S1, obtaining lesson preparation information and website information, extracting lesson preparation keywords based on the lesson preparation information, and matching the lesson preparation keywords with the websites to determine the websites to be searched; Step S2, calculating the anti-crawling friendliness based on the response delay tolerance, proxy IP compatibility, and verification complexity of the website to be retrieved, and planning several website crawling paths of the web crawler based on the anti-crawling friendliness and the resource update frequency of the website; Step S3: determining a primary crawling path and several backup crawling paths based on the average resource quality density and path independence of the several website crawling paths; Step S4: Monitor in real time the time interval between the web crawler sending a request to the primary crawling path and receiving data, as well as the percentage of requests in which the crawler successfully obtains valid data per unit time. Determine whether to switch from the primary crawling path to the backup crawling path based on whether the average response time of the primary crawling path fluctuates or whether the real-time success rate is abnormal. Step S5: determining the switched backup crawling path based on the coverage of the restricted nodes of the primary crawling path and the unblocked nodes of the backup crawling path and the resource association degree between the primary crawling path and the backup crawling path.

2. The AI-based auxiliary lesson preparation method according to claim 1 is characterized in that: In step S2, the anti-crawling friendliness is calculated based on the response delay tolerance, proxy IP compatibility and verification complexity of the website to be retrieved, including weighted summation of the response delay tolerance, proxy IP compatibility and verification complexity to obtain the anti-crawling friendliness.

3. The AI-based auxiliary lesson preparation method according to claim 2 is characterized in that: In step S2, planning a plurality of website crawling paths of a web crawler based on the anti-crawling friendliness and the resource update frequency of the website includes: The website score is calculated by weighting the anti-crawling friendliness and resource update frequency; Sort the searched websites based on their scores, and select the searched websites with scores higher than a preset threshold as target websites; Performing cluster analysis on the target website to form a number of path groups; The websites in each path group are sorted in descending order of scores to form several initial crawling paths.

4. The AI-based auxiliary lesson preparation method according to claim 3 is characterized in that: In step S3, determining the main crawling path and a plurality of backup crawling paths includes: Get the average resource quality density and path independence of several website crawling paths; The website crawling path with an average resource quality density greater than a preset average resource quality density and a path independence degree greater than a preset path independence degree is used as a backup crawling path; The backup crawling path with the maximum value obtained by weighted average sum of the average resource quality density and the path independence degree in the backup crawling path is used as the main crawling path.

5. The AI-based auxiliary lesson preparation method according to claim 4 is characterized in that: In step S3, the average resource quality density of several website crawling paths is determined based on the unit path resource quantity and the similarity between the unit path resource and the lesson preparation keywords, and the path independence degree is determined based on the path node overlap rate and the resource source difference.

6. The AI-based auxiliary lesson preparation method according to claim 5, characterized in that: In step S4, determining to switch from the primary crawling path to the backup crawling path includes: If the average response time of the main crawling path fluctuates or the real-time success rate is abnormal, it is determined to switch from the main crawling path to the backup crawling path.

7. The AI-based auxiliary lesson preparation method according to claim 6 is characterized in that: In step S4, the average response time of the main crawling path fluctuates, including the average response time of the main crawling path being greater than the preset response time, and the real-time success rate of the main crawling path is abnormal, including the real-time success rate being less than the preset success rate for three consecutive detections.

8. The AI-based auxiliary lesson preparation method according to claim 7, characterized in that: In step S5, determining the backup crawling path after switching includes: Calculate the coverage of the restricted nodes of the main crawling path and the unblocked nodes of the backup crawling path, as well as the resource correlation degree between the main crawling path and the backup crawling path; The weighted sum of coverage and resource association is used to obtain the priority score of the alternative crawling path; The backup crawling path with the highest backup crawling path priority score is used as the backup crawling path after switching.

9. The AI-based auxiliary lesson preparation method according to claim 8, characterized in that: In step S5, calculating the coverage of the restricted nodes of the primary crawling path and the unblocked nodes of the backup crawling path includes: Identify the main path restriction node; Identify available nodes on the backup path; The ratio of the coverable nodes of the backup crawling path in the main crawling path restriction nodes to the main crawling path restriction nodes is used as the coverage ratio of the main crawling path restriction nodes to the unblocked nodes of the backup crawling path.

10. An auxiliary lesson preparation system applied to the AI-based auxiliary lesson preparation method according to any one of claims 1 to 9, characterized in that: include: A data acquisition module is used to obtain lesson preparation information and website information, extract lesson preparation keywords based on the lesson preparation information, and match the lesson preparation keywords with websites to determine the websites to be searched; A website crawling path planning module, connected to the data acquisition module, is used to calculate the anti-crawling friendliness based on the response delay tolerance, proxy IP compatibility, and verification complexity of the website to be retrieved, and plan several website crawling paths for the web crawler based on the anti-crawling friendliness and the resource update frequency of the website; a crawling path division module, connected to the website crawling path planning module, for determining a main crawling path and a plurality of backup crawling paths based on the average resource quality density and path independence of a plurality of website crawling paths; A path switching module, connected to the crawling path division module, is used to monitor in real time the time interval between the web crawler sending a request to the primary crawling path and receiving data, as well as the proportion of requests in which the crawler successfully obtains valid data per unit time, and to determine whether to switch from the primary crawling path to the backup crawling path based on whether the average response time of the primary crawling path fluctuates or whether the real-time success rate is abnormal; The backup crawling path determination module is connected to the path switching module and is used to determine the switched backup crawling path based on the coverage of the main crawling path restriction nodes and the unblocked nodes of the backup crawling path and the resource association degree between the main crawling path and the backup crawling path.

Citation Information

Patent Citations

  • AI-Assisted Method and System for Pushing Lesson Preparation Materials

    CN117056612B