Webpage electronic file processing method based on artificial intelligence algorithm

Through the hybrid traversal model and structural entropy perception adjustment based on artificial intelligence algorithms, the problems of low collection efficiency and resource redundancy in complex website structures are solved, and efficient and accurate web electronic archive management is achieved.

CN120234462AActive Publication Date: 2025-07-01GANSU JIYOUPIN NETWORK TECH CO LTD

Patent Information

Application Number
CN202510696444.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-07-01
Estimated Expiration
2045-05-28

AI Technical Summary

Technical Problem

When facing complex target website structures, the existing technology has low collection efficiency, incomplete results and redundant resources, and lacks joint modeling of the complexity and semantic correlation of web page structure, resulting in limited intelligence level and operational efficiency of web page electronic archive management system.

Method used

Using a hybrid traversal model based on artificial intelligence algorithm, the initial feature vector is constructed through neural network training of structural entropy perception adjustment and policy deviation feedback, and combined with structural diagrams and aggregation intensity evaluation, automatic judgment and priority classification of web page tags and aggregation intensity are realized.

Benefits of technology

It improves the crawling coverage and content relevance of web electronic archives, avoids crawling redundancy and missing key resources, improves the accuracy and robustness of electronic archive collection, and optimizes the scheduling and response efficiency of high-value archive resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120234462A_ABST
    Figure CN120234462A_ABST
Patent Text Reader

Abstract

The invention discloses a webpage electronic archive processing method based on an artificial intelligence algorithm, and relates to the technical field of information collection, and the method comprises the following steps: obtaining a URL set Uof a target website, the URL set Ucomprises a plurality of seed URLs, and extracting metadata corresponding to each seed URL through a page analysis engine to generate a metadata set; according to the metadata set, constructing an initial feature vector; obtaining a metadata set of a new target website, generating a strategy probability for each hyperlink in the new target website by using the trained hybrid traversal model, and determining a traversal strategy based on the strategy probability; according to path information recorded in the traversal strategy execution process, webpage labels and aggregation strength are obtained, archiving priority classification is carried out based on the webpage labels and the aggregation strength, and fine recognition and dynamic adaptation of webpage structure complexity and hyperlink semantic value are achieved; the traversal path can be flexibly adjusted in webpage structures with different levels and different density distributions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of information collection, and particularly relates to a method for processing web page electronic files based on an artificial intelligence algorithm. Background Art

[0002] With the continuous deepening of informatization construction, the electronic file resources (such as announcement documents, statistical annual reports, policies and regulations, etc.) published on web pages by government affairs websites, enterprise and institution portals, and industry platforms are increasing day by day. Web page electronic files have become important information sources and data assets. In order to realize the automatic acquisition and archiving of these file resources scattered in different sites and pages with different structures, it is necessary to build an efficient web page collection and processing mechanism to ensure the comprehensiveness, timeliness, and structured quality of information collection.

[0003] Traditional web page collection methods mostly implement traversal control based on the breadth-first search (BFS) or depth-first search (DFS) strategy. The former is suitable for horizontally grabbing resources on the same layer of pages, and the latter is suitable for mining deep resources under directory nesting. In scenarios where the specific website structure is single or the link distribution is highly regular, such traversal strategies already have relatively high collection efficiency. However, in the face of the actual complex target website structure - such as features like multiple entrances, multiple levels, heterogeneous modules intertwined and nested, different sub-sites deployed under different domain names or paths, etc., a single strategy often has limited traversal scope, gets stuck in structural islands, or fails to identify high-value links, resulting in incomplete collection results or excessive redundant resources, seriously restricting the archiving quality and collection efficiency of web page electronic files.

[0004] Some current studies have introduced crawler scheduling models based on link analysis or page scoring, and tried to adjust the traversal path through link priorities or page scores to improve resource discovery capabilities. However, such methods usually lack the ability to jointly model the complexity of the web page DOM structure, document type characteristics, and page semantic association degree, and it is difficult to achieve a dynamic balance between structural adaptability and collection value orientation. At the same time, they lack a strategy feedback optimization mechanism, and thus cannot self-adjust strategy preferences when facing different types of target websites, resulting in insufficient generalization and robustness of the model.

[0005] In addition, the current web page collection and archiving processing process generally lacks the perception and analysis of "aggregation structure" and "resource value", and cannot effectively distinguish directory aggregation pages from actual file resource pages, nor can it prioritize the archiving or scheduling of resources according to the importance of the content. This leads to problems such as redundant occupation of storage space, delayed processing of high-value resources, or untimely scheduling response, restricting the intelligent level and operation efficiency of the web page electronic file management system. Summary of the Invention

[0006] To solve the above technical problems, the present invention provides a method for processing web page electronic files based on artificial intelligence algorithms, including: S1. Target website parsing: Obtain the URL set U0 of the target website. The URL set U0 contains multiple seed URLs. Extract the metadata corresponding to each seed URL through a page parsing engine to generate a metadata set. S2. Initial feature vector acquisition: Specifically, construct an initial feature vector according to the metadata set. S3. Hybrid traversal model training: Specifically, use a neural network based on a structural entropy perception adjustment strategy and policy deviation feedback as the hybrid traversal model. S4. Traversal strategy determination: Specifically, obtain the metadata set of the new target website, use the trained hybrid traversal model to generate policy probabilities for each hyperlink in the new target website, and determine the traversal strategy based on the policy probabilities. S5. Archiving priority classification: Specifically, according to the path information recorded during the execution of the traversal strategy, obtain the web page tags and aggregation strength, and perform archiving priority classification based on the web page tags and aggregation strength.

[0007] Further, the metadata includes the page main title text, the page body content, and the page hyperlink attribute set.

[0008] Further, the training process of the neural network based on the structural entropy perception adjustment strategy and policy deviation feedback is as follows: S31. Construct a three-layer feedforward neural network. The three-layer feedforward neural network includes an input layer, a hidden layer, and an output layer. Initialize the weight matrix and bias vector between each layer, and set the activation function and the initial learning rate. S32. Scale the initial feature vector through the structural entropy perception adjustment strategy to obtain an adjusted input vector. S33. Input the adjusted input vector into the feedforward neural network for forward propagation, and output the output result of the hidden layer and the policy probability ; S34. Calculate and obtain the loss function based on the preset policy label and policy probability. S35. Calculate and obtain the policy deviation according to the number of successful archivings corresponding to each round of policy probability and the preset expected number of archivings. S36. Update the neural network parameters based on the loss function and the policy deviation. The neural network parameters include the weight matrix, the bias vector, and the learning rate. S37. When the decrease amplitude of the loss function is less than for three consecutive rounds, deploy the corresponding neural network as the hybrid traversal model.

[0009] Further, obtain the adjustment input vector, expressed as: ; In the formula, is the adjustment input vector, is the structural entropy adjustment factor, is the initial feature vector, is the structural entropy.

[0010] Further, the policy label includes the BFS policy and the DFS policy. When the policy label is the BFS policy, the output value is 1; when the policy label is the DFS policy, the output value is 0. Calculate and obtain the loss function, expressed as: ; In the formula, is the loss function, is the total number of samples in the current training batch, is the output value of the policy label of the th sample, is the policy probability of the model of the th sample, is the regularization term for suppressing model overfitting, is the weight matrix from the input layer to the hidden layer, is the weight matrix from the hidden layer to the output layer, is the weight matrix the sum of the squares of all elements within, is the weight matrix the sum of the squares of all elements within.

[0011] Further, the logic for calculating the policy deviation is: Record the policy probability reached by the model training in the round of the number of successfully archived electronic files of the hyperlinks, is an integer greater than 0, which is used as the actual policy benefit , obtain the expected number of archives corresponding to the policy probability in this round as the expected policy benefit , and calculate the difference between the actual policy benefit and the expected policy benefit to obtain the policy deviation, expressed as: .

[0012] Further, update the neural network parameters based on the loss function and the policy deviation, expressed as: ; ; In the formula, is the learning rate of the current round, is the initial learning rate, is the error adjustment coefficient, is the local adjustment rate used to amplify and correct the weight of the high misjudgment area, is the bias vector from the hidden layer to the output layer; Calculate the gradient through the backpropagation algorithm and update the weight matrix using the Adam optimizer , ; and the bias vector , ; It is expressed as: ; ; ; ; In the formula, is the bias vector from the input layer to the hidden layer, , is the gradient of the loss function with respect to the weight matrix and the bias vector , , is the gradient of the loss function with respect to the weight matrix and the bias vector .

[0013] Furthermore, the traversal strategy is: when the output strategy probability ≥0.5, preferentially adopt the BFS traversal strategy, and when the output strategy probability <0.5, preferentially adopt the DFS traversal strategy.

[0014] Furthermore, the path information includes the set of traversed web pages and the set of jump edges of the traversed web pages. The logic for obtaining the web page labels and aggregation strength is as follows: S51, construct an attribution structure graph G=(X, E), where X is the set of traversed web pages, , E is the set of jump edges of the traversed web pages, ; Both and ζ are integers greater than 0; S52, based on each traversed web page in the structure graph, extract the structural attributes, and the structural attributes include page in-degree, page out-degree, structural level, and the proportion of pointed files; S53, based on the structural attributes, perform web page label determination on the traversed web pages, and the web page labels include directory aggregation pages and electronic file resources; S54. Calculate the corresponding aggregation strength based on the traversed web pages determined to be directory aggregation pages, and divide the subordinate document pages of the traversed web pages into priorities based on the aggregation strength.

[0015] Further, define each directory aggregation page as , ∈X, calculate and obtain the aggregation strength of the traversed web page, expressed as: ; is the aggregation strength of the directory aggregation page, is the set of all web pages directly pointed to by hyperlinks of the target aggregation page in the structure diagram, is the web page within, is the proportion of pointing-type files, is the directory aggregation page 's page out-degree.

[0016] Compared with the prior art, the beneficial effects of the present invention are as follows: By introducing a hybrid traversal model training mechanism, the present invention realizes the efficient understanding and intelligent scheduling of web pages with complex structures, multiple entrances, and multi-levels, improves the crawling coverage rate and content relevance matching degree of web page electronic files, and has the ability to adaptively identify different types of web page structures and link densities, avoiding problems such as crawling redundancy and missing key resources in traditional traversal methods in complex websites, thereby enhancing the accuracy and robustness of electronic file collection.

[0017] In addition, through the construction of the structure diagram and the aggregation strength evaluation mechanism, combined with the filing role determination strategy, the automatic judgment and priority division of the filing target value are realized, and then it is possible to accurately identify high-value file resources and directory aggregation pages, avoid the waste of storage resources on low-value content, and improve the scheduling response efficiency of high-value file content. Brief Description of the Drawings

[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required to be used in the embodiments. Obviously, the drawings described below are only some embodiments recorded in the present invention. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings.

[0019] Figure 1 is the step diagram of a method for processing web page electronic files based on artificial intelligence algorithms provided by an embodiment of the present invention; Figure 2 is the training step diagram of the neural network of a method for processing web page electronic files based on artificial intelligence algorithms provided by an embodiment of the present invention; Figure 3 It is a training flow chart of a neural network for a web electronic file processing method based on an artificial intelligence algorithm provided by an embodiment of the present invention; Figure 4 It is a classification step diagram of filing priorities for a web electronic file processing method based on an artificial intelligence algorithm provided by an embodiment of the present invention. Detailed implementation manners

[0020] The following describes the embodiments of the present disclosure in detail with reference to the accompanying drawings.

[0021] The following illustrates the embodiments of the present disclosure through specific specific examples. Those skilled in the art can easily understand other advantages and effects of the present disclosure from the content disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all of the embodiments. The present disclosure can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present disclosure. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present disclosure without creative efforts belong to the scope of protection of the present disclosure.

[0022] It should also be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present disclosure in a schematic manner. The diagrams only show the components related to the present disclosure rather than being drawn according to the number, shape, and size of the components in actual implementation. The types, quantities, and proportions of the components in actual implementation can be arbitrarily changed, and the component layout form may also be more complex.

[0023] Refer to Figure 1 , a web electronic file processing method based on an artificial intelligence algorithm, the method includes: S1, target website parsing Obtain a URL set U0 of the target website. The URL set U0 contains multiple seed URLs, and extract the metadata corresponding to each seed URL through a page parsing engine to generate a metadata set; It should be noted that: for a target website, such as a certain government affairs platform, enterprise archive network, or institutional announcement page, different business modules (such as announcements, policies, and annual statistical reports) of the target website correspond to different main entrances, and there may also be multiple "entrance paths" in the page paging, year directory, and column nesting within the target website; in addition, each sub-site (such as a city sub-portal) is deployed on different second-level domains or URL paths; Therefore, to cover the starting entry of the target website structure, it is necessary to extract multiple seed URLs to form a URL set U0 to prevent the crawler from getting stuck in a local area and having an incomplete collection scope; Among them, the seed URL set U0 can be generated by manually specifying the entry page or the sitemap file provided by the target website.

[0024] Specifically, the metadata includes the main page title text, the main body content of the page, and the set of page hyperlink attributes; It should be noted that: the main page title text is used to represent the core semantic information of the page; the main body content of the page is used to express the main content paragraphs of the page, and the content with the largest DOM density and the most stable text block length is preferentially extracted from the main body content of the page; the set of target page hyperlink attributes is the set of attributes of all hyperlinks in the target page, through the Extract the href attribute of the label, and the hyperlink points to other page resources or attachment files within the website.

[0025] S2. Construct an initial feature vector according to the metadata set; Specifically, the constructed initial feature vector is expressed as: ={ , , , , } In the formula, is the th hyperlink parsed from the web page, is an integer greater than 0; is the link depth, is the link density, is the document type code, is the metadata semantic matching degree, is the structural entropy; Specifically, the link depth is obtained by calculating the number of URL path levels in the page hyperlink attribute set, and is expressed as;

[0026] In the formula, means splitting the th hyperlink into multiple path segments according to the slash / , for example, / gov / policy / file.pdf will be split into 4 segments, expressed as ['', 'gov', 'policy', 'file.pdf']; () represents counting the number of path segments after splitting, c is a compensation constant, preferably, the compensation constant c is set to 2; It should be noted that: the compensation constant c comes from the empirical statistics of the actual website path depth distribution, and the value range is 1 to 3; Specifically, the link density is calculated by Obtained by dividing the number of nodes of the label by the total number of all nodes in the DOM tree, expressed as:

[0027] In the formula, is The number of nodes of the label, is the total number of all nodes in the DOM tree; It should be noted that: The label statistics scope is the link nodes directly parsed and visible through the standard HTML structure on the current page, excluding the virtual links generated by dynamic loading through JavaScript. All standard DOM element nodes (including text, div, table, etc.) are included; Specifically, the document type encoding According to the suffix name types of files pointed to in the page hyperlink attribute set, such as.pdf,.docx,.xls, etc., it is encoded and classified, expressed as:

[0028] For example, when the suffix name type in the u-th hyperlink is the output is 0.9; Preferably, to adapt to diverse file types, a complete document type encoding dictionary is constructed as follows:

[0029] Specifically, the metadata semantic matching degree It is obtained by calculating the cosine similarity, expressed as:

[0030] Among them, is the page semantic vector, and the page semantic vector is generated by splicing the page main title text and the page body content. is the preset keyword vector, and the keyword vector is constructed through the archive keyword set; If = 1, it means that the content of the web electronic archive captured on the page perfectly matches the preset keyword content; if = 0, it means that the content of the web electronic archive captured on the page has no relation to the preset keyword content; preferably, when > 0.6, the relevance between the content of the web electronic archive captured on the page and the preset keyword content is strong; It should be noted that: the vector and the vector are constructed by using the existing BERT model, and the BERT model uses the pre-trained language model bert-base-chinese; Specifically, the structural entropy is calculated by statistically analyzing the probabilities of various HTML tags (including but not limited to div, table, and a tags) in the DOM node distribution in the page hyperlink attribute set, and is calculated according to the structural entropy formula, expressed as:

[0031]

[0032] Wherein, is the proportion of HTML tags of the th type in the page, is the total number of tag types in the DOM, is the probability that HTML tags of the th type appear in the DOM structure, is the th type of tag in the number of nodes that appear in the page; It should be noted that: the tag types do not include <script> 与 <style> 标签,以避免非可见内容对结构分布的干扰。

[0033] S3,混合遍历模型训练利用步骤S2中构建的初始特征向量,输入至混合遍历模型中进行优化训练,本发明基于策略偏差反馈与结构熵感知调节策略的神经网络作为混合遍历模型,用于在广度优先搜索(BFS)与深度优先搜索(DFS)之间进行动态权衡,在面对结构复杂或链接密度极化的网页结构时,能够更加精准地分配抓取路径,以提高有价值电子档案资源的采集效率,确保混合遍历模型的适应性与鲁棒性;具体地,基于结构熵感知调节策略和策略偏差反馈的神经网络的训练流程为:S31,构建三层前馈神经网络,三层前馈神经网络包括输入层,隐藏层和输出层,初始化每层之间的权重矩阵与偏置向量,并设置激活函数与初始学习率;具体的,初始采用Xavier法进行初始化,表示为:;;式中,为正态分布,为输入层的神经元数量,为隐藏层的神经元数量,为输入层到隐藏层的权重矩阵,为隐藏层到输出层的权重矩阵,为输入层到隐藏层的偏置向量,为隐藏层到输出层的偏置向量,和的初始值设定为零;优选的,三层前馈神经网络中,输入层维度为5,隐藏层神经元数优选为64,输出层神经元为1;隐藏层采用ReLU激活函数;输出层采用Sigmoid激活函数。S32,通过结构熵感知调节策略对初始特征向量进行缩放,以获取调节输入向量;具体的,获取调节输入向量,表示为:

[0034] 式中,为调节输入向量,为结构熵调节因子,为初始特征向量,为结构熵,优选的,∈[0.2, 1.0];需要说明的是:当结构熵趋近于0时,对初始特征向量进行调节的幅度近似于1(无调节),结构熵越大,调节越显著;S33,将调节输入向量输入前馈神经网络进行前向传播,输出隐藏层输出结果和策略概率;具体的,输出隐藏层输出结果和策略概率,表示为:;;;;式中,为输入层到隐藏层的线性变换结果,为隐藏层到输出层的线性变换结果,为隐藏层激活函数,为输出层激活函数,为输入层到隐藏层的权重矩阵,为隐藏层到输出层的权重矩阵,为输入层到隐藏层的偏置向量,为隐藏层到输出层的偏置向量,优选的,策略概率∈(0,1),策略门限设为0.5;需要说明的是:前向传播顺序为输入层→隐藏层→输出层;BFS为广度优先搜索,DFS为深度优先搜索, 广度优先搜索(BFS)用于优先采集与当前页面同层的链接资源,适合获取页面横向分布的档案内容;而深度优先搜索(DFS)则用于优先深入访问当前页面中的下级链接,适合发掘目录型页面中的嵌套档案资源;S34,基于预设策略标签和策略概率,计算获取损失函数;具体的,策略标签包括BFS策略和DFS策略,当策略标签为BFS策略时,输出值为1,当策略标签为DFS策略时,输出值为0;计算获取损失函数,表示为:;其中,为损失函数,为当前训练批次中的样本总数,为第个样本的策略标签的输出值,为第条样本的模型的策略概率,为抑制模型过拟合的正则化项,为正则化系数,为输入层到隐藏层的权重矩阵,为隐藏层到输出层的权重矩阵,为权重矩阵内所有元素的平方和,为权重矩阵内所有元素的平方和,优选的,设为10⁻4;需要说明的是:设为10⁻4用于防止模型过拟合;S35,根据每轮策略概率对应的归档成功数和预设的期望归档数,计算获取策略偏差;具体的,计算获取策略偏差的逻辑为:记录第轮次训练的策略概率所触达超链接的电子档案的归档成功数,以作为实际策略收益,获取该轮次策略概率对应的期望归档数,以作为期望策略收益,将实际策略收益与期望策略收益进行差值计算,得到策略偏差,表示为:;需要说明的是:统计归档成功数时去除文件名重复项,期望策略收益通过策略概率的历史数据的平均归档数量进行设定。

[0035] S36,基于损失函数和策略偏差更新神经网络参数,神经网络参数包括权重矩阵,偏执向量和学习率;具体的,基于损失函数和策略偏差更新神经网络参数,表示为:;;式中,为当前轮次学习率,为初始学习率,为误差调整系数,为隐藏层到输出层的偏置向量,为用于对高误判区域权重放大修正的局部调整率,优选的,∈[0.001,0.01],∈[0.0032,0.013];通过反向传播算法计算梯度,采用Adam优化器更新权重矩阵,;以及偏置向量,;表示为:;;;;式中,为输入层到隐藏层的偏置向量,为隐藏层到输出层的偏置向量,,为损失函数关于权重矩阵和偏置向量的梯度,,为损失函数关于权重矩阵和偏置向量的梯度。

[0036] S37,当损失函数下降幅度连续三轮小于时,部署对应的神经网络作为混合遍历模型。

[0037] S4,网页遍历策略调整获取新的目标网站的元数据集合,利用已训练完成的混合遍历模型新的目标网站内各个超链接对应生成策略概率,基于策略概率确定遍历策略。

[0038] 具体的,遍历策略为:输出策略概率≥0.5时,优先采用BFS的遍历策略,输出策略概率<0.5时,优先采用DFS的遍历策略;S5,归档优先级分类根据遍历策略执行过程中记录的路径信息,获取网页标签和聚合强度,并基于网页标签和聚合强度进行归档优先级分类;具体的,获取网页标签和聚合强度的逻辑为:所述路径信息包括已遍历网页集合和已遍历网页的跳转边集合;需要说明的是:已遍历网页集合由多个已遍历的网页组成;已遍历网页的跳转边集合为已遍历网页的超链接指向的其他网页的集合;S51,构建归属结构图G=(X,E),其中,X为已遍历网页集合,, E为已遍历网页的跳转边集合,;和ζ皆为大于0的整数;S52,基于结构图中每个遍历网页,提取结构属性,所述结构属性包括页面入度,页面出度,结构层级,指向型档案比重;需要说明的是:页面入度为网页被其他网页超链接指向的次数,页面出度为超链接指向其他网页的数量,结构层级,基于URL中斜杠" / ”计数,指向型档案比重为网页指向PDF、Word 等电子档案格式资源的比例,通过网页中归档文件链接数和全部有效链接数计算获取;S53,基于结构属性,对已遍历网页执行网页标签判定,所述网页标签包括目录聚合页和电子档案资源;具体的,网页标签判定逻辑为:当页面出度大于页面入度,且指向型档案比重小于0.3时,判定网页为目录聚合页;当页面出度小于等于2,且指向型档案比重大于0.5时,判定网页为电子档案资源;其余网页则判定不具备聚合或归档意义;S54,基于判定为目录聚合页的已遍历网页,计算对应的聚合强度,并基于聚合强度对已遍历网页的下属文档页进行优先级划分;需要说明的是:下属文档页为判定为目录聚合页的已遍历网页内超链接所指向的网页;具体的,定义每个目录聚合页为,∈X,计算获取;为目录聚合页的聚合强度,为目标聚合页在结构图中通过超链接所直接指向的全部网页集合,为内的网页,为指向型档案比重,为目录聚合页的页面出度;若≥0.6,则将目标聚合页及其下属文档页划分为"高价值批注”,自动优先归档和调度;若<0.25,则将目标聚合页及其下属文档页划分为"低价值批注”,自动延后归档;若0.25≤<0.6,则通过人工策略进行处理;需要说明的是:0.25≤<0.6无法自动进行归档优先级判定,需要人工进行参与确定归档方案。

[0039] 在本实施例中, 首先对网页标签被判定为电子档案资源标签的网页中所包含的电子档案执行优先归档与调度操作;对于聚合强度满足≥0.6的目录聚合页,其所链接的下属文档页所包含的电子档案视为次级归档与调度对象,予以高价值电子档案内容补充;当目录聚合页的聚合强度满足小于<0.25时,其所链接的下属文档页中的电子档案被标记为延后归档与调度;若聚合强度落于区间0.25≤<0.6,则对应目录聚合结构无法直接归入自动归档策略范围,需由人工策略判定后再行归档处理。

[0040] 以上所述的实施例仅是对本发明的优选实施方式进行描述,并非对本发明的范围进行限定,在不脱离本发明设计精神的前提下,本领域普通技术人员对本发明的技术方案做出的各种变形和改进,均应落入本发明权利要求书确定的保护范围内。< / script>

Claims

1. A method for processing web electronic files based on artificial intelligence algorithms, characterized in that, The method includes: S1. Target website parsing: Obtain the URL set U0 of the target website. The URL set U0 contains multiple seed URLs. Extract the metadata corresponding to each seed URL through a page parsing engine to generate a metadata set. S2. Initial feature vector acquisition: Specifically, construct an initial feature vector according to the metadata set. S3. Hybrid traversal model training: Specifically, use a neural network based on a structure entropy perception adjustment strategy and policy deviation feedback as the hybrid traversal model. S4. Traversal strategy determination: Specifically, obtain the metadata set of the new target website. Use the trained hybrid traversal model to generate policy probabilities for each hyperlink in the new target website, and determine the traversal strategy based on the policy probabilities. S5. Archiving priority classification: Specifically, according to the path information recorded during the execution of the traversal strategy, obtain the web page labels and aggregation strengths, and perform archiving priority classification based on the web page labels and aggregation strengths.

2. The web electronic file processing method based on artificial intelligence algorithm according to claim 1, wherein The metadata includes the main page title text, the main body content of the page, and the set of page hyperlink attributes.

3. The web electronic file processing method based on artificial intelligence algorithm according to claim 2, wherein, The training process of the neural network based on the structure entropy perception adjustment strategy and policy deviation feedback is as follows: S31. Construct a three-layer feedforward neural network. The three-layer feedforward neural network includes an input layer, a hidden layer, and an output layer. Initialize the weight matrix and bias vector between each layer, and set the activation function and the initial learning rate. S32. Scale the initial feature vector through the structure entropy perception adjustment strategy to obtain an adjusted input vector. S33, input the adjusted input vector into the feedforward neural network for forward propagation, and output the output result of the hidden layer and the policy probability ; S34. Calculate and obtain the loss function based on the preset policy label and policy probability. S35. Calculate and obtain the policy deviation according to the number of successful archivings corresponding to each round of policy probability and the preset expected number of archivings. S36. Update the neural network parameters based on the loss function and the policy deviation. The neural network parameters include the weight matrix, the bias vector, and the learning rate. S37, when the decrease in the loss function is less than for three consecutive rounds, deploy the corresponding neural network as the hybrid traversal model.

4. The method for processing web electronic files based on artificial intelligence algorithm according to claim 3, characterized in that, Obtain the adjusted input vector, expressed as: ; In the formula, is the adjusted input vector, is the structural entropy adjustment factor, is the initial feature vector, is the structural entropy.

5. The method for processing web electronic files based on artificial intelligence algorithm according to claim 4, wherein, The policy labels include the BFS policy and the DFS policy. When the policy label is the BFS policy, the output value is 1; when the policy label is the DFS policy, the output value is 0. Calculate and obtain the loss function, expressed as: ; Wherein, is the loss function, is the total number of samples in the current training batch, is the output value of the policy label of the th sample, is the policy probability of the model of the th sample, is the regularization term for suppressing model overfitting, is the regularization coefficient, is the weight matrix from the input layer to the hidden layer, is the weight matrix from the hidden layer to the output layer, is the weight matrix and is the sum of the squares of all elements therein, is the weight matrix and is the sum of the squares of all elements therein.

6. The web electronic file processing method based on artificial intelligence algorithm according to claim 5, characterized in that, The logic for calculating and obtaining the policy deviation is: Record the policy probability of the number of successfully archived electronic files of the hyperlinks reached, which is an integer greater than 0 and used as the actual policy reward , and obtain the policy probability of this round corresponding expected number of archived files as the expected policy reward , and calculate the difference between the actual policy reward and the expected policy reward to obtain the policy deviation, expressed as: .

7. A method for processing web electronic files based on artificial intelligence algorithms according to claim 6, characterized in that, Update the neural network parameters based on the loss function and the policy deviation, expressed as: ; ; Wherein, is the learning rate of the current round, is the initial learning rate, is the error adjustment coefficient, is the local adjustment rate used to amplify and correct the weight of the high misjudgment area, is the bias vector from the hidden layer to the output layer; Calculate the gradient through the backpropagation algorithm and update the weight matrix using the Adam optimizer , ; And Bias vector , ; Expressed as: ; ; ; ; In the formula, is the bias vector from the input layer to the hidden layer, , is the gradient of the loss function with respect to the weight matrix and the bias vector . , is the gradient of the loss function with respect to the weight matrix and the bias vector .

8. A method for processing web electronic files based on artificial intelligence algorithms according to claim 7, characterized in that The traversal strategy is: output the strategy probability When it is ≥ 0.5, preferentially adopt the BFS traversal strategy and output the strategy probability When it is < 0.5, preferentially adopt the DFS traversal strategy.

9. A method for processing web electronic files based on an artificial intelligence algorithm according to claim 8, characterized in that The path information includes the set of traversed web pages and the set of jump edges of the traversed web pages. The logic for obtaining the web page labels and aggregation strengths is: S51. Construct an attribution structure graph G = (X, E), where X is the set of traversed web pages, , and E is the set of jump edges of the traversed web pages, ; Both and ζ are integers greater than 0; S52, based on each traversed web page in the structure diagram , extract structure attributes, where the structure attributes include page in-degree, page out-degree, structure level, and the proportion of pointed files; S53. Based on the structural attributes, perform web page label determination on the traversed web pages. The web page labels include the directory aggregation page and the electronic file resources. S54. Based on the traversed web pages determined to be directory aggregation pages, calculate the corresponding aggregation strengths, and perform priority division on the subordinate document pages of the traversed web pages based on the aggregation strengths.

10. A method for processing web electronic files based on artificial intelligence algorithms according to claim 9, characterized in that, Define each directory aggregation page as , ∈ X, calculate and obtain the aggregation strength of the traversed web pages, expressed as: ; is the aggregation strength of the directory aggregation page, is the set of all web pages directly pointed to by hyperlinks in the structure diagram of the target aggregation page, is the web pages within, is the proportion of pointed files, is the directory aggregation page 's page out-degree.

Citation Information

Patent Citations

  • Web vulnerability detection method and device, model training method and device and electronic equipment

    CN118199946A

  • Unmanned workshop equipment upgrading intelligent scheduling optimization system based on reinforcement learning

    CN119398568A

  • Network data adaptive acquisition method and system based on pre-training large model

    CN119884527A

  • Deep reinforcement learning for field development planning optimization

    US20220164657A1

Cited By

  • Self-adaptive retrieval strategy optimization method and system based on artificial intelligence

    CN122286003A

  • An adaptive retrieval strategy optimization method and system based on artificial intelligence

    CN122286003B