Webpage processing method and device based on mirror station cluster and hash solidification, equipment and medium

CN122817583APending Publication Date: 2026-09-25ZHEJIANG FUBO MEDIA TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611291300.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-25
Publication Date
2026-09-25

AI Technical Summary

Benefits of technology

[0015]综上可知,本申请基于取证指令访问疑似侵权网页,得到所述疑似侵权网页的网页信息,根据所述网页信息提取多维特征并构建各所述疑似侵权网页的目标页面特征;根据获取的所述目标页面特征,确定任意两个所述疑似侵权网页之间的镜像相似度,并根据所述镜像相似度对所述疑似侵权网页进行聚类,得到包括不同站群标签的各个镜像站群的站群聚类结果;基于预设规则确定各所述镜像站群的代表页面,对各所述代表页面执行第一证据采集操作得到第一证据,对各个所述镜像站群中非所述代表页面的页面执行第二证据采集操作得到第二证据,根据所述第一证据、所述第二证据以及所述站群聚类结果,生成初始站群关联证据包;对所述初始站群关联证据包进行哈希固化,输出经哈希固化的目标站群关联证据包至用户端,以根据所述目标站群关联证据包处理所述疑似侵权网页。由上可知,本申请首先通过预设浏览器访问携带特定标识的疑似侵权页面,采集其DOM快照、渲染截图和网络请求轨迹等网页信息。随后,利用网页信息得到结构指纹、基础设施特征、网络请求特征、截图感知哈希,以构建目标页面特征,进而计算任意两个疑似侵权网页间的镜像相似度,并据此聚类得到带有不同站群标签的镜像站群结果。此后,依据预设规则确定各站群的代表页面,对代表页面执行第一证据采集操作,对其余页面执行第二证据采集操作,结合聚类结果生成初始站群关联证据包,最终对该证据包进行哈希固化,输出经固化的目标站群关联证据包至用户端,以供后续对疑似侵权网页采取相应处理措施。这样一来,能够自动采集并固定疑似侵权网页的关键页面证据,识别关联镜像站群,减少重复取证,并生成便于复核的站群关联证据包。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122817583A_ABST
    Figure CN122817583A_ABST
Patent Text Reader

Abstract

The application discloses a webpage processing method and device based on mirror site cluster and hash fixation, equipment and medium, relates to the network forensics field, and comprises the following steps: accessing a suspected infringing webpage based on a forensic instruction, obtaining webpage information of the suspected infringing webpage, extracting multidimensional features based on the webpage information, and constructing target page features of each suspected infringing webpage; determining mirror similarity based on the target page features, clustering the suspected infringing webpage based on the mirror similarity to obtain a site cluster result; determining a representative page based on a preset rule, performing an evidence collection operation on the representative page and a non-representative page to obtain first evidence and second evidence, generating an initial site cluster associated evidence package based on the first evidence, the second evidence and the site cluster result; hash fixing the initial site cluster associated evidence package, and outputting a target site cluster associated evidence package to a user end. Key page evidence of the suspected infringing webpage can be automatically collected and fixed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of network forensics, and in particular to a webpage processing method, apparatus, device, and medium based on mirror site clustering and hash solidification. Background Technology

[0002] Existing web page forensics solutions typically target fixed URLs, providing screenshots, source code, logs, and timestamps to prove that a web page's content was collected and saved at a specific moment. However, in scenarios involving copyright protection, piracy monitoring, and processing of suspected infringing web pages, the objects to be addressed are often not single web pages, but rather a large number of suspected infringing web pages, pirated playback pages, mirror site pages, short-link redirect pages, and player-nested pages. Multiple pages may use the same video source, the same player code, similar DOM structures, or the same CDN nodes, but their domain names, paths, titles, and display templates differ.

[0003] In conclusion, how to automatically collect and secure key page evidence of suspected infringing web pages is an urgent problem to be solved. Summary of the Invention

[0004] In view of this, the purpose of this invention is to provide a webpage processing method, apparatus, device, and medium based on mirror site clustering and hash-based solidification, capable of automatically collecting and securing key page evidence of suspected infringing webpages. The specific solution is as follows: Firstly, this application provides a webpage processing method based on mirror site clustering and hash solidification, including: Based on the evidence collection instructions, access the suspected infringing web pages to obtain the web page information of the suspected infringing web pages, extract multi-dimensional features based on the web page information, and construct the target page features of each of the suspected infringing web pages; Based on the acquired target page features, determine the mirror similarity between any two suspected infringing web pages, and cluster the suspected infringing web pages according to the mirror similarity to obtain the clustering results of each mirror site group including different site group tags; Based on preset rules, representative pages of each of the mirror site groups are determined. A first evidence collection operation is performed on each of the representative pages to obtain first evidence. A second evidence collection operation is performed on pages in each of the mirror site groups that are not representative pages to obtain second evidence. An initial site group association evidence package is generated based on the first evidence, the second evidence, and the site group clustering results. The initial site group association evidence package is hashed and solidified, and the hashed and solidified target site group association evidence package is output to the user terminal to process the suspected infringing web pages based on the target site group association evidence package.

[0005] Optionally, the webpage information includes DOM snapshots, rendered screenshots, and network request traces, and the suspected infringing webpage is a webpage carrying a preset suspected identifier; Accordingly, the step of extracting multi-dimensional features based on the webpage information and constructing target page features for each of the suspected infringing webpages includes: Identify the player container and media interface in the suspected infringing webpage by using the DOM snapshot and / or the network request trajectory, so as to obtain the media content loaded or played by the suspected infringing webpage, generate a video clip hash based on the media content, and generate a player code fingerprint based on the player script in the suspected infringing webpage; The DOM snapshot is compressed to obtain the DOM structure fingerprint, the domain name associated with the suspected infringing webpage is parsed to obtain the corresponding domain name infrastructure characteristics, the network request trajectory is analyzed to generate network request characteristics, and the rendered screenshot is hashed to obtain the corresponding screenshot perception hash. Based on the DOM structure fingerprint, the screenshot-aware hash, the video clip hash, the player code fingerprint, the network request characteristics, and the domain name infrastructure characteristics, the target page characteristics of the suspected infringing webpage are constructed.

[0006] Optionally, generating a video clip hash based on the media content includes: Keyframe sampling is performed on the media content to obtain the corresponding keyframe summary; The media content is sampled to obtain corresponding audio segment summaries; The media content is segmented and sampled to obtain corresponding media segment summaries; The video segment hash is generated based on the obtained keyframe summary, audio segment summary or media segment summary, and the corresponding sampling time position.

[0007] Optionally, generating a player code fingerprint based on the player script in the suspected infringing webpage includes: The player script is parsed to obtain an abstract syntax tree summary; Extract the media interface path from the player script to obtain the interface path summary; Extract the key parameter names from the player script to obtain a parameter name summary; Extract the obfuscation features from the player script to obtain a script obfuscation feature summary; The player code fingerprint is generated based on the abstract syntax tree digest, the interface path digest, the parameter naming digest, and the script obfuscation feature digest.

[0008] Optionally, determining the mirror similarity between any two suspected infringing web pages based on the acquired target page features includes: Based on the target page features of the two suspected infringing web pages, the initial similarity is determined on the DOM structure fingerprint, screenshot-aware hash, video clip hash, player code fingerprint, network request features, and domain name infrastructure features, respectively. Count the number of evidence channels that meet the preset credibility conditions in each of the initial similarities; If the number of evidence channels meets a preset quantity condition, then a geometric consistency operation is performed based on the initial similarity that meets the preset credibility condition using a preset mirror similarity calculation formula to obtain the mirror similarity.

[0009] Optionally, the step of clustering the suspected infringing web pages based on the mirror similarity includes: The suspected infringing webpages are used as graph nodes. When the mirror similarity between two suspected infringing webpages is greater than or equal to a preset site group threshold, a similar edge is established between the two nodes to construct a page similarity graph. Perform connected component identification, label propagation, spectral clustering, or community detection processing on the page similarity graph to obtain the site clustering result.

[0010] Optionally, determining the representative page of each of the mirror site groups based on preset rules includes: Obtain the collection integrity score, evidence coverage, access stability, abnormal penalty items, and evidence collection cost for each of the suspected infringing web pages in the mirror site group; A lexicographical comparison operation is performed based on the order of the collection integrity score, the evidence coverage, the access stability, the anomaly penalty item, and the evidence collection cost to generate the corresponding comparison results. The suspected infringing webpages that meet the preset optimal conditions in the comparison results are identified as the representative pages.

[0011] Optionally, the step of hashing and solidifying the initial site group association evidence package to output the hash-solidified target site group association evidence package includes: For each evidence file in the initial site cluster association evidence package, determine the corresponding file-level hash value; According to the preset evidence category, the file-level hash values ​​of each evidence file belonging to the same preset evidence category are aggregated to generate the corresponding category hash value; The root hash value of the initial site group association evidence package is generated based on the category hash value, and the root hash value is used as a solidified identifier to determine the target site group association evidence package containing the solidified identifier.

[0012] Secondly, this application provides a webpage processing apparatus based on mirror site clustering and hash solidification, comprising: The feature construction module is used to access suspected infringing web pages based on evidence collection instructions, obtain web page information of the suspected infringing web pages, extract multi-dimensional features based on the web page information, and construct target page features for each of the suspected infringing web pages. The webpage clustering module is used to determine the mirror similarity between any two suspected infringing webpages based on the acquired target page features, and to cluster the suspected infringing webpages based on the mirror similarity to obtain the clustering results of each mirror site group including different site group tags. The evidence package generation module is used to determine the representative page of each of the mirror site groups based on preset rules, perform a first evidence collection operation on each of the representative pages to obtain first evidence, perform a second evidence collection operation on the pages in each of the mirror site groups that are not the representative pages to obtain second evidence, and generate an initial site group association evidence package based on the first evidence, the second evidence and the site group clustering results. The evidence package solidification module is used to hash and solidify the initial site group associated evidence package, and output the hash-solidified target site group associated evidence package to the user terminal, so as to process the suspected infringing webpage according to the target site group associated evidence package.

[0013] Thirdly, this application provides an electronic device, comprising: Memory, used to store computer programs; A processor is used to execute the computer program to implement the web page processing method based on mirror site clustering and hash solidification as described above.

[0014] Fourthly, this application provides a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the aforementioned web page processing method based on mirror site clustering and hash solidification.

[0015] In summary, this application accesses suspected infringing web pages based on evidence collection instructions to obtain web page information of the suspected infringing web pages. Multi-dimensional features are extracted based on the web page information, and target page features for each suspected infringing web page are constructed. Based on the obtained target page features, the mirror similarity between any two suspected infringing web pages is determined, and the suspected infringing web pages are clustered based on the mirror similarity to obtain clustering results for each mirror site group, including different site group tags. Representative pages for each mirror site group are determined based on preset rules. A first evidence collection operation is performed on each representative page to obtain first evidence, and a second evidence collection operation is performed on pages in each mirror site group that are not representative pages to obtain second evidence. An initial site group-related evidence package is generated based on the first evidence, the second evidence, and the site group clustering results. The initial site group-related evidence package is hashed and solidified, and the hashed and solidified target site group-related evidence package is output to the user terminal to process the suspected infringing web pages. As described above, this application first accesses suspected infringing pages carrying specific identifiers through a preset browser, collecting webpage information such as DOM snapshots, rendered screenshots, and network request trajectories. Subsequently, it uses the webpage information to obtain structural fingerprints, infrastructure features, network request features, and screenshot-aware hashes to construct target page features. Then, it calculates the mirror similarity between any two suspected infringing webpages and clusters them to obtain mirror site group results with different site group labels. Afterward, it determines representative pages for each site group according to preset rules, performs a first evidence collection operation on the representative pages, and a second evidence collection operation on the remaining pages. Combining the clustering results, it generates an initial site group-related evidence package. Finally, it hashes and solidifies this evidence package, outputting the solidified target site group-related evidence package to the user's end for subsequent processing of suspected infringing webpages. In this way, it can automatically collect and fix key page evidence of suspected infringing webpages, identify associated mirror site groups, reduce duplicate evidence collection, and generate site group-related evidence packages that are easy to verify. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0017] Figure 1 This is a system structure block diagram disclosed in this application; Figure 2 This is a flowchart of a webpage processing method based on mirror site clustering and hash solidification disclosed in this application; Figure 3This is a schematic diagram of multi-source page features and mirror site clustering disclosed in this application; Figure 4 This is a flowchart illustrating the interaction between the associated evidence package generation and hash solidification module disclosed in this application. Figure 5 This is a schematic diagram of a web page processing device based on mirror site clustering and hash solidification disclosed in this application; Figure 6 This is a structural diagram of an electronic device disclosed in this application. Detailed Implementation

[0018] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0019] Currently, webpage forensics solutions typically target fixed page screenshots, source code, logs, and timestamps for a single URL, proving that the content of a specific webpage was collected and saved at a particular moment. However, in scenarios involving copyright protection, piracy monitoring, and processing of suspected infringing webpage evidence, the objects to be processed are often not single webpages, but rather a large number of suspected infringing webpages, pirated playback pages, mirror site pages, short-link redirect pages, and player-nested pages. Multiple pages may use the same video source, the same player code, similar DOM structure, or the same CDN node, but their domain names, paths, titles, and display templates differ. To address the aforementioned technical problems, this application discloses a webpage processing method, apparatus, device, and medium based on mirror site clustering and hash-based solidification, capable of automatically collecting and securing key page evidence of suspected infringing webpages.

[0020] like Figure 1As shown, the webpage processing system based on mirror site clustering and hash solidification in this application includes a task scheduling module for receiving suspected infringing webpage evidence collection tasks, configuring collection strategies, allocating browser environments, and managing collection queues; a browser evidence collection module for collecting DOM snapshots, webpage source code, rendered screenshots, console logs, and network request traces; a media and player analysis module for identifying player containers, media interfaces, video clip hashes, and player code fingerprints; an infrastructure resolution module for resolving domain names, CDN nodes, IPs, ASNs, TLS certificates, and certificate fingerprints; a multi-source feature construction module for generating DOM structure fingerprints, screenshot-aware hashes, video clip hashes, player code fingerprints, network request features, and domain name infrastructure features; a mirror site clustering module for calculating mirror similarity, constructing page similarity graphs, generating site clustering results, and selecting representative pages; a site cluster association evidence package generation module for assembling page evidence, media evidence, network evidence, infrastructure evidence, clustering evidence, and evidence hash lists; and a hash solidification and verification module for generating file-level hashes, category hashes, evidence package root hashes, signatures, timestamps, and verification interfaces.

[0021] See Figure 2 As shown, this embodiment of the invention discloses a webpage processing method based on mirror site clustering and hash solidification, including: Step S11: Access the suspected infringing webpage based on the evidence collection instructions, obtain the webpage information of the suspected infringing webpage, extract multi-dimensional features based on the webpage information, and construct the target page features of each of the suspected infringing webpages.

[0022] In this embodiment, the system receives a task to collect evidence from a suspected infringing webpage. The evidence collection instruction (or task data) may include the URL of the suspected infringing webpage, task number, copyright work identifier, webpage type, browser configuration, collection region, maximum number of redirects, maximum playback duration, media sampling duration, evidence package storage path, and hash algorithm, etc. The system generates a collection session identifier (CID) for each task and records the session start time (T_start) so that all evidence generated within the same collection session can be linked together. The preset suspected identifier is used to mark the webpage as a target webpage for infringement evidence collection; for example, it may be a copyright work identifier, task number, or clue source identifier.

[0023] Furthermore, different data collection strategies can be configured based on page type. For example, for static image and text pages, the system prioritizes collecting DOM elements, screenshots, and source code; for video playback pages, the system adds player detection, media request monitoring, and video clip hashing; for short link redirect pages, the system adds redirect link collection; and for batch tasks involving multiple websites, a clustering queue and representative page selection strategy are added.

[0024] When accessing suspected infringing web pages, the system uses a default browser (either headless or with headers) to access the URL of the suspected infringing web page, simulating the user agent, language, resolution, cookie policy, and network environment of a real browser. It records the page state at various points in time, including the start of page loading, initial content rendering, network idle time, player loading, and the end of sampling. The system also records redirection chains, DOM snapshots, rendered screenshots, web page source code, console logs, and the initial network request trajectory. Specifically, the DOM snapshot saves the rendered DOM tree, key node paths, player container nodes, iframe nodes, and script nodes; the rendered screenshots save the first screen screenshot, full-page screenshot, player area screenshot, and screenshots of abnormal states; the web page source code saves the initial HTML and dynamically modified outer HTML digests; the console and error logs record script exceptions, player errors, cross-domain errors, and media loading failures; and the network request trajectory records the request URL, method, status code, redirection chain, request header digest, response header digest, resource type, timeline, and response body hash.

[0025] In addition, to measure the quality of a single page's data collection, the system can calculate a page collection integrity score, which can be expressed by the following formula: ; in, This represents the data collection integrity score for page p. This represents the data collection gating function, which takes the value 0 if any key evidence, such as DOM evidence, network request trajectory, or data collection timestamp, is missing, and otherwise takes the value 1. Indicates the completeness of DOM evidence. =Number of DOM elements already captured / Number of elements to be captured, i.e., DOM tree, key node paths, player container, iframe, script nodes; Indicates the completeness of the network request trajectory. = Recorded request trajectory fields / Fields to be recorded, i.e., URL, method, status code, redirection chain, request / response headers, timeline, response body hash; Indicating the completeness of media evidence, = Acquired media items / Expected items, namely keyframes, audio summaries, segment summaries, playlists, and sampling time positions. Unplayable pages are assigned a baseline value according to preset rules to avoid being always 0; Indicates the completeness of the collected logs. =Logged categories / Categories that should be logged; E represents the exception penalty item. p For normalization anomaly penalty, E p≥0, such that 1 / (1+E_p)∈(0,1). This formula, through gating and short-board constraints, avoids situations where excessively high scores for one type of evidence mask the absence of key evidence, thereby ensuring that the accepted page evidence has a high degree of completeness.

[0026] Next, the player container and media interface in the suspected infringing webpage are identified through the DOM snapshot and / or the network request trajectory to obtain the media content loaded or played by the suspected infringing webpage. A video clip hash is generated based on the media content, and a player code fingerprint is generated based on the player script in the suspected infringing webpage.

[0027] In this embodiment, keyframe sampling is performed on the media content to obtain corresponding keyframe summaries; audio segment sampling is performed on the media content to obtain corresponding audio segment summaries; media segment sampling is performed on the media content to obtain corresponding media segment summaries; and the video segment hash is generated based on the obtained keyframe summaries, audio segment summaries, or media segment summaries, and the corresponding sampling time positions. Specifically, the system identifies video tags, iframe players, canvas players, script loaders, and media interfaces within the page, and obtains media segment URLs, playlist URLs (e.g., m3u8 addresses), video file URLs (e.g., mp4 addresses), segment requests, or interface return summaries by clicking playback controls, waiting for media buffering, intercepting network requests, or parsing player configurations, thereby obtaining the media content loaded or played by the suspected infringing webpage. For playable media, the system extracts keyframes, short segments, or audio summaries according to the sampling strategy without saving the complete video, and generates a video segment hash. The video segment hash can be represented as: ; in, This represents the hash of the video clip on page p. Represents a hash function for video clips; Represents a keyframe summary; Indicates a summary of an audio segment; This indicates a summary of media segments; Indicates the sampling time position.

[0028] Furthermore, the player script is parsed to obtain an abstract syntax tree digest; media interface paths are extracted from the player script to obtain an interface path digest; key parameter names are extracted from the player script to obtain a parameter name digest; obfuscation features are extracted from the player script to obtain a script obfuscation feature digest; and the player code fingerprint is generated based on the abstract syntax tree digest, the interface path digest, the parameter name digest, and the script obfuscation feature digest. Specifically, an abstract syntax tree digest is generated from the program syntax of the player script, an interface path digest is extracted from the media interface paths, and a key parameter name digest is generated by analyzing the interface paths and key parameter names. The abstract syntax tree digest, interface path digest, parameter name digest, and script obfuscation features are then integrated to generate the player code fingerprint. The player code fingerprint can be represented as: ; in, This represents the fingerprint of the player code on page p; Indicates a code fingerprint function; Represents a summary of the player's script abstract syntax tree; This represents a summary of the media interface path. This indicates a naming summary for key parameters; This indicates script obfuscation features.

[0029] Understandably, by simultaneously collecting video clip hashes and player code fingerprints, the system can distinguish between situations such as different page outer templates loading the same video source, the same player code loading different video sources, and the same website group reusing the same playback infrastructure.

[0030] Subsequently, the DOM snapshot is compressed to obtain the DOM structure fingerprint, the domain name associated with the suspected infringing webpage is parsed to obtain the corresponding domain name infrastructure characteristics, the network request trajectory is analyzed to generate network request characteristics, and the rendered screenshot is hashed to obtain the corresponding screenshot perception hash.

[0031] In this embodiment, the DOM tree is compressed into a DOM structure fingerprint. This DOM structure fingerprint may include tag hierarchy sequence, node degree distribution, player container path, script reference location, key attribute name set, resource loading slots, and ad module location. To reduce the impact of page noise, before compressing the DOM snapshot to obtain the DOM structure fingerprint, noise information such as random ads, timestamps, session parameters, and irrelevant recommendation lists is removed according to preset noise identification rules. Furthermore, the main page domain, script domain, image domain, media domain, interface domain, and redirect domain are extracted from the network request trajectory. DNS resolution, CDN node identification, TLS certificate collection, certificate fingerprint calculation, and IP attribution analysis are performed on these domains to obtain domain infrastructure characteristics. These domain infrastructure characteristics can be represented as: ; in, This indicates the domain infrastructure characteristics of page p; Indicates domain name resolution characteristics; Indicates CDN node characteristics; Indicates certificate characteristics; Indicates IP address characteristics; This indicates the characteristics of an autonomous system.

[0032] It should be noted that analyzing network request trajectories generates network request characteristics, and the network request trajectories are saved as request chains in chronological order. Each node in the request chain can include the normalized result of the request URL, resource type, initiator, response status, content hash, and time offset. For redirection chains, the system records the URL, status code, and redirection time before and after each redirect to verify the actual page entry point and the final access address. Hash calculations are performed on rendered screenshots to obtain screenshot-aware hashes. For example, a perceptual hashing algorithm can be used to calculate perceptual hash values ​​resistant to slight changes for screenshots of the first screen or the player area.

[0033] Then, based on the DOM structure fingerprint, the screenshot-aware hash, the video clip hash, the player code fingerprint, the network request characteristics, and the domain infrastructure characteristics, the target page features of the suspected infringing webpage are constructed. The generated features are then transformed and combined into a unified target page feature, which can be represented by the following formula: ; Among them, F p D represents the multi-source page feature vector of page p, i.e., the target page feature; p Represents the fingerprint of the DOM structure; S p V represents a screenshot-aware hash; p Represents the hash of a video clip; C p Indicates the player's code fingerprint; R p Indicates network request characteristics; I p This represents the characteristics of the domain name infrastructure. The system can assign credibility and weight to each type of characteristic; for example, when the video clip hash is available, V... p The weights are relatively high; when the video is restricted from being played, the weights of the DOM structure fingerprint, player code fingerprint, and network request features are increased, thereby maintaining the stability and comparability of the target page features under different page collection conditions.

[0034] Step S12: Based on the acquired target page features, determine the mirror similarity between any two suspected infringing web pages, and cluster the suspected infringing web pages according to the mirror similarity to obtain the clustering results of each mirror site group including different site group tags.

[0035] In this embodiment, as Figure 3 As shown, firstly, based on the target page features of the two suspected infringing web pages, initial similarities are determined on the DOM structure fingerprint, screenshot-aware hash, video clip hash, player code fingerprint, network request features, and domain infrastructure features, respectively. The number of evidence channels satisfying the preset confidence conditions in each initial similarity is counted. If the number of evidence channels meets the preset quantity condition, a geometric consistency operation is performed based on the initial similarities satisfying the preset confidence conditions using a preset mirror similarity calculation formula to obtain the mirror similarity. Specifically, for two suspected infringing web pages p and q, six types of target page features—DOM structure fingerprint, screenshot-aware hash, video clip hash, player code fingerprint, network request features, and domain infrastructure features—are extracted, and the initial similarity between each pair is calculated to obtain a set of original evidence scores covering page structure, visual presentation, dynamic content, underlying communication, and domain ownership, i.e., the initial similarity. However, not all evidence channels have equal reference value; therefore, only when the similarity value of a channel reaches a preset confidence threshold is it counted as a valid evidence channel, and the number of channels satisfying the preset confidence threshold is counted. If the preset quantitative conditions are met, such as at least four channels being highly similar simultaneously, it indicates that the two web pages are consistent across multiple independent dimensions, sufficient to rule out mere coincidence. At this point, a preset mirror similarity calculation formula is invoked. By checking whether the numerical distribution of these similarities satisfies the proportionality identity or transformation invariance characteristic of mirror replication, a normalized similarity score reflecting the overall degree of mirroring is finally derived, i.e., the mirror similarity score. Mirror similarity can be expressed as: ; in, This indicates the mirror similarity between page p and page q; This represents the key evidence gating function, which is set to 0 when the two pages lack comparable key evidence or only have template-level similarity. Indicates DOM structure similarity; Indicates the hash similarity of video clips; Indicates the similarity of the player's code fingerprint; Indicates the similarity of network request trajectories; Indicates domain infrastructure similarity; SimS indicates screenshot-aware hash similarity. This represents the number of evidence channels that participate in the comparison and meet the minimum credibility requirement. p,q When =0, by Gate p,q =0, directly let S p,q =0, to avoid exponential divergence.

[0036] Furthermore, after completing the quantitative calculation of pairwise mirror similarity, the suspected infringing web pages are used as graph nodes. When the mirror similarity between two suspected infringing web pages is greater than or equal to a preset site group threshold, a similar edge is established between the two nodes to construct a page similarity graph. Connected component identification, label propagation, spectral clustering, or community detection processing operations are performed on the page similarity graph to obtain the site group clustering result.

[0037] Specifically, each suspected infringing webpage is mapped as an independent node in a graph structure. Then, a preset site group threshold is used as the judgment boundary. The infringing webpage is determined if and only if the mirror similarity between any two nodes is greater than or equal to the preset site group threshold. At this point, an undirected similarity edge is established between the corresponding two nodes p and q. After traversing all node pairs in this way, a page similarity graph that can comprehensively reflect the relationship between pages is generated. =( , ),in A collection of page nodes. This represents the set of similar edges on the pages. To divide the graph into densely connected subnets, algorithms such as connected component identification, label propagation, spectral clustering, or community detection can be used to separate highly coupled sets of nodes from the whole. The final output of the site clustering result is a set of web pages that are tightly coupled internally and relatively loosely connected externally. These sets often directly correspond to mirror site groups controlled by the same infringing entity, and can be used for subsequent batch evidence collection or traffic management operations. The determination of site group members can be expressed as: ; in, This indicates whether pages p and q are grouped into the same mirror site group; 1[·] represents an indicator function, which takes a value of 1 when the condition within the square brackets is true, and a value of 0 otherwise. That is, when the mirror similarity between pages p and q is not lower than the mirror similarity threshold, and the number of evidence channels participating in the comparison and meeting the minimum credibility is not lower than the key evidence consistency item threshold, =1 indicates that pages p and q belong to the same mirror site group; otherwise... =0 indicates that page p and page q are not in the same mirror site group; Indicates the mirror similarity threshold; This indicates the number of evidence channels that participate in the comparison and meet the minimum credibility requirement. This indicates the threshold for key evidence consistency. See also: Figure 4 The diagram shows the interaction between the website cluster association evidence package generation and hash solidification modules, illustrating the interaction between the task scheduling module, browser evidence collection module, media and player analysis module, mirror website cluster clustering module, website cluster association evidence package generation module, and hash solidification and verification module.

[0038] Step S13: Determine the representative page of each of the mirror site groups based on preset rules, perform a first evidence collection operation on each of the representative pages to obtain first evidence, perform a second evidence collection operation on the pages in each of the mirror site groups that are not the representative pages to obtain second evidence, and generate an initial site group association evidence package based on the first evidence, the second evidence and the site group clustering results.

[0039] In this embodiment, after obtaining the site clustering results, a representative page is selected for each mirror site cluster. This requires obtaining the collection integrity score, evidence coverage, access stability, anomaly penalty item, and evidence collection cost for each suspected infringing webpage in the mirror site cluster. A lexicographical comparison operation is performed based on the order of the collection integrity score, evidence coverage, access stability, anomaly penalty item, and evidence collection cost to generate corresponding comparison results. The suspected infringing webpage that meets the preset optimal conditions in the comparison results is determined as the representative page. Specifically, after obtaining the site clustering results, a representative page is selected for each mirror site cluster. The representative page prioritizes conditions such as high collection integrity score, usable video clip hash, complete network request trajectory, stable online page performance, and no anomalies in the evidence file. To avoid ordinary weighted summation causing pages that are "easy to collect but have weak evidence" to be incorrectly selected as representative pages, the system adopts a lexicographical order constraint selection method. The representative page can be represented as: ; in, This represents the representative page of the mirror site group g; Lex represents the lexicographical comparison function. Indicates the data collection integrity score; This indicates the coverage of key evidence from the website cluster on page p; Indicates page access stability; Indicates an abnormal penalty item; This represents the cost of collecting complete evidence. The system first compares the completeness of the evidence collected, then the coverage of key evidence, and subsequently compares page stability, anomalies, and evidence collection costs. This ensures that the representative page must first meet the requirement of sufficient evidence before considering efficiency.

[0040] Next, complete evidence collection is performed on the representative page. The resulting first evidence may include the complete DOM, full-page screenshots, player area screenshots, media samples, full network request records, certificate information, and a list of evidence hashes. Simplified evidence collection is performed on non-representative pages in each mirror site group. The resulting second evidence may include the entry URL, final URL, screenshot summary, core DOM summary, key request summary, similarity score, and member page hashes. This approach proves the relationships between site group members while reducing the cost of repeatedly downloading complete media and performing repeated complete evidence collection. Based on the first and second evidences and the site group clustering results, an initial site group association evidence package is generated.

[0041] Step S14: Hash and solidify the initial site group association evidence package, and output the hash-solidified target site group association evidence package to the user terminal to process the suspected infringing webpage according to the target site group association evidence package.

[0042] In this embodiment, hash solidification is performed on the initial site group associated evidence package. First, the corresponding file-level hash value is determined for each evidence file in the initial site group associated evidence package. Based on a preset evidence category, the file-level hash values ​​of evidence files belonging to the same preset evidence category are aggregated to generate a corresponding category hash value. A root hash value of the initial site group associated evidence package is generated based on the category hash value, and this root hash value is used as a solidification identifier to identify the target site group associated evidence package containing the solidification identifier. In this way, the user can process suspected infringing web pages based on the target site group associated evidence package, such as verifying evidence, initiating takedown notices, filing complaints, or initiating lawsuits. Simultaneously, the root hash of the evidence package can be submitted to a trusted timestamp service, blockchain, electronic signature service, or internal audit system, thereby further enhancing the tamper resistance and verifiability of the evidence. The root hash of the evidence package is generated using a hierarchical Merkle tree, which can be represented as: ; in, Indicates the root hash of the evidence; This represents the leaf hash of the i-th evidence file or clustering result entry; Indicate the type of evidence; Indicates the path within the evidence package; This represents a hash of the file content or a hash of the structured result; This indicates the time when the evidence item was collected or generated. Using a Merkle tree structure, it is possible to verify whether a single evidence file, a single site cluster entry, or a set of evidence has been replaced without re-expanding the entire evidence package.

[0043] As described above, this embodiment first accesses a suspected infringing page carrying a specific identifier through a preset browser, collecting webpage information such as DOM snapshots, rendered screenshots, and network request trajectories. Then, it uses the DOM snapshot or network request trajectory to identify the player container and media interface, extracts the loaded or played media content, and generates video clip hashes. Simultaneously, it generates code fingerprints based on the player script, and calculates the hash value of the webpage information to obtain structural fingerprints, infrastructure features, network request features, and screenshot-aware hashes. Combining the above DOM structural fingerprints, screenshot-aware hashes, video clip hashes, player code fingerprints, network request features, and domain infrastructure features, it constructs target page features, calculates the mirror similarity between any two suspected infringing webpages, and clusters them to obtain mirror site group results with different site group labels. Subsequently, it determines representative pages for each site group according to preset rules, performs a first evidence collection operation on the representative pages, performs a second evidence collection operation on the remaining pages, and generates an initial site group-related evidence package based on the clustering results. Finally, it hashes and solidifies the evidence package, outputting the solidified target site group-related evidence package to the user terminal for subsequent processing of suspected infringing webpages. This allows for the automatic collection and preservation of key page evidence from suspected infringing web pages, identification of associated mirror site groups, reduction of redundant evidence collection, and generation of site group-related evidence packages that are easy to verify.

[0044] See Figure 5 As shown, this embodiment of the invention discloses a webpage processing device based on mirror site clustering and hash solidification, comprising: The feature construction module 11 is used to access suspected infringing web pages based on evidence collection instructions, obtain web page information of the suspected infringing web pages, extract multi-dimensional features based on the web page information, and construct target page features for each of the suspected infringing web pages. The webpage clustering module 12 is used to determine the mirror similarity between any two suspected infringing webpages based on the acquired target page features, and to cluster the suspected infringing webpages based on the mirror similarity to obtain the clustering results of each mirror site group including different site group tags. The evidence package generation module 13 is used to determine the representative page of each of the mirror site groups based on preset rules, perform a first evidence collection operation on each of the representative pages to obtain first evidence, perform a second evidence collection operation on the pages in each of the mirror site groups that are not the representative pages to obtain second evidence, and generate an initial site group association evidence package based on the first evidence, the second evidence and the site group clustering results. The evidence package solidification module 14 is used to perform hash solidification on the initial site group associated evidence package and output the hash solidified target site group associated evidence package to the user terminal so as to process the suspected infringing webpage according to the target site group associated evidence package.

[0045] As described above, this application first accesses suspected infringing pages carrying specific identifiers through a preset browser, collecting webpage information such as DOM snapshots, rendered screenshots, and network request trajectories. Subsequently, it uses the DOM snapshots or network request trajectories to identify the player container and media interface, extracts the loaded or played media content, and generates video clip hashes. Simultaneously, it generates code fingerprints based on the player scripts, and calculates the hash values ​​of the webpage information to obtain structural fingerprints, infrastructure features, network request features, and screenshot-aware hashes. Combining the aforementioned DOM structural fingerprints, screenshot-aware hashes, video clip hashes, player code fingerprints, network request features, and domain infrastructure features, it constructs target page features, calculates the mirror similarity between any two suspected infringing webpages, and clusters them to obtain mirror site group results with different site group labels. Then, according to preset rules, it determines representative pages for each site group, performs a first evidence collection operation on the representative pages, and a second evidence collection operation on the remaining pages. Combining the clustering results, it generates an initial site group-related evidence package. Finally, it hashes and solidifies this evidence package, outputting the solidified target site group-related evidence package to the user terminal for subsequent processing of suspected infringing webpages. This allows for the automatic collection and preservation of key page evidence from suspected infringing web pages, identification of associated mirror site groups, reduction of redundant evidence collection, and generation of site group-related evidence packages that are easy to verify.

[0046] In some specific implementations, the webpage information includes DOM snapshots, rendered screenshots, and network request traces, and the suspected infringing webpage is a webpage carrying a preset suspected identifier; Accordingly, the feature construction module 11 may specifically include: A webpage information generation unit is used to identify the player container and media interface in the suspected infringing webpage through the DOM snapshot and / or the network request trajectory, so as to obtain the media content loaded or played by the suspected infringing webpage, generate a video clip hash based on the media content, and generate a player code fingerprint based on the player script in the suspected infringing webpage. The hash value acquisition unit is used to compress the DOM snapshot to obtain the DOM structure fingerprint, parse the domain name associated with the suspected infringing webpage to obtain the corresponding domain name infrastructure characteristics, analyze the network request trajectory to generate network request characteristics, and perform hash calculation on the rendered screenshot to obtain the corresponding screenshot perception hash. The feature construction unit is used to construct the target page features of the suspected infringing webpage based on the DOM structure fingerprint, the screenshot-aware hash, the video clip hash, the player code fingerprint, the network request features, and the domain name infrastructure features.

[0047] In some specific implementations, the webpage information generation unit may specifically include: The first summary acquisition subunit is used to perform keyframe sampling on the media content to obtain the corresponding keyframe summary. The second summary acquisition subunit is used to sample audio segments from the media content to obtain corresponding audio segment summaries. The third summary acquisition subunit is used to perform media segmentation sampling on the media content to obtain the corresponding media segment summary; The hash generation subunit is used to generate the video segment hash based on the obtained keyframe summary, audio segment summary or media segment summary, and the corresponding sampling time position.

[0048] In some specific implementations, the webpage information generation unit may specifically include: The fourth summary acquisition subunit is used to perform syntax parsing on the player script to obtain an abstract syntax tree summary; The fifth summary acquisition subunit is used to extract the media interface path in the player script and obtain the interface path summary; The sixth summary acquisition subunit is used to extract the key parameter names in the player script and obtain a parameter name summary; The seventh summary acquisition subunit is used to extract obfuscation features from the player script and obtain a script obfuscation feature summary; The fingerprint generation subunit is used to generate the player code fingerprint based on the abstract syntax tree digest, the interface path digest, the parameter naming digest, and the script obfuscation feature digest.

[0049] In some specific implementations, the webpage clustering module 12 may specifically include: The similarity determination unit is used to determine the initial similarity on the DOM structure fingerprint, screenshot-aware hash, video clip hash, player code fingerprint, network request features and domain name infrastructure features respectively based on the target page features of the two suspected infringing web pages. The quantity statistics unit is used to count the number of evidence channels that meet the preset credibility conditions in each of the initial similarities; The similarity acquisition unit is used to perform geometric consistency calculation based on the initial similarity that meets the preset credibility condition, using a preset mirror similarity calculation formula, to obtain the mirror similarity if the number of evidence channels meets a preset quantity condition.

[0050] In some specific implementations, the webpage clustering module 12 may specifically include: The similarity graph construction unit is used to take the suspected infringing web pages as graph nodes. When the mirror similarity between two suspected infringing web pages is greater than or equal to a preset site group threshold, a similar edge is established between the two nodes to construct a page similarity graph. The result unit is used to perform connected component identification, label propagation, spectral clustering, or community detection processing operations on the page similarity graph to obtain the site clustering result.

[0051] In some specific implementations, the evidence package generation module 13 may specifically include: The result generation unit is used to obtain the collection integrity score, evidence coverage, access stability, anomaly penalty item, and evidence collection cost corresponding to each suspected infringing webpage in the mirror site group; and to perform a lexicographical comparison operation based on the order of the collection integrity score, evidence coverage, access stability, anomaly penalty item, and evidence collection cost to generate the corresponding comparison results. The representative page determination unit is used to determine the suspected infringing webpage that meets the preset optimal conditions in the comparison results as the representative page.

[0052] Furthermore, embodiments of this application also disclose an electronic device, Figure 6 This is a structural diagram of an electronic device 20 according to an exemplary embodiment. The content of the diagram should not be construed as limiting the scope of this application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 stores a computer program, which is loaded and executed by the processor 21 to implement the relevant steps in the webpage processing method based on mirror site clustering and hash solidification disclosed in any of the foregoing embodiments. Furthermore, the electronic device 20 in this embodiment may specifically be a computer.

[0053] In this embodiment, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of this application, and is not specifically limited here; the input / output interface 25 is used to acquire external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs, and is not specifically limited here.

[0054] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, random access memory, disk or optical disk, etc. The resources stored thereon can include operating system 221, computer program 222, etc., and the storage method can be temporary storage or permanent storage.

[0055] The operating system 221 is used to manage and control the various hardware devices on the electronic device 20 and the computer program 222, which may be Windows Server, Netware, Unix, Linux, etc. In addition to including a computer program capable of performing the webpage processing method based on mirror site clustering and hash solidification disclosed in any of the foregoing embodiments, the computer program 222 may further include computer programs capable of performing other specific tasks.

[0056] Furthermore, this application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the aforementioned webpage processing method based on mirror site clustering and hash solidification. Specific steps of this method can be found in the corresponding content disclosed in the foregoing embodiments, and will not be repeated here.

[0057] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to in the method section.

[0058] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0059] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0060] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0061] The technical solutions provided in this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A webpage processing method based on mirror site clustering and hash-based data solidification, characterized in that, include: Based on the evidence collection instructions, access the suspected infringing web pages to obtain the web page information of the suspected infringing web pages, extract multi-dimensional features based on the web page information, and construct the target page features of each of the suspected infringing web pages; Based on the acquired target page features, determine the mirror similarity between any two suspected infringing web pages, and cluster the suspected infringing web pages according to the mirror similarity to obtain the clustering results of each mirror site group including different site group tags; Based on preset rules, representative pages of each of the mirror site groups are determined. A first evidence collection operation is performed on each of the representative pages to obtain first evidence. A second evidence collection operation is performed on pages in each of the mirror site groups that are not representative pages to obtain second evidence. An initial site group association evidence package is generated based on the first evidence, the second evidence, and the site group clustering results. The initial site group association evidence package is hashed and solidified, and the hashed and solidified target site group association evidence package is output to the user terminal to process the suspected infringing web pages based on the target site group association evidence package.

2. The webpage processing method based on mirror site clustering and hash solidification according to claim 1, characterized in that, The webpage information includes DOM snapshots, rendered screenshots, and network request traces; the suspected infringing webpage is a webpage carrying a preset suspected identifier. Accordingly, the step of extracting multi-dimensional features based on the webpage information and constructing target page features for each of the suspected infringing webpages includes: Identify the player container and media interface in the suspected infringing webpage by using the DOM snapshot and / or the network request trajectory, so as to obtain the media content loaded or played by the suspected infringing webpage, generate a video clip hash based on the media content, and generate a player code fingerprint based on the player script in the suspected infringing webpage; The DOM snapshot is compressed to obtain the DOM structure fingerprint, the domain name associated with the suspected infringing webpage is parsed to obtain the corresponding domain name infrastructure characteristics, the network request trajectory is analyzed to generate network request characteristics, and the rendered screenshot is hashed to obtain the corresponding screenshot perception hash. Based on the DOM structure fingerprint, the screenshot-aware hash, the video clip hash, the player code fingerprint, the network request characteristics, and the domain name infrastructure characteristics, the target page characteristics of the suspected infringing webpage are constructed.

3. The webpage processing method based on mirror site clustering and hash solidification according to claim 2, characterized in that, The step of generating a video clip hash based on the media content includes: Keyframe sampling is performed on the media content to obtain the corresponding keyframe summary; The media content is sampled to obtain corresponding audio segment summaries; The media content is segmented and sampled to obtain corresponding media segment summaries; The video segment hash is generated based on the obtained keyframe summary, audio segment summary or media segment summary, and the corresponding sampling time position.

4. The webpage processing method based on mirror site clustering and hash solidification according to claim 2, characterized in that, The process of generating a player code fingerprint based on the player script in the suspected infringing webpage includes: The player script is parsed to obtain an abstract syntax tree summary; Extract the media interface path from the player script to obtain the interface path summary; Extract the key parameter names from the player script to obtain a parameter name summary; Extract the obfuscation features from the player script to obtain a script obfuscation feature summary; The player code fingerprint is generated based on the abstract syntax tree digest, the interface path digest, the parameter naming digest, and the script obfuscation feature digest.

5. The webpage processing method based on mirror site clustering and hash solidification according to claim 2, characterized in that, The step of determining the mirror similarity between any two suspected infringing web pages based on the acquired target page features includes: Based on the target page features of the two suspected infringing web pages, the initial similarity is determined on the DOM structure fingerprint, screenshot-aware hash, video clip hash, player code fingerprint, network request features, and domain name infrastructure features, respectively. Count the number of evidence channels that meet the preset credibility conditions in each of the initial similarities; If the number of evidence channels meets a preset quantity condition, then a geometric consistency operation is performed based on the initial similarity that meets the preset credibility condition using a preset mirror similarity calculation formula to obtain the mirror similarity.

6. The webpage processing method based on mirror site clustering and hash solidification according to claim 1, characterized in that, The step of clustering the suspected infringing web pages based on the mirror similarity includes: The suspected infringing webpages are used as graph nodes. When the mirror similarity between two suspected infringing webpages is greater than or equal to a preset site group threshold, a similar edge is established between the two nodes to construct a page similarity graph. Perform connected component identification, label propagation, spectral clustering, or community detection processing on the page similarity graph to obtain the site clustering result.

7. The webpage processing method based on mirror site clustering and hash solidification according to any one of claims 1 to 6, characterized in that, The step of determining the representative page for each of the mirror site groups based on preset rules includes: Obtain the collection integrity score, evidence coverage, access stability, abnormal penalty items, and evidence collection cost for each of the suspected infringing web pages in the mirror site group; A lexicographical comparison operation is performed based on the order of the collection integrity score, the evidence coverage, the access stability, the anomaly penalty item, and the evidence collection cost to generate the corresponding comparison results. The suspected infringing webpages that meet the preset optimal conditions in the comparison results are identified as the representative pages.

8. The webpage processing method based on mirror site clustering and hash solidification according to claim 7, characterized in that, The step of hashing and solidifying the initial site group association evidence package to output the hash-solidified target site group association evidence package includes: For each evidence file in the initial site cluster association evidence package, determine the corresponding file-level hash value; According to the preset evidence category, the file-level hash values ​​of each evidence file belonging to the same preset evidence category are aggregated to generate the corresponding category hash value; The root hash value of the initial site group association evidence package is generated based on the category hash value, and the root hash value is used as a solidified identifier to determine the target site group association evidence package containing the solidified identifier.

9. A webpage processing device based on mirror site clustering and hash solidification, characterized in that, include: The feature construction module is used to access suspected infringing web pages based on evidence collection instructions, obtain web page information of the suspected infringing web pages, extract multi-dimensional features based on the web page information, and construct target page features for each of the suspected infringing web pages. The webpage clustering module is used to determine the mirror similarity between any two suspected infringing webpages based on the acquired target page features, and to cluster the suspected infringing webpages based on the mirror similarity to obtain the clustering results of each mirror site group including different site group tags. The evidence package generation module is used to determine the representative page of each of the mirror site groups based on preset rules, perform a first evidence collection operation on each of the representative pages to obtain first evidence, perform a second evidence collection operation on the pages in each of the mirror site groups that are not the representative pages to obtain second evidence, and generate an initial site group association evidence package based on the first evidence, the second evidence and the site group clustering results. The evidence package solidification module is used to hash and solidify the initial site group associated evidence package, and output the hash-solidified target site group associated evidence package to the user terminal, so as to process the suspected infringing webpage according to the target site group associated evidence package.

10. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the web page processing method based on mirror site clustering and hash solidification as described in any one of claims 1 to 8.

11. A computer-readable storage medium, characterized in that, Used to store computer programs; wherein, when the computer programs are executed by a processor, they implement the web page processing method based on mirror site clustering and hash solidification as described in any one of claims 1 to 8.