Webpage archiving method and device, equipment and medium
By performing content recognition and adaptive algorithm processing on web pages, the problem of preserving dynamic content during web page archiving is solved, and efficient and reliable web page archiving is achieved, which is suitable for approval process management in the financial and medical fields.
Patent Information
- Application Number
- CN202510687309.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-26
- Publication Date
- 2025-09-05
AI Technical Summary
Existing web page archiving methods are difficult to save dynamic web page content efficiently and reliably, and are difficult to closely integrate with specific application processes, resulting in inconvenient archiving processes and complex operations.
By performing content recognition on the web pages to be archived, the dynamic and static web page data are extracted respectively, and after being processed using Base64 and QP encoding, an adaptive algorithm is used to determine the appropriate web page archive format, such as MHTML or WARC format, to generate and store the target archive file.
It achieves accurate archiving and efficient storage of dynamic and static web page content, improves archiving efficiency and information integrity, ensures the transparency and verifiability of the approval process, meets regulatory compliance requirements, and is suitable for approval tracking and historical data backtracking in the financial and medical fields.
Smart Images

Figure CN120596445A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology and can be applied to the financial and medical fields, and in particular to a web page archiving method, device, equipment and medium. Background Art
[0002] Amidst the accelerating digitalization of the financial and healthcare industries, a significant number of business processes have migrated to internet platforms. Web pages have become a crucial vehicle for core processes such as approvals, reimbursements, and compliance reviews. Credit approvals and expense reimbursements in the financial sector, as well as medical procedures and medical insurance settlements, all involve large amounts of dynamic web content and sensitive data. Efficiently and reliably archiving this web content to ensure process integrity, traceability, and compliance has become a pressing challenge.
[0003] Traditional web archiving methods, such as screenshot archiving, can preserve the appearance of web pages, but they have significant shortcomings when dealing with complex web content, especially dynamically generated content. Furthermore, performance issues with screenshot generation can affect archiving efficiency and user experience. Alternatively, web page archiving can be accomplished independently through the browser. However, this approach makes it difficult to closely integrate the archiving of approved content with specific application processes, leading to inconvenience and operational complexity during the archiving process.
[0004] Therefore, it is necessary to propose a web page archiving method that can ensure the integrity of web page content while enhancing the relevance of the archiving process with applications in the financial and medical fields, improving archiving efficiency and ensuring the long-term preservation and compliance of data. Summary of the Invention
[0005] The present invention provides a web page archiving method, apparatus, computer equipment and medium to solve the technical problem that the archiving of approval content is difficult to be closely integrated with the specific application process, which easily causes inconvenience and operational complexity in the archiving process.
[0006] In a first aspect, a webpage archiving method is provided, comprising:
[0007] Acquire a web page to be archived, and perform content recognition on the web page to be archived to obtain target web page data; wherein the target web page data includes dynamic web page data and static web page data;
[0008] Independently encoding the dynamic web page data and the static web page data to obtain dynamic encoded data and static encoded data;
[0009] Analyzing the dynamic coded data and the static coded data by an adaptive algorithm, and determining the webpage archive format corresponding to the target webpage data according to the analysis result;
[0010] According to the web page archive format, a target archive file corresponding to the target web page data is generated, and the target archive file is stored.
[0011] In a second aspect, a web page archiving device is provided, comprising:
[0012] An acquisition module is used to acquire a web page to be archived and perform content recognition on the web page to be archived to obtain target web page data; wherein the target web page data includes dynamic web page data and static web page data;
[0013] An encoding module, configured to encode the dynamic web page data and the static web page data respectively to obtain dynamic encoded data and static encoded data;
[0014] An analysis module, configured to analyze the dynamic coded data and the static coded data using an adaptive algorithm, and determine a web page archive format corresponding to the target web page data based on the analysis result;
[0015] The archiving module is used to generate a target archive file corresponding to the target web page data according to the web page archiving format, and store the target archive file.
[0016] In a third aspect, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above-mentioned web page archiving method when executing the computer program.
[0017] In a fourth aspect, a computer-readable storage medium is provided, the computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the web page archiving method described above. In the solution implemented based on the web page archiving method, apparatus, computer device, and storage medium described above, a web page to be archived can be obtained and content identification performed on the web page to be archived to obtain target web page data; the target web page data includes dynamic web page data and static web page data. Furthermore, the dynamic web page data and static web page data can be independently encoded to obtain dynamic encoded data and static encoded data, and the dynamic encoded data and static encoded data can be analyzed using an adaptive algorithm. The web page archiving format corresponding to the target web page data is determined based on the analysis results. Thus, a target archive file corresponding to the target web page data can be generated based on the web page archiving format and stored. In the present invention, by performing content identification and analysis on the approval web page to be archived, dynamic web page data (such as approval status and real-time data updates) and static web page data (such as fixed policy documents and approval rules) can be accurately extracted. These data are independently encoded and processed using an adaptive algorithm to ensure accurate archiving and efficient storage of dynamic and static content. By determining the appropriate archiving format, approval web pages can be saved in a standardized and structured manner, making key data in the approval process traceable and easily retrievable over the long term. For financial institutions, this not only helps meet regulatory compliance requirements, but also provides important support in future audits and data analysis. In particular, for cross-cycle approval tracking and historical data backtracking, this solution improves archiving efficiency and information integrity, ensuring the transparency and verifiability of the approval process. In the medical field, this solution is also applicable to scenarios such as medical insurance review and treatment plan approval, helping to restore the complete decision-making process and improve the security and compliance management capabilities of medical data. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments of the present invention. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0019] Figure 1 This is a schematic diagram of an application environment of a web page archiving method according to an embodiment of the present invention;
[0020] Figure 2 This is a flow chart of a web page archiving method according to an embodiment of the present invention;
[0021] Figure 3 yes Figure 1 A schematic flow chart of a specific implementation of step S10;
[0022] Figure 4 is a structural diagram of a web page archiving device according to an embodiment of the present invention;
[0023] Figure 5 is a structural diagram of a computer device in one embodiment of the present invention;
[0024] Figure 6 FIG. 2 is another structural diagram of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0025] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0026] The web page archiving method provided by the embodiment of the present invention can be applied in the following situations: Figure 1In an application environment, a client communicates with a server via a network. The server can obtain a web page to be archived through the client, perform content recognition on the web page to be archived, and obtain target web page data; wherein the target web page data includes dynamic web page data and static web page data; independently encode the dynamic web page data and the static web page data to obtain dynamic encoded data and static encoded data; analyze the dynamic encoded data and the static encoded data using an adaptive algorithm, and determine a web page archive format corresponding to the target web page data based on the analysis results; generate a target archive file corresponding to the target web page data according to the web page archive format, store the target archive file, and feed the target archive file back to the client. In the present invention, by performing content recognition and analysis on the approval web page to be archived, dynamic web page data (such as approval status, real-time data updates) and static web page data (such as fixed policy documents, approval rules, etc.) can be accurately extracted. These data are independently encoded and processed by the adaptive algorithm to ensure accurate archiving and efficient storage of dynamic and static content. By determining a suitable archive format, the approval web page can be saved in a standardized and structured manner, making key data in the approval process traceable and easy to retrieve over the long term. For financial institutions, this not only helps to meet regulatory compliance requirements, but also provides important support in future audits and data analysis. In particular, for cross-cycle approval tracking and historical data backtracking, this solution improves archiving efficiency and information integrity, and ensures the transparency and verifiability of the approval process. In the medical field, this solution is also applicable to scenarios such as medical insurance review and treatment plan approval, which helps to restore the complete decision-making process and improve the security and compliance management capabilities of medical data. Among them, the client can be but is not limited to various personal computers, laptops, smart phones, tablets and portable wearable devices. The server can be implemented with an independent server or a server cluster consisting of multiple servers. The present invention is described in detail below through specific embodiments.
[0027] See also Figure 2 As shown, Figure 2 A flowchart of a web page archiving method provided by an embodiment of the present invention includes the following steps:
[0028] S10: Acquire a web page to be archived, and perform content recognition on the web page to be archived to obtain target web page data.
[0029] The target web page data includes dynamic web page data and static web page data.
[0030] First, a target web page may be obtained, where the target web page may include a web page extracted from the Internet through user request, scheduled crawling, or automated means; the web page to be archived may be a news website, a financial website, a government announcement page, etc., which usually contains important information and needs to be preserved for a long time, and this application does not limit this.
[0031] Furthermore, after acquiring the web page, the target web page can be analyzed and identified for content. Specifically, the various types of information in the target web page can be classified, located, and identified to distinguish which are dynamic web page data (which may change over time or due to user operations) and which are static web page data (information that remains unchanged).
[0032] It should be noted that dynamic web page data is used to represent web page data that changes with time, user behavior, or background updates. For example, real-time stock market quotes, commodity prices, user account balances, transaction records, etc. on financial websites. This data is usually loaded dynamically through technologies such as JavaScript and is time-sensitive and changeable. Static web page data is used to represent data that is fixed and unchanging, and usually contains information that already exists when the web page is loaded. For example, text content on a web page (such as articles, announcements), images, tables, fixed links or files, etc. This data does not change with user operations or changes over time.
[0033] Furthermore, in the medical field, dynamic web data can include real-time updates of patient medical records, examination report upload status, medical insurance settlement progress, and physician approval opinions; while static web data can include disease catalogs, medical insurance policy descriptions, standard medication lists, and fixed approval forms. Accurately identifying and classifying this type of information not only improves the record integrity of the medical approval process but also enables efficient information archiving and precise traceability, facilitating future medical dispute resolution, data auditing, and the development of intelligent decision-making support systems.
[0034] By identifying the target web page content, all dynamic web page data and static web page data will be extracted from the target web page to form the target web page data, which provides a basis for subsequent archiving, management and analysis. Figure 3 As shown, in step S10, that is, performing content recognition on the target webpage to obtain corresponding target webpage data, the following steps are included:
[0035] S11: parsing the document object model structure of the target webpage, separating the boundary between the dynamic webpage data and the static webpage data in the target webpage, and obtaining the target webpage after boundary separation.
[0036] S12: Analyze the script and interface of the target web page after the separation boundary to extract the dynamic web page data; and analyze the web page structure of the target web page after the separation boundary to extract the static web page data.
[0037] The dynamic web page data at least includes multimedia data and form data; the static web page data at least includes font files, embedded images and web page text.
[0038] S13: Obtaining the target web page data based on the dynamic web page data and the static web page data.
[0039] In steps S11-S12, by analyzing the document object model structure, it is possible to identify the dynamically loaded and static portions of the target web page, namely, the dynamic web page data and static web page data. Dynamic web page data is typically dynamically loaded via scripts such as JavaScript, while static web page data is content that already exists at the time of loading. This allows for a clear distinction between the two within the target web page, identifying which portions are static and which are dynamically loaded or interactively generated content.
[0040] It should be noted that the Document Object Model (DOM) is a structured representation of a web page, including the hierarchical relationships of web page elements (such as HTML tags, attributes, text nodes, etc.). By parsing the DOM structure of a target web page, the location and organization of each element on the web page can be understood.
[0041] Furthermore, dynamic web page data and static web page data can be extracted. Among them, dynamic web page data is usually loaded through scripts (such as JavaScript) or obtained by communicating with the back-end server through an API interface. For example, real-time data on a web page, form submission, user interaction, etc. These scripts and interface requests can be analyzed to extract dynamic web page data. For example, real-time stock market quotes or form data entered by the user. Static web page data refers to content that already exists when loading and does not rely on user interaction or background updates. Therefore, the HTML structure of the target web page can be analyzed to extract fixed elements. These static web page data include font files, embedded images (such as PNG, JPG files) and web page text (such as article content, titles, paragraphs, etc.). Static web page data is usually directly embedded in HTML or loaded through resources such as CSS files and image files.
[0042] In step S13, the dynamic web page data and static web page data extracted in step S12 can be merged to form complete target web page data. Ultimately, the target web page data obtained will include the key information of all web pages: dynamic web page data (such as multimedia and form content) and static web page data (such as text, images, font files, etc.). This data will serve as the core content of the archive and can be used to generate archive files or perform other data analysis.
[0043] S20: independently encoding the dynamic web page data and the static web page data to obtain dynamic encoded data and static encoded data.
[0044] In some embodiments, the dynamic web page data and the static web page data are independently encoded to obtain dynamic encoded data and static encoded data, including: encoding the non-text content in the dynamic web page data through Base64 encoding to obtain the dynamic encoded data; and encoding the text content in the static web page data through QP encoding to obtain the static encoded data.
[0045] It should be noted that Base64 is an encoding method that converts binary data (such as images or audio files) into printable ASCII characters. It encodes data as a set of characters (including letters, numbers, and symbols), allowing files that originally only stored and transmitted binary data to be processed in text format. Therefore, Base64 encoding can be used to encode non-text content (such as multimedia data such as images and audio) in dynamic web pages, converting this binary data into text format, making it suitable for storage and transmission. QP (Quoted-Printable) encoding is a common text encoding method, commonly used in email and HTTP protocols, and is particularly suitable for encoding text that can contain non-printable characters. The basic concept of QP encoding is to convert non-ASCII characters into the form "=XX", where "XX" is the hexadecimal representation of the character. For example, the character "&" might be encoded as "=26". This method ensures compatibility and readability of text data across different platforms.
[0046] For example, dynamic web page data usually includes real-time updated content, user interaction data, form data, pictures, videos, audio, etc. Since these data are dynamically generated when the web page is loaded, and may contain binary files (such as pictures, videos, etc.), they need to be processed using Base64 encoding to ensure that they can be correctly stored and transmitted during the archiving process. Static web page data includes text content, HTML structure, CSS style sheets, font files, etc. These contents already exist when the web page is loaded, and are generally text or other non-binary data. Plain text in the web page (such as web page titles, paragraph text, etc.) or special characters in HTML tags (such as "&") will be converted into a transmittable format through QP encoding to ensure that they can be correctly interpreted and displayed in different storage or transmission environments.
[0047] It should be understood that the encoding process described above ensures that data is not corrupted during transmission and storage. In particular, when transmitting over a network, text encoding ensures that data is not subject to the limitations of the transmission protocol. Furthermore, encoding can mitigate security issues caused by special characters, particularly preventing the injection of illegal characters during archiving or transmission.
[0048] Therefore, the encoding process in S20 processes dynamic web page data and static web page data separately by using encoding methods suitable for different types of data, making data storage and transmission more stable, secure and reliable.
[0049] S30: Analyzing the dynamic coding data and the static coding data by using an adaptive algorithm, and determining the webpage archive format corresponding to the target webpage data according to the analysis result.
[0050] In some embodiments, the web page archive format includes MHTML format and WARC format, and the dynamic coding data and the static coding data are analyzed by an adaptive algorithm, and the web page archive format corresponding to the target web page data is determined based on the analysis results, including: performing type analysis on the dynamic coding data by the adaptive algorithm to obtain corresponding target type information; wherein the target type information includes multimedia data type and form data type; and performing complexity analysis on the static coding data by the adaptive algorithm to obtain a corresponding complexity value; if the target type information is the multimedia data type, and the complexity data is greater than or equal to the preset complexity value, determining that the web page archive format corresponding to the target web page data is the WARC format; if the target type information is the form data type, and the complexity data is less than the preset complexity value, determining that the web page archive format corresponding to the target web page data is the MHTML format.
[0051] For example, adaptive algorithms can automatically select the most appropriate archiving format based on the characteristics of dynamic and static data by analyzing it. This process can make the best archiving choice based on the specific content and complexity of the target web page, thereby improving archiving efficiency and usability.
[0052] For example, dynamically encoded data usually involves multimedia data, form data, etc. These data are usually loaded through scripts and may change according to user interaction or page changes. The target type information is the data type information obtained after analyzing the dynamically encoded data through an adaptive algorithm. Specifically including: Multimedia data type, which refers to multimedia files such as audio, video, and pictures contained in web pages. This type of data usually needs to be processed separately because they may be very large and in various formats, and the storage requirements are different from other text data. Form data type refers to user input content in web pages, such as search box input, registration form, comment form, etc. This type of data is usually text information and has a relatively fixed structure.
[0053] Static encoded data includes fixed content on web pages, such as text, images, and HTML tags. This data typically does not change over time or with user interaction, making archiving relatively simple. The adaptive algorithm also performs a complexity analysis on static data. This complexity value measures the complexity of static web page content, such as the number of images, the amount of text, and the complexity of the HTML structure. The greater the complexity of the static page, the more complex the archiving method may be, which will also affect the size and structure of the archived file.
[0054] Furthermore, based on the analysis results (type information of dynamically encoded data and complexity value of statically encoded data), the adaptive algorithm can decide which web page archiving format to use. Specifically: WARC format is a format commonly used for web page archiving, suitable for storing large-scale and diverse web page data, including a large amount of dynamic content, images, videos, etc. WARC format archive files can accommodate complex data structures and multimedia files. Therefore, if the dynamically encoded data is of multimedia type and the complexity of the statically encoded data is high (exceeding the preset complexity threshold), the WARC format will be selected for archiving. MHTML is a format that packages web pages and their related resources (such as images, CSS files, etc.) into a single file. It is suitable for storing relatively simple web pages, usually including less dynamic content or form data. If the dynamically encoded data belongs to the form data type and the complexity of the statically encoded data is low (lower than the preset complexity threshold), the MHTML format is selected for archiving.
[0055] Based on the above embodiment, when the target type information is the form data type, it also includes: performing interactivity analysis on the form data in the dynamically encoded data to determine whether it contains interactive functions, the interactive functions including submit buttons and validation rules; if it contains the interactive functions, optimizing the form data through the adaptive algorithm to obtain optimized dynamically encoded data, and updating the target web page data based on the optimized dynamically encoded data to obtain updated target web page data.
[0056] It should be noted that the form data type refers to user-entered content contained in a web page, such as a registration form, login form, search box, etc. Form data can be interactive, allowing users to perform operations such as input, modification, and submission. Adaptive algorithms can determine whether these forms contain interactive functions based on the form data in the dynamically encoded data. Interactive functions refer to the interactive operations between users and web forms, such as the submit button, which means that after the user fills out the form, they click the submit button to submit the data to the server. Validation rules, that is, the data validation function in the form, are used to ensure that the data entered by the user conforms to the predetermined format or requirements (such as email format, password strength, etc.).
[0057] It should be understood that the purpose of interactivity analysis is to identify whether form data contains these interactive features. If a form is interactive, it means that the form is not just a simple text input box, but also involves user interaction with the page (such as submitting data, validating input, etc.). This type of form data is more complex than form data without interactive features and requires special processing when archiving.
[0058] For example, if the form data includes interactive features (such as submit buttons or validation rules), the adaptive algorithm will optimize this data. The purpose of this step is to ensure that interactive features continue to work after archiving by adjusting the storage method of the data. For example, the user input content, buttons, validation rules, etc. in the form are structured to ensure that the behavior and logic of the form data can be effectively preserved and reproduced even in an environment without real-time interaction.
[0059] Furthermore, the optimized dynamically encoded data can update the target web page data in the archive to include the optimized interactive form data. This allows the archive file to more accurately reflect the interactive functionality of the form, ensuring that users can still understand and restore the original interactive behavior of the form when viewing or replaying the page in the future. The optimized dynamically encoded data will be integrated into the target web page data to form the updated target web page data. This step ensures that all form interactivity, logic, and data processing (such as submission and validation rules) have been correctly preserved and updated.
[0060] Furthermore, generating a target archive file corresponding to the target web page data according to the web page archive format and storing the target archive file includes: generating an updated archive file corresponding to the updated target web page data according to the web page archive format and storing the updated archive file.
[0061] For example, based on the content of the target web page, especially the optimized dynamic encoding data and static encoding data, the adaptive algorithm can select a suitable archive format (such as MHTML or WARC). This step ensures that the appropriate format is selected to store the target web page, and the archive format takes into account the complexity, interactivity and data type of the web page. Based on the updated target web page data (including optimized form data), a new archive file will be generated. This file will include the final version of all dynamic content, interactive functions and static content, ensuring that the archived web page can accurately and completely reflect the content and functionality of the original web page. The generated updated archive file will be stored in an appropriate location (such as an archive server, cloud storage or local storage). This step is the final link in web page archiving. The storage process ensures that the data of the target web page can be securely accessed and retrieved in the future.
[0062] S40: generating a target archive file corresponding to the target web page data according to the web page archive format, and storing the target archive file.
[0063] In some embodiments, generating a target archive file corresponding to the target web page data according to the web page archive format and storing the target archive file includes: organizing the dynamic web page data and the static web page data into a first data module and a second data module respectively according to the web page archive format, and creating a first identifier corresponding to the first data module and a second identifier corresponding to the second data module; generating the target archive file according to the first data module and the first identifier, and the second data module and the second identifier; compressing the target archive file, and storing the compressed target archive file.
[0064] For example, during the archive generation process, the target web page data (including dynamic web page data and static web page data) will first be organized and processed to adapt to the selected archive format. Among them, the first data module contains dynamic web page data, such as multimedia files, form data, scripts and API requests. These data are usually more complex and require special storage methods, such as Base64 encoding. The second data module contains static web page data, such as web page text, pictures, CSS, font files, etc. These data structures are relatively fixed and can be stored directly. Through this modular approach, dynamic and static data can be processed separately, so that when archiving, the organizational structure of the archive file is clearer and more scalable.
[0065] Furthermore, each data module (first data module and second data module) will be assigned a unique identifier (such as UUID). These identifiers ensure that the data of different modules can be correctly associated in the archive file and facilitate subsequent data retrieval and access. Based on the modules and identifiers organized as described above, a complete target archive file can be generated. This file contains all the content of the web page, organizes dynamic data and static data in a structured manner, and each module can be accurately referenced by an identifier. The generated target archive file is usually compressed to reduce the file size for easy storage and transmission. Compression can use common compression formats (such as ZIP, GZ, etc.) to reduce the storage space of the archive file and improve storage efficiency.
[0066] In other embodiments, storing the target archive file further includes: encrypting the target archive file to obtain an encrypted archive file; and storing the encrypted archive file in at least one of a cloud storage system, a local storage medium, and a database.
[0067] For example, archive files are encrypted to enhance data security. Archive file encryption is usually performed using common encryption algorithms (such as AES, RSA, etc.) to ensure that even if the archive files are stolen or leaked during storage, they cannot be illegally accessed or tampered with. Encryption ensures the confidentiality and integrity of the target archive files, preventing sensitive data from being leaked or maliciously tampered with. This is especially important for storing web data involving sensitive information (such as financial data, personal information, etc.). The encrypted archive files are stored in a suitable storage medium, such as cloud storage, local storage, or database, to ensure reliable storage and secure access to the archive files.
[0068] As can be seen, in the above solution, by performing content recognition and analysis on archived approval web pages, dynamic web page data (such as approval status and real-time data updates) and static web page data (such as fixed policy documents and approval rules) can be accurately extracted. This data is independently encoded and processed using adaptive algorithms to ensure accurate archiving and efficient storage of both dynamic and static content. By determining an appropriate archiving format, approval web pages can be stored in a standardized and structured manner, making key data from the approval process traceable and easily retrievable over the long term. For financial institutions, this not only helps meet regulatory compliance requirements but also provides important support for future audits and data analysis. In particular, for cross-cycle approval tracking and historical data backtracking, this solution improves archiving efficiency and information integrity, ensuring transparency and verifiability of the approval process. In the healthcare sector, this solution is also applicable to scenarios such as medical insurance review and treatment plan approval, helping to restore the complete decision-making process and enhance medical data security and compliance management capabilities.
[0069] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0070] In one embodiment, a web page archiving device is provided, which corresponds one-to-one to the web page archiving method in the above embodiment. As shown in the figure, the web page archiving device includes an acquisition module 101, an encoding module 102, an analysis module 103, and an archiving module 104. The functional modules are described in detail as follows:
[0071] The acquisition module 101 is used to acquire a web page to be archived and perform content recognition on the web page to be archived to obtain target web page data; wherein the target web page data includes dynamic web page data and static web page data;
[0072] The encoding module 102 is used to encode the dynamic web page data and the static web page data respectively to obtain dynamic encoded data and static encoded data;
[0073] An analysis module 103 is configured to analyze the dynamic encoding data and the static encoding data using an adaptive algorithm, and determine a web page archive format corresponding to the target web page data based on the analysis result;
[0074] The archiving module 104 is configured to generate a target archive file corresponding to the target web page data according to the web page archive format, and store the target archive file.
[0075] The acquisition module 101 is used to parse the document object model structure of the target web page, separate the boundary between the dynamic web page data and the static web page data in the target web page, and obtain the target web page after the boundary is separated; analyze the script and interface of the target web page after the boundary is separated to extract the dynamic web page data; and analyze the web page structure of the target web page after the boundary is separated to extract the static web page data; wherein the dynamic web page data includes at least multimedia data and form data; the static web page data includes at least font files, embedded images and web page text; and obtain the target web page data based on the dynamic web page data and the static web page data.
[0076] The encoding module 102 is used to encode the non-text content in the dynamic web page data using Base64 encoding to obtain the dynamic encoded data; and to encode the text content in the static web page data using QP encoding to obtain the static encoded data.
[0077] The analysis module 103 is used to perform type analysis on the dynamically encoded data through the adaptive algorithm to obtain corresponding target type information; wherein, the target type information includes multimedia data type and form data type; and, perform complexity analysis on the statically encoded data through the adaptive algorithm to obtain a corresponding complexity value; if the target type information is the multimedia data type, and the complexity data is greater than or equal to the preset complexity value, determine that the web page archive format corresponding to the target web page data is the WARC format; if the target type information is the form data type, and the complexity data is less than the preset complexity value, determine that the web page archive format corresponding to the target web page data is the MHTML format.
[0078] The analysis module 103 is further used to: perform an interactive analysis on the form data in the dynamically encoded data to determine whether it contains interactive functions, the interactive functions including a submit button and validation rules; if it contains the interactive functions, optimize the form data through the adaptive algorithm to obtain optimized dynamically encoded data, and update the target web page data based on the optimized dynamically encoded data to obtain updated target web page data; generate a target archive file corresponding to the target web page data according to the web page archive format, and store the target archive file, including: generating an updated archive file corresponding to the updated target web page data according to the web page archive format, and storing the updated archive file.
[0079] The archiving module 104 is configured to organize the dynamic web page data and the static web page data into a first data module and a second data module respectively according to the web page archiving format, and to create a first identifier corresponding to the first data module and a second identifier corresponding to the second data module;
[0080] generating the target archive file according to the first data module and the first identifier, and the second data module and the second identifier;
[0081] The target archive file is compressed, and the compressed target archive file is stored.
[0082] The archiving module 104 is further configured to: encrypt the target archive file to obtain an encrypted archive file; and store the encrypted archive file in at least one of a cloud storage system, a local storage medium, and a database.
[0083] The present invention provides a web page archiving device, which can accurately extract dynamic data (such as approval status, real-time data updates) and static data (such as fixed policy documents, approval rules, etc.) by performing content identification and analysis on the approval web pages to be archived. These data are independently encoded and processed by adaptive algorithms to ensure accurate archiving and efficient storage of dynamic and static content. By determining the appropriate archiving format, the approval web pages can be saved in a standardized and structured manner, so that key data in the approval process can be traced and easily retrieved over the long term. For financial institutions, this not only helps to meet regulatory compliance requirements, but also provides important support in future audits and data analysis. In particular, for cross-cycle approval tracking and historical data backtracking, the solution improves archiving efficiency and information integrity, and ensures the transparency and verifiability of the approval process. In the medical field, the solution is also applicable to scenarios such as medical insurance review and treatment plan approval, which helps to restore the complete decision-making process and improve the security and compliance management capabilities of medical data.
[0084] The specific definition of the web page archiving device can be found in the definition of the web page archiving method above and will not be repeated here. Each module in the above-mentioned web page archiving device can be implemented in whole or in part through software, hardware, or a combination thereof. Each of the above-mentioned modules can be embedded in or independent of the processor in the computer device in hardware form, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the corresponding operations of each of the above modules.
[0085] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 5As shown. The computer device includes a processor, a memory, a network interface and a database connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external client via a network connection. When the computer program is executed by the processor, it realizes a function or step on the server side based on the web page archiving method.
[0086] In one embodiment, a computer device is provided, which may be a client, and its internal structure diagram may be shown in Figure 6. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it realizes a function or step on the client side of a web page archiving method.
[0087] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are performed:
[0088] Acquire a web page to be archived, and perform content recognition on the web page to be archived to obtain target web page data; wherein the target web page data includes dynamic web page data and static web page data;
[0089] Independently encoding the dynamic web page data and the static web page data to obtain dynamic encoded data and static encoded data;
[0090] Analyzing the dynamic coded data and the static coded data by an adaptive algorithm, and determining the webpage archive format corresponding to the target webpage data according to the analysis result;
[0091] According to the web page archive format, a target archive file corresponding to the target web page data is generated, and the target archive file is stored.
[0092] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0093] Acquire a web page to be archived, and perform content recognition on the web page to be archived to obtain target web page data; wherein the target web page data includes dynamic web page data and static web page data;
[0094] Independently encoding the dynamic web page data and the static web page data to obtain dynamic encoded data and static encoded data;
[0095] Analyzing the dynamic coded data and the static coded data by an adaptive algorithm, and determining the webpage archive format corresponding to the target webpage data according to the analysis result;
[0096] According to the web page archive format, a target archive file corresponding to the target web page data is generated, and the target archive file is stored.
[0097] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or computer device can be found in the relevant descriptions of the server side and the client side in the aforementioned method embodiment. To avoid repetition, they will not be described one by one here.
[0098] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchl ink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0099] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0100] The embodiments described above are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention, and should all be included in the scope of protection of the present invention.
Claims
1. A web page archiving method, characterized in that: include: Acquire a web page to be archived, and perform content recognition on the web page to be archived to obtain target web page data; wherein the target web page data includes dynamic web page data and static web page data; Independently encoding the dynamic web page data and the static web page data to obtain dynamic encoded data and static encoded data; Analyzing the dynamic coded data and the static coded data by an adaptive algorithm, and determining the webpage archive format corresponding to the target webpage data according to the analysis result; According to the web page archive format, a target archive file corresponding to the target web page data is generated, and the target archive file is stored.
2. The method according to claim 1, characterized in that The performing content recognition on the target webpage to obtain corresponding target webpage data includes: Parsing the document object model structure of the target webpage, separating the boundary between the dynamic webpage data and the static webpage data in the target webpage, and obtaining the target webpage after the boundary is separated; Analyzing the scripts and interfaces of the target web page after the separation boundary to extract the dynamic web page data; and analyzing the web page structure of the target web page after the separation boundary to extract the static web page data; wherein the dynamic web page data includes at least multimedia data and form data; and the static web page data includes at least font files, embedded images, and web page text; The target web page data is obtained based on the dynamic web page data and the static web page data.
3. The method according to claim 1, characterized in that The independently encoding the dynamic web page data and the static web page data to obtain dynamic encoded data and static encoded data includes: Encoding the non-text content in the dynamic web page data by Base64 encoding to obtain the dynamic encoded data; and The text content in the static web page data is encoded using a QP encoding method to obtain the static encoded data.
4. The method according to claim 1, wherein The web page archive format includes MHTML format and WARC format. The adaptive algorithm is used to analyze the dynamic coding data and the static coding data, and the web page archive format corresponding to the target web page data is determined according to the analysis result, including: Performing type analysis on the dynamically encoded data using the adaptive algorithm to obtain corresponding target type information; wherein the target type information includes multimedia data type and form data type; and performing complexity analysis on the statically encoded data using the adaptive algorithm to obtain a corresponding complexity value; If the target type information is the multimedia data type, and the complexity data is greater than or equal to a preset complexity value, determining that the webpage archive format corresponding to the target webpage data is the WARC format; If the target type information is the form data type, and the complexity data is less than the preset complexity value, it is determined that the web page archive format corresponding to the target web page data is the MHTML format.
5. The method according to claim 4, characterized in that When the target type information is the form data type, the following further includes: Performing an interactive analysis on the form data in the dynamically encoded data to determine whether it contains interactive functions, wherein the interactive functions include a submit button and a validation rule; If the interactive function is included, the form data is optimized by the adaptive algorithm to obtain optimized dynamic coding data, and the target webpage data is updated based on the optimized dynamic coding data to obtain updated target webpage data; Generating a target archive file corresponding to the target web page data according to the web page archive format and storing the target archive file includes: An updated archive file corresponding to the updated target web page data is generated according to the web page archive format, and the updated archive file is stored.
6. The method according to claim 1, wherein Generating a target archive file corresponding to the target web page data according to the web page archive format and storing the target archive file includes: According to the webpage archive format, the dynamic webpage data and the static webpage data are respectively organized into a first data module and a second data module, and a first identifier corresponding to the first data module and a second identifier corresponding to the second data module are created; generating the target archive file according to the first data module and the first identifier, and the second data module and the second identifier; The target archive file is compressed, and the compressed target archive file is stored.
7. The method according to claim 1, characterized in that The storing of the target archive file further includes: Encrypting the target archive file to obtain an encrypted archive file; The encrypted archive file is stored in at least one of a cloud storage system, a local storage medium, and a database.
8. A web page archiving device, characterized in that: include: An acquisition module is used to acquire a web page to be archived and perform content recognition on the web page to be archived to obtain target web page data; wherein the target web page data includes dynamic web page data and static web page data; An encoding module, configured to encode the dynamic web page data and the static web page data respectively to obtain dynamic encoded data and static encoded data; An analysis module, configured to analyze the dynamic coded data and the static coded data using an adaptive algorithm, and determine a web page archive format corresponding to the target web page data based on the analysis result; The archiving module is used to generate a target archive file corresponding to the target web page data according to the web page archiving format, and store the target archive file.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the web page archiving method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the web page archiving method according to any one of claims 1 to 7 are implemented.