A page adaptation method and system

By dividing form areas and extracting feature of pages in the stock system, and obtaining data using positioning identifiers, the problems of high labor costs and low efficiency in the prior art are solved, and efficient page adaptation and mobile device adaptation are achieved.

CN119003917BActive Publication Date: 2025-07-22GUANGDONG YUEDIAN INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411192841.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-28
Publication Date
2025-07-22
Estimated Expiration
2044-08-28

AI Technical Summary

Technical Problem

When the prior art adapts the stock system to a mobile device, the labor cost is high, and the characteristics need to be adjusted manually after the page elements change, resulting in inefficient adaptation.

Method used

By obtaining multiple pages of the same business, dividing form areas and non-form areas, extracting form field features and read-only business fields, using positioning identifiers to obtain page data, and performing rendering adaptations, including preprocessing, area division, feature extraction and data rendering.

Benefits of technology

Improve the efficiency and accuracy of page adaptation, reduce labor costs, and the adapted page conforms to mobile device usage habits.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119003917B_ABST
    Figure CN119003917B_ABST
Patent Text Reader

Abstract

The present invention discloses a page adaptation method and system, including: obtaining a plurality of first pages belonging to a first service, and extracting a form area and a non-form area corresponding to each of the first pages; performing feature analysis on the fields located in the form area to obtain form field features corresponding to the first service; traversing the plurality of first pages based on the non-form area to obtain read-only service fields corresponding to the first service; performing comparison analysis on the plurality of first pages to obtain positioning identifiers corresponding to the first service; obtaining a plurality of second pages corresponding to the first service based on the positioning identifiers; extracting page data corresponding to each of the second pages based on the form field features and the read-only service fields, and rendering the page data to complete page adaptation, thereby improving the efficiency of page adaptation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of web design, and in particular, to a page adaptation method and system. Background Art

[0002] With the popularization of mobile devices, more and more enterprises need to adapt their legacy systems to mobile devices to meet the needs of users. Among them, a large number of systems lack maintenance or are technologically outdated due to the limitations of past technologies, resulting in the inability to adapt the mobile pages through secondary development.

[0003] For these systems, the prior art usually uses a headless browser to simulate the crawler process. Taking the element hierarchy, style, etc. as features, after positioning through an element selector, each field is extracted and then converted into the required format of the interface and returned, so as to perform page adaptation.

[0004] However, this technical means requires manual analysis of each page structure. When there are many pages or elements to be adapted, the labor cost is high. And the features of each field are relatively fixed. When the page changes due to later development or style adjustment, it is necessary to manually adjust the features of the corresponding fields, resulting in extremely low adaptation efficiency. Summary of the Invention

[0005] In order to solve the above technical problems, the present invention discloses a page adaptation method and system for improving the efficiency of page adaptation.

[0006] To achieve the above object, the present invention discloses a page adaptation method, including:

[0007] Obtain a plurality of first pages belonging to a first service, and extract the form area and non-form area corresponding to each of the first pages;

[0008] Perform feature analysis on the fields located in the form area to obtain the form field features corresponding to the first service;

[0009] Traverse the plurality of first pages based on the non-form area to obtain the read-only service fields corresponding to the first service;

[0010] Perform comparison and analysis on the plurality of first pages to obtain the positioning identifiers corresponding to the first service;

[0011] Obtain a plurality of second pages corresponding to the first service based on the positioning identifiers;

[0012] Extract the page data corresponding to each of the second pages based on the form field features and the read-only service fields, and render the page data to complete the adaptation of the page.

[0013] A page adaptation method disclosed by the present invention first processes multiple pages belonging to the same service to extract the common features shared by the multiple pages in the first service, and then extracts the data of the pages according to the common features to complete the adaptation of the pages and improve the efficiency of page adaptation. Among them, first divide the areas in the obtained first page into form areas and non-form areas, find the form fields directly corresponding to the first service according to the form areas, and then extract the form field features of the form fields. Further, since the service also includes fields that are not related to the form but are related to the service, in order to ensure the integrity of page data extraction, traverse all the first pages based on the non-form areas to obtain the read-only service fields corresponding to the first service. Then, in order to adapt all the pages belonging to the same service to the terminal to be adapted, by comparing the first page, obtain the identifier corresponding to the first service, and then find all the second pages according to the identifier, and then use the extracted form field features and the read-only service fields to extract each of the second pages to complete the adaptation of the pages and improve the adaptation efficiency.

[0014] As a preferred example, the obtaining of several first pages belonging to the first service includes:

[0015] Obtain several initial pages belonging to the first service;

[0016] Traverse each of the initial pages and parse the structure of the initial page to delete the useless data in the initial page according to the structure;

[0017] Convert each initial page after deleting the useless data into a DOM tree structure to obtain the first page corresponding to each initial page.

[0018] The present invention deletes the useless data in each first page to reduce the amount of data to be processed subsequently and improve the efficiency of feature extraction. Further, convert each page into a DOM tree structure to analyze the page according to the DOM tree structure and improve the accuracy of analysis.

[0019] As a preferred example, the extraction of the form area and the non-form area corresponding to each of the first pages includes:

[0020] Traverse the DOM tree structure and convert the bottom - layer nodes of the DOM tree structure into corresponding first - weight arrays; wherein, the first - weight array includes the values pointed to by each of the bottom - layer nodes to the tree nodes and the weights of each of the bottom - layer nodes.

[0021] Add adjacent elements in the first - weight array to obtain a second - weight array.

[0022] Obtain the regions where high - weight values continuously appear in the second - weight array and determine these regions as the form regions.

[0023] Based on the DOM tree structure, the present invention uses the bottom - layer nodes, that is, the bottom - layer data, to merge upward one by one, gradually check the form regions in the page, improve the accuracy of form - region division, and further improve the accuracy of page - feature extraction.

[0024] As a preferred example, the analysis of the first fields located in the form region to obtain the form - field features corresponding to the first page includes:

[0025] Obtain the forms in each form region to extract the hierarchical structure corresponding to the fields in the form.

[0026] Parse the name and value of the field according to the hierarchical structure; wherein, the form - field features include the hierarchical structure of each form, the name and value of the field.

[0027] Based on the fact that forms of the same business have the same fields, the present invention parses the fields in the form, obtains the hierarchical structure corresponding to each form and the name and value of the fields in the form, and determines them as form - field features, so as to extract the data of all pages according to the form - field features subsequently and improve the efficiency of page adaptation.

[0028] As a preferred example, the traversal of several first pages based on the non - form region to obtain the read - only service fields corresponding to the first service includes:

[0029] According to the DOM tree structure, map the bottom - layer nodes in each first page to a first array; wherein, the value of each element in the first array is the length of the node text content.

[0030] Calculate the sliding value of the text length corresponding to each non - form region, and successively divide the first array into several intervals according to the sliding value of the text length.

[0031] Calculate the average text length corresponding to each interval and extract the first interval when the average text length is less than a preset length threshold.

[0032] Obtain the common parent node corresponding to the first interval in the DOM tree structure, so as to obtain the read-only service field from the first page according to the common parent node.

[0033] The present invention divides a number of text intervals one by one by using the non-form area and the DOM tree structure, calculates the average text length of each text interval, extracts the area where read-only service fields may exist according to a preset threshold and the average text length of the text, and then extracts the read-only service fields in the area, ensuring the integrity and accuracy of page data extraction.

[0034] As a preferred example, the comparing and analyzing the several first pages to obtain the positioning identifier corresponding to the first service includes:

[0035] Compare and analyze the uniform resource locators of the several first pages;

[0036] Obtain the identifier of the uniform resource locator by using a preset longest prefix matching method, and determine the identifier as the positioning identifier corresponding to the first service.

[0037] The present invention performs the longest prefix matching on the uniform resource locator of the first page, and can obtain the unified identifier of the pages belonging to the first service, so that all the pages belonging to the same service can be obtained according to the identifier subsequently, ensuring the integrity of page adaptation of the service.

[0038] As a preferred example, the rendering the page data to complete the page adaptation includes:

[0039] Perform streaming formatting conversion on the page data;

[0040] Render the page data after conversion according to the engine of the terminal for displaying the page data to complete the adaptation of the page to the terminal.

[0041] The present invention performs streaming formatting conversion on the page data so that the data can be adapted to different terminals, improving the universality of page adaptation. Further, the page data is rendered by using the engine of the terminal, so that the page data is adapted to the terminal to complete the adaptation of the page to the terminal.

[0042] On the other hand, the present invention discloses a page adaptation system, including a region division module, a form extraction module, a field extraction module, a page positioning module, a service query module and a page adaptation module;

[0043] The area division module is used to obtain a plurality of first pages belonging to the first service, and extract the form area and non-form area corresponding to each of the first pages;

[0044] The form extraction module is used to perform feature analysis on the fields located in the form area to obtain the form field features corresponding to the first service;

[0045] The field extraction module is used to traverse a plurality of the first pages based on the non-form area to obtain the read-only service fields corresponding to the first service;

[0046] The page positioning module is used to perform comparison and analysis on a plurality of the first pages to obtain the positioning identifier corresponding to the first service;

[0047] The service query module is used to obtain a plurality of second pages corresponding to the first service based on the positioning identifier;

[0048] The page adaptation module is used to extract the page data corresponding to each of the second pages based on the form field features and the read-only service fields, and render the page data to complete the page adaptation.

[0049] A page adaptation system disclosed by the present invention first obtains multiple pages belonging to the same service for processing to extract the common features of multiple pages in the first service, and then extracts the data of the pages according to the common features to complete the page adaptation and improve the efficiency of page adaptation. Among them, first, the area in the obtained first page is divided into a form area and a non-form area to find the form fields directly corresponding to the first service according to the form area, and then the form field features of the form fields are extracted. Further, since there are also fields in the service that are irrelevant to the form but related to the service, for this, in order to ensure the integrity of page data extraction, all the first pages are traversed based on the non-form area to obtain the read-only service fields corresponding to the first service. Then, in order to adapt all the pages belonging to the same service to the terminal to be adapted, the identifier corresponding to the first service is obtained by comparing the first pages, and then all the second pages are found according to the identifier. Furthermore, each of the second pages is extracted by using the extracted form field features and the read-only service fields to complete the page adaptation and improve the adaptation efficiency.

[0050] As a preferred example, the area division module includes a preprocessing unit and an area screening unit;

[0051] The preprocessing unit is used to obtain a number of initial pages belonging to the first service; traverse each of the initial pages, and parse the structure of the initial pages to delete the useless data in the initial pages according to the structure; convert each initial page after deleting the useless data into a DOM tree structure to obtain the first page corresponding to each initial page.

[0052] The area screening module is used to traverse the DOM tree structure and convert the bottom-layer nodes of the DOM tree structure into corresponding first weight arrays; wherein, the first weight array includes the values pointed to by the tree nodes of each of the bottom-layer nodes and the weights of each of the bottom-layer nodes; add adjacent elements in the first weight array to obtain a second weight array; obtain the areas where high weight values continuously appear in the second weight array, and determine the areas as the form areas.

[0053] The present invention deletes the useless data in each first page to reduce the amount of data to be processed subsequently and improve the efficiency of feature extraction. Further, each page is converted into a DOM tree structure to analyze the page according to the DOM tree structure, improving the accuracy of analysis. Further, based on the DOM tree structure, the bottom-layer nodes, i.e., the bottom-layer data, are merged upward one by one to gradually check the form areas in the page, improving the accuracy of form area division and further improving the accuracy of page feature extraction.

[0054] As a preferred example, the form extraction module includes a hierarchical parsing unit and a field extraction unit;

[0055] The hierarchical parsing unit is used to obtain the forms within each of the form areas to extract the hierarchical structures corresponding to the fields within the forms;

[0056] The field extraction unit is used to parse the names and values of the fields according to the hierarchical structures; wherein, the form field features include the hierarchical structures of each form, the names and values of the fields.

[0057] Based on the fact that the forms of the same service have the same fields, the present invention parses the fields within the forms to obtain the hierarchical structures corresponding to each form and the names and values of the fields within the forms, which are determined as form field features, so as to extract the data of all pages according to the form field features subsequently and improve the efficiency of page adaptation. Description of the Drawings

[0058] Figure 1 : It is a schematic flow chart of a page adaptation method disclosed in an embodiment of the present invention;

[0059] Figure 2: Schematic structural diagram of a page adaptation system disclosed in an embodiment of the present invention;

[0060] Figure 3 : Schematic flow diagram of a page adaptation method disclosed in another embodiment of the present invention;

[0061] Figure 4 : Schematic structural diagram of a form area extraction based on a DOM tree disclosed in another embodiment of the present invention;

[0062] Figure 5 : Schematic structural diagram of a form field extraction based on a DOM tree disclosed in another embodiment of the present invention;

[0063] Figure 6 : Schematic structural diagram of a read-only field extraction based on a DOM tree disclosed in another embodiment of the present invention. Detailed implementation manners

[0064] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0065] Embodiment 1

[0066] This embodiment discloses a page adaptation method. For the specific implementation process of the page adaptation method, please refer to Figure 1 , which mainly includes steps 101 to 106. The steps are mainly as follows:

[0067] Step 101: Obtain a plurality of first pages belonging to a first service, and extract the form area and non-form area corresponding to each of the first pages.

[0068] In this embodiment, this step mainly includes: obtaining a plurality of initial pages belonging to the first service; traversing each of the initial pages and parsing the structure of the initial page to delete the useless data in the initial page according to the structure; converting each initial page after deleting the useless data into a DOM tree structure to obtain the first page corresponding to each of the initial pages.

[0069] Further, traverse the DOM tree structure and convert the bottom - most nodes of the DOM tree structure into corresponding first - weight arrays; wherein, the first - weight array includes the values pointed to by each of the bottom - most nodes to tree nodes and the weights of each of the bottom - most nodes; add adjacent elements in the first - weight array to obtain a second - weight array; obtain the regions where high - weight values continuously appear in the second - weight array, and determine the regions as the form regions.

[0070] In this embodiment, this step deletes the useless data in each first page to reduce the amount of data to be processed subsequently and improve the efficiency of feature extraction. Further, convert each page into a DOM tree structure to analyze the page according to the DOM tree structure and improve the accuracy of analysis. Further, based on the DOM tree structure, use the bottom - most nodes, that is, the bottom - most data, to merge upward one by one, gradually check the form regions in the page, improve the accuracy of form - region division, and thus improve the accuracy of page - feature extraction.

[0071] Step 102: Perform feature analysis on the fields within the form region to obtain the form - field features corresponding to the first service.

[0072] In this embodiment, this step mainly includes: obtaining the forms within each form region to extract the hierarchical structure corresponding to the fields within the forms; parsing the names and values of the fields according to the hierarchical structure; wherein, the form - field features include the hierarchical structure of each form, the names and values of the fields.

[0073] In this embodiment, since the forms of the same service have the same fields, by parsing the fields within the forms, obtain the hierarchical structure corresponding to each form and the names and values of the fields within the forms, and determine them as form - field features, so as to extract the data of all pages according to the form - field features subsequently and improve the efficiency of page adaptation.

[0074] Step 103: Traverse several first pages based on the non - form region to obtain the read - only service fields corresponding to the first service.

[0075] In this embodiment, this step mainly includes: mapping the bottommost nodes in each of the first pages to a first array according to the DOM tree structure; wherein the value of each element in the first array is the length of the node text content. Calculating the text length sliding value corresponding to each of the non-form regions, and sequentially dividing the first array into several intervals according to the text length sliding value; calculating the average text length corresponding to each of the intervals, and extracting the first interval corresponding to when the average text length is less than a preset length threshold; obtaining the common parent node corresponding to the first interval in the DOM tree structure, and obtaining the read-only service fields from the first page according to the common parent node.

[0076] In this embodiment, this step uses the non-form regions and the DOM tree structure to divide several text intervals one by one, and calculates the average text length of each of the text intervals, so as to extract the regions where read-only service fields may exist according to the preset threshold and the average text length, and then extract the read-only service fields in the regions, ensuring the integrity and accuracy of page data extraction.

[0077] Step 104: Compare and analyze several of the first pages to obtain the positioning identifier corresponding to the first service.

[0078] In this embodiment, this step mainly includes: comparing and analyzing the uniform resource locators of several of the first pages; obtaining the identifier of the uniform resource locator through a preset longest prefix matching method, and determining the identifier as the positioning identifier corresponding to the first service.

[0079] In this embodiment, this step performs the longest prefix matching on the uniform resource locators of the first pages, and can obtain the unified identifier of the pages belonging to the first service, so that all the pages belonging to the same service can be obtained according to the identifier later, ensuring the integrity of the page adaptation of the service.

[0080] Step 105: Obtain several second pages corresponding to the first service based on the positioning identifier.

[0081] Step 106: Extract the page data corresponding to each of the second pages based on the form field features and the read-only service fields, and render the page data to complete the page adaptation.

[0082] In this embodiment, this step mainly includes: performing a streaming formatting conversion on the page data; rendering the page data after conversion according to the engine of the terminal for displaying the page data to complete the adaptation between the page and the terminal.

[0083] In this embodiment, this step performs a streaming formatting conversion on the page data so that the data can be adapted to different terminals, improving the universality of page adaptation. Further, the engine of the terminal is used to render the page data, so that the page data is adapted to the terminal, completing the adaptation of the page and the terminal.

[0084] On the other hand, this embodiment also discloses a page adaptation system. For the specific structural composition of the page adaptation system, please refer to Figure 2 , including a region division module 201, a form extraction module 202, a field extraction module 203, a page positioning module 204, a service query module 205, and a page adaptation module 206;

[0085] The region division module 201 is used to obtain a plurality of first pages belonging to the first service, and extract the form region and non-form region corresponding to each of the first pages;

[0086] The form extraction module 202 is used to perform feature analysis on the fields located in the form region to obtain the form field features corresponding to the first service;

[0087] The field extraction module 203 is used to traverse a plurality of the first pages based on the non-form region to obtain the read-only service fields corresponding to the first service;

[0088] The page positioning module 204 is used to perform comparison and analysis on a plurality of the first pages to obtain the positioning identifier corresponding to the first service;

[0089] The service query module 205 is used to obtain a plurality of second pages corresponding to the first service based on the positioning identifier;

[0090] The page adaptation module 206 is used to extract the page data corresponding to each of the second pages based on the form field features and the read-only service fields, and render the page data to complete the adaptation of the page.

[0091] In this embodiment, the region division module 201 includes a preprocessing unit and a region screening unit;

[0092] The preprocessing unit is used to obtain a plurality of initial pages belonging to the first service; traverse each of the initial pages, and parse the structure of the initial page to delete the useless data in the initial page according to the structure; convert each initial page after deleting the useless data into a DOM tree structure to obtain the first page corresponding to each of the initial pages;

[0093] The area screening module is used to traverse the DOM tree structure and convert the bottom - most nodes of the DOM tree structure into corresponding first weight arrays; wherein, the first weight array includes the values pointed to by each of the bottom - most nodes to tree nodes and the weights of each of the bottom - most nodes; add adjacent elements in the first weight array to obtain a second weight array; obtain the areas where high - weight values continuously appear in the second weight array, and determine the areas as the form areas.

[0094] In this embodiment, the form extraction module 202 includes a hierarchical parsing unit and a field extraction unit;

[0095] The hierarchical parsing unit is used to obtain the forms within each of the form areas to extract the hierarchical structures corresponding to the fields within the forms;

[0096] The field extraction unit is used to parse the names and values of the fields according to the hierarchical structures; wherein, the form field features include the hierarchical structures of each form, the names of the fields, and the values of the fields.

[0097] Embodiment Two

[0098] This embodiment further provides a page adaptation method. The specific implementation process of the page adaptation method is as follows with reference to Figure 3 , mainly including steps 301 to 306, and the steps are mainly as follows:

[0099] Step 301: Obtain at least two first pages belonging to the first service, and pre - process the first pages to delete the useless information in the first pages.

[0100] In this embodiment, this step is mainly as follows: Obtain at least two initial pages belonging to the first service; traverse each of the initial pages, and parse the structure of the initial pages, so as to delete the useless data in the initial pages according to the structure, and convert each of the initial pages after deleting the useless data into a DOM tree structure to obtain the first page corresponding to each of the initial pages.

[0101] Specifically, based on form feature parsing, only 2 pages are required to obtain, while for read - only service field feature parsing, since content comparison is involved, at least 2 pages are required. Thus, at least 2 initial pages belonging to the same service are obtained. Further, the page structures belonging to the same service are the same, while the field values in different forms are different. Thus, use a headless browser to open the page, start traversing from the root node, parse the structure of the web page, remove the parts irrelevant to the service, including scripts, comments, navigation areas, etc., and convert the web page into a DOM tree structure.

[0102] Step 302: Analyze each of the first pages from which the useless information has been deleted, and extract the form area of each of the first pages and the field features of the forms within each of the form areas.

[0103] In this embodiment, this step mainly includes: traversing the DOM tree structure to obtain the form area of the first page. Further, obtain the forms within each of the form areas to extract the hierarchical structure corresponding to the fields within the forms, and then parse the names and values of the fields according to the hierarchical structure; wherein, the form field features include the hierarchical structure of each form, the name and value of the field.

[0104] Specifically, analyze a single first page. According to the distribution of form features, find the areas of all forms, and analyze the field features within a single form to find the complete hierarchical structure of the form fields, and further find the parsing methods of the field names and field values. The field features of the form include the hierarchical relationship, tag name, and class. Using these features, the corresponding fields can be directly selected and the content can be extracted.

[0105] Optionally, refer to Figure 4 , which is a schematic structural diagram of a DOM tree corresponding to a first page provided in this embodiment. As Figure 4 shown, starting from the root node, depth-first traverse the entire DOM tree structure until the bottommost node, and then convert the bottommost node into a weight array w[n]. Each array contains 2 elements, one is the value pointing to the tree node, and the other is the weight of the node. If the node is a form element, such as input, select, textarea, the weight of the node is 1, otherwise it is 0.

[0106] Secondly, add adjacent elements of the array to obtain a new array w'[n]. This step is considered because the form contains two areas of name and value. In some designs, the value and name fields appear alternately in the structure. Finally, find the areas where high-weight values continuously appear in w'[n]. These areas often correspond to the business forms in the page.

[0107] In some embodiments of this embodiment, the process of determining the service form is as follows: Define the start subscript ns, end subscript ne, interval length l, cumulative weight sum_w, and average weight w' of the start area. Then, start traversing the array from array subscript 0, find the first element with a non-zero weight and record the start subscript ns, ne. The interval length = ne - ns + 1, and sum_w = w'[ns]. Then continue traversing the array, accumulate the array weight values to sum_w, and calculate the average weight w' until w' < 0.2. At this time, the interval pointed to by ns and ne is the area where the form is located. Repeat the above process of determining the form until all form areas are obtained. After finding all form areas, for each found area, take 3 non-consecutive points from it, generally take ns, ne, and the midpoint of the interval [(ne - ns) / 2], find the corresponding elements through the node pointer, and take the parent nodes pnode_s, pnode_e, pnode_m of these elements respectively until the first common parent node is found, that is, pnode_s = pnoden = pnode_m. And the common parent node of each area is the flag node of the form, that is Figure 4 the form node shown.

[0108] Based on the determined service form, find the corresponding name of the form field. Among them, for the array w[n] generated by traversing the DOM tree, traverse each element. When traversing the element, if the corresponding node is not a form field and the node includes text information (after converting the node to pure text, the length is not 0), record the node subscript value as t and find all text nodes. Further, if there is a DOM structure where S(t2) - S(t1) = S(t3) - S(t2), as Figure 5 shown, that is, the text nodes appear regularly, then the text nodes are the description information of the nearest form node (as shown in the figure, the label pointed to by t1 is the description information of the right input).

[0109] Then, based on the name, find the x-path representation forms of the form field and the name field. Specifically, in a single form, calculate the number of form fields or name fields, denoted as n_field. Add the parent field of the field to the x-path selector. For example, Figure 5 the selector for input in is div.col>input. Continue to count the number of elements that meet this selector format until it is less than n_field. Take the form node found in the form area determination process as the outermost layer to obtain the final x-path format, such as form>div.col>input.

[0110] Step 303: Analyze the first page based on the non-form area of the first page, and extract the read-only fields corresponding to the first page.

[0111] In this embodiment, this step mainly includes: traversing several first pages based on the non-form area to obtain the read-only service fields corresponding to the first service.

[0112] Specifically, since in a business document, in addition to the form that needs to be filled in on the mobile device, there are also parts that do not need to be filled in, such as read-only fields and remarks information, etc. Therefore, it is necessary to find the read-only fields related to the business. During the process of finding the read-only fields, the bottom-level nodes in the DOM tree structure converted from each first page are mapped to an array t[n], and the value of each element is the length of the node text content. Among them, the mapping of the bottom-level nodes to the array can refer to Figure 6 .

[0113] As Figure 6 shown, for the non-form area, calculate the moving average of the text length. Assume the moving length is l, generally l = 5, and calculate the average text length l_avg in the interval from t[n] to t[n + 5]. If l_avg is greater than the threshold, record the point where n + 2 is located as the starting point, and continue to calculate the average length along the array direction until l_avg < the threshold, and record the end point as node_end.

[0114] Find the first common parent node corresponding to n_start, n_md, and n_mid. This node is the area where the read-only content is located; repeat the above steps for each first page. If there are the same feature areas, that is, the x-path representations are the same, but the text contents are different, then the corresponding areas may be the read-only contents related to the business.

[0115] Step 304: Analyze the URL structure of the first page to obtain the URL features corresponding to the first page.

[0116] In this embodiment, this step mainly includes: obtaining the identifiers of the uniform resource locators of all the first pages through a preset longest prefix matching method, and determining the identifiers as the location identifiers corresponding to the first service.

[0117] Specifically, compare and analyze the page URLs of multiple services of the same type, and use the longest prefix matching method to find the identifiers of the URLs, so as to use the field features (forms, read-only fields) for other pages of the same type of service.

[0118] Step 305: Obtain several second pages of the first service based on the URL features, and extract the data of each second page based on the field features and the URL features.

[0119] In this embodiment, this step mainly includes: performing streaming formatting conversion on the page data.

[0120] Specifically, based on the form field features and read-only field features corresponding to the page extracted in the above step, other pages of the same business type are extracted and converted into a streaming JSON format.

[0121] Step 306: Convert the extracted data in format and render the data after format conversion to complete the adaptation of the page.

[0122] Specifically, in this embodiment, the page data after conversion is rendered according to the engine of the terminal for displaying the page data to complete the adaptation of the page to the terminal.

[0123] A page adaptation method disclosed in this embodiment can automatically extract page key fields according to the page structure in combination with the built-in feature library. Only a small amount of manual verification is required to accurately adapt a traditional PC-side page to a mobile-side page. At the same time, compared with directly modifying the style to generate an HTML5 page, since the field extraction in this embodiment is more accurate and a specially customized component library is used, it can be more in line with the usage habits of mobile devices in terms of page style and operation habits.

[0124] The above specific embodiments have further elaborated on the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the protection scope of the present invention. It is particularly pointed out that for those skilled in the art, any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A page adaptation method, characterized in that, Including: Obtain a number of first pages belonging to the first service, and extract the form area and non-form area corresponding to each of the first pages; Conduct feature analysis on the fields located within the form area to obtain the form field features corresponding to the first service; wherein, obtain the forms within each of the form areas to extract the hierarchical structure corresponding to the fields within the forms; parse the names and values of the fields according to the hierarchical structure; wherein, the form field features include the hierarchical structure of each form, the names and values of the fields; Traverse the number of first pages based on the non-form area to obtain the read-only service fields corresponding to the first service; wherein, according to the DOM tree structure corresponding to the first page, map the bottommost nodes in each of the first pages to a first array; wherein, the value of each element in the first array is the length of the node text content; calculate the text length sliding value corresponding to each of the non-form areas, and sequentially divide the first array into a number of intervals according to the text length sliding value; calculate the average text length corresponding to each of the intervals, and extract the first interval when the average text length is less than a preset length threshold; obtain the common parent node corresponding to the first interval in the DOM tree structure, and obtain the read-only service fields from the first page according to the common parent node; Conduct comparison and analysis on the number of first pages to obtain the positioning identifier corresponding to the first service; Obtain a number of second pages corresponding to the first service based on the positioning identifier; Extract the page data corresponding to each of the second pages based on the form field features and the read-only service fields, and render the page data to complete page adaptation.

2. The page adaptation method according to claim 1, characterized in that, The obtaining a number of first pages belonging to the first service includes: Obtain a number of initial pages belonging to the first service; Traverse each of the initial pages and parse the structure of the initial pages to delete the useless data in the initial pages according to the structure; Convert each of the initial pages after deleting the useless data into a DOM tree structure to obtain the first page corresponding to each of the initial pages.

3. The page adaptation method according to claim 2, wherein The extracting the form area and non-form area corresponding to each of the first pages includes: Traverse the DOM tree structure and convert the bottommost nodes of the DOM tree structure into a corresponding first weight array; wherein, the first weight array includes the value of each of the bottommost nodes pointing to the tree nodes and the weight of each of the bottommost nodes; Add adjacent elements in the first weight array to obtain a second weight array; Obtain the area where high weight values continuously appear in the second weight array, and determine the area as the form area.

4. A page adaptation method according to claim 1, characterized in that, The conducting comparison and analysis on the number of first pages to obtain the positioning identifier corresponding to the first service includes: Conduct comparison and analysis on the uniform resource locators of the number of first pages; Obtain the identifier of the uniform resource locator through a preset longest prefix matching method, and determine the identifier as the location identifier corresponding to the first service.

5. A page adaptation method according to claim 1, characterized in that, The rendering of the page data to complete the adaptation of the page includes: Perform a streaming formatting conversion on the page data; Render the converted page data according to the engine of the terminal for displaying the page data to complete the adaptation of the page to the terminal.

6. A page adaptation system, characterized in that, It includes a region division module, a form extraction module, a field extraction module, a page location module, a service query module, and a page adaptation module; The region division module is used to obtain a plurality of first pages belonging to the first service, and extract the form region and non-form region corresponding to each of the first pages; The form extraction module is used to perform feature analysis on the fields located in the form region to obtain the form field features corresponding to the first service; wherein, the form extraction module includes a hierarchical parsing unit and a field extraction unit; the hierarchical parsing unit is used to obtain the forms in each form region to extract the hierarchical structure corresponding to the fields in the form; the field extraction unit is used to parse the name and value of the field according to the hierarchical structure; wherein, the form field features include the hierarchical structure of each form, the name and value of the field; The field extraction module is used to traverse a plurality of the first pages based on the non-form region to obtain the read-only service fields corresponding to the first service; wherein, according to the DOM tree structure corresponding to the first page, map the bottommost nodes in each of the first pages to a first array; wherein, the value of each element in the first array is the length of the node text content; calculate the text length sliding value corresponding to each non-form region, and divide the first array into several intervals in sequence according to the text length sliding value; calculate the average text length corresponding to each interval, and extract the first interval when the average text length is less than a preset length threshold; obtain the common parent node corresponding to the first interval in the DOM tree structure, and obtain the read-only service fields from the first page according to the common parent node; The page location module is used to perform comparison and analysis on a plurality of the first pages to obtain the location identifier corresponding to the first service; The service query module is used to obtain a plurality of second pages corresponding to the first service based on the location identifier; The page adaptation module is used to extract the page data corresponding to each of the second pages based on the form field features and the read-only service fields, and render the page data to complete the adaptation of the page.

7. A page adaptation system according to claim 6, characterized in that, The region division module includes a preprocessing unit and a region screening unit; The preprocessing unit is used to obtain a plurality of initial pages belonging to the first service; traverse each of the initial pages, and parse the structure of the initial pages to delete the useless data in the initial pages according to the structure; Convert each initial page after deleting the useless data into a DOM tree structure to obtain a first page corresponding to each of the initial pages; The area screening module is used to traverse the DOM tree structure and convert the bottom-layer nodes of the DOM tree structure into corresponding first weight arrays; wherein, the first weight array includes the values of each of the bottom-layer nodes pointing to tree nodes and the weights of each of the bottom-layer nodes; add adjacent elements in the first weight array to obtain a second weight array; obtain the area where high weight values continuously appear in the second weight array, and determine the area as the form area.

Citation Information

Patent Citations

  • Page loading method and device, computer equipment and storage medium

    CN109582899A

  • Business data configuration processing method and device, computer equipment and storage medium

    CN111310427A