Webpage data visualization method, device, equipment, medium and program product

By using different styles to display data asset identifiers and metadata information in web documents, the problem of low efficiency in displaying data asset information in web documents is solved, enabling cross-platform metadata display without page switching and improving the visualization efficiency of data asset information.

CN122132639APending Publication Date: 2026-06-02BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING ZITIAO NETWORK TECH CO LTD
Filing Date
2026-02-11
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

In enterprise-level data analytics and collaboration scenarios, the display efficiency of data asset information in web documents is low. Existing technologies require switching pages to obtain data asset information, resulting in low efficiency.

Method used

By displaying the identifiers and other information of data assets in different styles in web page documents, and displaying their metadata information when the interaction point is located in the data asset identifier area, metadata can be viewed without switching pages, thus achieving cross-platform metadata information display.

Benefits of technology

It improves the efficiency of displaying data asset information in web documents, enabling users to intuitively understand the location and related information of data assets on the web, and reducing page switching operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122132639A_ABST
    Figure CN122132639A_ABST
Patent Text Reader

Abstract

This invention relates to the field of data visualization technology and discloses a method, apparatus, device, medium, and program product for visualizing web page data. The method displays a web page document containing at least one identifier for a data asset and other information. The identifier for the data asset is displayed using a first style, while the other information is displayed using a second style, which differs from the first style. If the interactive point is located within the area containing the identifier of the first data asset, then the metadata information of the first data asset is displayed. This method improves the efficiency of displaying data asset information in web page documents.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This relates to the field of data visualization technology, specifically to methods, devices, equipment, media, and program products for visualizing web page data. Background Technology

[0002] In enterprise-level data analytics and collaboration scenarios, data analysts, data developers, and business analysts frequently need to view or reference data asset information in various web documents. Related technologies typically require switching to the corresponding page to access the data asset information. This approach results in low efficiency in displaying data asset information within web documents. Summary of the Invention

[0003] A method, apparatus, device, medium, and program product for visualizing web page data, to solve the problem of low display efficiency of data asset information in web page documents.

[0004] Firstly, a method for visualizing web page data includes:

[0005] Display a webpage document, which includes at least one identifier of a data asset and other information. The identifier of the data asset is displayed in a first style, and the other information is displayed in a second style, which is different from the first style. If the location of the interaction point is within the area where the identifier of the first data asset is located, then the metadata information of the first data asset will be displayed.

[0006] Secondly, a device for visualizing web page data includes: A webpage document display module is used to display a webpage document, which includes at least one identifier of a data asset and other information. The identifier of the data asset is displayed in a first style, and the other information is displayed in a second style, which is different from the first style. The metadata information display module is used to display the metadata information of the first data asset if the location of the interaction point is located in the area where the identifier of the first data asset is located.

[0007] Thirdly, an electronic device includes: a memory and a processor, which are communicatively connected to each other. The memory stores computer instructions, and the processor executes the computer instructions to perform the web page data visualization method of the first aspect or any corresponding embodiment described above.

[0008] Fourthly, a computer-readable storage medium storing computer instructions for causing a computer to perform the web page data visualization method of the first aspect or any corresponding embodiment thereof.

[0009] Fifthly, a computer program product includes computer instructions for causing a computer to execute the web page data visualization method described in the first aspect or any corresponding embodiment thereof.

[0010] In some cases, webpage data visualization methods involve displaying a webpage document containing at least one data asset identifier and other information. The data asset identifier is displayed using a first style, while the other information is displayed using a second style, which differs from the first. If the interactive point is located within the area containing the first data asset identifier, then the metadata information of the first data asset is displayed. This method displays at least one data asset identifier in the webpage document using the first style, intuitively representing the location of the data asset identifier. When the interactive point is located within the area containing the first data asset identifier, the metadata information of the first data asset is displayed, allowing the metadata information to be viewed directly within the webpage document without requiring page switching. In other words, this method integrates the metadata information of the webpage document and the data asset, enabling the webpage document to understand the data asset. This achieves cross-platform metadata information display on the webpage, improving the efficiency of displaying data asset information in webpage documents. Attached Figure Description

[0011] To more clearly illustrate the technical solutions in the specific embodiments or related technologies, the accompanying drawings used in the description of the specific embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0012] Figure 1 These are schematic diagrams based on application scenarios under certain conditions; Figure 2 This is a flowchart illustrating methods for visualizing web page data under certain circumstances; Figure 3 This is a first-type illustration based on browser pages under certain circumstances; Figure 4 This is a diagram based on the first window in some situations; Figure 5 This is a second type of illustration based on browser pages under certain circumstances; Figure 6 This is a diagram illustrating methods for visualizing web page data under certain circumstances; Figure 7 This is a schematic diagram illustrating data asset identification under certain circumstances; Figure 8This is a schematic diagram illustrating the batch asynchronous processing of data asset identification under certain circumstances; Figure 9 It is a flowchart illustrating the visualization of web page data under certain circumstances; Figure 10 It is a structural block diagram of a visualization device for web page data under certain circumstances; Figure 11 These are schematic diagrams of the hardware structure of electronic devices under certain conditions. Detailed Implementation

[0013] To make the objectives, technical solutions, and advantages clearer in some cases, the technical solutions in some cases will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments, not all embodiments. Based on the embodiments in some cases, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this solution.

[0014] It is understood that before using the technical solutions disclosed in the various embodiments in certain situations, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in certain situations and their authorization should be obtained in accordance with relevant laws and regulations through appropriate means.

[0015] For example, upon receiving a user's proactive request, a prompt message can be sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware such as electronic devices, applications, servers, or storage media that perform the operation based on the prompt message.

[0016] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0017] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the specific implementation method. Other methods that comply with relevant laws and regulations may also be applied to this implementation method.

[0018] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.

[0019] The terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this document, "multiple" means two or more, unless otherwise explicitly specified.

[0020] As an optional application scenario in some situations, such as Figure 1 As shown, application 101 is installed in terminal device 110, and user 130 can interact with application 101 through terminal device 110 and / or access device of terminal device 110. In some cases, application 101 is a browser application.

[0021] exist Figure 1 In the application scenario shown, if application 101 is active, the terminal device 110 can display the interface 102 of application 101. The interface 102 may include various pages that application 101 can provide, such as interactive pages, settings pages, query pages, etc.

[0022] In some embodiments, terminal device 110 is communicatively connected to server 120 to provide services to application 101. Terminal device 110 may be a mobile terminal, fixed terminal, or portable terminal, etc., including but not limited to mobile phones, desktop computers, laptop computers, multimedia tablets, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. In some embodiments, terminal device 110 may also support any type of interface, and server 120 may be various types of computing systems or servers capable of providing computing power, including but not limited to mainframes, edge computing nodes, computing devices in cloud environments, etc.

[0023] It should be noted that, Figure 1 This is merely an example of an application scenario and does not limit the scope of protection.

[0024] The following description, in conjunction with the accompanying drawings, outlines some embodiments. It should be understood that the pages shown in the drawings are merely examples, and various page designs are possible in practice. The graphic elements on the page may have different arrangements and visual representations; one or more elements may be omitted or replaced, and one or more other elements may also be present, without any limitation herein. Furthermore, the embodiments described below primarily pertain to terminal device 110. It should be understood that the actions described relative to terminal device 110 can be performed by application 101 on terminal device 110, or by application 101 in collaboration with its server (e.g., server 120).

[0025] For example, users view web page data in application 101, including but not limited to online documents, data requirement documents, etc. The web page data contains identifiers for multiple data assets, which are highlighted for easy user understanding. Furthermore, if the interaction point is located in the area where a data asset's identifier is located, the metadata information of that data asset is also displayed, allowing users to view the metadata information of data assets on the web page without switching pages.

[0026] It should be understood that, in some cases, the visualization of web page data can be achieved by installing a plugin in a browser application, enabling the browser application to have the visualization function of web page data as described in some cases; or, it can be achieved by providing a new browser application that integrates the visualization function of web page data as described in some cases.

[0027] This document does not impose any restrictions on the methods for visualizing web page data in certain situations; the specific settings can be configured according to actual needs.

[0028] In some cases, an embodiment of a method for visualizing web page data is provided. It should be noted that the steps shown in the flowcharts in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowcharts, in some cases, the steps shown or described may be performed in a different order than that shown here.

[0029] In some cases, web page data visualization methods can be used on the aforementioned terminal devices. Figure 2 It is a flowchart based on methods for visualizing web page data in certain situations, such as Figure 2 As shown, the process includes the following steps: Step S201: Display the webpage document.

[0030] The web page document includes at least one identifier for a data asset and other information. The identifier for the data asset is displayed in a first style, and the other information is displayed in a second style, which is different from the first style.

[0031] A web page document is used to represent a document displayed in a browser application, such as an online document. A web page document includes at least one identifier for a data asset, as well as other information. The identifier for the data asset can be related to the type of the data asset; for example, if the data asset is a table, the identifier could be the table name; if the data asset is text, the identifier could be the file name; and so on.

[0032] The web page document displays the identifier of the data asset. For example, if the web page document involves the description of multiple data tables, then the identifiers of these multiple data tables are displayed in the web page document. The web page document also displays other information that characterizes information other than the data asset identifier. To distinguish the data asset from other information, a first style is used to display the data asset identifier, and a second style is used to display the other information, with the first style and the second style being different.

[0033] The first style includes, but is not limited to, bolding, highlighting, and color filling. There are no restrictions on the specific form of the first and second styles; they can be set according to actual needs.

[0034] For example, when an online document is opened in a browser application, the identifiers of data assets involved in the online document are highlighted. Based on this, users can intuitively understand the identifiers of data assets in the online document.

[0035] In step S202, if the location of the interaction point is in the area where the identifier of the first data asset is located, then the metadata information of the first data asset is displayed.

[0036] Interaction points can be used to represent the current position of the mouse, and are not limited to the position point corresponding to the interaction operation. That is, the position of the interaction point represents the position covered (hovered) by the mouse.

[0037] Specifically, as described above, the web page document includes the identifier of at least one data asset. By detecting the location of the interaction point, if the location of the interaction point is detected to be in the area where the identifier of a certain data asset is located, then the identifier of that data asset is referred to as the identifier of the first data asset.

[0038] The area where the identifier of the first data asset is located can be the rest of the smallest bounding box corresponding to the identifier of the first data asset; it can also be an area formed by extending outward from the center position of the identifier of the first data asset by a predetermined length; of course, other methods can also be used to define the area where the identifier of the data asset is located, and no restrictions are placed on them here.

[0039] If the location of the interaction point is detected to be within the area identified by the first data asset, the metadata information of the first data asset will be displayed. This metadata information includes, but is not limited to, the generation time of the first data asset, its data source, and a data summary.

[0040] By generating metadata for data assets, for example, a metadata generation model is used. Information related to each data asset and prompts are input into the metadata generation model to obtain the metadata information of the data assets. This data asset-related information includes, but is not limited to, descriptions of the data assets, associations of the data assets, and original information about the data assets.

[0041] The above is just an example of how metadata information is generated, and it does not limit the scope of protection. The specific settings can be made according to actual needs.

[0042] For example, by associating and storing the identifier of a data asset with its metadata information, after determining the identifier of the first data asset, the metadata information corresponding to the identifier of the first data asset can be retrieved using the associated stored content. Accordingly, the metadata information of the identifier of the first data asset is displayed.

[0043] In some alternative implementations, the metadata information of the first data asset may be displayed in a floating window, a sidebar, or a pop-up window, etc. The specific display method is not limited here.

[0044] In some cases, web page data visualization methods display the identifier of at least one data asset in the web page document using a first style, which can intuitively represent the location of the data asset identifier. When the interaction point is located in the area where the first data asset identifier is located, the metadata information of the first data asset is displayed. The metadata information can be viewed in the web page document without switching pages. That is, the metadata information of the web page document and the data asset is integrated, so that the web page document can understand the data asset. This enables cross-platform metadata information display on the web page, improving the display efficiency of data asset information in the web page document.

[0045] In some optional implementations, in step S202 above, displaying the metadata information of the first data asset includes: Step a1: Display the first window in the web page document.

[0046] Step a2: Display the metadata information of the first data asset in the first window.

[0047] Step a3: In response to interaction with the data analysis controls in the first window, display a text description of the first data asset.

[0048] The first window is used to display the metadata information of the data asset. If an interaction point is detected to be located in the area where the identifier of a data asset is located, the first window will be displayed; if an interaction point is detected to be located outside the area where the identifier of the data asset is located, the first window will be hidden.

[0049] For example, such as Figure 3 As shown, a webpage document is displayed on browser page 301. The webpage document includes other information 302 and the identifier of the data asset 303. Figure 3 The display style of the identifier 303 for data assets is different from the display style of other information 302.

[0050] If the location of the interaction point is detected to be in the area where the identifier 303 of the data asset is located, the first window 304 is displayed, and the metadata information of the data asset is displayed in the first window 304.

[0051] Furthermore, the first window also displays data analysis controls, which are used to trigger analysis of the first data asset and obtain a text description of it. This text description can be used to characterize the first data asset, such as its function, origin, and target audience. Optionally, users can interact with the text description to view more details.

[0052] For example, such as Figure 4 As shown, the first window 401 displays the identifier 402 of the first data asset and the data analysis control 403. By interacting with the data analysis control 403, the text description of the first data asset is triggered to be displayed.

[0053] The first window in the web page document displays the metadata information of the first data asset; by setting up a data analysis control in the first window, the analysis of the first data asset is triggered by interacting with the data analysis control, and the corresponding text description is displayed, so as to facilitate the user's understanding of the first data asset.

[0054] In some optional implementations, step S202 above, which involves displaying the metadata information of the first data asset, further includes: Step a4: Display the data source switching control in the first window.

[0055] Step a5: In response to the interaction with the data source switching control, determine the target data source and display the metadata information of the first data asset under the target data source.

[0056] Since data assets with the same identifier may exist in different data sources, a data source switching control is also displayed in the first window to switch the display of metadata information of the first data asset under different data sources.

[0057] The data source switching control can be a drop-down list displaying data sources associated with the primary data asset, or it can be multiple tabs, each corresponding to a data source. Of course, other methods can also be used to display the data source switching control; there are no restrictions on these methods, and the specific settings should be tailored to actual needs.

[0058] Users determine the target data source by interacting with the data source switching control, and accordingly, the metadata information of the first data asset under the target data source is displayed. That is, different target data sources will display different metadata information.

[0059] Since the same data asset identifier may exist in different data sources, the first window displays a data source switching control to switch the display of metadata information under different data sources, making it easy for users to compare metadata information from different data sources without jumping to another page.

[0060] In some alternative implementations, the above-mentioned method for visualizing web page data further includes: Step b1: Display the webpage document in the browser.

[0061] Step b2, in response to the interaction with the first control in the browser page, displays the metadata information of all data assets in the second window of the browser page.

[0062] The webpage document is displayed in the browser. A first control is set up in the browser to trigger the display of metadata information for all data assets within the displayed webpage document. This first control can be displayed in the browser's toolbar or at a specified location on the browser page; no specific restrictions are placed on its location here.

[0063] Users interact with the first control, triggering the display of metadata information for all data assets in the webpage document. Following interaction with the first control, the metadata information for all data assets is then displayed sequentially in a second window of the browser page.

[0064] For example, such as Figure 5 As shown, a first control is displayed on browser page 501. Interacting with the first control triggers the display of a second window, which displays metadata information for two data assets: window 502 corresponding to the metadata information of data asset 1 and window 503 corresponding to the metadata information of data asset 2.

[0065] It should be understood that, Figure 5The amount of metadata information for data assets displayed in the second window is merely an example and does not limit the scope of protection. As described above, the second window displays the metadata information for all data assets in the webpage document. Due to the limited window size of the second window, only a portion of the data asset metadata information can be displayed at a time. A slider can be displayed on the side of the second window, allowing users to interact with the slider to switch the content displayed in the second window, thereby enabling the display of metadata information for each data asset.

[0066] A primary control is set up in the browser page. Interacting with the primary control can trigger the batch display of metadata information of all data assets, further improving the visualization efficiency of web page data.

[0067] In some alternative implementations, the above-mentioned method for visualizing web page data further includes: Step c1, in response to the selection instruction for the identifier of the second data asset in the web page document, displays the data requirement control for the second data asset.

[0068] Step c2: In response to the selection instruction of the sub-control generated for the query statement in the data requirement control, the query statement corresponding to the second data asset is displayed in the third window of the browser page. The query statement is configured to be generated based on the metadata information of the second data asset.

[0069] The system provides the ability to select and process data requirements for data assets within a webpage document. Specifically, the identifier of the selected data asset in the webpage document is referred to as the identifier of the second data asset. After selecting the identifier of the second data asset, the display of the data requirement control for the second data asset is triggered.

[0070] For example, selecting the identifier of the second data asset and right-clicking triggers the display of a control window. This window contains multiple interactive controls, including a data requirement control. Of course, other controls may also be included among these interactive controls; no specific limitations are imposed here, and they can be set according to actual needs.

[0071] The data requirement control also has several subordinate sub-controls, including a query statement generation sub-control. Interacting with the query statement generation sub-control triggers the generation of a query statement for the second data asset. Correspondingly, the query statement for the second data asset is displayed in the third window of the browser page. As described above, the identifier of a data asset is related to its metadata information. Therefore, once the identifier of the second data asset is determined, its metadata information can be obtained. Based on this metadata information, a query statement for the second data asset is generated and displayed in the third window of the browser page.

[0072] It should be understood that the generated query statement can be copied to other documents for code writing in those documents.

[0073] Since the identification of data assets and their corresponding metadata information are integrated on the web page, after selecting the identification of the second data asset, the generation of the query statement can be triggered by the query statement generation sub-control, which facilitates the subsequent writing of query code for the second data asset.

[0074] In some alternative implementations, the above-mentioned method for visualizing web page data further includes: Step d1: Identify and display the identifier of the data asset of the first displayed content in the web page document. The first displayed content is used to represent the content in the web page document that is being displayed in the browser page.

[0075] Step d2: Identify and store the identifier of the data asset of the second displayed content in the web page document. The position of the second displayed content in the web page document is after the position of the first displayed content in the web page document.

[0076] The identification of data assets in web page documents is processed using a segmented asynchronous scanning and recognition method. Since the length of a web page document may exceed the display size of the browser window, a slider is needed to display the entire text in batches. Based on this, a web page document includes a first display content and a second display content, and the position of the second display content in the web page document is the same as the position of the first display content in the web page document.

[0077] The first set of displayed content represents the content of the webpage document currently being displayed in the browser. This content is scanned and identified first to reduce user waiting time. For the second set of displayed content, which is not yet displayed, a separate thread can be used to scan and identify it when the browser is idle, and the identifiers of the identified data assets are stored after identification.

[0078] For example, when displaying web page documents, they are generally browsed from front to back. Accordingly, the data asset identifiers of the currently displayed parts can be identified first; the data asset identifiers of the undisplayed parts can also be identified in parallel and stored after identification.

[0079] Through a segmented recognition mechanism, asynchronous recognition and rendering are achieved in the context of large-scale web page data, so that the web page scanning process does not hinder user operation and reduces the time spent waiting for metadata information to be displayed.

[0080] In some alternative implementations, the above-described method for visualizing web page data further includes: in response to a display instruction for the second display content, retrieving and displaying the identifier of the data asset of the second display content from storage space.

[0081] When browsing the second display content, the identifier of the data asset of the second display content is directly retrieved from the storage space and rendered in the web page document in the first style, so that the identifier of the data asset of the second display content can be displayed in the first style.

[0082] Since the identifiers of the data assets in the second display content have been cached, they can be read directly from the storage space when the second display content is displayed, which improves the display efficiency of the metadata information of the data assets in the second display content.

[0083] In some alternative implementations, the identification of data assets is achieved through the following methods: Step e1: Determine the first data asset identifier based on the node attributes of each node in the document object model of the web page document.

[0084] Step e2: Use more than one filtering method to filter the first data asset identifier to obtain the second data asset identifier.

[0085] Step e3: Complete the identifier of the second data asset to obtain the identifier of at least one data asset.

[0086] The Document Object Model (DOM) of a web page document is used to represent and manipulate the content and structure of documents in markup languages ​​such as Hypertext Markup Language (HTML) and Extensible Markup Language (XML). The DOM parses the document into a tree structure composed of nodes and objects. Based on the node attributes of each node in the DOM, it can initially filter out potential data asset identifiers and obtain the first data asset identifier.

[0087] Based on this, the first data asset identifier is filtered using more than one filtering method to obtain the second data asset identifier. These multiple filtering methods include, but are not limited to, rule matching, whitelist filtering, and model filtering, etc., and are set according to actual needs; no restrictions are placed on them here.

[0088] The selected second data asset identifiers are then augmented to obtain a complete identifier, meaning at least one identifier for a data asset is obtained. This augmentation can be achieved by comparing the identifier's configuration rules to identify any missing parts in each second data asset identifier and then supplementing them according to the rules.

[0089] By identifying data asset identifiers through multiple methods, we can achieve high recall and high accuracy compatibility of data asset identifiers in web page data.

[0090] For example, in some cases, the visualization method for web page data relies on plugins installed in the browser, i.e., a browser plugin system. For example... Figure 6 As shown, the browser plugin system includes a content script module, a background thread module, and a sidebar module. The content script module is used for DOM traversal, identifier recognition, identifier highlighting, and event detection in the webpage document; the background thread is used for backend source data application programming interface (API) requests, cache management, and debouncing synchronization; the sidebar module displays complete metadata information and task links, etc. The task links represent links to tasks belonging to the corresponding data assets, and interacting with these links allows users to view task details.

[0091] Furthermore, the content script module requests identifier verification by interacting with the background worker thread, identifies data asset identifiers by interacting with DOM text nodes, and obtains rendering results by interacting with the webpage user interface (UI) highlighting module, and displays complete metadata information in conjunction with the sidebar module.

[0092] The background worker thread connects to the backend metadata service API, which is used for verifying the authenticity of data asset identifiers and querying metadata information. The backend metadata service API also connects to the database metadata storage, which stores the cached results of verified data asset identifiers to reduce duplicate requests.

[0093] In some alternative implementations, step e2 above includes: Step e21: Perform regular expression matching on the first data asset identifier to obtain candidate data asset identifiers.

[0094] Step e22: Compare the candidate data asset identifier with the identifier database to obtain the second data asset identifier.

[0095] The filtering for the first data asset includes regular expression matching and identification database comparison. Specifically, regular expression matching is performed on the identification of the first data asset to obtain candidate identifications; based on this, these candidate identifications are compared with the identifications in the identification database to filter the candidate identifications and obtain the second data asset identifications.

[0096] For example, if the data asset is a data table, such as Figure 7As shown, a four-layer identification mechanism is used to balance recall and precision. The four-layer identification mechanism includes regular expression identification, whitelist filtering, label completion, and authenticity verification.

[0097] Specifically, the web page document undergoes DOM scanning and text extraction. Candidate table names are identified using regular expressions, resulting in a set of candidate table names. Regular expression identification uses multi-pattern regular expression matching to identify strings that conform to database naming conventions, thus obtaining the candidate table name set.

[0098] Based on this, a whitelist comparison is performed, that is, the whitelist is compared with the built-in database or table list to eliminate common misidentified words and obtain a set of reliable candidates; further, table name completion and multi-database matching are performed, that is, the database name is completed for the table name without a database prefix to obtain a complete set of table names.

[0099] For the completed table names, the existence and status of the tables are verified by calling the authenticity verification API to obtain the final set of valid tables.

[0100] After merging and deduplicating the results from the valid table set, the results are highlighted. If a user hovers over the data and triggers a notification window, the metadata information corresponding to the relevant data asset's identifier is displayed after the notification window.

[0101] Furthermore, non-blocking scanning and piecewise traversal algorithms are used for DOM scanning of web page documents to achieve large-scale DOM text scanning without affecting web page rendering. For example, such as... Figure 8 As shown, the user opens a web page document through a browser application. The content script module scans the web page DOM in chunks. For example, after processing 500 DOM nodes each time, it yields the main thread and then uses a callback function to continue scanning and storing the results when the browser is idle. The scan results are asynchronously sent back to the background worker thread in batches for merging and processing.

[0102] By using a multi-level combined identification method, the accuracy of data asset identification has been improved.

[0103] As a specific application example in some situations, an online document is displayed on a browser page, and the online document involves multiple data tables, such as... Figure 9 As shown, when a user's mouse hovers over a highlighted table name for a stable, preset duration, the cache is checked and the backend is requested to obtain metadata information, including but not limited to primary key, number of fields, average daily row count, latest partition, and latency. The metadata information is then displayed by rendering the content of a floating window. For example, an independent floating window rendering can be created using the Shadow DOM to avoid style conflicts. Furthermore, the floating window contains a multi-database switching control, allowing real-time switching of metadata sources. If the user's mouse leaves the window, the floating window is closed.

[0104] If the web page document is updated, only the newly added nodes are scanned to reduce redundant calculations and keep the page recognition status synchronized in real time.

[0105] Furthermore, after identifying the data asset's identifier, the identifier and its metadata are stored as key-value pairs. The front-end scan results are then merged and sent together to reduce communication overhead.

[0106] It should be understood that, in some cases, the visualization methods for web page data, by establishing a relationship between the identification of data assets and metadata information, can be extended to a wide range of downstream applications, including but not limited to field identification, indicator interpretation, and query statement generation. The specific downstream applications involved are not limited here, and all downstream applications developed on this basis fall within the scope of protection under certain circumstances.

[0107] In some cases, a web page data visualization device is also provided, which is used to implement the above embodiments and preferred embodiments, and will not be repeated as already described. As used below, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.

[0108] A device for visualizing web page data in some situations, such as Figure 10 As shown, it includes: The webpage document display module 1001 is used to display a webpage document, which includes at least one identifier of a data asset and other information. The identifier of the data asset is displayed in a first style, and the other information is displayed in a second style, which is different from the first style.

[0109] The metadata information display module 1002 is used to display the metadata information of the first data asset if the location of the interaction point is located in the area where the identifier of the first data asset is located.

[0110] In some optional implementations, the metadata information display module 1002 includes: The first display unit is used to display the first window in a web page document.

[0111] The second display unit is used to display the metadata information of the first data asset in the first window.

[0112] The third display unit is used to display a text description of the first data asset in response to interaction with the data analysis controls in the first window.

[0113] In some optional implementations, the metadata information display module 1002 further includes: The fourth display unit is used to display the data source switching control in the first window.

[0114] The fifth display unit is used to respond to the interaction with the data source switching control, determine the target data source, and display the metadata information of the first data asset under the target data source.

[0115] In some alternative implementations, the web page data visualization device further includes: The webpage document display module is used to display webpage documents in a browser.

[0116] The first interaction module is used to respond to interactions with the first control on the browser page and display metadata information of all data assets in the second window of the browser page.

[0117] In some alternative implementations, the web page data visualization device further includes: The first response module is used to respond to the selection instruction for the identifier of the second data asset in the web page document and display the data requirement control for the second data asset.

[0118] The second response module is used to respond to the query statement in the data requirement control to generate the selection instruction of the sub-control, and to display the query statement corresponding to the second data asset in the third window of the browser page. The query statement is configured to be generated based on the metadata information of the second data asset.

[0119] In some alternative implementations, the web page data visualization device further includes: The first identification module is used to identify and display the identifier of the data asset of the first displayed content in the web page document. The first displayed content is used to represent the content in the web page document that is being displayed in the browser page.

[0120] The second identification module is used to identify and store the identifier of the data asset of the second displayed content in the web page document. The position of the second displayed content in the web page document is after the position of the first displayed content in the web page document.

[0121] In some alternative implementations, the web page data visualization device further includes: The third response module is used to respond to the display instruction for the second display content by retrieving the identifier of the data asset of the second display content from the storage space and displaying it.

[0122] In some alternative implementations, the identification of data assets is achieved through the following modules: The first identification module is used to determine the first data asset identifier based on the node attributes of each node in the document object model of the web page document.

[0123] The second identification module is used to filter the first data asset identifier using more than one filtering method to obtain the second data asset identifier.

[0124] The identifier completion module is used to complete the identifier of the second data asset to obtain the identifier of at least one data asset.

[0125] In some alternative implementations, the second identification module includes: The matching unit is used to perform regular expression matching on the first data asset identifier to obtain candidate data asset identifiers.

[0126] The comparison unit is used to compare the candidate data asset identifier with the identifier database to obtain the second data asset identifier.

[0127] In some cases, the provided web page data visualization device can execute the above-described web page data visualization method, possessing the corresponding functional modules and beneficial effects for executing the method. Further functional descriptions of the various modules and units described above are the same as in the corresponding embodiments described above, and will not be repeated here.

[0128] Figure 11 This is a schematic diagram of the structure of an electronic device provided in certain situations.

[0129] The following is a detailed reference. Figure 11 This diagram illustrates a structural schematic suitable for implementing an electronic device in certain situations. The electronic device may include a processor (e.g., a central processing unit, graphics processor, etc.) 1101, which can perform various appropriate actions and processes based on a program stored in read-only memory (ROM) 1102 or a program loaded from memory 1108 into random access memory (RAM) 1103. The RAM 1103 also stores various programs and data required for the operation of the electronic device. The processor 1101, ROM 1102, and RAM 1103 are interconnected via a bus 1104. An input / output interface 1105 is also connected to the bus 1104.

[0130] Typically, the following devices can be connected to the input / output interface 1105: input devices 1106 including, for example, a touchscreen, touchpad, keyboard, mouse, camera, microphone, accelerometer, gyroscope, etc.; output devices 1107 including, for example, a liquid crystal display, speaker, vibrator, etc.; memory devices 1108 including, for example, magnetic tape, hard disk, etc.; and communication devices 1109. Communication device 1109 allows electronic devices to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 11 Electronic devices with various devices are shown, but it should be understood that it is not required to implement or have all of the devices shown, and more or fewer devices may be implemented or have instead.

[0131] Specifically, the processes described in the flowchart above can be implemented as computer software programs. For example, some cases include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowchart. In such embodiments, the computer program can be downloaded and installed from a network via communication device 1109, or installed from memory 1108, or installed from ROM 1102. When the computer program is executed by processor 1101, it performs the functions defined in some cases of methods for visualizing web page data.

[0132] Figure 11 The electronic devices shown are merely examples and should not be construed as limiting their functionality or scope of use in any situation.

[0133] In some cases, a computer-readable storage medium is also provided, in which the above-described methods can be implemented in hardware or firmware, or implemented as computer code that can be recorded on a storage medium, or implemented as computer code downloaded via a network and originally stored on a remote storage medium or a non-transitory machine-readable storage medium and then stored on a local storage medium. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium can also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code, which, when accessed and executed by the computer, processor, or hardware, implements the web page data visualization method shown in the above embodiments.

[0134] Some of the above solutions can be applied as computer program products, such as computer program instructions. When executed by a computer, these instructions, through the operation of the computer, can invoke or provide the aforementioned methods and / or technical solutions. Those skilled in the art should understand that the forms in which computer program instructions exist in computer-readable media include, but are not limited to, source files, executable files, and installation package files. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executing the instruction; the computer compiling the instruction and then executing the corresponding compiled program; the computer reading and executing the instruction; or the computer reading and installing the instruction and then executing the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to a computer.

[0135] Although embodiments in some cases have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the above description, and such modifications and variations all fall within the scope defined by the appended claims.

Claims

1. A method for visualizing web page data, comprising: Display a webpage document, which includes at least one identifier of a data asset and other information. The identifier of the data asset is displayed in a first style, and the other information is displayed in a second style, which is different from the first style. If the location of the interaction point is within the area where the identifier of the first data asset is located, then the metadata information of the first data asset will be displayed.

2. The method according to claim 1, wherein displaying the metadata information of the first data asset includes: The first window is displayed in the web page document; The metadata information of the first data asset is displayed in the first window; In response to interaction with the data analysis controls in the first window, a text description of the first data asset is displayed.

3. The method according to claim 2, wherein displaying the metadata information of the first data asset further includes: The data source switching control is displayed in the first window; In response to the interaction with the data source switching control, a target data source is determined, and the metadata information of the first data asset under the target data source is displayed.

4. The method according to claim 1, further comprising: The webpage document is displayed on the browser page; In response to an interaction with a first control on the browser page, metadata information of all the data assets is displayed in a second window of the browser page.

5. The method according to claim 1, further comprising: In response to a selection instruction for the identifier of a second data asset in the web page document, a data request control for the second data asset is displayed; In response to a selection instruction to generate a sub-control for a query statement in the data requirement control, a query statement corresponding to the second data asset is displayed in a third window of the browser page. The query statement is configured to be generated based on the metadata information of the second data asset.

6. The method according to claim 1, further comprising: Identify and display the identifier of the data asset of the first displayed content in the web page document, wherein the first displayed content is used to represent the content of the web page document being displayed in the browser page; Identify and store the identifier of the data asset for the second displayed content in the webpage document, wherein the second displayed content is located after the first displayed content in the webpage document.

7. The method according to claim 6, further comprising: In response to a display instruction for the second display content, the identifier of the data asset of the second display content is retrieved from the storage space and displayed.

8. The method according to claim 1, wherein the identifier of the data asset is identified in the following manner: Based on the node attributes of each node in the document object model of the web page document, the first data asset identifier is determined. The first data asset identifier is filtered using more than one filtering method to obtain the second data asset identifier; The identifier of the second data asset is completed to obtain the identifier of the at least one data asset.

9. The method according to claim 8, wherein filtering the first data asset identifier using more than one filtering method to obtain the second data asset identifier includes: Perform regular expression matching on the first data asset identifier to obtain candidate data asset identifiers; The candidate data asset identifier is compared with the identifier database to obtain the second data asset identifier.

10. A device for visualizing web page data, comprising: A webpage document display module is used to display a webpage document, which includes at least one identifier of a data asset and other information. The identifier of the data asset is displayed in a first style, and the other information is displayed in a second style, which is different from the first style. The metadata information display module is used to display the metadata information of the first data asset if the location of the interaction point is located in the area where the identifier of the first data asset is located.

11. An electronic device, comprising: A memory and a processor are communicatively connected, the memory storing computer instructions, and the processor executing the computer instructions to perform the web page data visualization method according to any one of claims 1 to 9.

12. A computer-readable storage medium storing computer instructions for causing a computer to perform the method for visualizing web page data according to any one of claims 1 to 9.

13. A computer program product comprising computer instructions for causing a computer to perform the method for visualizing web page data according to any one of claims 1 to 9.