HTML-based webpage word cutting annotation generation method, system and equipment

By capturing the DOM node path data and combining the three-dimensional node fingerprint model, the problems of low storage efficiency and insufficient cross-browser compatibility in web word annotations are solved, and an efficient and accurate web page annotation solution is realized to adapt to dynamic page changes.

CN120541324APending Publication Date: 2025-08-26TONGFANG KNOWLEDGE DIGITAL PUBLISHING TECH CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510655659.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-21
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

The existing web page marking annotation methods have problems such as low storage efficiency, poor adaptability of dynamic web pages, lack of fault tolerance mechanisms, and insufficient cross-browser compatibility, which is difficult to meet the actual needs of users.

Method used

By capturing the DOM node path data of the user's word selection, combining the three-dimensional node fingerprint model to generate a unique fingerprint identifier, and storing it in JSON format, it uses the intelligent offset compensation algorithm to achieve high-precision positioning and dynamic adjustment to ensure the accuracy of the label and cross-browser compatibility.

Benefits of technology

It realizes efficient and accurate web page labeling, adapts to dynamic page changes, ensures two-way rendering of the label and the original text, and improves storage efficiency and cross-browser compatibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120541324A_ABST
    Figure CN120541324A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, in particular to an HTML (Hypertext Markup Language)-based webpage word cutting annotation generation method, system and equipment, and aims to solve the problems of low storage efficiency, poor dynamic webpage adaptability, lack of fault tolerance and insufficient cross-browser compatibility of an existing word cutting annotation method. The method comprises the steps of capturing a selected area; obtaining DOM node path data and text offset, and generating positioning data; recursively storing the path data to generate a unique fingerprint identifier; storing the fingerprint data as historical fingerprint data in a JSON format; performing matching and positioning based on the historical fingerprint data; and mapping information in the historical fingerprint data to an HTML page to realize bidirectional rendering. According to the method, the label name, the class name and the hierarchical relationship of the DOM node are recorded through the node path to construct the fingerprint identifier, and cross-platform data exchange is realized through the JSON structure, so that webpage labeling which is efficient, accurate, high in fault tolerance and adaptive to a dynamic page is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data processing, and in particular relates to a method, system and device for generating web page word annotations based on HTML. Background Art

[0002] With the rapid development and widespread application of internet technology, users are increasingly demanding the ability to mark, annotate, or take notes on web content while browsing. These operations allow users to more efficiently organize and mark key information for later review and review, and they can also communicate and collaborate with others by sharing annotated content. However, current traditional web markup and annotation methods have numerous problems and struggle to meet users' actual needs.

[0003] Existing methods for web page word annotation mainly include three methods: screenshot or picture saving, plain text storage, and DOM node path storage. The method based on screenshot or picture saving is easy to operate. Users only need to take a screenshot of the selected area of ​​the web page to save the annotated content. However, this method has obvious disadvantages. First, high-resolution screenshots will take up a lot of storage space. For users who frequently use this function, the storage cost will increase significantly. Second, because the annotated content exists in the form of pictures, text retrieval and editing are impossible, which is not conducive to users' in-depth processing of the annotated content. In addition, under this method, the annotations are independent of the original content and lack dynamic association. Once the web page content is updated, the annotations will not be able to correspond to the new original text location, resulting in invalid annotations.

[0004] The plain text storage method only saves the text content of the selected word, without recording its specific location in the HTML document. Although this method saves storage space, it has serious drawbacks. On the one hand, when users re-view the annotation, it is difficult to accurately locate the original text based on the plain text content, and the context cannot be quickly matched, which affects the user experience. On the other hand, in a dynamic web page environment, such as a web page that uses AJAX to load content, the page structure and content may change in real time, and the plain text annotation is very easy to lose.

[0005] Compared to the first two methods, storage methods based on DOM node paths (such as XPath / CSS Path) achieve persistent annotation by recording the node path of the selected word in the HTML document. However, for dynamic web pages (such as single-page applications (SPAs), their DOM structure frequently changes due to user operations. Relying solely on static node paths can easily lead to invalid annotations. Furthermore, different browsers differ in how they handle selections (Selection API), making the accuracy of such annotations unreliable when used across browsers.

[0006] To sum up, the existing web page word segmentation and tagging technology has problems such as low storage efficiency, poor adaptability to dynamic web pages, lack of fault tolerance mechanism and insufficient cross-browser compatibility. There is an urgent need for a more efficient, stable and compatible web page word segmentation and tagging solution. Summary of the Invention

[0007] In order to solve the above-mentioned problems in the prior art, namely, the problems of low storage efficiency, poor adaptability to dynamic web pages, lack of fault tolerance mechanism and insufficient cross-browser compatibility in the existing web page word marking and annotation technology, the first aspect of the present invention proposes a method for generating web page word marking and annotation based on HTML, which comprises the following steps: S1. Capture the user's word selection on the web page and record the text information of each node; S2. Obtain the DOM node path data of the selected area and the text offset relative to the parent node, dynamically associate the timestamp information, and establish a spatiotemporal association between the content and the page state; S3, recursively storing the DOM node path data, and generating a unique fingerprint identifier in combination with a three-dimensional node fingerprint model; S4. The text information, text offset, timestamp and unique fingerprint identifier of the word selection area are associated and stored as historical fingerprint data in JSON format; S5, matching the current DOM node based on the tag name in the historical fingerprint data, performing multi-dimensional similarity calculation, and then locating the text node of the word selection area; S6. Dynamically adjust the text offset based on an intelligent offset compensation algorithm, and map the information in the historical fingerprint data to the HTML page to achieve two-way rendering of the notes and the original text.

[0008] In some preferred embodiments, the DOM node path data of the selected area is obtained by: S21, determining the XPath location path of the DOM node of the selected area by a document traversal method; S22. Determine the CSS selector path of the DOM node of the selected area through a recursive parent node hierarchy structure; S23. Record the index position of the DOM node of the selected area in the parent container; S24: Combine the XPath location path, CSS selector path and index position, and store them as DOM node path data of the word selection area.

[0009] In some preferred embodiments, a unique fingerprint identifier is generated by: S31, extracting the tag name of the DOM node of the selected area; S32, traversing the attribute set of the DOM node of the selected area, and converting each attribute into a key-value pair structure including a name and a value; S33. Calculate the position index of the target node by traversing the parent container child node list; S34. Generate a unique fingerprint identifier based on the three-dimensional node fingerprint model and in combination with the tag name, attribute set, and location index.

[0010] In some preferred embodiments, the three-dimensional node fingerprint model is: FP 3D =(Tag,Attributes,Structure); Among them, Tag is the tag feature, Attributes is the attribute feature, Structure is the structure feature; fail is the failure probability of the three-dimensional node fingerprint model; ΔD1 is the variation of label features, ΔD2 is the variation of attribute features, and ΔD3 is the variation of structural features; D1 is the sensitivity of label features, D2 is the sensitivity of attribute features, and D3 is the sensitivity of structural features; Structure=(ParentPath,ChildIndex,SiblingHash); Among them, ParentPath is the parent node path, ChildIndex is the child node index, and SiblingHash is the sibling hash.

[0011] In some preferred embodiments, the text information includes the original text value, user annotation information, note content, and original text link, which is used to fully record the context information of the annotated note.

[0012] In some preferred implementations, when storing historical fingerprint data in JSON format, a multi-level hash index is constructed as follows: S41. Use the page URL hash value as the primary index key; S42. Use the prefix of the XPath location path / CSS selector path as the secondary index key; S43. Associate the index key with an annotation identification list; wherein the annotation identification list includes the text information of the word selection area, the text offset, the timestamp and the unique fingerprint identification.

[0013] In some preferred embodiments, multi-dimensional similarity calculation is performed as follows: S51. Obtain a set of nodes of the same type through a document query interface; S52, performing similarity matching on candidate nodes in the node set of the same type with nodes in the historical fingerprint data one by one, and calculating weighted similarity scores of the nodes; The similarity score includes label consistency, attribute similarity and position consistency; S53 , sorting the calculated results in descending order, and selecting nodes with scores greater than a preset similarity threshold as matching results.

[0014] In some preferred embodiments, the text offset is dynamically adjusted based on an intelligent offset compensation algorithm, and the method is as follows: S61: local anchor point matching to determine the offset of the DOM node of the selected area in the sliding window; S62, comparing the sequence with the set of nodes of the same type through a sequence alignment algorithm to determine semantic difference compensation; S63: Determine linear scale compensation, and perform weighted summation with the offset of the text window and the semantic difference compensation to determine the text offset.

[0015] A second aspect of the present invention provides an HTML-based web page annotation generation system, comprising: A data collection module is configured to capture a user's word selection on a web page, record the text information of each node, obtain the DOM node path data of the selection and the text offset relative to the parent node, dynamically associate the timestamp information, and establish a spatiotemporal association between the content and the page state; A data processing module is configured to recursively store the DOM node path data and generate a unique fingerprint identifier in combination with a three-dimensional node fingerprint model; A data storage module configured to associate and store the text information, text offset, timestamp and unique fingerprint identifier of the word selection area into historical fingerprint data in JSON format; A positioning module is configured to match the current DOM node based on the tag name in the historical fingerprint data, perform multi-dimensional similarity calculation, and then locate the text node of the word selection area; The rendering module is configured to dynamically adjust the text offset based on an intelligent offset compensation algorithm, and map the information in the historical fingerprint data to the HTML page to achieve two-way rendering of the notes and the original text.

[0016] A third aspect of the present invention provides an electronic device, comprising: at least one processor; and a memory communicatively connected to at least one of the processors; wherein, The memory stores instructions that can be executed by the processor, and the instructions are used to be executed by the processor to implement the above-mentioned HTML-based web page word annotation generation method.

[0017] Beneficial effects of the present invention: 1) This invention combines DOM node path recursive storage, HTML node fingerprint fault tolerance mechanism and JSON structured storage to ensure high-precision positioning of word selection areas, and realizes an efficient, accurate and dynamic web page annotation solution; 2) The tag name, class name, and hierarchical relationship of DOM nodes are recorded through the node path to construct a fingerprint identifier, and cross-platform data exchange is achieved through a standardized JSON structure. The storage structure adopts a JSON structure and combines it with index storage to reduce memory usage while improving query efficiency. The storage structure and compensation algorithm do not rely on specific browser APIs. Even if the page structure changes, the target area can still be quickly located, and cross-browser compatibility is guaranteed. 3) The text offset is combined with text fingerprint and intelligent offset compensation algorithm to build a fault-tolerant mechanism, which can accurately detect content modification and intelligently correct the selection position; the three-level compensation system forms a complete error correction chain. The first level compensation is local anchor point matching; the second level compensation is for semantic difference compensation; the third level compensation performs linear proportional compensation, and the offset is obtained by weighting the three-level results, and the selection position is dynamically adjusted, which comprehensively improves the system's error correction capabilities in complex scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Other features, objects and advantages of the present application will become more apparent upon reading the detailed description of non-limiting embodiments made with reference to the following drawings: Figure 1 This is a flow chart of a method for generating HTML-based web page annotations in an embodiment of the present invention; Figure 2 is a flow chart of a three-path generation strategy according to an embodiment of the present invention; Figure 3 This is a structural diagram of a unique fingerprint identifier in an embodiment of the present invention; Figure 4 This is a flow chart of marking academic documents in an embodiment of the present invention; Figure 5 It is an architectural diagram of an HTML-based web page word annotation generation system in an embodiment of the present invention. DETAILED DESCRIPTION

[0019] The present application will be further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are intended only to illustrate the relevant invention and are not intended to limit the invention. It should also be noted that, for ease of description, only portions relevant to the relevant invention are shown in the accompanying drawings.

[0020] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in this application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0021] The present invention provides a method for generating web page word marking and annotation based on HTML. This method achieves breakthroughs in data acquisition accuracy, dynamic adaptability, storage efficiency and rendering performance, establishes a complete technical system in the field of web page word marking and annotation, and realizes a web page annotation solution that is efficient, accurate and adaptable to dynamic pages.

[0022] The present invention provides a method for generating web page annotations based on HTML, comprising the following steps: S1. Capture the user's word selection on the web page and record the text information of each node; S2. Obtain the DOM node path data of the selected area and the text offset relative to the parent node, dynamically associate the timestamp information, and establish a spatiotemporal association between the content and the page state; S3, recursively storing the DOM node path data, and generating a unique fingerprint identifier in combination with a three-dimensional node fingerprint model; S4. The text information, text offset, timestamp and unique fingerprint identifier of the word selection area are associated and stored as historical fingerprint data in JSON format; S5, matching the current DOM node based on the tag name in the historical fingerprint data, performing multi-dimensional similarity calculation, and then locating the text node of the word selection area; S6. Dynamically adjust the text offset based on an intelligent offset compensation algorithm, and map the information in the historical fingerprint data to the HTML page to achieve two-way rendering of the notes and the original text.

[0023] In order to more clearly illustrate the HTML-based web page annotation generation method of the present invention, the following Figure 1 Each step in the embodiment of the present invention is described in detail.

[0024] The HTML-based webpage annotation generation method of the first embodiment of the present invention includes steps S1 to S6, each of which is described in detail as follows: S1. Capture the user's word selection on the web page and record the text information of each node.

[0025] Preferably, in this embodiment, the user's word selection on the web page is captured, and based on the DOM change observation interface (such as MutationObserver) and selection operation interface (such as Selection API) provided by the browser, real-time monitoring of user interface interaction events triggering data collection, including but not limited to mouse release events and touch screen end events.

[0026] Preferably, in this embodiment, the selection operation interface supported by the browser is automatically selected through a feature detection mechanism, and the standard interface (such as window.getSelection) is used first, and the backup interface (such as document.selection) is fallen back to for legacy browsers to achieve cross-browser compatibility.

[0027] Preferably, a global DOM tree change listener (such as MutationObserver) is registered. When a change in the child node list is detected, the validity verification process of the annotation data is automatically triggered, forming a DOM change monitoring mechanism to achieve dynamic page processing.

[0028] Preferably, the text information includes the original text value, user annotation information, note content and original text link, which is used to fully record the context information of the annotated note.

[0029] S2. Obtain the DOM node path data of the selected area and the text offset relative to the parent node, dynamically associate the timestamp information, and establish a spatiotemporal association between the content and the page state.

[0030] Preferably, the DOM node path data of the selected area is obtained, and a three-path generation strategy is adopted, such as Figure 2 As shown, it includes the following steps: S21, determining the XPath location path of the DOM node of the selected area by a document traversal method; S22. Determine the CSS selector path of the DOM node of the selected area through a recursive parent node hierarchy structure; S23. Record the index position of the DOM node of the selected area in the parent container; S24: Combine the XPath location path, CSS selector path and index position, and store them as DOM node path data of the word selection area.

[0031] Further preferably, when capturing the target selection, contextual text content of a predetermined length before and after the selection is synchronously extracted to ensure positioning robustness when dynamic content changes.

[0032] As an option, in this embodiment, the predetermined length of the front and back characters is 50 characters.

[0033] S3. Recursively store the DOM node path data and generate a unique fingerprint identifier in combination with a three-dimensional node fingerprint model.

[0034] Preferably, a unique fingerprint identifier is generated, such as Figure 3 As shown, the method is: S31, extracting the tag name (tagName attribute) of the DOM node of the selected area; S32, traversing the attribute set of the DOM node of the selected area, and converting each attribute into a key-value pair structure including a name (name) and a value (value); S33. Calculate the position index of the target node by traversing the parent container child node list; S34. Generate a unique fingerprint identifier based on the three-dimensional node fingerprint model and in combination with the tag name, attribute set, and location index.

[0035] Further preferably, a traditional two-dimensional fingerprint (a two-dimensional model only includes labels + attributes); FP 2D =(Tag,Attributes); The three-dimensional node fingerprint model is: FP 3D =(Tag,Attributes,Structure); Among them, Tag is the tag feature, Attributes is the attribute feature, Structure is the structure feature; fail is the failure probability of the three-dimensional node fingerprint model; ΔD1 is the label feature change, that is, the ratio of the node label name / type to change, ranging from [0, 1]; D1 is the label feature sensitivity, that is, the importance of the label name in matching; ΔD2 is the attribute feature change, that is, the ratio of the node attribute (id / class, etc.) to change, ranging from [0, 1], and D2 is the attribute feature sensitivity, that is, the importance of the attribute in matching; ΔD3 is the structural feature change, that is, the ratio of the change in the parent-child hierarchical relationship / sibling node order; D3 is the structural feature sensitivity, that is, the weight of the node position relationship in matching; Structure=(ParentPath,ChildIndex,SiblingHash); Among them, ParentPath is the parent node path, ChildIndex is the child node index, and SiblingHash is the sibling hash.

[0036] S4. The text information, text offset, timestamp and unique fingerprint identifier of the word selection area are associated and stored as historical fingerprint data in JSON format.

[0037] Preferably, when storing historical fingerprint data in JSON format, a multi-level hash index is constructed as follows: S41. Use the page URL hash value as the primary index key; S42. Use the prefix of the XPath location path / CSS selector path as the secondary index key; S43. Associate the index key with an annotation identification list; wherein the annotation identification list includes the text information of the word selection area, the text offset, the timestamp and the unique fingerprint identification.

[0038] Preferably, the historical fingerprint data in JSON format is stored in a hierarchical structure, including: Target positioning data layer: XPath path, CSS selector path, text offset (starting position and ending position); Annotation content layer: user-entered note text and style configuration (including color, underline type, line width, etc.); Metadata layer: contains unique fingerprint identifier, creation timestamp, and operator information.

[0039] Cross-platform data exchange is achieved through standardized JSON structure, and index optimization strategies (such as URL hash mapping and path prefix index) are supported to improve query efficiency.

[0040] S5. Match the current DOM node based on the tag name in the historical fingerprint data, perform multi-dimensional similarity calculation, and then locate the text node of the word selection area.

[0041] Preferably, multi-dimensional similarity calculation is performed by: S51. Obtain a set of nodes of the same type through the document query interface and filter candidate nodes: C={c|c∈currentDom∧tagName(c)=originalFp.tagName}; Where: currentDom represents the current document DOM tree, tagName(c) represents the tag name of the node on the current document DOM tree; originalFp.tagName is the tag name of the original node (the text node of the word selection area); S52: perform similarity matching on candidate nodes in the node set of the same type with the nodes in the historical fingerprint data one by one, and calculate the similarity score of the nodes by weight: S total =W tag ·S tag +W attr ·S attr +W pos ·S pos ; The similarity score includes the label consistency S tag , attribute similarity S attr and position consistency S pos ;W tag 、W attr 、W pos Represent the weight of each attribute respectively; S53 , sorting the calculated results in descending order, and selecting nodes with scores greater than a preset similarity threshold as matching results.

[0042] Preferably, in this embodiment, the similarity score Score 3D The calculation method is: When the label consistency is completely matched, a 40% basic score is obtained, otherwise it is 0; when calculating attribute similarity, the score is accumulated with a weight of 30% based on the maximum matching ratio of shared attribute values; when calculating position consistency, a 30% weight score is obtained when the index position of the node in the parent container is completely matched, otherwise it is 0; Score 3D =0.4·S tag +0.3·S attr +0.3·S struct ; The structural similarity is: Among them, I (parentMatch) is the parent node matching indicator function, that is, judging whether the candidate node matches the direct parent node of the original node; It is a position matching indicator function, which determines whether the index position of the candidate node in the parent container is consistent with the original node; It is an adjacent node matching indicator function, which determines whether the previous sibling node of the candidate node matches the original node.

[0043] In this embodiment, the similarity threshold is 0.85.

[0044] Further preferably, the tag consistency (40%) is calculated as follows: Earn the full 40% weighted points or 0 points.

[0045] In this embodiment, the example is as follows: the original node is , the candidate nodes are also →Score 40 points.

[0046] Further preferably, the attribute similarity (30%) is calculated as follows: a) Find the common attribute name set Acommon; b) Compare the values ​​of these attributes to see if they are the same; c) Calculate the matching ratio:

[0047] In this embodiment, an example is given below: the original node has three attributes (id, class, data-x); the candidate node matches the id and class → gets (2 / 3)×30=20 points.

[0048] Further preferably, the position consistency (30%) is calculated as follows: Earn the full 30% weighted points or 0 points.

[0049] In this embodiment, an example is given as follows: the original node is the second child element of the parent node; the candidate node is the second child element of the parent node → 30 points.

[0050] S6. Dynamically adjust the text offset based on an intelligent offset compensation algorithm, and map the information in the historical fingerprint data to the HTML page to achieve two-way rendering of the notes and the original text.

[0051] Preferably, the text offset is: Original text: T o ={t 01 ,t 02 ,...,t 0n }; Current text: T c ={t c1 ,t c2 ,...,t cm }; Original offset: δ o ∈[0,n]; Offset after compensation: δ c ∈[0,m].

[0052] Further preferably, the text offset is dynamically adjusted based on an intelligent offset compensation algorithm, and a three-level compensation system is adopted, the method of which is: S61: Local anchor point matching to determine the offset of the DOM node of the selected area in the sliding window Among them, S(T o ,δ o ,w) represents the text T o At position δ o The sliding window with width w is at the offset δ o The window feature vector with width w at S(T c ,j,w) represents the current text T c The feature vector of the window with width w at position j; ∈ is the matching threshold, which can balance the matching accuracy and robustness; ||·||2 indicates that the L2 norm (Euclidean distance) is the core indicator for measuring the similarity of text windows; S62: Compare the sequence with the node set of the same type through a sequence alignment algorithm to determine semantic difference compensation. Among them, A is the set of matching regions found by the sequence alignment algorithm; the op interval object is the matching interval (start, end, length) of the original text; the oc interval object is the matching interval (start, end, length) of the current text; δ o is the original offset; A=align(T o ,T c )={(op1,oc1),...,(op k ,oc k )}; S63: Determine a linear scale compensation, and perform a weighted summation on the linear scale compensation, the text window offset, and the semantic difference compensation to determine a text offset. Among them, n is the original text length; m is the current text length; δ o is the original offset; The three-level results are weighted to obtain: Among them, the weight coefficient α is calculated dynamically: Among them, τ is the empirical threshold; sim(T o ,T c ) is the text similarity calculation (such as cosine similarity or edit distance normalization result).

[0053] Furthermore, in this embodiment, the three-level compensation system is optimized through performance indicators in the following manner: The accuracy is determined as: Where N is the total number of test samples; is the algorithm prediction offset of the i-th sample; is the true offset of the i-th sample; Is an indicator function (when the condition is met = 1, otherwise = 0; Experimental results show that: The time complexity is: T(n,m)=O(min(w 2 ,n)m)+O(nm)+O(1)≈O(nm); Among them, n is the original text length; m is the current text length; w is the sliding window width; after sliding window optimization, the actual complexity is reduced to O(logw·m)).

[0054] In this embodiment, the sequence alignment algorithm is Needleman-Wunsch.

[0055] In this embodiment, the sliding window width w=10, the optimal ∈=2, and τ=0.6.

[0056] Preferably, when rendering a note, if the original HTML node cannot be matched due to a page update, an approximate node search is performed based on the unique fingerprint identifier, and the offset is adjusted to restore the annotation position.

[0057] Preferably, in this embodiment, Figure 4 As shown, based on the above HTML-based web page annotation generation method, annotating an academic document is done as follows: The user selects certain key data in the "Experimental Results" paragraph; the system obtains the original text value and original link of the selected area, and obtains the corresponding DOM node path data and text offset relative to the parent node, dynamically associates the timestamp information, establishes a spatiotemporal association between the content and the page state, and generates a unique fingerprint identifier based on the three-dimensional node fingerprint model; System record: XPath: / / div[@class='article'] / section[3] / p[5]; text offset: 32-45; The user adds a note: "There is a significant difference compared to the control group (p < 0.01)"; the system obtains the annotation information and note content, associates all text information, text offset, timestamp and unique fingerprint identifier of the selected area and stores it as historical fingerprint data in JSON format; Two weeks later, when the page was revised, the system matched the current DOM node based on the tag name in the historical fingerprint data and was still able to accurately locate the corresponding position of the new version of the page; the text offset was dynamically adjusted based on the intelligent offset compensation algorithm, and the "significant difference from the control group (p<0.01)" in the historical fingerprint data was mapped to certain key data in the "Experimental Results" paragraph in the literature.

[0058] Preferably, in this embodiment, based on the above-mentioned HTML-based web page annotation generation method, price comparison of product parameters annotated on different e-commerce websites is performed, and the process is as follows: A user selects "Processor: i7-11800H" on website A. The system obtains the original text value, context, and original link of the selected area, obtains the corresponding DOM node path data, dynamically associates the timestamp information, establishes a spatiotemporal association between the content and the page state, and generates a unique fingerprint identifier based on a three-dimensional node fingerprint model. The system associates and stores all the text information, timestamp, and unique fingerprint identifier of the selected area as historical fingerprint data in JSON format. The system automatically records the CSS path: #specs>ul>li:nth-child(2); When a user opens website B, the system matches the content based on the tag names in the historical fingerprint data, automatically matches the same configuration parameter location on website B, and generates a price comparison note: "The price on website A is 15% lower than that on website B."

[0059] Preferably, in this embodiment, the process of taking notes on video subtitles based on the above-mentioned HTML-based web page word annotation generation method is as follows: The user selects a key concept in the video subtitles; the system obtains the timestamp (01:23:45) and text content of the selected area, obtains the corresponding DOM node path data, dynamically associates the timestamp information, establishes a spatiotemporal association between the content and the page state, automatically associates the video progress bar position, and generates a unique fingerprint identifier based on a three-dimensional node fingerprint model; all text information, timestamps, and unique fingerprint identifiers of the selected area are associated and stored as historical fingerprint data in JSON format; When the user clicks on the notes during review, the system matches based on the historical fingerprint data and jumps directly to the corresponding position in the video.

[0060] Although the various steps in the above embodiment are described in the above-mentioned order, those skilled in the art will understand that in order to achieve the effect of this embodiment, different steps do not have to be executed in such an order. They can be executed simultaneously (in parallel) or in a reverse order. These simple changes are within the scope of protection of the present invention.

[0061] The second embodiment of the present invention is a web page word annotation generation system based on HTML, such as Figure 5 Shown, including: A data collection module is configured to capture a user's word selection on a web page, record the text information of each node, obtain the DOM node path data of the selection and the text offset relative to the parent node, dynamically associate the timestamp information, and establish a spatiotemporal association between the content and the page state; A data processing module is configured to recursively store the DOM node path data and generate a unique fingerprint identifier in combination with a three-dimensional node fingerprint model; A data storage module configured to associate and store the text information, text offset, timestamp and unique fingerprint identifier of the word selection area into historical fingerprint data in JSON format; A positioning module is configured to match the current DOM node based on the tag name in the historical fingerprint data, perform multi-dimensional similarity calculation, and then locate the text node of the word selection area; The rendering module is configured to dynamically adjust the text offset based on an intelligent offset compensation algorithm, and map the information in the historical fingerprint data to the HTML page to achieve two-way rendering of the notes and the original text.

[0062] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working process and related instructions of the system described above can refer to the corresponding process in the aforementioned method embodiment and will not be repeated here.

[0063] It should be noted that the HTML-based web page word annotation generation system provided in the above embodiment is only illustrated by the division of the above functional modules. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the modules or steps in the embodiments of the present invention can be decomposed or combined. For example, the modules in the above embodiment can be combined into one module, or further divided into multiple sub-modules to complete all or part of the functions described above. The names of the modules and steps involved in the embodiments of the present invention are only for distinguishing the modules or steps and are not regarded as improper limitations on the present invention.

[0064] An electronic device according to a third embodiment of the present invention includes: at least one processor; and a memory communicatively connected to at least one of the processors; wherein, The memory stores instructions that can be executed by the processor, and the instructions are used to be executed by the processor to implement the above-mentioned HTML-based web page word annotation generation method.

[0065] A fourth embodiment of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to be executed by the computer to implement the above-mentioned method for generating HTML-based web page word annotations.

[0066] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes and related instructions of the electronic device and computer-readable storage medium described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0067] Those skilled in the art should be able to appreciate that, in conjunction with the modules and method steps of each example described in the embodiments disclosed herein, it is possible to implement them with electronic hardware, computer software, or a combination of the two, and the programs corresponding to the software modules and method steps can be placed in random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disks, removable disks, CD-ROMs, or any other form of storage medium known in the art. In order to clearly illustrate the interchangeability of electronic hardware and software, the composition and steps of each example have been generally described in terms of function in the above description. Whether these functions are performed in electronic hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0068] Computer program code for performing the operations of the present application can be written in one or more programming languages, or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0069] The flow charts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, program segment or a part of code, and the module, program segment or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the specified function or operation, or can be implemented by a combination of dedicated hardware and computer instructions.

[0070] The terms "first", "second", etc. are used to distinguish similar objects, rather than to describe or indicate a specific order or sequence.

[0071] The term "comprise," "comprising," or any other similar term is intended to cover non-exclusive inclusion such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed or inherent to such process, method, article, or apparatus.

[0072] Thus far, the technical solutions of the present invention have been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art may make equivalent changes or substitutions to the relevant technical features, and the technical solutions after such changes or substitutions will fall within the scope of protection of the present invention.

Claims

1. A method for generating web page annotations based on HTML, characterized in that: The following steps are involved: S1. Capture the user's word selection on the web page and record the text information of each node; S2. Obtain the DOM node path data of the selected area and the text offset relative to the parent node, dynamically associate the timestamp information, and establish a spatiotemporal association between the content and the page state; S3, recursively storing the DOM node path data, and generating a unique fingerprint identifier in combination with a three-dimensional node fingerprint model; S4. The text information, text offset, timestamp and unique fingerprint identifier of the word selection area are associated and stored as historical fingerprint data in JSON format; S5, matching the current DOM node based on the tag name in the historical fingerprint data, performing multi-dimensional similarity calculation, and then locating the text node of the word selection area; S6. Dynamically adjust the text offset based on an intelligent offset compensation algorithm, and map the information in the historical fingerprint data to the HTML page to achieve two-way rendering of the notes and the original text.

2. The method for generating HTML-based web page annotations according to claim 1, characterized in that: Get the DOM node path data of the selected area using the following method: S21, determining the XPath location path of the DOM node of the selected area by a document traversal method; S22. Determine the CSS selector path of the DOM node of the selected area through a recursive parent node hierarchy structure; S23. Record the index position of the DOM node of the selected area in the parent container; S24: Combine the XPath location path, CSS selector path and index position, and store them as DOM node path data of the word selection area.

3. The method for generating HTML-based web page annotations according to claim 1, wherein: Generate a unique fingerprint identifier by: S31, extracting the tag name of the DOM node of the selected area; S32, traversing the attribute set of the DOM node of the selected area, and converting each attribute into a key-value pair structure including a name and a value; S33. Calculate the position index of the target node by traversing the parent container child node list; S34. Generate a unique fingerprint identifier based on the three-dimensional node fingerprint model and in combination with the tag name, attribute set, and location index.

4. The method for generating HTML-based web page annotations according to claim 3, characterized in that: The three-dimensional node fingerprint model is: FP 3D =(Tag,Attributes,Structure); Among them, Tag is the tag feature, Attributes is the attribute feature, Structure is the structure feature; fail is the failure probability of the three-dimensional node fingerprint model; ΔD1 is the variation of label features, ΔD2 is the variation of attribute features, and ΔD3 is the variation of structural features; D1 is the sensitivity of label features, D2 is the sensitivity of attribute features, and D3 is the sensitivity of structural features; Structure=(ParentPath,ChildIndex,SiblingHash); Among them, ParentPath is the parent node path, ChildIndex is the child node index, and SiblingHash is the sibling hash.

5. The method for generating HTML-based web page annotations according to claim 1, wherein: The text information includes the original text value, user annotation information, note content and original text link, which is used to fully record the context information of the annotated note.

6. The method for generating HTML-based web page annotations according to claim 1, wherein: When storing historical fingerprint data in JSON format, a multi-level hash index is constructed as follows: S41. Use the page URL hash value as the primary index key; S42. Use the prefix of the XPath location path / CSS selector path as the secondary index key; S43. Associate the index key with an annotation identification list; wherein the annotation identification list includes the text information of the word selection area, the text offset, the timestamp and the unique fingerprint identification.

7. The method for generating HTML-based web page annotations according to claim 1, wherein: The method for multi-dimensional similarity calculation is: S51. Obtain a set of nodes of the same type through a document query interface; S52, performing similarity matching on candidate nodes in the node set of the same type with nodes in the historical fingerprint data one by one, and calculating weighted similarity scores of the nodes; The similarity score includes label consistency, attribute similarity and position consistency; S53 , sorting the calculated results in descending order, and selecting nodes with scores greater than a preset similarity threshold as matching results.

8. The method for generating HTML-based web page annotations according to claim 7, characterized in that: The text offset is dynamically adjusted based on the intelligent offset compensation algorithm. The method is as follows: S61: local anchor point matching to determine the offset of the DOM node of the selected area in the sliding window; S62, comparing the sequence with the set of nodes of the same type through a sequence alignment algorithm to determine semantic difference compensation; S63: Determine linear scale compensation, and perform weighted summation with the offset of the text window and the semantic difference compensation to determine the text offset.

9. A web page annotation generation system based on HTML, characterized in that: The system includes: A data collection module is configured to capture a user's word selection on a web page, record the text information of each node, obtain the DOM node path data of the selection and the text offset relative to the parent node, dynamically associate the timestamp information, and establish a spatiotemporal association between the content and the page state; A data processing module is configured to recursively store the DOM node path data and generate a unique fingerprint identifier in combination with a three-dimensional node fingerprint model; A data storage module configured to associate and store the text information, text offset, timestamp and unique fingerprint identifier of the word selection area into historical fingerprint data in JSON format; A positioning module is configured to match the current DOM node based on the tag name in the historical fingerprint data, perform multi-dimensional similarity calculation, and then locate the text node of the word selection area; The rendering module is configured to dynamically adjust the text offset based on an intelligent offset compensation algorithm, and map the information in the historical fingerprint data to the HTML page to achieve two-way rendering of the notes and the original text.

10. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to at least one of the processors; wherein, The memory stores instructions that can be executed by the processor, and the instructions are used to be executed by the processor to implement the HTML-based web page word annotation generation method according to any one of claims 1-8.

Citation Information

Cited By

  • Directory information extraction method and device, equipment and medium

    CN121145797A

  • Webpage element positioning method and device, electronic equipment and storage medium

    CN121542527A

  • Webpage element positioning method and device, electronic equipment, and storage medium

    CN121542527B