Multi-keyword accurate matching and marking line drawing method and system based on DOM (Document Object Model)
Through the DOM-based multi-keyword precise matching and identification line drawing method, the problem of precise matching and highlighting of key information in big data articles is solved, and the effect of efficiently positioning and displaying keywords is achieved while retaining the original format of the article.
Patent Information
- Application Number
- CN202411773362.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-04
- Publication Date
- 2025-05-06
AI Technical Summary
In a big data environment, how to accurately match and identify key information for large articles while retaining the original format of the article, especially when facing articles with complex formats, it is difficult for the existing technology to achieve efficient keyword matching and highlighting.
The DOM-based multi-keyword exact matching and identification line drawing method is adopted. By converting the article content into a DOM tree structure, the DOM tree is traversed in depth first, the plain text content is extracted, keyword preprocessing and matching is performed, the string matching is used for precise matching, and the highlighting effect and identification line are achieved by segmenting text nodes and Bezier curve drawing.
It realizes the precise matching of multiple groups of keywords. On the basis of retaining the original format of the article, it accurately locates and highlights the keywords, providing intuitive information positioning methods, and enhancing users' fast positioning capabilities in massive information.
Smart Images

Figure CN119940294A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of front-end development, and in particular relates to a DOM-based multi-keyword precise matching and identification line drawing method and system. Background Art
[0002] With the advent of the big data era, article retrieval and extraction have gradually become an indispensable part of various fields. In various applications, the display of article content has become more and more common. As the amount of information in articles has become huge with the development of big data, the identification, extraction and positioning of key information in articles are extremely important. Faced with the growth of Internet information and the personalization of user needs, how to annotate key information in large articles has become a problem that must be solved.
[0003] In the Internet field under the big model environment, article data will have its own corresponding tags, which are marked by the big model. These tags will have a series of specific keyword groups inside them. We need to solve the problem of accurately matching the corresponding content from the article based on this group of keywords, highlighting the matched content, and marking and connecting it with the tags.
[0004] In the processing of large-scale data, text search is a very common operation method. JavaScript language technology has developed to date and has also provided many methods for text retrieval. If the main body of the article is plain text, then the simplest text replacement method can achieve content highlighting, but big data articles are not as simple as imagined. Big data articles come from a variety of sources, and the articles themselves may be formatted text. For articles with complex formats, such as those containing HTML tags, accurate matching becomes very difficult.
[0005] The biggest difference between searching for plain text and searching for HTML strings is that when the matching keyword is longer than two words, such as "advertisement", for plain text it is "advertisement", but for HTML strings, it may be "advertisement" or "advertisement". tell ”, or it could be “ wide tell "... There are endless possibilities, but no matter what the original string is, it will be "advertisement" when it is presented on the page. As long as it is "advertisement", we need to match it and highlight it, and ensure that its formatting style is not lost.
[0006] In view of this, it is very meaningful to propose a DOM-based multi-keyword precise matching and identification line drawing method and system. Summary of the invention
[0007] In order to better mark the key content of the article and retain the original format of the article to the greatest extent, the present invention provides a DOM-based multi-keyword precise matching and marking line drawing method and system to solve the above-mentioned technical defects.
[0008] In a first aspect, the present invention proposes a DOM-based multi-keyword accurate matching and identification line drawing method, the method comprising the following steps:
[0009] S1, in response to converting the article content to be processed into a DOM tree structure, constructing a DOM tree by creating HTML node elements using a DOM API of JavaScript, so that the article content is represented in a tree structure in memory;
[0010] S2. Depth-first traversal of the constructed DOM tree, starting from the root node using a recursive algorithm, sequentially accessing each node and its child nodes, and when a text node is accessed, recording relevant information of the text node;
[0011] S3, extracting the plain text content from the DOM tree, traversing the text nodes in the DOM tree, splicing the text contents of all the text nodes in order, ignoring HTML tags and other non-text elements in the splicing process, and obtaining a plain text representation of the article for keyword matching operation;
[0012] S4, pre-processing the input keyword group, including removing repeated keywords and ignoring special symbols in the keywords;
[0013] S5. For each keyword preprocessed in step S4, traverse the DOM node content obtained in step S2. During the traversal process, for each DOM node, determine whether it contains the keyword to be matched;
[0014] S6, when a keyword is found in the process of traversing the DOM node in step S5, a string matching algorithm is used to perform an exact match, determine the starting position and the ending position of the keyword in the text node, store the position information, and convert the text match into a position match;
[0015] S7, traversing the matching results obtained in step S6 and the text nodes obtained in step S2, processing them by using a method of splitting the text nodes, and performing highlighting effect processing on each matching result and text node;
[0016] S8, for the text node segmented and processed in step S7, use the replaceChild method of the Node interface in the DOM to replace the original text node with a new node with a highlight style attribute, and during the replacement process, record the ID information of the new node;
[0017] S9, traverse each highlighted node obtained in step S8, and use the bezierCurveTo method of canvas drawing to draw a Bezier curve.
[0018] Preferably, in step S7, a method of splitting text nodes is used for processing, and each matching result and text node is subjected to highlighting effect processing, including: using the splitText method of the Text node in the DOM to split the text node containing the keyword into two or more independent brother nodes, during the segmentation process, accurately locating the position of the keyword in the text node according to the position information of the matching result, and performing highlighting effect processing on the text node portion corresponding to each matching result, without destroying the original DOM tree structure and text format during the processing, and retaining HTML tags and related style attributes.
[0019] Further preferably, the method further includes performing a node splitting operation, specifically including:
[0020] S71, starting from the first matching result, sequentially obtain each matching result and its corresponding text node;
[0021] S72, determine whether the index interval of the current text node and the index interval of the current matching result have an intersection, if there is an intersection, record the start and end indexes of the intersection, and proceed to step S73; if there is no intersection, proceed to step S75;
[0022] S73, judging whether the current text node has been completely matched, that is, judging whether the end point index of the matching result is greater than or equal to the end point index of the text node, if so, proceeding to step S75; if not, proceeding to step S74;
[0023] S74, judging whether the current matching result has been completely matched, that is, judging whether the index of the text node is greater than or equal to the end index of the matching result, if so, the matching result points to the next one, and enters step S76; if not, the matching result points to the next one, and returns to step S72;
[0024] S75, judging whether the current matching result has been completely matched, that is, judging whether the index of the text node is greater than or equal to the end index of the matching result, if so, the matching result points to the next one and enters sub-step S76; if not, returning to sub-step S72;
[0025] S76: Based on all the matching intervals of the current text node obtained, traverse the intersection and start splitting the nodes, including:
[0026] S761, determine whether the current intersection item is the first intersection, if so, proceed to step S72; if not, directly split the remaining nodes and proceed to step S73;
[0027] S762, determine whether the current intersection item hits the first character of the node, if so, do not process; if not, split the node from the starting point of the intersection and go to step S73;
[0028] S763. Determine whether the length of the current intersection item is less than the length of the current text node text. If so, split the node using the length of the intersection; if not, do nothing.
[0029] Further preferably, before step S71, the method further includes:
[0030] The matching results are obtained from step S6, and each matching item records the corresponding starting index in the plain text; the node information is obtained from step S2, and each node records the text content of the node and its corresponding starting and ending indexes in the plain text.
[0031] Preferably, in step S9, drawing a Bezier curve using the bezierCurveTo method of canvas drawing includes:
[0032] The bezierCurveTo method receives three point coordinate parameters, where the first point coordinate is the position coordinate of the highlighted node on the page (x1, y1), the second point coordinate is the position coordinate of the tag associated with the keyword on the page (x2, y2), and the third point coordinate (x, y) is calculated by the following formula:
[0033] x=|(x2-x1)|*4 / 5+min(x1,x2)
[0034] y=x1 <x2?y1∶y2
[0035] By drawing Bezier curves, intuitive identification connections between keywords and tags are achieved, and the drawing parameters of the Bezier curves can be dynamically adjusted according to the page layout and node distribution to ensure that the curves are clear and beautiful and do not obscure important content on the page.
[0036] Preferably, step S6 also includes deduplication and sorting of the matching results: for the same keyword that appears multiple times, the position information of each appearance is recorded; and all matching results are deduplicated and sorted according to certain rules, and the rules include sorting according to the order in which the keywords appear in the article.
[0037] Preferably, after step S9, it also includes: a step of visually displaying the processing results, displaying the highlight effect of the keywords and the results of the marking line drawing in an intuitive manner on the page, and providing user interaction functions, such as users clicking on keywords to view related detailed information or perform further operations, to enhance the user experience and the practicality of the method.
[0038] Preferably, it also includes: a step of logging the processing process, the logging content includes the execution time of each step, the amount of data processed, the problems encountered and the processing results, so as to debug and analyze when problems arise, and also provide data support for performance optimization of the method.
[0039] In a second aspect, an embodiment of the present invention provides a DOM-based multi-keyword accurate matching and identification line drawing system, including:
[0040] A DOM tree building module is configured to convert the article content to be processed into a DOM tree structure, and to build the DOM tree by creating HTML node elements using the DOM API of JavaScript, so that the article content is represented in a tree structure in memory;
[0041] A text node information acquisition module is configured to traverse the constructed DOM tree in depth first, using a recursive algorithm starting from the root node, and sequentially accessing each node and its child nodes. When a text node is accessed, the relevant information of the text node is recorded, including the text content, the hierarchical position in the DOM tree, and the start and end indexes in the plain text content. The index calculation is based on a predefined character encoding standard to ensure accuracy.
[0042] A plain text extraction module is configured to extract plain text content from the DOM tree, traverse the text nodes in the DOM tree, and splice the text contents of all text nodes in order, ignoring HTML tags and other non-text elements during the splicing process, to obtain a plain text representation of the article, so as to perform a keyword matching operation;
[0043] A keyword preprocessing module is configured to perform preprocessing operations on the input keyword group, including removing repeated keywords and ignoring special symbols in the keywords;
[0044] The keyword matching module is configured to traverse the DOM node content obtained by the text node information obtaining module for each keyword preprocessed by the keyword preprocessing module, and during the traversal process, for each DOM node, determine whether it contains the keyword to be matched;
[0045] The matching result processing module is configured to use a string matching algorithm to perform precise matching when a keyword is found in the process of the keyword matching module traversing the DOM node, determine the starting position and ending position of the keyword in the text node, store the position information, and convert the text match into a position match; store the matching result in a specific data format, and perform duplicate removal and sorting operations on all matching results. The duplicate removal and sorting rules are based on a predefined priority strategy, taking into account the first appearance position and keyword length factors;
[0046] A node segmentation module is configured to traverse the matching results obtained by the matching result processing module and the text nodes obtained by the text node information obtaining module, and process them by using a method of segmenting the text nodes;
[0047] The node highlighting module is configured to replace the original text node with a new node with highlighting style attributes using the replaceChild method of the Node interface in the DOM for the text node segmented and processed by the node segmentation module. During the replacement process, the ID information of the new node is recorded and associated with the node text content and the DOM tree position.
[0048] The line drawing module is configured to traverse each highlighted node obtained in the node highlighting module and draw the Bezier curve using the bezierCurveTo method of the canvas drawing.
[0049] Preferably, the node segmentation module also includes: being configured to use the splitText method of the Text node in the DOM to segment the text node containing the keyword into two or more independent brother nodes; during the segmentation process, accurately locating the position of the keyword in the text node according to the position information of the matching result, and performing highlighting effect processing on the text node portion corresponding to each matching result; during the processing, the original DOM tree structure and text format are not destroyed, and the HTML tags and related style attributes are retained.
[0050] Further preferably, the node segmentation module further includes:
[0051] The node matching interval judgment submodule is configured to obtain the matching results and their corresponding text nodes in sequence starting from the first matching result, and judge whether there is an intersection between the current text node index interval and the matching result index interval. Based on a strict mathematical interval comparison algorithm, if there is an intersection, the start and end indexes are recorded;
[0052] The node matching status judgment submodule is configured to judge whether all matching of the current text node has been completed, by comparing the matching result end index with the text node end index, and judging whether all matching of the current matching result has been completed, and comparing the text node index with the matching result end index, and deciding the next operation according to the judgment result, such as pointing to the next matching result, continuing to judge the interval intersection, or performing node segmentation;
[0053] The node segmentation execution submodule is configured to traverse the intersection segmentation nodes based on all matching intervals of the obtained text nodes, including determining whether the current intersection item is the first intersection. If so, determining whether the first character of the node is hit. Deciding whether to segment the node based on the determination result. If not, directly segmenting the remaining nodes. It is also necessary to determine the relationship between the length of the intersection item and the length of the text node to determine the segmentation method.
[0054] Preferably, the line drawing module further includes:
[0055] The configuration is used to receive three point coordinate parameters using the bezierCurveTo method, dynamically adjust the curve drawing parameters according to the page layout and node distribution, ensure that the curve is smooth and beautiful without blocking important content on the page, and realize the connection between keywords and tags;
[0056] The first point coordinates are the position coordinates of the highlighted node on the page (x1, y1), the second point coordinates are the position coordinates of the tag associated with the keyword on the page (x2, y2), and the third point coordinates (x, y) are calculated by the following formula: x = |(x2-x1)|*4 / 5+min(x1, x2), y = x1 <x2?y1∶y2。
[0057] In a third aspect, an embodiment of the present invention provides an electronic device, comprising: one or more processors; a storage device for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation manner in the first aspect.
[0058] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described in any implementation manner in the first aspect.
[0059] Compared with the prior art, the beneficial results of the present invention are:
[0060] The present invention provides a method and system for accurately marking big data articles, which can accurately match multiple groups of keywords, accurately locate the positions of these keywords in the article on the basis of retaining the original format of the article, and highlight them, providing users with a very intuitive information positioning method. Combined with the method of marking and drawing lines, users can quickly find key information in massive information. At the same time, based on the processing method of DOM tree, the technical solution can be compatible with diverse article content structures and has good scalability and adaptability.
[0061] In the big data environment, efficient marking and drawing methods will be widely used in major websites and applications. It can help users better understand and explore the content of articles and help researchers better extract valuable information from large paragraphs. In the field of education, educational platforms can highlight important knowledge points in textbooks to help students remember and understand key knowledge points. In the academic field, researchers can use the present invention to quickly locate key concepts in documents and improve the efficiency of retrieval. With the continuous development and improvement of technology, the present invention will play an important role in more fields. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] The accompanying drawings are included to provide a further understanding of the embodiments and are incorporated into and constitute a part of this specification. The accompanying drawings illustrate the embodiments and are used together with the description to explain the principles of the present invention. It will be easy to recognize other embodiments and many expected advantages of the embodiments because they become better understood by reference to the following detailed description. The elements of the drawings are not necessarily to scale with each other. The same reference numerals refer to corresponding similar parts.
[0063] Figure 1 It is a flow chart of a DOM-based multi-keyword accurate matching and identification line drawing method according to an embodiment of the present invention;
[0064] Figure 2 A schematic diagram of a method implementation flow of a specific embodiment of the present invention;
[0065] Figure 3 A schematic diagram of a flowchart of node segmentation in an embodiment of the present invention;
[0066] Figure 4 It is a basic architecture diagram of a DOM-based multi-keyword accurate matching and identification line drawing system according to an embodiment of the present invention;
[0067] Figure 5 This is a keyword highlighting effect diagram of an embodiment of the present invention;
[0068] Figure 6 This is a logo connection effect diagram of an embodiment of the present invention;
[0069] Figure 7 It is a schematic diagram of the structure of a computer device suitable for implementing an electronic device of an embodiment of the present invention. DETAILED DESCRIPTION
[0070] The present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments. It is to be understood that the specific embodiments described herein are only used to explain the relevant invention, rather than to limit the invention. It is also necessary to explain that, for ease of description, only the parts related to the relevant invention are shown in the accompanying drawings.
[0071] It should be noted that, in the absence of conflict, the embodiments of the present invention and the features in the embodiments can be combined with each other. The present invention will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0072] First, Figure 1 It is shown that the embodiment of the present invention discloses a DOM-based multi-keyword accurate matching and identification line drawing method, such as Figure 1 As shown, the method comprises the following steps:
[0073] S1, in response to converting the article content to be processed into a DOM tree structure, constructing a DOM tree by creating HTML node elements using a DOM API of JavaScript, so that the article content is represented in a tree structure in memory;
[0074] S2. Depth-first traversal of the constructed DOM tree, starting from the root node using a recursive algorithm, sequentially accessing each node and its child nodes, and when a text node is accessed, recording relevant information of the text node;
[0075] S3, extracting the plain text content from the DOM tree, traversing the text nodes in the DOM tree, splicing the text contents of all the text nodes in order, ignoring HTML tags and other non-text elements in the splicing process, and obtaining a plain text representation of the article for keyword matching operation;
[0076] S4, pre-processing the input keyword group, including removing repeated keywords and ignoring special symbols in the keywords;
[0077] S5. For each keyword preprocessed in step S4, traverse the DOM node content obtained in step S2. During the traversal process, for each DOM node, determine whether it contains the keyword to be matched;
[0078] S6, when a keyword is found in the process of traversing the DOM node in step S5, a string matching algorithm is used to perform an exact match, determine the starting position and the ending position of the keyword in the text node, store the position information, and convert the text match into a position match;
[0079] Specifically, it also includes deduplication and sorting of matching results: for the same keyword that appears multiple times, the position information of each appearance is recorded; and all matching results are deduplicated and sorted according to certain rules, and the rules include sorting according to the order in which the keywords appear in the article.
[0080] S7, traversing the matching results obtained in step S6 and the text nodes obtained in step S2, processing them by using a method of splitting the text nodes, and performing highlighting effect processing on each matching result and text node;
[0081] Specifically, in this embodiment, a method of splitting text nodes is adopted for processing, and highlighting effect processing is performed on each matching result and text node, including: using the splitText method of the Text node in the DOM to split the text node containing the keyword into two or more independent brother nodes. During the segmentation process, the position of the keyword in the text node is accurately located according to the position information of the matching result, and the text node part corresponding to each matching result is highlighted. The original DOM tree structure and text format are not destroyed during the processing, and the HTML tags and related style attributes are retained.
[0082] Perform node splitting operations, including:
[0083] S70, obtaining the matching results from step S6, where each matching item records the corresponding start index in the plain text; obtaining the node information from step S2, where each node records the text content of the node and its corresponding start and end indexes in the plain text;
[0084] S71, starting from the first matching result, sequentially obtain each matching result and its corresponding text node;
[0085] S72, determine whether the index interval of the current text node and the index interval of the current matching result have an intersection, if there is an intersection, record the start and end indexes of the intersection, and proceed to step S73; if there is no intersection, proceed to step S75;
[0086] S73, judging whether the current text node has been completely matched, that is, judging whether the end point index of the matching result is greater than or equal to the end point index of the text node, if so, proceeding to step S75; if not, proceeding to step S74;
[0087] S74, judging whether the current matching result has been completely matched, that is, judging whether the index of the text node is greater than or equal to the end index of the matching result, if so, the matching result points to the next one, and enters step S76; if not, the matching result points to the next one, and returns to step S72;
[0088] S75, judging whether the current matching result has been completely matched, that is, judging whether the index of the text node is greater than or equal to the end index of the matching result, if so, the matching result points to the next one and enters sub-step S76; if not, returning to sub-step S72;
[0089] S76: Based on all the matching intervals of the current text node obtained, traverse the intersection and start splitting the nodes, including:
[0090] S761, determine whether the current intersection item is the first intersection, if so, proceed to step S72; if not, directly split the remaining nodes and proceed to step S73;
[0091] S762, determine whether the current intersection item hits the first character of the node, if so, do not process; if not, split the node from the starting point of the intersection and go to step S73;
[0092] S763. Determine whether the length of the current intersection item is less than the length of the current text node text. If so, split the node using the length of the intersection; if not, do nothing.
[0093] S8, for the text node segmented and processed in step S7, use the replaceChild method of the Node interface in the DOM to replace the original text node with a new node with a highlight style attribute, and during the replacement process, record the ID information of the new node;
[0094] S9, traverse each highlighted node obtained in step S8, and use the bezierCurveTo method of canvas drawing to draw a Bezier curve.
[0095] Furthermore, the bezierCurveTo method receives three point coordinate parameters, where the first point coordinate is the position coordinate of the highlighted node on the page (x1, y1), the second point coordinate is the position coordinate of the tag associated with the keyword on the page (x2, y2), and the third point coordinate (x, y) is calculated by the following formula:
[0096] x=|(x2-x1)|*4 / 5+min(x1,x2)
[0097] y=x1 <x2?y1∶y2
[0098] By drawing Bezier curves, intuitive identification connections between keywords and tags are achieved, and the drawing parameters of the Bezier curves can be dynamically adjusted according to the page layout and node distribution to ensure that the curves are clear and beautiful and do not obscure important content on the page.
[0099] This embodiment also includes: a step of visually displaying the processing results, displaying the highlight effect of keywords and the results of marking lines in an intuitive manner on the page, and providing user interaction functions, such as users clicking on keywords to view related detailed information or perform further operations, to enhance the user experience and the practicality of the method.
[0100] It also includes the step of logging the processing process, which includes the execution time of each step, the amount of data processed, the problems encountered and the processing results, so as to facilitate debugging and analysis when problems arise, and also provide data support for performance optimization of the method.
[0101] Specifically, the technical implementation scheme of the embodiment of the present invention is mainly:
[0102] First, treat the article as an HTML element and build a DOM tree. Then, traverse the DOM tree in depth first to obtain text node information, and take out all text content to splice into the plain text content of the article. Then, perform text search and matching on the keyword group and the plain text content, store the start and end index of each match, and convert from text matching to position matching. The matching results are split into text nodes, adding highlight effects to the nodes without destroying the original format. Finally, use the API provided by Canvas to customize the parameters and draw a smooth Bezier curve.
[0103] The method steps disclosed in the present invention are described in detail below using a specific embodiment. Figure 2 As shown, the specific steps are as follows:
[0104] Step 1: Use JavaScript's DOM API to create HTML node elements and build the article content into a DOM tree;
[0105] Step 2: Use recursive method to implement depth-first traversal of DOM tree and obtain text node information;
[0106] Step 3: Extract the plain text content from the complex DOM structure tree;
[0107] Step 4: Pre-process the keyword groups to remove duplicates and ignore special symbols;
[0108] Step 5: Traverse each keyword and traverse the DOM node content;
[0109] Step 6: Use the string matching algorithm to perform text search and matching on each keyword and the plain text content, store the start and end indexes of each match, and convert from text matching to position matching; remove duplicates and sort the matching results;
[0110] Step 7: Traverse the matching results and text nodes; Use the method of splitting text nodes. Use the splitText method of the Text node in DOM to split the node into two independent brother nodes, and accurately highlight each matching result and text node;
[0111] Example 1: For the HTML string: "Dragon Boat Festival Boat I haven't seen enough of the race yet", keyword: "Dragon Boat Race", its DOM tree corresponds to 3 text nodes, namely "Dragon Boat Festival Dragon", "Boat", "I haven't seen enough of the race yet". Through keyword matching and node segmentation, the text nodes corresponding to its DOM tree will be split into 5: "Dragon Boat Festival", "Dragon", "Boat", "Race", "I haven't seen enough".
[0112] Example 2: For the pure text: "I haven't had enough of watching the Dragon Boat Race during the Dragon Boat Festival", with the keyword: "Dragon Boat Race", the DOM tree corresponding text node has only 1. Through keyword matching and node splitting, the DOM tree corresponding text node will be split into 3: "Dragon Boat Festival", "Dragon Boat Race", "I haven't had enough of watching".
[0113] As Figure 3 shown, the detailed process of this step is as follows:
[0114] S101: Obtain the matching results from Step 6. Each matching item records the corresponding starting index on the pure text. Example 1 can be converted to: [[2,4]];
[0115] S102: Obtain the node information from Step 2. Each node records the text content of the node and its corresponding start and end indexes on the pure text. Example 1 can be converted to: [{text:'Dragon Boat Festival',range:[0,3]},{text:'Dragon',range:[3,4]},{text:'Race I haven't had enough of watching',range:[4,10]}];
[0116] S103: Start from the first matching item and traverse all nodes;
[0117] S104: Determine whether there is an intersection between the index range of the current node and the index range of the current matching item. If there is an intersection, record the start and end indexes of the intersection and continue with Step S105; otherwise, enter Step S106;
[0118] S105: Determine whether the current node has been fully matched (i.e., the end index of the matching item is greater than or equal to the end index of the node). If so, continue with Step S106; otherwise, enter Step S107;
[0119] S106: Determine whether the current matching item has been fully matched (i.e., the index of the node is greater than or equal to the end index of the matching item). If so, increment the matching item by 1; enter Step S108;
[0120] S107: Determine whether the current matching item has been fully matched (i.e., the index of the node is greater than or equal to the end index of the matching item). If so, increment the matching item by 1; continue with Step S104;
[0121] S108: At this time, all the matching intervals of the current node have been obtained, and traverse the intersection to start splitting the node;
[0122] S109: Determine whether the current intersection item is the first intersection. If so, enter Step S110; otherwise, directly split the remaining nodes and enter Step S111;
[0123] S110: Determine whether the first character of the current intersection item hits a node. If so, do nothing; otherwise, split the node from the starting point of the intersection and proceed to step S111.
[0124] S111: Determine whether the length of the current intersection item is less than the length of the current node text. If so, split the node with the length of the intersection; otherwise, do nothing.
[0125] S112: End.
[0126] Step Eight: Use the replaceChild method of the Node interface in the DOM for the split text nodes to replace a child node in the DOM tree, so that the node has a highlighted style attribute. During the replacement process, record the ID of the node for the generation of subsequent line drawing markings.
[0127] In Example 1, the replaced text nodes will become: "Dragon Boat Festival", " <span style="color:red"> dragon ", " <span style="color:red"> Boat ", " <span style="color:red"> Race ", "
[0128] Step Nine: Traverse each highlighted node and use the bezierCurveTo method of canvas drawing to draw a Bezier curve. The bezierCurveTo method receives three point coordinate parameters. The first point is the coordinate (x1, y1) of the highlighted node, the second point is the coordinate (x2, y2) of the label, and the third point coordinate is unknown. Here, a method needs to be defined to calculate the coordinate of the third point so that the drawn curve is smooth and beautiful.
[0129] The calculation formula for the third point coordinate is as follows:
[0130] x = |(x2 - x1)| * 4 / 5 + min(x1, x2)
[0131] y = x1 < x2? y1 : y2
[0132] For further reference Figure 4 , as an implementation of the methods shown in the above figures, the present application provides an embodiment of a system. This system embodiment corresponds to Figure 1 the method embodiment shown, and this system can be specifically applied to various electronic devices.
[0133] In a second aspect, the embodiments of the present invention also disclose a DOM-based multi-keyword exact matching and marking line drawing system, as shown in Figure 4 , including:
[0134] A DOM tree construction module 41 is configured to convert the article content to be processed into a DOM tree structure, and to construct the DOM tree by creating HTML node elements using a DOM API of JavaScript, so that the article content is represented in a tree structure in memory;
[0135] The text node information acquisition module 42 is configured to traverse the constructed DOM tree in depth first, using a recursive algorithm starting from the root node, and sequentially accessing each node and its child nodes. When a text node is accessed, the relevant information of the text node is recorded, including the text content, the hierarchical position in the DOM tree, and the start and end indexes in the plain text content. The index calculation is based on a predefined character encoding standard to ensure accuracy.
[0136] A plain text extraction module 43 is configured to extract plain text content from the DOM tree, by traversing the text nodes in the DOM tree, splicing the text contents of all text nodes in order, ignoring HTML tags and other non-text elements during the splicing process, and obtaining a plain text representation of the article for keyword matching operations;
[0137] The keyword preprocessing module 44 is configured to perform preprocessing operations on the input keyword group, including removing repeated keywords and ignoring special symbols in the keywords;
[0138] The keyword matching module 45 is configured to traverse the DOM node content obtained by the text node information obtaining module for each keyword preprocessed by the keyword preprocessing module, and during the traversal process, for each DOM node, determine whether it contains the keyword to be matched;
[0139] The matching result processing module 46 is configured to perform precise matching using a string matching algorithm when a keyword is found in the process of the keyword matching module traversing the DOM node, determine the starting position and the ending position of the keyword in the text node, store the position information, and convert the text match into a position match; store the matching result in a specific data format, and perform duplicate removal and sorting operations on all matching results. The duplicate removal and sorting rules are based on a predefined priority strategy, taking into account the first appearance position and the keyword length factors;
[0140] The node segmentation module 47 is configured to traverse the matching results obtained by the matching result processing module and the text nodes obtained by the text node information obtaining module, and process them by using a method of segmenting the text nodes;
[0141] The node highlighting module 48 is configured to replace the original text node with a new node having a highlighting style attribute by using the replaceChild method of the Node interface in the DOM for the text node segmented and processed by the node segmentation module. During the replacement process, the ID information of the new node is recorded and associated with the node text content and the DOM tree position.
[0142] The line drawing module 49 is configured to traverse each highlighted node obtained in the node highlighting module and draw a Bezier curve using the bezierCurveTo method of canvas drawing.
[0143] Furthermore, the node segmentation module 47 also includes: being configured to use the splitText method of the Text node in the DOM to segment the text node containing the keyword into two or more independent brother nodes, during the segmentation process, accurately locating the position of the keyword in the text node according to the position information of the matching result, performing highlighting effect processing on the text node portion corresponding to each matching result, and not destroying the original DOM tree structure and text format during the processing process, and retaining HTML tags and related style attributes;
[0144] Further, such as Figure 4 As shown, the node segmentation module 47 also includes:
[0145] The node matching interval judgment submodule 471 is configured to obtain the matching results and their corresponding text nodes in sequence starting from the first matching result, and judge whether there is an intersection between the current text node index interval and the matching result index interval, and record the start and end indexes if there is an intersection based on a strict mathematical interval comparison algorithm;
[0146] The node matching state judgment submodule 472 is configured to judge whether all matching of the current text node has been completed, by comparing the matching result end index with the text node end index, and judging whether all matching of the current matching result has been completed, and comparing the text node index with the matching result end index, and determining the next operation according to the judgment result, such as pointing to the next matching result, continuing to judge the interval intersection, or performing node segmentation;
[0147] The node segmentation execution submodule 473 is configured to traverse the intersection segmentation nodes based on all matching intervals of the acquired text nodes, including determining whether the current intersection item is the first intersection. If so, determining whether the first character of the node is hit. Deciding whether to segment the node is based on the determination result. If not, directly segmenting the remaining nodes. It is also necessary to determine the relationship between the length of the intersection item and the length of the text node to determine the segmentation method.
[0148] The line drawing module 49 also includes: a configuration for receiving three point coordinate parameters using the bezierCurveTo method, dynamically adjusting the curve drawing parameters according to the page layout and node distribution, ensuring that the curve is smooth and beautiful and does not block important content of the page, and realizing the identification connection between keywords and tags;
[0149] The first point coordinates are the position coordinates of the highlighted node on the page (x1, y1), the second point coordinates are the position coordinates of the tag associated with the keyword on the page (x2, y2), and the third point coordinates (x, y) are calculated by the following formula: x = |(x2-x1)|*4 / 5+min(x1, x2), y = x1 <x2?y1∶y2。
[0150] The functions of the above modules correspond to the methods and will not be repeated here.
[0151] Figure 5 The keyword highlighting effect diagram of this embodiment is shown; Figure 6 The effect diagram of the marking connection of this embodiment is shown in FIG. Figure 5 and 6 shown.
[0152] The embodiment of the present invention provides a method and system for accurately marking big data articles. Through the above steps, multiple groups of keywords can be accurately matched. On the basis of retaining the original format of the article, the positions of these keywords in the article can be accurately located and highlighted, providing users with a very intuitive information positioning method. Combined with the method of marking and drawing lines, users can quickly find key information in massive amounts of information. At the same time, based on the processing method of the DOM tree, the technical solution can be compatible with diverse article content structures and has good scalability and adaptability.
[0153] In the big data environment, efficient marking and drawing methods will be widely used in major websites and applications. It can help users better understand and explore the content of articles and help researchers better extract valuable information from large paragraphs. In the field of education, educational platforms can highlight important knowledge points in textbooks to help students remember and understand key knowledge points. In the academic field, researchers can use the present invention to quickly locate key concepts in documents and improve the efficiency of retrieval. With the continuous development and improvement of technology, the present invention will play an important role in more fields.
[0154] Reference below Figure 7 , which shows a schematic structural diagram of a computer device 600 of an electronic device suitable for implementing an embodiment of the present invention. Figure 7 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present invention.
[0155] like Figure 7 As shown, the computer device 600 includes a central processing unit (CPU) 601 and a graphics processing unit (GPU) 602, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 603 or a program loaded from a storage part 609 to a random access memory (RAM) 604. In the RAM 604, various programs and data required for the operation of the device 600 are also stored. The CPU 601, GPU 602, ROM 603 and RAM 604 are connected to each other through a bus 605. An input / output (I / O) interface 606 is also connected to the bus 605.
[0156] The following components are connected to the I / O interface 606: an input section 607 including a keyboard, a mouse, etc.; an output section 608 including a display such as a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 609 including a hard disk, etc.; and a communication section 610 including a network interface card such as a LAN card, a modem, etc. The communication section 610 performs communication processing via a network such as the Internet. A drive 611 may also be connected to the I / O interface 606 as needed. A removable medium 612, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 611 as needed so that a computer program read therefrom is installed into the storage section 609 as needed.
[0157] In particular, according to the embodiments disclosed in the present invention, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the embodiments disclosed in the present invention include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 610, and / or installed from the removable medium 612. When the computer program is executed by the central processing unit (CPU) 601 and the graphics processing unit (GPU) 602, the above-mentioned functions defined in the method of the present invention are executed.
[0158] It should be noted that the computer-readable medium described in the present invention may be a computer-readable signal medium or a computer-readable medium or any combination of the above two. The computer-readable medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor device, apparatus or device, or any combination of the above. More specific examples of computer-readable media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, a computer-readable medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution device, apparatus or device. In the present invention, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable program code. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium may also be any computer-readable medium other than a computer-readable medium that can send, propagate or transmit a program for use by or in conjunction with an instruction execution device, apparatus or device. The program code contained on the computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.
[0159] Computer program code for performing the operations of the present invention may be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0160] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the device, method and computer program product according to various embodiments of the present invention. In this regard, each box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based device that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0161] The modules involved in the embodiments of the present invention may be implemented in software or hardware, and the modules described may also be arranged in a processor.
[0162] As another aspect, the present invention further provides a computer-readable medium, which may be included in the electronic device described in the above embodiment; or may exist independently without being assembled into the electronic device. The above computer-readable medium carries one or more programs, and when the above one or more programs are executed by the electronic device, the electronic device: executes the method steps described in the first aspect.
[0163] The above description is only a preferred embodiment of the present invention and an explanation of the technical principles used. Those skilled in the art should understand that the scope of the invention involved in the present invention is not limited to the technical solution formed by a specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the above features are replaced with the technical features with similar functions disclosed in the present invention (but not limited to) to form a technical solution.
Claims
1. A DOM-based multi-keyword accurate matching and marking line drawing method, characterized in that: The method comprises the following steps: S1. In response to converting the article content to be processed into a DOM tree structure, a DOM tree is constructed by creating HTML node elements using a DOM API of JavaScript, so that the article content is represented in a tree structure in memory; S2. Depth-first traversal of the constructed DOM tree, starting from the root node using a recursive algorithm, sequentially accessing each node and its child nodes, and when a text node is accessed, recording relevant information of the text node; S3, extracting the plain text content from the DOM tree, traversing the text nodes in the DOM tree, splicing the text contents of all the text nodes in order, ignoring HTML tags and other non-text elements in the splicing process, and obtaining a plain text representation of the article for keyword matching operation; S4, preprocessing the input keyword group, including removing repeated keywords and ignoring special symbols in the keywords; S5. For each keyword preprocessed in step S4, traverse the DOM node content obtained in step S2. During the traversal process, for each DOM node, determine whether it contains the keyword to be matched; S6, when a keyword is found in the process of traversing the DOM node in step S5, a string matching algorithm is used to perform an exact match, determine the starting position and the ending position of the keyword in the text node, store the position information, and convert the text match into a position match; S7, traversing the matching results obtained in step S6 and the text nodes obtained in step S2, processing them by using a method of segmenting the text nodes, and performing highlighting effect processing on each matching result and text node; S8, for the text node segmented and processed in step S7, use the replaceChild method of the Node interface in the DOM to replace the original text node with a new node with a highlight style attribute, and during the replacement process, record the ID information of the new node; S9, traverse each highlighted node obtained in step S8, and use the bezierCurveTo method of canvas drawing to draw a Bezier curve.
2. The DOM-based multi-keyword accurate matching and labeling line drawing method according to claim 1, characterized in that: In step S7, a method of segmenting text nodes is used to perform a highlighting effect on each matching result and text node, including: Use the splitText method of the Text node in DOM to split the text node containing keywords into two or more independent brother nodes. During the splitting process, the position of the keyword in the text node is accurately located according to the position information of the matching result, and the text node part corresponding to each matching result is highlighted. The original DOM tree structure and text format are not destroyed during the processing, and the HTML tags and related style attributes are retained.
3. The DOM-based multi-keyword accurate matching and labeling line drawing method according to claim 2, characterized in that: It also includes performing node splitting operations, including: S71, starting from the first matching result, sequentially obtain each matching result and its corresponding text node; S72, determine whether the index interval of the current text node and the index interval of the current matching result have an intersection, if there is an intersection, record the start and end indexes of the intersection, and proceed to step S73; if there is no intersection, proceed to step S75; S73, judging whether the current text node has been completely matched, that is, judging whether the end point index of the matching result is greater than or equal to the end point index of the text node, if so, proceeding to step S75; if not, proceeding to step S74; S74, judging whether the current matching result has been completely matched, that is, judging whether the index of the text node is greater than or equal to the end index of the matching result, if so, the matching result points to the next one, and enters step S76; if not, the matching result points to the next one, and returns to step S72; S75, judging whether the current matching result has been completely matched, that is, judging whether the index of the text node is greater than or equal to the end index of the matching result, if so, the matching result points to the next one and enters sub-step S76; if not, returning to sub-step S72; S76: Based on all the matching intervals of the current text node obtained, traverse the intersection and start splitting the nodes, including: S761, determine whether the current intersection item is the first intersection, if so, proceed to step S72; if not, directly split the remaining nodes and proceed to step S73; S762, determine whether the current intersection item hits the first character of the node, if so, do not process; if not, split the node from the starting point of the intersection and go to step S73; S763. Determine whether the length of the current intersection item is less than the length of the current text node text. If so, split the node using the length of the intersection; if not, do nothing.
4. The DOM-based multi-keyword accurate matching and labeling line drawing method according to claim 3, characterized in that: Before step S71, the method further includes: The matching results are obtained from step S6, and each matching item records the corresponding starting index in the plain text; The node information is obtained from step S2, and each node records the text content of the node and its corresponding start and end indexes in the plain text.
5. The DOM-based multi-keyword accurate matching and labeling line drawing method according to claim 1, characterized in that: In step S9, the BezierCurveTo method of canvas drawing is used to draw the Bezier curve, including: The bezierCurveTo method receives three point coordinate parameters, where the first point coordinate is the position coordinate of the highlighted node on the page (x1, y1), the second point coordinate is the position coordinate of the tag associated with the keyword on the page (x2, y2), and the third point coordinate (x, y) is calculated by the following formula: x=|(x2-x1)|*4 / 5+min(x1,x2) y=x1 <x2?y1∶y2 By drawing Bezier curves, intuitive identification connections between keywords and tags are achieved, and the drawing parameters of the Bezier curves can be dynamically adjusted according to the page layout and node distribution to ensure that the curves are clear and beautiful and do not obscure important content on the page.
6. The DOM-based multi-keyword accurate matching and labeling line drawing method according to claim 1, characterized in that: Step S6 also includes sorting and ranking the matching results: For the same keyword that appears multiple times, the location information of each occurrence is recorded; and all matching results are deduplicated and sorted according to certain rules, including sorting in the order in which the keywords appear in the article.
7. The DOM-based multi-keyword accurate matching and labeling line drawing method according to claim 1, characterized in that: After step S9, the method further includes: The steps of visually displaying the processing results are to display the highlighting effect of keywords and the results of marking lines in an intuitive way on the page, and provide user interaction functions, such as users clicking on keywords to view related detailed information or perform further operations, to enhance the user experience and the practicality of the method.
8. The DOM-based multi-keyword accurate matching and labeling line drawing method according to claim 1, characterized in that: Also includes: The steps of logging the processing process include the execution time of each step, the amount of data processed, the problems encountered and the processing results, so as to facilitate debugging and analysis when problems arise, and also provide data support for performance optimization of the method.
9. A DOM-based multi-keyword accurate matching and marking line drawing system, characterized in that: include: A DOM tree building module is configured to convert the article content to be processed into a DOM tree structure, and to build the DOM tree by creating HTML node elements using the DOM API of JavaScript, so that the article content is represented in a tree structure in memory; A text node information acquisition module is configured to traverse the constructed DOM tree in depth first, using a recursive algorithm starting from the root node, and sequentially accessing each node and its child nodes. When a text node is accessed, the relevant information of the text node is recorded, including the text content, the hierarchical position in the DOM tree, and the start and end indexes in the plain text content. The index calculation is based on a predefined character encoding standard to ensure accuracy. A plain text extraction module is configured to extract plain text content from the DOM tree, traverse the text nodes in the DOM tree, and splice the text contents of all text nodes in order, ignoring HTML tags and other non-text elements during the splicing process, to obtain a plain text representation of the article, so as to perform a keyword matching operation; A keyword preprocessing module is configured to perform preprocessing operations on the input keyword group, including removing repeated keywords and ignoring special symbols in the keywords; The keyword matching module is configured to traverse the DOM node content obtained by the text node information obtaining module for each keyword preprocessed by the keyword preprocessing module, and during the traversal process, for each DOM node, determine whether it contains the keyword to be matched; The matching result processing module is configured to use a string matching algorithm to perform precise matching when a keyword is found in the process of the keyword matching module traversing the DOM node, determine the starting position and ending position of the keyword in the text node, store the position information, and convert the text match into a position match; store the matching result in a specific data format, and perform duplicate removal and sorting operations on all matching results. The duplicate removal and sorting rules are based on a predefined priority strategy, taking into account the first appearance position and keyword length factors; A node segmentation module is configured to traverse the matching results obtained by the matching result processing module and the text nodes obtained by the text node information obtaining module, and process them by using a method of segmenting the text nodes; The node highlighting module is configured to replace the original text node with a new node with highlighting style attributes using the replaceChild method of the Node interface in the DOM for the text node segmented and processed by the node segmentation module. During the replacement process, the ID information of the new node is recorded and associated with the node text content and the DOM tree position. The line drawing module is configured to traverse each highlighted node obtained in the node highlighting module and draw the Bezier curve using the bezierCurveTo method of the canvas drawing.
10. The DOM-based multi-keyword accurate matching and labeling line drawing method according to claim 9, characterized in that: The node segmentation module also includes: a configuration for using the splitText method of the Text node in the DOM to segment the text node containing the keyword into two or more independent brother nodes. During the segmentation process, the position of the keyword in the text node is accurately located according to the position information of the matching result, and the text node part corresponding to each matching result is highlighted. The original DOM tree structure and text format are not destroyed during the processing, and the HTML tags and related style attributes are retained.
11. The DOM-based multi-keyword accurate matching and labeling line drawing method according to claim 10, characterized in that: The node segmentation module also includes: The node matching interval judgment submodule is configured to obtain the matching results and their corresponding text nodes in sequence starting from the first matching result, and judge whether there is an intersection between the current text node index interval and the matching result index interval. Based on a strict mathematical interval comparison algorithm, if there is an intersection, the start and end indexes are recorded; The node matching status judgment submodule is configured to judge whether all matching of the current text node has been completed, by comparing the matching result end index with the text node end index, and judging whether all matching of the current matching result has been completed, and comparing the text node index with the matching result end index, and deciding the next operation according to the judgment result, such as pointing to the next matching result, continuing to judge the interval intersection, or performing node segmentation; The node segmentation execution submodule is configured to traverse the intersection segmentation nodes based on all matching intervals of the obtained text nodes, including determining whether the current intersection item is the first intersection. If so, determining whether the first character of the node is hit. Deciding whether to segment the node based on the determination result. If not, directly segmenting the remaining nodes. It is also necessary to determine the relationship between the length of the intersection item and the length of the text node to determine the segmentation method.
12. The DOM-based multi-keyword accurate matching and labeling line drawing method according to claim 9, characterized in that: The line drawing module also includes: The configuration is used to receive three point coordinate parameters using the bezierCurveTo method, dynamically adjust the curve drawing parameters according to the page layout and node distribution, ensure that the curve is smooth and beautiful without blocking important content on the page, and realize the connection between keywords and tags; The first point coordinates are the position coordinates of the highlighted node on the page (x1, y1), the second point coordinates are the position coordinates of the tag associated with the keyword on the page (x2, y2), and the third point coordinates (x, y) are calculated by the following formula: x = |(x2-x1)|*4 / 5+min(x1, x2), y = x1 <x2?y1∶y2。 13. An electronic device, comprising: one or more processors; A storage device for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 8.
14. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.