Data labeling method and device, electronic equipment and computer readable storage medium
By converting lightweight markup language text into hypertext markup language documents and generating tags, the problems of low efficiency and low accuracy of lightweight markup language text data annotation are solved, and more efficient and accurate data annotation is achieved.
Patent Information
- Application Number
- CN202510487551.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-09-19
AI Technical Summary
In the prior art, data annotation of lightweight markup language text has low efficiency and low accuracy, mainly because the labelers need to understand the grammatical rules of the lightweight markup language and cannot directly view the actual rendering results.
The lightweight markup language text is converted into a hypertext markup language document. By displaying the rendering result of the hypertext markup language document, the selected operation is detected and the corresponding label is generated, thereby realizing the refined data annotation of the lightweight markup language text.
It reduces the difficulty of understanding lightweight markup language text and the complexity of data annotation, improves the efficiency and accuracy of data annotation, and provides a flexible annotation method.
Smart Images

Figure CN120671665A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and more specifically, to a data annotation method, device, electronic device, and computer-readable storage medium. Background Art
[0002] Lightweight markup languages (LMLs) offer advantages such as high readability, simple syntax, cross-platform compatibility, and easy conversion. Consequently, they are widely used in natural language processing (NLP) and related fields, particularly large language models (LLMs). Data annotation is a crucial step in the training of large language processing models, and its quality impacts their ability to understand and accurately interpret natural language.
[0003] However, existing technologies usually directly mark lightweight markup language texts. This marking method requires the labeler to have a certain understanding of the grammatical rules of the lightweight markup language text and is prone to marking errors. Therefore, the efficiency of data labeling is low and the accuracy is not high. Summary of the Invention
[0004] In view of this, embodiments of the present invention provide a data labeling method, device, electronic device and computer-readable storage medium, which reduce the difficulty of understanding lightweight markup language text and the complexity of data labeling by converting lightweight markup language text into hypertext markup language documents, thereby improving the efficiency and accuracy of data labeling.
[0005] In a first aspect, an embodiment of the present invention provides a data labeling method, the method comprising:
[0006] Displaying a rendering result of a hypertext markup language document, the hypertext markup language document being obtained by converting a lightweight markup language text, the rendering result including a plurality of selectable elements;
[0007] In response to detecting a selection operation on the selectable element, determining a selected element corresponding to the selection operation;
[0008] Determining a target element in the lightweight markup language text according to the selected element;
[0009] In response to detecting a marking operation on the selected element, a first label corresponding to the target element is generated according to the marking operation, and a second label of the selected element is displayed according to a display position of the selected element.
[0010] Optionally, the hypertext markup language document is converted in the following manner:
[0011] Parsing the lightweight markup language text to obtain a corresponding first abstract syntax tree;
[0012] Converting the first abstract syntax tree into a second abstract syntax tree, where the second abstract syntax tree is an abstract syntax tree of the hypertext markup language document;
[0013] The second abstract syntax tree is converted into the hypertext markup language document.
[0014] Optionally, the rendering result includes text information and non-text information, and the selectable elements are characters in the text information or the non-text information.
[0015] Optionally, the first abstract syntax tree includes a text node for each of the text information;
[0016] The converting the second abstract syntax tree into the hypertext markup language document comprises:
[0017] Decomposing each text node into a character node according to characters in each text message, and adding a positioning node for each character node in the second abstract syntax tree to obtain a processed second abstract syntax tree, wherein the positioning node is used to store a positioning element for each character, and the positioning element is determined according to a position of the character in the text message;
[0018] The processed second abstract syntax tree is converted into the hypertext markup language document.
[0019] Optionally, the method further includes:
[0020] A markup file of the lightweight markup language text is generated according to the target element and the tag attribute information of the corresponding first tag.
[0021] Optionally, the method further includes:
[0022] In response to detecting a viewing operation on a lightweight markup language text, the lightweight markup language text and the first tag of the target element are displayed.
[0023] In a second aspect, an embodiment of the present invention provides a data labeling device, the device comprising:
[0024] A result display unit, configured to display a rendering result of a hypertext markup language document, the hypertext markup language document being obtained by converting a lightweight markup language text, the rendering result including a plurality of selectable elements;
[0025] a first element determining unit, configured to, in response to detecting a selection operation on the selectable element, determine a selected element corresponding to the selection operation;
[0026] A second element determining unit, configured to determine a target element in the lightweight markup language text according to the selected element;
[0027] A generating and displaying unit is configured to, in response to detecting a marking operation on the selected element, generate a first label corresponding to the target element according to the marking operation, and display a second label of the selected element according to a display position of the selected element.
[0028] In a third aspect, an embodiment of the present invention provides an electronic device comprising a memory and a processor, wherein the memory is used to store one or more computer program instructions, wherein the one or more computer program instructions are executed by the processor to implement a method as described in any one of the first aspects.
[0029] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method as described in any one of the first aspects is implemented.
[0030] In a fifth aspect, an embodiment of the present invention provides a computer program product, which includes a computer program / instructions, and when the computer program / instructions are executed by a processor, implements the method as described in any one of the first aspects.
[0031] The embodiment of the present invention displays the rendering result of a hypertext markup language document including multiple selectable elements, and when a selection operation for the selectable elements is detected, determines the selected elements among the selectable elements, and then determines the target element in the lightweight markup language text corresponding to the hypertext markup language document based on the selected elements, so that when a marking operation for the selected elements is detected, a label of the target element is generated according to the marking operation, and the label of the selected element is displayed according to the display position of the selected element. The embodiment of the present invention reduces the difficulty of understanding the lightweight markup language text and the complexity of data marking by converting the lightweight markup language text into a hypertext markup language document, thereby effectively improving the efficiency and accuracy of data marking. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] The above and other objects, features and advantages of the present invention will become more apparent through the following description of the embodiments of the present invention with reference to the accompanying drawings, in which:
[0033] Figure 1 is a flow chart of a data labeling method according to an embodiment of the present invention;
[0034] Figure 2 This is a flow chart of obtaining a hypertext markup language document in an optional implementation of an embodiment of the present invention;
[0035] Figure 3 is a schematic diagram of the structure of an abstract syntax tree according to an embodiment of the present invention;
[0036] Figure 4 This is a flow chart of obtaining a hypertext markup language document in an optional implementation of an embodiment of the present invention;
[0037] Figure 5 is a schematic diagram of the structure of an abstract syntax tree according to an embodiment of the present invention;
[0038] Figure 6 This is a schematic diagram of an interface according to an embodiment of the present invention;
[0039] Figure 7-Figure 8 is another interface schematic diagram of an embodiment of the present invention;
[0040] Figure 9 is another interface schematic diagram of an embodiment of the present invention;
[0041] Figure 10 is a schematic diagram of a data labeling device according to an embodiment of the present invention;
[0042] Figure 11 is a schematic diagram of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0043] The present application is described below based on the following embodiments, but the present application is not limited to these embodiments. In the detailed description of the present application below, certain specific details are described in detail. Those skilled in the art can fully understand the present application without the description of these details. To avoid obscuring the essence of the present application, well-known methods, processes, procedures, components, and circuits are not described in detail.
[0044] Furthermore, persons of ordinary skill in the art will appreciate that the figures provided herein are for illustration purposes only and are not necessarily drawn to scale.
[0045] Unless the context clearly requires otherwise, words like “include”, “comprising” and the like throughout this application should be interpreted as including rather than exclusive or exhaustive; that is, as meaning “including but not limited to”.
[0046] In the description of this application, it should be understood that the terms "first", "second", etc. are used for descriptive purposes only and should not be understood to indicate or imply relative importance. In addition, in the description of this application, unless otherwise specified, "plurality" means two or more.
[0047] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0048] Lightweight Markup Language (LML) is a text language that uses simple syntax to describe simple formats. Its native format is close to natural language, and can include multimedia content such as text and images, as well as reflect rich text layout information. Common LMLs include Markdown, reStructuredText, AsciiDoc, Textile, and Org-mode. Taking Markdown as an example, in Markdown, the "#" symbol indicates a title, and the number of "#" symbols indicates the title's level, such as "#" for a first-level title and "##" for a second-level title. Italics are indicated by "*" or "_" symbols placed on either side of the text, such as "*text*" or "_text_" for "text." Bold text is indicated by "**" or "__" (double underscores), such as "**text**" or "__text__" for "text."
[0049] However, there are still certain differences between lightweight markup languages and natural languages, and there are also differences between the grammatical rules of different lightweight markup languages. Therefore, when performing data annotation on lightweight markup language text, if the labeler does not have a high level of understanding of the grammatical rules of the lightweight markup language, it will take extra time to understand the meaning expressed by each symbol in the grammatical rules, which will reduce the efficiency of data annotation. At the same time, the content of the lightweight markup language may differ from the actual rendering result, and the labeler is usually unable to directly view the actual rendering result. For example, images are mainly represented by the uniform resource locator (URL) of the image in the lightweight markup language text, which makes it impossible for the labeler to directly view the rendered image when viewing the lightweight markup language text. Therefore, the accuracy of existing data annotation is usually not high.
[0050] In order to solve the above problems, the embodiments of the present invention propose a data labeling method, device, electronic device and computer-readable storage medium, which reduce the difficulty of understanding lightweight markup language text and the complexity of data labeling by converting lightweight markup language text into hypertext markup language document, thereby improving the efficiency and accuracy of data labeling.
[0051] The embodiments of the present invention are mainly described using markdown as a lightweight markup language as an example. It should be understood that this embodiment is not limited to this. Existing lightweight markup languages that can support corresponding functions or lightweight markup languages that can support corresponding functions with future technological developments are all within the scope of protection of the embodiments of the present invention.
[0052] The following is an explanation of the method by way of an embodiment. Figure 1 Flowchart of the data labeling method according to an embodiment of the present invention. Figure 1 As shown, the method of this embodiment includes the following steps:
[0053] Step S101: displaying the rendering result of the Hypertext Markup Language document.
[0054] In this step, the Hyper Text Markup Language (HTML) text may be rendered through various existing methods, such as a browser, a server, etc., and the rendered result of the HTML document may be displayed.
[0055] The data annotation of this embodiment is performed based on lightweight markup language text. Therefore, after obtaining the original lightweight markup language text, the original lightweight markup language text can be converted to obtain a hypertext markup language document, and the rendering result of the hypertext markup language document can be displayed.
[0056] Figure 2 FIG. 1 is a flow chart of obtaining a hypertext markup language document in an optional implementation of an embodiment of the present invention. Figure 2 As shown, in an optional implementation of this embodiment, the lightweight markup language text can be converted and a hypertext markup language document can be obtained by the following method:
[0057] Step S201: Parse the lightweight markup language text to obtain a corresponding first abstract syntax tree.
[0058] In this step, the original lightweight markup language text can be parsed using various existing methods to obtain the abstract syntax tree (AST) of the lightweight markup language text, that is, the first abstract syntax tree. For example, the open source remark tool can be called to parse the markdown text and obtain the abstract syntax tree of the markdown text.
[0059] Among them, the abstract syntax tree of the lightweight markup language text is a tree representation of the abstract syntax structure of the lightweight markup language text. It is a data structure commonly used by development tools such as compilers and interpreters when processing lightweight markup language text, and is used to analyze lightweight markup language text.
[0060] Step S202: Convert the first abstract syntax tree into a second abstract syntax tree.
[0061] To ensure that the lightweight markup language text can be correctly displayed in a browser or rich text environment, in this step, the abstract syntax tree of the lightweight markup language text can be converted into the abstract syntax tree of the hypertext markup language document, that is, the second abstract syntax tree, through various existing methods. For example, the open source unified tool can be used to convert the abstract syntax tree of the markdown text into the abstract syntax tree of the hypertext markup language document.
[0062] Figure 3 Schematic diagram of the abstract syntax tree of the embodiment of the present invention. Figure 3 As shown, the abstract syntax tree 31 is the abstract syntax tree of the markdown text, and the abstract syntax tree 31 includes nodes 311, 312, 313, and 314. Nodes 311 and 312 are used to store the types of information such as words and images in the markdown text. Specifically, the type stored in node 311 is the first-level title (heading(1)), and the type stored in node 311 is the paragraph (paragraph); nodes 313 and 314 are used to store the content of the information in the markdown text. Specifically, node 313 is a child node of node 311, and the content stored in node 313 is "hello". Node 314 is a child node of node 312, and the content stored in node 314 is an image (image).
[0063] Abstract syntax tree 32 is the abstract syntax tree of the HTML document obtained by converting abstract syntax tree 31. Abstract syntax tree 32 includes nodes 321, 322, 323, 324, and 325. Nodes 321, 322, and 323 are used to store the types of various pieces of information in the HTML document. Specifically, node 321 stores the type of a first-level heading (h1), node 322 stores the type of a newline character (\n), and node 323 stores the type of a paragraph (p). Nodes 324 and 325 are used to store the content of various pieces of information in the HTML document. Specifically, node 324 is a child node of node 321 and stores the content "hello." Node 325 is a child node of node 323 and stores the content of an image (img).
[0064] Step S203: convert the second abstract syntax tree into a hypertext markup language document.
[0065] The second AST cannot be directly recognized by a browser or rendering engine. Therefore, in this step, the second AST can be converted into a Hypertext Markup Language document using various existing methods. For example, the open-source unified tool can be used to convert the AST of the Hypertext Markup Language document into a Hypertext Markup Language document. Depending on actual needs, the Hypertext Markup Language document can be an Hypertext Markup Language tree or an Hypertext Markup Language document.
[0066] Figure 4 FIG. 1 is a flow chart of obtaining a hypertext markup language document in an optional implementation of an embodiment of the present invention. Figure 4 As shown, in an optional implementation of this embodiment, step S203 may include the following steps:
[0067] In step S401 , each text node is decomposed into character nodes according to characters in each text information, and a positioning node of each character node is added to a second abstract syntax tree to obtain a processed second abstract syntax tree.
[0068] According to the type of information, the nodes in the second abstract syntax tree can be roughly divided into two types: one is the text node of text information, and the other is the non-text node of non-text information. Among them, text information is a linear sequence of readable characters, that is, phrases, words, sentences, etc.; non-text information is structured data or binary content in non-character form, mainly including images, and in some cases may also include charts, formulas, links, etc. In the existing data annotation methods for lightweight markup language text, the smallest unit of text information that can be selected and marked is usually a word or phrase. In actual applications, the annotator sometimes does not need to annotate the entire word or phrase, but only needs to annotate some characters in the word or phrase. Therefore, the existing data annotation methods lack flexibility.
[0069] In order to achieve refined text manipulation and control, and thus enhance the flexibility of data annotation, in this step, the text nodes of the text information in the second abstract syntax tree are further broken down into character nodes, so that the selectable elements of the text information class in the rendered result of the hypertext markup language document are updated from words or phrases to characters, and the selectable elements of the non-text information class remain non-text information. At the same time, to ensure that the annotation personnel's data annotation results for the lightweight markup language text through the rendered result of the hypertext markup language document are accurate, this step also adds a positioning node for each character node in the second abstract syntax tree, thereby obtaining a processed second abstract syntax tree.
[0070] In this embodiment, the positioning node of a character node is used to store the positioning element of the corresponding character, and the positioning element of each character is determined based on the position of the character in the text node of the original text information. The positioning element can be a string generated based on the position of the character in the text information and through various existing methods. For example, the positioning element of the character can be generated based on the position of the text information in a lightweight markup language text and the position of the character in the text information and according to predetermined encoding rules. The position of the character in the text information can also be directly determined as the positioning element of the character, etc. This embodiment does not impose any restrictions on this.
[0071] Figure 5 Schematic diagram of the structure of the abstract syntax tree of the embodiment of the present invention. Figure 3 , node 313 is the text node of the text information "hello" in the abstract syntax tree 32. Figure 5 As shown, the node 313 can be disassembled into five character nodes, namely the "h" node (also known as node 511), the "e" node (also known as node 512), the "l" node (also known as node 513), the "l" node (also known as node 514) and the "o" node (also known as node 515), and the position of each character in the text information "hello" can be determined, namely "h" is the first position in "hello", "e" is the second position in "hello", "l" is the third and fourth positions in "hello", and "o" is the fifth position in "hello". Furthermore, according to the position of each character in the text information "hello", the positioning nodes of nodes 511 to 515 can be added to the abstract syntax tree 32, namely node 521 corresponding to node 511, node 522 corresponding to node 512, node 523 corresponding to node 513, node 524 corresponding to node 514, and node 525 corresponding to node 515. Thus, it can be obtained Figure 5 The processed second abstract syntax tree shown is abstract syntax tree 32 ′.
[0072] Step S402: converting the processed second abstract syntax tree into a hypertext markup language document.
[0073] In this step, the processed second abstract syntax tree can be converted into a hypertext markup language document through various existing methods. Specifically, the implementation of this step can refer to step S203 and will not be repeated here.
[0074] Step S102 : In response to detecting a selection operation on a selectable element, determining a selected element corresponding to the selection operation.
[0075] After the lightweight markup language text is processed as described above and a rendered result of a hypertext markup language document is generated, the marker can select an element to be annotated from the selectable elements. Therefore, this embodiment can detect the marker's selection operation on each selectable element in the rendered result of the hypertext markup language document. If a selection operation is detected on any selectable element, this embodiment can determine the selectable element corresponding to the selection operation as a selected element.
[0076] Optionally, for text information, the selection operation of this embodiment can be various existing text selection operations, such as a full selection operation for the display page of the rendering result, a long press and drag operation with the position of any selectable element in the display page of the rendering result as the operation starting point and the position of another selectable element in the display page of the rendering result as the operation ending point, a double-click operation generated at the position of any selectable element in the display page of the rendering result, etc. This embodiment does not impose any restrictions on this.
[0077] Optionally, for non-text information, in the process of rendering a hypertext markup language document, each non-text information is usually treated as a whole. Therefore, the selection operation of this embodiment can be various existing non-text selection operations, such as a single-click operation generated in the display area of any selectable element in the display page of the rendering result, etc. This embodiment does not impose any restrictions on this.
[0078] Figure 6 This is a schematic diagram of an interface according to an embodiment of the present invention. Figure 6 Page 60 is a display page of the rendered result of a hypertext markup language document. Page 60 displays text information and non-text information, wherein the non-text information is image P6, and the display area of image P6 is area 62. When a long press and drag operation is detected with the position of the selectable element "t" in the text information "the" in page 60 as the operation starting point and the position of the selectable element "o" in the text information "Markdown" in page 60 as the operation ending point, the terminal device can determine that the text interval corresponding to the long press and drag operation is text interval 61, and then determine that each selectable element in text interval 61 is a selected element. When it is detected that the cursor hovers in area 62 for more than a preset time (such as 1 second), the terminal device can display a selection control 63 of image P6 in the upper right corner of area 62. If it is detected that the selection control 63 is triggered, such as a single click operation on the selection control 63, the terminal device can determine that the selectable element image P6 in area 62 is a selected element.
[0079] Step S103: determining a target element in the lightweight markup language text according to the selected element.
[0080] In this step, the positioning element of each selected element can be determined in the second abstract syntax tree, and the element with a mapping relationship with the selected element can be determined in the first abstract syntax tree through the positioning element, and then the element can be determined as the target element in the lightweight markup language text.
[0081] Step S104 , in response to detecting a marking operation on the selected element, generating a first label corresponding to the target element according to the marking operation, and displaying a second label of the selected element according to the display position of the selected element.
[0082] In this step, the annotator can annotate the selected element, so this embodiment can detect the annotation operation on the selected element, and when the annotation operation is detected, generate a label corresponding to the target element (i.e., the first label) according to the annotation operation, and display the label of the selected element (i.e., the second label) according to the display position of the selected element. It is easy to understand that the first label and the second label can be the same label or different labels, and this embodiment does not limit this.
[0083] In this embodiment, the labeler can select a labeling template for a lightweight markup language text, and the labeling template is provided with problem description information and corresponding problem options for each problem, and the problem options can be single-choice options or multiple-choice options. The labeling template can be created by the labeler according to business needs, or it can be pre-set by the system, and this embodiment does not limit this. For example, the problem description information can include the problem types existing in the text, the problem types existing in the non-text, and the associated problem types between the text and the non-text, wherein the problem types existing in the text can include unclear semantics, non-factual, spelling errors, etc., and the problem types existing in the non-text can include image logic errors (such as human body structure errors, object relationship contradictions, perspective errors, light and shadow errors, etc.), inconsistent multi-image styles, blurred images, garbled charts, non-existent mathematical formulas, etc.; the associated problem types between text and non-text can include irrelevant images and texts, inaccurate or missing image details, and mismatch between image and text styles.
[0084] Therefore, this embodiment can optionally display a label setting page according to the annotation template configured by the annotator, and display option controls and confirmation controls for the question options of each question in the annotation template on the label setting page. If at least one option control is detected to be selected and the confirmation control is triggered, it can be considered that an annotation operation for the selected element has been detected. Therefore, a label for the selected element can be generated based on the question option corresponding to the selected option control, and the label of the selected element can be displayed based on the display position of the selected element.
[0085] Figure 7-Figure 8 This is another interface schematic diagram of an embodiment of the present invention. Figure 7The interface 70 shown is a label setting page. Figure 6 After the terminal device determines that the selected element is image P6, if Figure 7 As shown, the annotator selects the option control 71 of the question option "Blurred Picture" corresponding to the question description information "Known Picture Question Type" and clicks the "Next Question" control 72 (i.e., the confirmation control). The terminal device can detect that the option control 71 is selected and the "Next Question" control 72 is triggered, and generates a label corresponding to the target element image P6 according to the question option "Blurred Picture", i.e., "Blurred Picture". At the same time, Figure 8 As shown, the terminal device displays a corresponding label 81, "Image Blur," based on the display position of the selected element image P6 on page 60, for example, in the upper left corner of image P6. Similarly, a label 82 corresponding to text section 61 is generated based on the question options "123456." The specific implementation method can be determined by referring to the above method and will not be repeated here.
[0086] Through the data annotation method of this embodiment, the annotator can not only implement data annotation of the original lightweight markup language text in a more refined manner in the rendering result of the converted hypertext markup document, but can also view the actual annotation results in the rendering result of the hypertext markup document. Therefore, it can reduce the time spent by the annotator in understanding the lightweight markup language text and reduce the deviation between the content of the lightweight markup language text and the actual rendering result, thereby improving the efficiency, accuracy and flexibility of data annotation.
[0087] In an optional implementation, the method of this embodiment may further include the following steps:
[0088] Step S105 : generating a lightweight markup language text annotation file according to the target element and the tag attribute information of the corresponding first tag.
[0089] In this optional implementation, to prevent tags generated by data annotation from modifying the original lightweight markup language text, a separate annotation file is generated in this step to store the metadata of the first tag. The metadata of the first tag refers to the tag attribute information of the target element and the corresponding first tag. Alternatively, the annotation file can store the tag metadata in various formats, such as a list of annotations.
[0090] The metadata of the first tag may include at least one of the identifier of the first tag, question option information corresponding to the first tag, scope information of the first tag, color of the first tag, name of the first tag, text content and tag type of the first tag. The identifier of the first tag may be a string generated according to a predetermined encoding rule; the question option information corresponding to the first tag may include classification information of the first tag and identifiers of the corresponding question options, wherein the classification information of the first tag is used to indicate whether the first tag is single-choice or multiple-choice, and the identifiers of the question options are identifiers of the question options used to generate the first tag; the range information of the first tag is used to indicate the target element and may be represented by the first character and the last character of the target element; the color of the first tag may be represented by a number in the HSL color representation model, where H, or hue, represents hue and ranges between 0° and 360°; S, or saturation, represents saturation and ranges between 0% and 100%; and L, or lightness, represents brightness and ranges between 0% and 100%; the name of the first tag may be the same as the identifier of the first tag or may be a string generated by other means, which is not limited in this embodiment; the text content is also lightweight markup language text; the tag type of the first tag is used to indicate the type of the target element, that is, text information or non-text information, which can be further divided into multiple types such as text, image, and link.
[0091] In an optional implementation, the method of this embodiment may further include the following steps:
[0092] Step S106 : in response to detecting a viewing operation on the lightweight markup language text, displaying the lightweight markup language text and the first tag of the target element.
[0093] In this optional implementation, if the labeler needs to view the labeling results of the original lightweight markup language text, the labeler can receive a viewing operation for the lightweight markup language text through various existing controls. After detecting this operation, the embodiment can also display the lightweight markup language text and display the corresponding first label according to the display position of each target element. In this way, the labeler can simultaneously view the labeling results in the original lightweight markup language text and the converted hypertext markup language document, thereby effectively improving the labeler's user experience.
[0094] It is easy to understand that in order to further improve the convenience for labelers to view the marking results and real marking results of lightweight markup language text, and to improve the convenience of data annotation, the display page of the above-mentioned lightweight markup language text, the display page of the rendering results of the hypertext markup language document, and the tag setting page can be displayed as sub-pages side by side on the same page.
[0095] Figure 9 This is another interface diagram of an embodiment of the present invention. Figure 8 The terminal device can also display a control 83 in the interface, such as in page 60, and the control 83 is used to receive a viewing operation for the lightweight markup language text. If the marker triggers the control 83 and turns the control 83 on, such as Figure 9 As shown, the terminal device can display the original lightweight markup language text in page 90, and display the corresponding first label 91 according to the display position of the target element corresponding to each selected element in the text interval 61 in page 90, and can also display the corresponding first label 92 according to the display position of the target element corresponding to image P6 in page 90.
[0096] The embodiment of the present invention displays the rendering result of a hypertext markup language document including multiple selectable elements, and when a selection operation for the selectable elements is detected, determines the selected elements among the selectable elements, and then determines the target element in the lightweight markup language text corresponding to the hypertext markup language document based on the selected elements, so that when a marking operation for the selected elements is detected, a label of the target element is generated according to the marking operation, and the label of the selected element is displayed according to the display position of the selected element. The embodiment of the present invention reduces the difficulty of understanding the lightweight markup language text and the complexity of data marking by converting the lightweight markup language text into a hypertext markup language document, thereby effectively improving the efficiency and accuracy of data marking.
[0097] Figure 10 Schematic diagram of a data tagging device according to an embodiment of the present invention. Figure 10 As shown, the data annotation device of this embodiment includes a result display unit 1001 , a first element determination unit 1002 , a second element determination unit 1003 and a generation display unit 1004 .
[0098] Among them, the result display unit 1001 is used to display the rendering result of the hypertext markup language document, which is obtained by converting the lightweight markup language text, and the rendering result includes multiple selectable elements; the first element determination unit 1002 is used to determine the selected element corresponding to the selection operation in response to detecting a selection operation for the selectable element; the second element determination unit 1003 is used to determine the target element in the lightweight markup language text according to the selected element; the generation display unit 1004 is used to generate a first label corresponding to the target element according to the marking operation in response to detecting a marking operation for the selected element, and display the second label of the selected element according to the display position of the selected element.
[0099] Furthermore, the hypertext markup language document is converted by a text parsing unit, a syntax tree conversion unit and a document conversion unit.
[0100] Among them, the text parsing unit is used to parse the lightweight markup language text to obtain the corresponding first abstract syntax tree; the syntax tree conversion unit is used to convert the first abstract syntax tree into a second abstract syntax tree, and the second abstract syntax tree is the abstract syntax tree of the hypertext markup language document; the document conversion unit is used to convert the second abstract syntax tree into the hypertext markup language document.
[0101] Furthermore, the rendering result includes text information and non-text information, and the selectable elements are characters in the text information or the non-text information.
[0102] Furthermore, the second abstract syntax tree includes text nodes of each of the text information;
[0103] The document conversion unit includes a syntax tree processing subunit and a document conversion subunit.
[0104] Among them, the syntax tree processing subunit is used to decompose each text node into character nodes according to the characters in each text information, and add a positioning node of each character node in the second abstract syntax tree to obtain the processed second abstract syntax tree, and the positioning node is used to store the positioning element of each character, and the positioning element is determined according to the position of the character in the text information; the document conversion subunit is used to convert the processed second abstract syntax tree into the hypertext markup language document.
[0105] Furthermore, the device also includes a file generating unit.
[0106] The file generating unit is configured to generate a markup file of the lightweight markup language text according to the target element and the tag attribute information of the corresponding first tag.
[0107] Furthermore, the device also includes an information display unit.
[0108] The information display unit is configured to display the lightweight markup language text and the first tag of the target element in response to detecting a viewing operation on the lightweight markup language text.
[0109] The embodiment of the present invention displays the rendering result of a hypertext markup language document including multiple selectable elements, and when a selection operation for the selectable elements is detected, determines the selected elements among the selectable elements, and then determines the target element in the lightweight markup language text corresponding to the hypertext markup language document based on the selected elements, so that when a marking operation for the selected elements is detected, a label of the target element is generated according to the marking operation, and the label of the selected element is displayed according to the display position of the selected element. The embodiment of the present invention reduces the difficulty of understanding the lightweight markup language text and the complexity of data marking by converting the lightweight markup language text into a hypertext markup language document, thereby effectively improving the efficiency and accuracy of data marking.
[0110] Figure 11 Schematic diagram of an electronic device according to an embodiment of the present invention. In this embodiment, the electronic device 11 includes a server, a terminal, etc. Figure 11 As shown, the electronic device 11: includes at least one processor 1101; and a memory 1102 communicatively connected to the at least one processor 1101; and a communication component 1103 communicatively connected to the scanning device, and the communication component 1103 receives and sends data under the control of the processor 1101; wherein the memory 1102 stores instructions that can be executed by the at least one processor 1101, and the instructions are executed by the at least one processor 1101 to implement the above-mentioned arbitration method.
[0111] Specifically, the electronic device includes: one or more processors 1101 and a memory 1102, Figure 11 A processor 1101 is taken as an example. The processor 1101 and the memory 1102 may be connected via a bus or other means. Figure 11 In the example above, a bus connection is used. Memory 1102, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer executable programs, and modules. Processor 1101 executes the non-volatile software programs, instructions, and modules stored in memory 1102 to execute various functional applications and data processing of the device, thereby implementing the aforementioned arbitration method.
[0112] The memory 1102 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and application programs required for at least one function; the data storage area may store a list of options, etc. In addition, the memory 1102 may include a high-speed random access memory and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some embodiments, the memory 1102 may optionally include a memory remotely located relative to the processor 1101, and these remote memories may be connected to an external device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0113] One or more modules are stored in the memory 1102 , and when executed by one or more processors 1101 , perform the arbitration method in any of the above method embodiments.
[0114] The above-mentioned product can execute the method provided in the embodiment of this application, and has the functional modules and beneficial effects corresponding to the execution method. For technical details not fully described in this embodiment, please refer to the method provided in the embodiment of this application.
[0115] This embodiment of the present invention determines the channel that issued the request in the current arbitration round, obtains the arbitration status of each of the channels in the current arbitration cycle of the current arbitration round and the arbitration decision result of the previous arbitration cycle, and then determines the channel that obtained authorization in the current arbitration cycle from among the channels based on this information. The arbitration decision result of the previous arbitration cycle includes the arbitration status of the authorized channel in the previous arbitration cycle. The arbitration status of each channel is determined based on its weight and authorization count. Therefore, this embodiment of the present invention can avoid frequent switching of arbitration states by the arbitration unit, thereby effectively reducing the overhead caused by arbitration switching and improving arbitration efficiency and flexibility.
[0116] Another embodiment of the present invention relates to a non-volatile storage medium for storing a computer-readable program, wherein the computer-readable program is used to enable a computer to execute part or all of the above method embodiments.
[0117] That is, those skilled in the art will understand that all or part of the steps in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a program, which is stored in a storage medium and includes a number of instructions for causing a device (which may be a single-chip microcomputer, chip, etc.) or a processor to execute all or part of the steps of the methods described in the embodiments of the present application. The aforementioned storage medium includes: a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc., various media that can store program code.
[0118] The foregoing is merely a preferred embodiment of the present application and is not intended to limit the present application. Persons skilled in the art will readily appreciate that various modifications and variations are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present application are intended to be within the scope of protection of the present application.
Claims
1. A data annotation method, characterized in that: The method comprises: Displaying a rendering result of a hypertext markup language document, the hypertext markup language document being obtained by converting a lightweight markup language text, the rendering result including a plurality of selectable elements; In response to detecting a selection operation on the selectable element, determining a selected element corresponding to the selection operation; Determining a target element in the lightweight markup language text according to the selected element; In response to detecting a marking operation on the selected element, a first label corresponding to the target element is generated according to the marking operation, and a second label of the selected element is displayed according to a display position of the selected element.
2. The method according to claim 1, characterized in that The hypertext markup language document is obtained by converting: Parsing the lightweight markup language text to obtain a corresponding first abstract syntax tree; Converting the first abstract syntax tree into a second abstract syntax tree, where the second abstract syntax tree is an abstract syntax tree of the hypertext markup language document; The second abstract syntax tree is converted into the hypertext markup language document.
3. The method according to claim 2, characterized in that The rendering result includes text information and non-text information, and the selectable elements are characters in the text information or the non-text information.
4. The method according to claim 3, characterized in that The second abstract syntax tree includes a text node for each of the text information; The converting the second abstract syntax tree into the hypertext markup language document comprises: Decomposing each text node into a character node according to characters in each text message, and adding a positioning node for each character node in the second abstract syntax tree to obtain a processed second abstract syntax tree, wherein the positioning node is used to store a positioning element for each character, and the positioning element is determined according to a position of the character in the text message; The processed second abstract syntax tree is converted into the hypertext markup language document.
5. The method according to claim 1, wherein The method further comprises: A markup file of the lightweight markup language text is generated according to the target element and the tag attribute information of the corresponding first tag.
6. The method according to claim 1, characterized in that The method further comprises: In response to detecting a viewing operation on a lightweight markup language text, the lightweight markup language text and the first tag of the target element are displayed.
7. A data labeling device, characterized in that: The device comprises: A result display unit, configured to display a rendering result of a hypertext markup language document, the hypertext markup language document being obtained by converting a lightweight markup language text, the rendering result including a plurality of selectable elements; a first element determining unit, configured to, in response to detecting a selection operation on the selectable element, determine a selected element corresponding to the selection operation; A second element determining unit, configured to determine a target element in the lightweight markup language text according to the selected element; A generating and displaying unit is configured to, in response to detecting a marking operation on the selected element, generate a first label corresponding to the target element according to the marking operation, and display a second label of the selected element according to a display position of the selected element.
8. An electronic device comprising a memory and a processor, characterized in that: The memory is configured to store one or more computer program instructions, wherein the one or more computer program instructions are executed by the processor to implement the method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.
10. A computer program product, characterized in that The computer program product comprises a computer program / instructions, which implement the method according to any one of claims 1 to 6 when executed by a processor.