Rich text processing method, device, equipment and storage medium

By replacing and fusion rich texts with labels, the problem of format and content integrity when rich text is passed in different systems and multi-language environments in the prior art is solved, and more efficient and accurate rich text processing is achieved.

CN118607478BActive Publication Date: 2025-05-16BEIJING YOUZHUJU NETWORK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410781728.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-17
Publication Date
2025-05-16
Estimated Expiration
2044-06-17

AI Technical Summary

Technical Problem

When processing rich text, it is difficult for the prior art to maintain the consistency of data between different systems and platforms. Especially in multi-language environments, conventional solutions are difficult to ensure the integrity of format and content, resulting in errors and loss during data transmission.

Method used

By replacing the tags in the source rich text with predefined tags and fusing them, the fused tags are obtained, and the to be processed based on these tags.

Benefits of technology

This method can retain the rich text structure, reduce the number of text characters that the model needs to process, ensure the complete structure of the sentence and the associated context information, improve processing accuracy and efficiency, and have strong adaptability and flexibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118607478B_ABST
    Figure CN118607478B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide a rich text processing method, apparatus, device and storage medium. In the method, at least one first tag in the source rich text is replaced with at least one predefined second tag. Each first tag corresponds to one or more second tags, and the source rich text includes at least one first tag and a text to be processed associated with the at least one first tag. In the method, at least one third tag is obtained by performing a fusion operation on at least one second tag, and the text to be processed is processed based on the at least one third tag. The number of at least one third tag is less than or equal to the number of at least one second tag. In this way, the number of tags related to the text to be processed is reduced, thereby improving the accuracy and efficiency of rich text processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Example embodiments of the present disclosure generally relate to the field of computer technology, and more particularly, to a method, apparatus, device, and computer-readable storage medium for processing rich text. Background Art

[0002] With the rapid development of computer technology, the demand for the transmission and processing of various types of data on digital platforms is increasing. These data may contain a variety of information, such as text, pictures, links, tables, etc., which increases the richness and complexity of information display. In actual operation, how to maintain consistency between different systems and platforms, and accurately transmit information in a multilingual environment, has become an important challenge. Conventional solutions often find it difficult to ensure the integrity of the format and content when processing data and multilingual conversion, resulting in errors and loss of data during the transmission process, which is undesirable. Therefore, the effective transmission of data needs to be further improved.

[0003] With the increase of information volume and the diversification of application scenarios, the difficulty of data processing is also increasing. Therefore, it is of great practical significance to develop more intelligent and efficient data processing technology to cope with various complex application requirements. Summary of the invention

[0004] In a first aspect of the present disclosure, a rich text processing method is provided. The method comprises: replacing at least one first tag in a source rich text with at least one predefined second tag, wherein each first tag corresponds to one or more second tags, and the source rich text comprises at least one first tag and text to be processed associated with the at least one first tag; obtaining at least one third tag by performing a fusion operation on the at least one second tag, wherein the number of the at least one third tag is less than or equal to the number of the at least one second tag; and processing the text to be processed based on the at least one third tag.

[0005] In a second aspect of the present disclosure, a rich text processing device is provided. The device includes: a replacement module, configured to replace at least one first tag in a source rich text with at least one predefined second tag, wherein each first tag corresponds to one or more second tags, and the source rich text includes at least one first tag and a text to be processed associated with the at least one first tag; a fusion module, configured to obtain at least one third tag by performing a fusion operation on at least one second tag, wherein the number of the at least one third tag is less than or equal to the number of the at least one second tag; and a processing module, configured to process the text to be processed based on the at least one third tag.

[0006] In a third aspect of the present disclosure, an electronic device is provided. The device includes: at least one processing unit; and at least one memory, the at least one memory is coupled to the at least one processing unit and stores instructions for execution by the at least one processing unit, and when the instructions are executed by the at least one processing unit, the device executes the method of the first aspect.

[0007] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided, wherein a computer program is stored on the medium, and when the computer program is executed by a processor, the method of the first aspect is implemented.

[0008] It should be understood that the content described in this section is not intended to limit the key features or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, wherein:

[0010] Figure 1 A schematic diagram showing an example rich text processing in which embodiments of the present disclosure can be implemented;

[0011] Figure 2 A flowchart showing a process for rich text processing according to some embodiments of the present disclosure is shown;

[0012] Figure 3A and Figure 3B Schematic diagrams of label replacement processes according to some embodiments of the present disclosure are respectively shown;

[0013] Figure 4 A flowchart for a label fusion process according to some embodiments of the present disclosure is shown;

[0014] Figure 5 A schematic diagram of a tag fusion process according to some embodiments of the present disclosure is shown;

[0015] Figure 6 A schematic diagram showing a label fusion process according to other embodiments of the present disclosure is shown;

[0016] Figure 7 A schematic diagram showing a process of rich text processing according to some other embodiments of the present disclosure;

[0017] Figure 8 A block diagram showing a rich text processing apparatus according to some embodiments of the present disclosure; and

[0018] Fig. 9 A block diagram of an electronic device is shown in which one or more embodiments of the present disclosure may be implemented. DETAILED DESCRIPTION

[0019] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments set forth herein. On the contrary, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not intended to limit the scope of protection of the present disclosure.

[0020] In the description of the embodiments of the present disclosure, the term "including" and similar terms should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "based at least in part on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may also be included below.

[0021] Herein, unless explicitly stated, executing a step “in response to A” does not mean executing the step immediately after “A” but may include one or more intermediate steps.

[0022] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and relevant provisions.

[0023] It is understandable that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, scope of use, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0024] For example, in response to receiving an active request from a user, a prompt message is sent to the user to clearly prompt the user that the operation requested to be performed will require obtaining and using the user's personal information, so that the user can independently choose whether to provide personal information to software or hardware such as electronic devices, applications, servers or storage media that execute operations of the technical solution of the present disclosure based on the prompt message.

[0025] As an optional but non-limiting implementation, in response to receiving an active request from the user, the prompt information is sent to the user in a manner such as a pop-up window, in which the prompt information can be presented in text form. In addition, the pop-up window can also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0026] It is understandable that the above notification and the process of obtaining user authorization are merely illustrative and do not constitute a limitation on the implementation of the present disclosure. Other methods that meet relevant laws and regulations may also be applied to the implementation of the present disclosure.

[0027] As used herein, the term "rich text" may refer to text that contains multiple formatting elements, such as text, images, links, tables, etc., to enhance the expression and presentation of information. Rich text can be used in various documents and web pages, and is implemented using markup languages ​​such as HTML. It should be understood that rich text is not limited to HTML, but may also include a variety of other formats and implementations, such as Markdown, RTF, etc.

[0028] The term "tag" may refer to a code tag used to define an element in a markup language, and is widely used in documents such as HTML and XML. It should be understood that tags are not limited to HTML or XML, but may also be used in other markup languages ​​and application scenarios.

[0029] The term "self-closing tag" may refer to a type of tag. A self-closing tag completes its definition and ends within a tag, without requiring a separate end tag. Self-closing tags are often used for elements that do not contain content, such as images, line breaks, input fields, and so on. In HTML, self-closing tags (also known as empty elements, single tags) are, for example, tags that have no content and do not require a matching end tag to close. Such tags are used, for example, to introduce media, define controls, or serve as layout separators. They are indicated as self-closing by adding a slash " / " at the end of the content of the start tag, for example, , etc.

[0030] The term "non-self-closing tag" may refer to another type of tag, which may also include two types of tags, namely, a start tag (also called an "open tag") and an end tag (also called a "closed tag"). Start tags and end tags are used in pairs, for example, to enclose certain content. These tags define the element of the enclosed content and mark the end of the element through a closing tag. The content of the start tag and the corresponding end tag are the same, the difference is, for example, that the corresponding closing tag is indicated by adding a slash " / " at the beginning of the content of the start tag. For example, if the start tag is , then the corresponding end tag is .

[0031] The term "model" can learn the association between the corresponding input and output from the training data, so that the corresponding output can be generated for the given input after the training is completed. The generation of the model can be based on machine learning technology. Deep learning is a machine learning algorithm that processes inputs and provides corresponding outputs by using multiple layers of processing units. A neural network model is an example of a model based on deep learning. In this article, "model" may also be referred to as "machine learning model", "learning model", "machine learning network" or "learning network", and these terms are used interchangeably in this article.

[0032] Generally, machine learning can be roughly divided into three stages, namely the training stage, the testing stage, and the application stage (also called the inference stage). In the training stage, a given model can be trained using a large amount of training data, and the parameter values ​​are continuously updated iteratively until the model can obtain consistent inferences that meet the expected goals from the training data. Through training, the model can be considered to be able to learn the association between input and output (also called input-to-output mapping) from the training data. The parameter values ​​of the trained model are determined. In the testing stage, the test input is applied to the trained model to test whether the model can provide the correct output, thereby determining the performance of the model. The testing stage can sometimes be integrated into the training stage. In the application or inference stage, the trained model can be used to process the actual model input based on the parameter values ​​obtained through training to determine the corresponding model output.

[0033] Current machine learning models are learning complex language patterns by training on large-scale text data to perform language tasks such as translation and summarization. They capture detailed language patterns through a huge number of parameters (usually in the billions).

[0034] Traditionally, there are several ways to translate rich text: manual translation, plain text translation replacement, automatic translation software translation, etc. In order to maintain the accuracy and format of rich text, manual translation is usually used. This method takes into account both text content and structure, but requires a higher cost. Plain text translation replacement refers to extracting the text content of rich text, then translating and filling it into the original format. This method takes a shorter time, but usually loses context due to structural breaks. Automatic translation software usually cannot handle complex formats, resulting in formatting and semantic errors.

[0035] There are currently many technical problems that need to be solved urgently. For example, after the text content and labels are separated, the text content becomes semantically fragmented after processing and loses contextual association, resulting in reduced accuracy of the processing results. For example, during the translation process, due to the significant differences in grammatical structure and vocabulary between different languages, the format of the processing results does not match and the word order is chaotic, resulting in changes in the original layout and structure.

[0036] According to an embodiment of the present disclosure, a scheme for rich text processing is proposed. In the scheme, the original rich text (also referred to as "source rich text") includes tags and text to be processed. The tags of the source rich text are first replaced with new tags. Then, the new tags are fused to obtain fused tags. The fused tags can then be processed together with the text content to be processed. Thereafter, the tags of the source rich text can be restored based on the fused tags and the mapping relationship between various tags, and the target rich text can be obtained based on the restored tags and the processed text.

[0037] The solution disclosed in the present invention can retain the structure of rich text and reduce the number of text characters that the model needs to process. Compared with conventional solutions, the solution disclosed in the present invention can retain the complete structure of the sentence and the associated context information, thereby overcoming the content fragmentation caused by label segmentation in conventional methods and achieving a more coherent and consistent output. In addition, through automated label processing and model processing procedures, compared with conventional manual processing, the solution achieves a significant improvement in the efficiency of rich text processing. At the same time, since the machine learning model has a wide range of knowledge and context understanding capabilities, the solution can quickly process different types of rich text and has strong adaptability and flexibility.

[0038] Some example embodiments of the present disclosure will be described below with reference to the accompanying drawings.

[0039] Figure 1 Schematic diagram of an example rich text process 100 in which embodiments of the present disclosure can be implemented. Figure 1 As shown, exemplarily, the source rich text 110 may include one or more tags and text to be processed associated with the tags. In some embodiments, the rich text may include structured text with tags, and the tags indicate the format or structure of the corresponding text.

[0040] The rich text processing device 120 is used to process the source rich text 110, for example, to translate, generate text, perform sentiment analysis, classify the text to be processed in the source rich text 110. It should be understood that the above examples of processing the text to be processed are merely illustrative and not restrictive. In the embodiments of the present disclosure, other appropriate or necessary processing may be performed on the text to be processed. The embodiments of the present disclosure are not limited to this.

[0041] After the source rich text 110 passes through the rich text processing device 120, a target rich text 130 is obtained. The target rich text 130 also includes one or more tags and processed text. In some embodiments, the processing performed on the text may be, for example, translating the text. In this case, the source rich text 110 may include tags and text in a first language associated with the tags, for example, English text. The target rich text 130 may include tags and text in a second language, for example, Chinese text. The second language text in the target rich text 130 has a word order that conforms to the usage habits of the second language, and the association relationship between the second language text and the tag is the same as that in the source rich text 110. In other words, the target rich text 130 has the same layout and structure as the source rich text 110. It should be understood that in Figure 1 The rich text processing 120 shown in FIG. 1 is exemplary only and is not intended to be limiting in any way.

[0042] Figure 1 The rich text processing device 120 shown can be implemented with any type of electronic device, such as a terminal device, a server device or other device. Examples of terminal devices include mobile terminals, fixed terminals or portable terminals, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, personal communication systems (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio broadcast receivers, e-book devices, gaming devices or any combination of the foregoing, including accessories and peripherals of these devices or any combination thereof. In some embodiments, the electronic device can also support any type of interface for the user (such as "wearable" circuits, etc.). Examples of server devices include, but are not limited to, mainframes, edge computing nodes, computing devices in cloud environments, and the like.

[0043] Figure 2 FIG. 2 is a flowchart of a process 200 for rich text processing according to some embodiments of the present disclosure. The process 200 may be performed by, for example Figure 1 The rich text processing device 120 or other suitable device is executed. In the following, for the purpose of explanation, reference is made to Figure 1 The process 200 is described below.

[0044] In box 210, the rich text processing device 120 replaces the first tag in the source rich text with a predefined second tag. In some embodiments, the first tag in the source rich text may have corresponding content and type, and may be extracted. If the type of the extracted first tag is a self-closing type, the first tag may be replaced with a paired second tag. In this case, one of the paired second tags is an open tag, the other is a closed tag, and the contents of the paired second tags are the same. Figure 3A and Figure 3B Detailed description.

[0045] Figure 3A and Figure 3B Schematic diagrams of the rich text processing process according to some embodiments of the present disclosure are respectively shown. Figure 3A As shown, the first tag 310 in the source rich text 305 has content and a type, wherein the type is a self-closing type. The first tag 310 is replaced by a second tag 330 and a second tag 340, wherein the second tag 330 is an open tag, the second tag 340 is a closed tag, and the second tag 330 has the same content as the second tag 340, and the content of the paired second tags 330 and 340 corresponds to the content of the first tag 310. In other embodiments, if the type of the extracted first tag is a non-self-closing type, the first tag may be replaced by a second tag, wherein the content of the second tag corresponds to the content of the first tag.

[0046] Unlike the self-closing tags above, Figure 3B A non-self-closing type of label is shown. Specifically, Figure 3B As shown, the first tag 350 in the source rich text 350 is a non-self-closing type, which is replaced by a second tag 365, and the content of the second tag 365 corresponds to the content of the first tag 350. The first tag 360 in the source rich text 350 is a non-self-closing type, which is replaced by a second tag 370, and the content of the second tag 370 corresponds to the content of the first tag 360. In the process of box 210, a mapping relationship between the first tag and the second tag will be established, and used to restore the processed rich text to the target rich text. The process of restoring using the mapping relationship between the first tag and the second tag will be described in detail below.

[0047] Now return to Figure 2 In block 220, the rich text processing device 120 performs a fusion operation on the second tag in the rich text and obtains a third tag. The specific method and embodiment of the fusion process will be described below in conjunction with Figures 4 to 7 Further description: After the fusion process, the number of third tags in the rich text is less than or equal to the number of second tags.

[0048] In block 230 , the rich text processing device 120 processes the text to be processed based on the third tag. In some embodiments, processing the text to be processed may include translation, text generation, sentiment analysis, text classification, and the like.

[0049] In some embodiments, optionally, before frame 210, the rich text processing device 120 may perform text preprocessing on the source rich text. The text preprocessing may remove any zero-width spaces or other special characters in the source rich text. Through preprocessing, the influence of zero-width spaces and other special characters on parsing may be eliminated to prepare for the processing in frames 210 to 230.

[0050] Compared with conventional solutions, by replacing and merging the tags in the rich text, the number of tags related to the text to be processed is reduced, effectively avoiding the interference of tags in the processing of the text to be processed, thereby improving the accuracy and efficiency of rich text processing.

[0051] As mentioned above, we will then combine Figures 4 to 7 The specific methods and embodiments of the fusion process are further described. In some embodiments, reference Figure 4 , shows a flowchart for a tag fusion process 400 according to some embodiments of the present disclosure. It should be understood that Figure 4 The embodiments shown in are exemplary only and are not intended to be limiting in any way.

[0052] In block 410, N pairs of second tags in the rich text are identified, where N ≥ 1. One of each pair of second tags is an open tag, and the other is a closed tag. In some embodiments, a node list can be constructed based on the second tags and the text to be processed, each of the second tags can be determined as a node, and the text to be processed between two adjacent second tags is determined as a node.

[0053] By traversing the node list, all N pairs of second labels can be found. Specifically, if a first node corresponding to an open label in the node list is detected, a second node can be found from the node list. In this case, the second node corresponds to a closed label that matches the open label.

[0054] Next, the steps in the following blocks 415-450 are iteratively performed.

[0055] In block 415 , determine whether there are any undetermined pairs of second tags in the N pairs of second tags. When block 415 is executed for the first time, since N is greater than or equal to 1, there must be undetermined pairs of second tags, so process 400 proceeds to block 420 .

[0056] In block 420, for the i-th pair of second labels, where the initial value of i is set to 1, and i≤N, if there is a second label between the i-th pair of second labels, i is set to i+1, and the process proceeds to block 415. In block 415, it is determined whether there are any undetermined pairs of second labels in the N pairs of second labels. If so, there are still second labels that need to be merged, and the process 400 proceeds to block 420 again; if not, that is, all second labels are converted to third labels, and there are no second labels that need to be merged, the process 400 ends in block 416. If there is no second label between the i-th pair of second labels, the process 400 proceeds to block 430.

[0057] Figure 5 A schematic diagram of a tag fusion process according to some embodiments of the present disclosure is shown. In this embodiment, 5 pairs of second tags have been identified from rich text 501, that is, N=5. In this case, the first pair of second tags is second tag 505 and second tag 523, the second pair of second tags is second tag 507 and second tag 521, the third pair of second tags is second tag 509 and second tag 515, the fourth pair of second tags is second tag 511 and second tag 513, and the fifth pair of second tags is second tag 517 and second tag 519.

[0058] Continue to refer Figure 4 The frame 420 is Figure 5 Taking the illustrated embodiment as an example, when i is set to 1, there are other second tags between the first pair of second tags 505 and 523, so i is set to 2, and block 420 is repeated until i is set to 4. Since there is no second tag between the fourth pair of tags, second tag 511 and second tag 513, process 400 proceeds to block 430.

[0059] In block 430, for the i-th pair of second tags, if the i-th pair of tags is preceded and followed by a pair of second tags, the process 400 proceeds to block 440, the i-th pair of tags is determined to be a fusible tag, and the fusible tag does not belong to the second tag, then i is reset to 1 and returns to block 420. If the i-th pair of tags is not preceded and followed by a pair of second tags, the process 400 proceeds to block 450. In this case, the i-th pair of tags may be preceded and followed by second tags, but the two second tags are not a pair; or there is no tag at least one of the front and back of the i-th pair of tags. In block 450, based on the i-th pair of second tags and the fusible tags associated with the i-th pair of tags, a pair of third tags is determined, then i is reset to 1 and returns to block 420.

[0060] Specifically, Figure 5As a reference to the embodiment, in box 430, the fourth pair of tags, the second tag 511 and the second tag 513, are preceded and followed by the second tag 509 and the second tag 515, respectively, and the second tag 509 and the second tag 515 are paired, at which time the process 400 proceeds to box 440, and the fourth pair of tags is determined to be fusible tags. Next, i is reset to 1, and box 420 is continued until i is set to 3. The third pair of second tags, i.e., the second tag 509 and the second tag 515, only have the fusible tags 511 and 513 and text between them, and there is no second tag. Therefore, the process 400 proceeds to box 430. Since the third pair of second tags 509 and 515 are preceded and followed by the second tag 507 and the second tag 517, respectively, the two second tags are not paired, so the process 400 proceeds to box 450.

[0061] At block 450 , the third pair of second tags 509 and 515 and the fusible tags 511 and 513 are determined as a third tag 529 . Similarly, the fifth pair of second tags 517 and 519 are determined as third tags 533 and 535 .

[0062] When the second pair of tags 507 and 521 are determined to be fusible tags, i is set to 1, and the process returns to 420. At this time, there are only fusible tags, third tags, and text between the first pair of second tags 505 and 523, so the process 400 proceeds to 430. At 430, since there are no tags before and after the first pair of second tags 505 and 523, the process 400 proceeds to box 450. At box 450, the first pair of second tags 505 and 523 and the fusible tags 507 and 521 are determined to be third tags 527 and 537. At this point, there are no more pairs of second tags that need to be determined, and the iterative process ends.

[0063] As an alternative, Figure 6 FIG. 2 shows a schematic diagram of a label fusion process according to some other embodiments of the present disclosure. Figure 6 As shown in , 5 pairs of second tags have been identified for rich text 501. Based on the text to be processed, a group of consecutive second tags in the second tags are merged, and there is no text to be processed between a group of consecutive second tags. In this case, second tags 505, 507, 509 and 511 will be merged, second tags 513, 515 and 517 will be merged, and second tags 519, 521 and 523 will be merged. Based on the fusion results, 3 third tags 603, 605 and 607 are obtained. It should be understood that in Figure 6 The embodiments shown in are exemplary only and are not intended to be limiting in any way.

[0064] In other embodiments, reference Figure 7, which shows a schematic diagram of the process of rich text processing according to some other embodiments of the present disclosure. It should be understood that Figure 7 The embodiment shown in is only exemplary and is not intended to be limiting. Before block 710, the source rich text is subjected to a text preprocessing stage to remove any zero-space or other special characters in the source rich text that may affect parsing.

[0065] After the processing between frame 710 and frame 720, tokenization and label classification are performed, that is, corresponding to Figure 2 The replacement process of the middle frame 210. The purpose of this step is to tokenize the cleaned rich text and split it into a node array, where each node may be an HTML opening tag, a closing tag, a self-closing tag or a plain text.

[0066] In some specific implementations, you can check each character from the beginning of the text. When a < character is detected, identify the beginning of the tag, and search backward from that position until the next > character to determine a complete tag. Then for each detected tag, first determine whether it is a legal rich text tag. If not, treat it as plain text. For legal HTML tags, then determine whether it is a self-closing tag, for example or These self-closing tags do not need to match the end tag and will be immediately processed as a pair of start and close virtual nodes to simplify subsequent translation operations and ensure that the self-closing structure does not break the node processing flow. For non-self-closing tags, their indexes are recorded in the stack so that when the corresponding closing tag is searched later, it can be removed from the stack and matched as a complete tag pair.

[0067] In this way, the tags in the text are decomposed into nodes with unique indexes and classified according to their types (start tags, end tags and plain text), wherein the text in the middle of the tag is stored as a plain text node, excluding HTML tags. In this process, if the operation of the stack does not match the tag, for example, a closed tag without a start tag is found, or there are still unclosed tags in the stack after searching the entire text, an error message will be given indicating that the tags do not match. Through these processes, the rich text in frame 720 can be obtained, and the number of invalid characters that are meaningless to translation can be greatly shortened when there are a large number of attributes in the HTML tag.

[0068] Table 1 shows an example of the mapping relationship between tags before and after replacement. This mapping relationship will be used in the process of generating the target rich text, which will be described in further detail below.

[0069] Table 1

[0070]

[0071]

[0072]

[0073] In the processing from block 720 to block 730, the labels corresponding to each node are merged, for example, corresponding to Figure 2 The fusion process of the middle frame 220. Through the node array of the previous step, the rich text has a clear segmentation mark, and the adjacent label nodes that can be merged are merged here. It should be understood that in the embodiments of the present disclosure, the term "fusion" is sometimes also referred to as "merging", both of which are ablation operations performed on labels, and there is no intention to limit the embodiments of the present disclosure in any way.

[0074] Specifically, in this fusion process, the node list is traversed to identify the opening and closing tags that form a pair, and to check whether there are other label nodes between them. If it is found that there are no other nested labels between the continuous nodes between a pair of opening and closing tags, it is regarded as a label paragraph that can be fused. Through this step of processing, the fusion simplification of complex nested label structures can be achieved, and finally a sparse structure with the minimum necessary number of labels is generated, thereby simplifying the subsequent processing process.

[0075] An exemplary pseudo code for the specific implementation of the above process is shown in Table 2. It should be understood that the pseudo code shown in Table 2 is only exemplary and is not intended to be limiting.

[0076] Table 2

[0077]

[0078]

[0079]

[0080] In some embodiments, the label fusion process in Table 2 can be further explained by the following process. First, input verification, that is, checking whether the number of input nodes is less than 2, if so, an error is returned; and verifying whether the type and value of the first and last nodes match, otherwise an error is returned. Next, node processing is performed, including initializing the result list and the start index, and traversing from the second node to process the left label, text node and its combination. Then recursively process the sublist and merged results. If a left label node is encountered, find a matching right label node, recursively merge the sublist, and merge the subresults into the result list to process the mixed conditions. Next, process the start and end nodes. If the result list is empty, create a new node to merge the first and last texts; if the result list length is 1, decide whether to mix the first and last texts into one node or create a new node; and if the result list length is greater than 1, process the text merge of the first and last nodes with the first and last nodes of the result list. Finally, return the result, that is, finally return the merged node list.

[0081] After block 730, the tag mapping text as described above can be generated. This step performs a reorganization operation to construct the merged node array into a new text with a tag mapping. In this step, the original HTML tags are replaced with temporary placeholders, such as <d0>The third tag may be in various forms, such as having different contents or different formats. For example, the third tag may be recorded as <d0> 、 <d1> 、 <d2>... and so on, other characters or letters may be used to replace D. In addition, the third label may also be named in other appropriate ways, such as <f0> 、 <f1> 、 <f2>......etc., the embodiments of the present disclosure do not impose any limitation on this.

[0082] In some embodiments, a mapping relationship table may be created for these tags to record the corresponding relationship between the placeholders and the original tags, as shown in Table 3.

[0083] Table 3

[0084]

[0085] Next, perform tag recovery and translated text recovery. In some embodiments, after translation, the temporary placeholders in the translated text will be converted back to the original HTML tags according to the tag mapping table to achieve the combination and recovery of the translated content and the original format. In some embodiments, based on the processing result of the third tag and the text to be processed and the correspondence between the second tag and the third tag, that is, as described in Table 3 above, an intermediate rich text including the second tag and the processing result can be obtained. And based on the intermediate rich text and the correspondence between the first tag and the second tag, that is, as described in Table 1 above, a processed target rich text can be generated.

[0086] In some embodiments, processing the text to be processed may include processing the fused rich text including the third label and the text to be processed using a pre-trained text processing model. In this case, the text processing model may be trained using the fused reference label and the reference text to be processed as input and using the fused reference label and the processing result for the reference text to be processed as output.

[0087] In some embodiments, the above text processing model can be implemented by a machine learning model. For example, the translation capability of the machine learning model can be used to translate the text to be processed after label replacement and fusion. Before using the text processing model, alignment training is required, such as fine-tuning the model to better process rich text content with simplified label sequences.

[0088] The purpose of the above training is to adapt the model to process texts with special marks. In some embodiments, the specific steps of training are as follows. The first is the data annotation and sample preparation process, which requires manual annotation of a large number of translation samples, which contain texts that have been processed by the aforementioned label fusion and replacement. Each sample consists of a replaced and fused version of the source rich text and the corresponding target language version. In the target language version, the text content has been translated, but the replaced label sequence is retained so that the model can learn how to complete the translation task while maintaining the text structure. Next is the model fine-tuning process. After the training data is prepared, the model is fine-tuned to enable it to learn the rich text samples after label fusion and replacement.

[0089] During the fine-tuning process, the model input is the source language text after replacement and fusion, and the output is the corresponding target language translation text, and the replaced structural labels are retained. It should be understood that the above-mentioned model fine-tuning training embodiment is only exemplary and is not intended to be limiting.

[0090] In this way, by using a machine learning model that has been fine-tuned and trained, the solution of the present disclosure achieves format and word order integrity after rich text processing, effectively avoiding the format confusion problem caused by standard extraction and backfilling strategies.

[0091] Figure 8 800 according to some embodiments of the present disclosure. The apparatus 800 may be implemented as or included in Figure 1 The rich text processing device 120 of the apparatus 800. Each module / component in the apparatus 800 may be implemented by hardware, software, firmware or any combination thereof.

[0092] As shown in the figure, the device 800 includes a replacement module 810, which is configured to replace at least one first tag in the source rich text with at least one predefined second tag, wherein each first tag corresponds to one or more second tags, and the source rich text includes at least one first tag and a text to be processed associated with the at least one first tag. The device 800 also includes a fusion module 820, which is configured to obtain at least one third tag by performing a fusion operation on at least one second tag, and the number of at least one third tag is less than or equal to the number of at least one second tag. The device 800 also includes a processing module 830, which is configured to process the text to be processed based on the at least one third tag.

[0093] In some embodiments, the replacement module 810 is also configured to: extract a first tag in the source rich text, the first tag having corresponding content and type; in response to the first tag being a self-closing type, replace the first tag with a paired second tag, one of the paired second tags being an open tag and the other being a closed tag, and the paired second tags have the same content, and the content of the paired second tags corresponds to the content of the first tag; and in response to the first tag being a non-self-closing type, replace the first tag with a second tag, and the content of the second tag corresponds to the content of the first tag.

[0094] In some embodiments, the fusion module 820 is also configured to: identify at least one pair of second tags from at least one second tag, one of the second tags in each pair is an open tag and the other is a closed tag, and the second tags of each pair have the same content; for each pair of at least one pair of second tags, iteratively perform the following steps: in response to determining that there are no other second tags between a pair of second tags in at least one pair of second tags, and the previous second tag of the previous second tag in a pair of second tags is paired with the next second tag in the pair of second tags, then determine the pair of second tags as fusible tags, wherein the fusible tags do not belong to at least one second tag; and in response to determining that there are no other second tags between a pair of second tags in at least one pair of second tags, and the previous second tag of the previous second tag in a pair of second tags is not paired with the next second tag in the pair of second tags, then determine a pair of third tags based on the pair of second tags and / or the fusible tags associated with the pair of second tags.

[0095] In some embodiments, the fusion module 820 is further configured as follows: at least one second tag and the text to be processed are used to construct a node list, wherein each of the at least one second tag is determined as a node, and the text to be processed between two adjacent second tags is determined as a node, and wherein at least one pair of second tags is identified by traversing the node list.

[0096] In some embodiments, the fusion module 820 is further configured to: in response to detecting a first node in the node list corresponding to an open tag, search for a second node in the node list, wherein the second node corresponds to a close tag matching the open tag.

[0097] In other embodiments, the fusion module 820 is further configured to: based on the text to be processed, fuse a group of continuous second tags in at least one second tag, where there is no text to be processed between the group of continuous second tags; and obtain at least one third tag based on the fusion result.

[0098] In some embodiments, the processing module 830 is further configured to: obtain an intermediate rich text including at least one second tag and the processing result based on at least one third tag and the processing result for the text to be processed and the correspondence between at least one second tag and at least one third tag; and generate a processed target rich text based on the intermediate rich text and the correspondence between at least one first tag and at least one second tag.

[0099] In some embodiments, the processing module 830 is further configured to: use a pre-trained text processing model to process a fused rich text comprising at least one third tag and a text to be processed, wherein the text processing model is trained using at least one fused reference tag and a reference text to be processed as input, and using at least one fused reference tag and a processing result for the reference text to be processed as output.

[0100] Fig. 9 900 is a block diagram of an electronic device in which one or more embodiments of the present disclosure may be implemented. It should be understood that Fig. 9 The electronic device 900 shown is merely exemplary and should not constitute any limitation on the functionality and scope of the embodiments described herein. Fig. 9 The electronic device 900 shown can be used to implement Figure 1 Rich text processing device 120 and Figure 8 Device 800, etc.

[0101] like Fig. 9 As shown, the electronic device 900 is in the form of a general-purpose computing device. The components of the electronic device 900 may include, but are not limited to, one or more processors or processing units 910, a memory 920, a storage device 930, one or more communication units 940, one or more input devices 950, and one or more output devices 960. The processing unit 910 may be an actual or virtual processor and is capable of performing various processes according to a program stored in the memory 920. In a multi-processor system, multiple processing units execute computer executable instructions in parallel to improve the parallel processing capability of the electronic device 900.

[0102] The electronic device 900 typically includes a plurality of computer storage media. Such media can be any available media accessible to the electronic device 900, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 920 can be a volatile memory (e.g., register, cache, random access memory (RAM)), a non-volatile memory (e.g., a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), flash memory) or some combination thereof. The storage device 930 can be a removable or non-removable medium, and can include a machine-readable medium, such as a flash drive, a disk, or any other medium, which can be used to store information and / or data and can be accessed within the electronic device 900.

[0103] The electronic device 900 may further include additional removable / non-removable, volatile / non-volatile storage media. Fig. 9 As shown in , a disk drive for reading or writing from a removable, non-volatile disk (e.g., a "floppy disk") and an optical drive for reading or writing from a removable, non-volatile optical disk may be provided. In these cases, each drive may be connected to the bus (not shown) by one or more data media interfaces. Memory 920 may include a computer program product 925 having one or more program modules that are configured to perform various methods or actions of various embodiments of the present disclosure.

[0104] The communication unit 940 implements communication with other electronic devices through a communication medium. Additionally, the functions of the components of the electronic device 900 can be implemented with a single computing cluster or multiple computing machines that can communicate through a communication connection. Therefore, the electronic device 900 can operate in a networked environment using a logical connection with one or more other servers, a network personal computer (PC), or another network node.

[0105] The input device 950 may be one or more input devices, such as a mouse, a keyboard, a tracking ball, etc. The output device 960 may be one or more output devices, such as a display, a speaker, a printer, etc. The electronic device 900 may also communicate with one or more external devices (not shown) through the communication unit 940 as needed, such as a storage device, a display device, etc., communicate with one or more devices that allow a user to interact with the electronic device 900, or communicate with any device that allows the electronic device 900 to communicate with one or more other electronic devices (e.g., a network card, a modem, etc.). Such communication may be performed via an input / output (I / O) interface (not shown).

[0106] According to an exemplary implementation of the present disclosure, a computer-readable storage medium is provided, on which computer-executable instructions are stored, wherein the computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, and the computer-executable instructions are executed by a processor to implement the method described above.

[0107] Various aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of the methods, devices, equipment, and computer program products implemented according to the present disclosure. It should be understood that each box in the flowchart and / or block diagram and the combination of each box in the flowchart and / or block diagram can be implemented by computer-readable program instructions.

[0108] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, so that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device that implements the functions / actions specified in one or more boxes in the flowchart and / or block diagram is generated. These computer-readable program instructions can also be stored in a computer-readable storage medium, and these instructions cause the computer, programmable data processing device, and / or other equipment to work in a specific manner, so that the computer-readable medium storing the instructions includes a manufactured product, which includes instructions for implementing various aspects of the functions / actions specified in one or more boxes in the flowchart and / or block diagram.

[0109] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, so that the instructions executed on the computer, other programmable data processing apparatus, or other device implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.

[0110] The flow chart and block diagram in the accompanying drawings show the possible architecture, function and operation of the system, method and computer program product according to multiple implementations of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a part of a module, program segment or instruction, and a part of a module, program segment or instruction includes one or more executable instructions for realizing the logical function of the specification. In some implementations as replacements, the function marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two continuous square boxes can actually be executed substantially in parallel, and they can sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be realized by a special hardware-based system that performs the function or action of the specification, or can be realized by a combination of special hardware and computer instructions.

[0111] The above descriptions of various implementations of the present disclosure are exemplary, non-exhaustive, and not limited to the disclosed implementations. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described implementations. The selection of terms used herein is intended to best explain the principles of the implementations, practical applications, or improvements to the technology in the market, or to enable other persons of ordinary skill in the art to understand the various implementations disclosed herein. < / f1> < / f0> < / d1> < / d0>

Claims

1. A rich text processing method, comprising: Replacing at least one first tag in a source rich text with at least one predefined second tag, wherein each first tag corresponds to one or more second tags, and the source rich text includes the at least one first tag and text to be processed associated with the at least one first tag; By converting the at least one second tag, at least one third tag is obtained, and the number of the at least one third tag is less than or equal to the number of the at least one second tag; as well as The text to be processed is processed based on the at least one third tag.

2. The method according to claim 1, wherein replacing the at least one first tag in the source rich text with at least one predefined second tag comprises: Extracting a first tag from the source rich text, where the first tag has corresponding content and type; In response to the first tag being a self-closing type, replacing the first tag with a paired second tag, one of the paired second tags being an open tag and the other being a closed tag, and the paired second tags having the same content, the content of the paired second tags corresponding to the content of the first tag; as well as In response to the first tag being a non-self-closing type, the first tag is replaced with a second tag, wherein content of the second tag corresponds to content of the first tag.

3. The method according to claim 1, wherein converting the at least one second tag to obtain the at least one third tag comprises: identifying at least one pair of second tags from the at least one second tag, one of the second tags of each pair being an open tag and the other being a closed tag, and the second tags of each pair having the same content; For each pair of the at least one pair of second tags, iteratively perform the following steps: In response to determining that no other second tags exist between a pair of second tags in the at least one pair of second tags, and a first second tag of a first second tag and a second tag of a second tag in the pair of second tags are paired, determining the pair of second tags as fusible tags, wherein the fusible tags do not belong to the at least one second tag; as well as In response to determining that there are no other second tags between a pair of second tags in the at least one pair of second tags, and that the previous second tag of the previous second tag and the next second tag of the next second tag in the pair of second tags are not paired, a pair of third tags is determined based on the pair of second tags and / or the fusible tags associated with the pair of second tags.

4. The method according to claim 3, wherein the at least one second tag and the to-be-processed text are used to construct a node list, wherein each of the at least one second tag is determined as a node, and the to-be-processed text between two adjacent second tags is determined as a node, and The at least one pair of second labels is identified by traversing the node list.

5. The method of claim 4, wherein identifying the at least one pair of second tags comprises: In response to detecting a first node in the node list corresponding to an open tag, looking up a second node from the node list, wherein the second node corresponds to a close tag that matches the open tag.

6. The method according to claim 1, wherein converting the at least one second tag to obtain at least one third tag comprises: Based on the text to be processed, merging a group of continuous second tags in the at least one second tag, wherein there is no text to be processed between the group of continuous second tags; and Based on the result of the fusion, the at least one third tag is obtained.

7. The method according to claim 1, further comprising: Based on the at least one third tag and the processing result for the to-be-processed text and the correspondence between the at least one second tag and the at least one third tag, obtaining an intermediate rich text including the at least one second tag and the processing result; as well as Based on the intermediate rich text and the corresponding relationship between the at least one first tag and the at least one second tag, a processed target rich text is generated.

8. The method according to claim 1, wherein processing the to-be-processed text based on the at least one third tag comprises: Using a pre-trained text processing model to process the fused rich text including the at least one third tag and the text to be processed, The text processing model is trained using at least one fused reference label and a reference text to be processed as input, and using the at least one fused reference label and a processing result of the reference text to be processed as output.

9. The method according to claim 1, wherein the processing of the to-be-processed text comprises at least one of the following: Translation, text generation, sentiment analysis, text classification. 10 . The method according to claim 1 , wherein the rich text comprises structured text with tags, and the tags indicate a format or structure of the corresponding text.

11. A rich text processing device, comprising: A replacement module, configured to replace at least one first tag in a source rich text with at least one predefined second tag, wherein each first tag corresponds to one or more second tags, and the source rich text includes the at least one first tag and a text to be processed associated with the at least one first tag; a fusion module configured to obtain at least one third label by converting the at least one second label, wherein the number of the at least one third label is less than or equal to the number of the at least one second label; as well as The processing module is configured to process the text to be processed based on the at least one third tag.

12. An electronic device comprising: at least one processing unit; as well as At least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions causing the device to perform the method according to any one of claims 1 to 10 when executed by the at least one processing unit.

13. A computer-readable storage medium having a computer program stored thereon, wherein the computer program implements the method according to any one of claims 1 to 10 when executed by a processor.

Citation Information

Patent Citations

  • Label conversion processing method and device, electronic equipment and readable storage medium

    CN111967274A

  • Text processing method and device, electronic equipment and storage medium

    CN112035408A