Text rewriting method and device and electronic equipment

By matching the format attributes of the original text block and the rewritten text block during text rewriting, the format inconsistency caused by AI rewriting tools is solved, and the practicality and reliability of the text rewriting tools are improved, ensuring the format consistency and readability of the rewritten text and the original text.

CN120542384APending Publication Date: 2025-08-26ZHUHAI KINGSOFT OFFICE SOFTWARE +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510439776.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

AI rewriting tools are difficult to maintain the consistency of text format in text processing, resulting in obvious differences in format between the rewritten text and the original text, weakening the aesthetics and readability of the text, and possibly destroying the logical structure and information emphasis effect of the text content.

Method used

By rewriting the content of the original text segment, the initial rewrite text segment is generated, and the original text block and the corresponding rewrite text block are matched therein, ensuring that the format attributes of each original text block are applied to the corresponding rewrite text block, and the target rewrite text segment is generated.

Benefits of technology

It effectively solves the problem of format inconsistency caused by text rewriting tools, improves the practicality and reliability of text rewriting tools, enhances the overall aesthetics and readability of text, and ensures that the rewritten text is highly consistent with the original text in terms of content, logic and information emphasis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120542384A_ABST
    Figure CN120542384A_ABST
Patent Text Reader

Abstract

The invention relates to a text rewriting method and device and electronic device.The method comprises the steps that content rewriting processing is conducted on an original text segment, an initial rewritten text segment is generated, and the original text segment comprises at least one original text block; determining a rewritten text block corresponding to each original text block in the initial rewritten text segment; and respectively applying the format attribute of each original text block to the corresponding rewritten text block to obtain a target rewritten text segment. Therefore, the format attribute in the original text segment is reserved after the text is rewritten, and the technical problem that text formats are inconsistent before and after the text is rewritten by using a text rewriting tool is effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of text processing, and in particular to a text rewriting method, device, and electronic device. Background Art

[0002] With the continuous development of artificial intelligence (AI) technology, AI (artificial intelligence) rewriting tools are becoming increasingly widely used in text processing, greatly improving the efficiency and quality of text processing. However, in actual application, there is a significant problem: when using AI technology to rewrite, polish, expand, or abbreviate the original text content, it is often difficult to maintain the consistency of the text's original formatting, which undoubtedly brings new challenges to text processing.

[0003] Taking a typical application scenario of daily text processing as an example, users often need to adjust the format of specific text in a sentence or paragraph to highlight keywords or important information, thereby enhancing the text's expressive effect. For example, users may carefully add bold or underline marks to keywords in a sentence to emphasize the keyword's central position in the text.

[0004] However, during the AI ​​rewriting process, these carefully set formatting marks are often overlooked or lost, resulting in noticeable formatting differences between the rewritten text and the original. This formatting inconsistency not only diminishes the overall aesthetics and readability of the text, but more importantly, it can seriously disrupt the logical structure and emphasis of the text content, causing unnecessary difficulties in interpreting and understanding the text. Summary of the Invention

[0005] The present application provides a text rewriting method, device, and electronic device to solve the technical problem that the text rewritten using a text rewriting tool has obvious differences in format from the original text, thereby weakening the overall aesthetics and readability of the text, and is likely to seriously damage the logical structure of the text content and the emphasis effect of information, causing unnecessary trouble in the interpretation and understanding of the text.

[0006] In a first aspect, the present application provides a text rewriting method, the method comprising:

[0007] Performing content rewriting processing on the original text segment to generate an initial rewritten text segment, wherein the original text segment includes at least one original text block;

[0008] Determining, in the initial rewritten text segment, a rewritten text block corresponding to each of the original text blocks;

[0009] The format attributes of each original text block are applied to the corresponding rewritten text block to obtain a target rewritten text segment.

[0010] In one possible implementation, rewriting the original text segment to generate the initial rewritten text segment includes:

[0011] Perform text block processing on the original text segment to obtain multiple original text blocks;

[0012] Adding an identification tag to each of the original text blocks in the original text segment to obtain a labeled text segment;

[0013] The labeled text segment is input into a text rewriting model to obtain an initial rewritten text segment containing the identification tag.

[0014] In a possible implementation, determining the rewritten text block corresponding to each of the original text blocks in the initial rewritten text segment includes:

[0015] The following processing is performed for each original text block:

[0016] The text blocks in the initial rewritten text segment that have the same identification tag as the original text block are determined as the rewritten text blocks corresponding to the original text blocks.

[0017] In a possible implementation, the text segment is segmented into multiple blocks to obtain a plurality of original text blocks, including:

[0018] Reading each text character in the original text segment in sequence, and when a first text character is read, recording a first tag for the first text character;

[0019] Starting from the second text character read, when the format attribute of the currently read text character is different from the format attribute of the text character read previously, recording the first tag for the currently read text character, and recording the second tag for the text character read previously;

[0020] The text characters between the adjacent first markers and the second markers are merged into the same original text block to obtain multiple original text blocks.

[0021] In a possible implementation, the text segmentation is performed on the original text segment to obtain a plurality of original text blocks, including:

[0022] The original text segment is segmented, and each word obtained is determined as an original text block, thereby obtaining a plurality of original text blocks.

[0023] In a possible implementation, applying the formatting attributes of each original text block to the corresponding rewritten text block includes:

[0024] Identifying a first text block having a specific formatting attribute from a plurality of original text blocks of the original text segment;

[0025] If it is determined that the first text block is semantically incomplete, a semantically complete second text block is determined from a text character set consisting of the first text block and its adjacent text characters, the specific format attributes are applied to the second text block, and the format attributes of the text characters in the text character set other than the second text block are modified to the first format attributes; wherein the second text block and the first text block have overlapping text characters; and the first format attributes are format attributes other than the specific format attributes that appear in the text character set.

[0026] Based on the original text segment after the above processing, the step of applying the format attributes of each original text block to the corresponding rewritten text block is performed.

[0027] In a possible implementation, determining the rewritten text block corresponding to each of the original text blocks in the initial rewritten text segment includes:

[0028] Performing word segmentation on the initial rewritten text segment, and determining each word obtained by the word segmentation as a rewritten text block, thereby obtaining a plurality of rewritten text blocks;

[0029] The following processing is performed for each rewritten text block:

[0030] An original text block that satisfies a preset semantic consistency condition with the rewritten text block is determined in the original text segment, and the original text block that satisfies the preset semantic consistency condition is determined as the original text block corresponding to the rewritten text block.

[0031] In a possible implementation, determining the rewritten text block corresponding to each of the original text blocks in the initial rewritten text segment includes:

[0032] identifying content differences between the original text segment and the initial rewritten text segment;

[0033] Based on the content difference portion, a plurality of original text blocks are defined in the original text segment, and a rewritten text block corresponding to each of the original text blocks is defined in the initial rewritten text segment.

[0034] In one possible implementation, identifying the content difference between the original text segment and the initially rewritten text segment includes:

[0035] Using a preset difference comparison algorithm, identifying a set of text locations in the original text segment that differ from the initial rewritten text segment;

[0036] The step of defining a plurality of original text blocks in the original text segment based on the content difference portion, and defining a rewritten text block corresponding to each of the original text blocks in the initial rewritten text segment, includes:

[0037] Based on the text position set, defining original text blocks corresponding to text positions with different contents in the original text segment, and defining rewritten text blocks corresponding to the original text blocks for portions corresponding to the text position set in the initial rewritten text segment;

[0038] Furthermore, continuous text intervals in which the text characters in the original text segment and the initial rewritten text segment are completely identical are defined as original text blocks and corresponding rewritten text blocks, respectively.

[0039] In a possible implementation manner, before identifying the content difference between the original text segment and the initially rewritten text segment, the method further includes:

[0040] identifying corresponding original text ranges and rewritten text ranges in the original text segment and the initial rewritten text segment;

[0041] The following processing is performed for each group of corresponding original text ranges and rewritten text ranges: the original text range is treated as the original text segment, and the rewritten text range is treated as the rewritten text segment, and the steps of identifying the content difference between the original text segment and the initial rewritten text segment and subsequent steps are performed.

[0042] In one possible implementation, the method further includes:

[0043] Obtaining the paragraph attributes of the original text segment;

[0044] Apply the paragraph attributes of the original text segment to the target rewritten text segment.

[0045] In a second aspect, the present application provides a text rewriting device, the device comprising:

[0046] a rewriting module, configured to perform content rewriting processing on an original text segment to generate an initial rewritten text segment, wherein the original text segment includes at least one original text block;

[0047] A mapping module, configured to determine, in the initial rewritten text segment, a rewritten text block corresponding to each of the original text blocks;

[0048] The format inheritance module is used to apply the format attributes of each original text block to the corresponding rewritten text block to obtain a target rewritten text segment.

[0049] In a possible implementation, the rewriting module includes:

[0050] A block division unit is used to perform text block processing on the original text segment to obtain multiple original text blocks;

[0051] a label adding unit, configured to add an identification label to each of the original text blocks in the original text segment to obtain a label text segment;

[0052] The text rewriting unit is used to input the label text segment into a text rewriting model to obtain an initial rewritten text segment containing the identification label.

[0053] In a possible implementation, the mapping module is specifically configured to:

[0054] The following processing is performed for each original text block:

[0055] The text blocks in the initial rewritten text segment that have the same identification tag as the original text block are determined as the rewritten text blocks corresponding to the original text blocks.

[0056] In a possible implementation, the block division unit is specifically configured to:

[0057] Reading each text character in the original text segment in sequence, and when a first text character is read, recording a first tag for the first text character;

[0058] Starting from the second text character read, when the format attribute of the currently read text character is different from the format attribute of the text character read previously, recording the first tag for the currently read text character, and recording the second tag for the text character read previously;

[0059] The text characters between the adjacent first markers and the second markers are merged into the same original text block to obtain multiple original text blocks.

[0060] In a possible implementation, the block division unit is specifically configured to:

[0061] The original text segment is segmented, and each word obtained is determined as an original text block, thereby obtaining a plurality of original text blocks.

[0062] In a possible implementation, the format inheritance module is specifically configured to:

[0063] Identifying a first text block having a specific formatting attribute from a plurality of original text blocks of the original text segment;

[0064] If it is determined that the first text block is semantically incomplete, a semantically complete second text block is determined from a text character set consisting of the first text block and its adjacent text characters, the specific format attributes are applied to the second text block, and the format attributes of the text characters in the text character set other than the second text block are modified to the first format attributes; wherein the second text block and the first text block have overlapping text characters; and the first format attributes are format attributes other than the specific format attributes that appear in the text character set.

[0065] Based on the original text segment after the above processing, the step of applying the format attributes of each original text block to the corresponding rewritten text block is performed.

[0066] In a possible implementation, the mapping module includes:

[0067] A word segmentation unit is used to perform word segmentation on the initial rewritten text segment, and determine each word obtained by word segmentation as a rewritten text block, thereby obtaining a plurality of rewritten text blocks;

[0068] A determination unit, configured to perform the following processing on each of the rewritten text blocks:

[0069] An original text block that satisfies a preset semantic consistency condition with the rewritten text block is determined in the original text segment, and the original text block that satisfies the preset semantic consistency condition is determined as the original text block corresponding to the rewritten text block.

[0070] In a possible implementation, the mapping module includes:

[0071] an identification unit, configured to identify a content difference between the original text segment and the initial rewritten text segment;

[0072] The defining unit is configured to define a plurality of original text blocks in the original text segment based on the content difference portion, and define a rewritten text block corresponding to each of the original text blocks in the initial rewritten text segment.

[0073] In a possible implementation, the identification unit is specifically configured to:

[0074] Using a preset difference comparison algorithm, identifying a set of text locations in the original text segment that differ from the initial rewritten text segment;

[0075] The defining unit is specifically used to:

[0076] Based on the text position set, defining original text blocks corresponding to text positions with different contents in the original text segment, and defining rewritten text blocks corresponding to the original text blocks for portions corresponding to the text position set in the initial rewritten text segment;

[0077] Furthermore, continuous text intervals in which the text characters in the original text segment and the initial rewritten text segment are completely identical are defined as original text blocks and corresponding rewritten text blocks, respectively.

[0078] In a possible implementation, the device further includes:

[0079] a coarse screening module, configured to identify the original text range and the rewritten text range corresponding to each other in the original text segment and the initial rewritten text segment before identifying the content difference between the original text segment and the initial rewritten text segment;

[0080] The recognition unit is used to perform the following processing for each group of corresponding original text ranges and rewritten text ranges: taking the original text range as the original text segment and the rewritten text range as the rewritten text segment, and performing the step of identifying the content difference between the original text segment and the initial rewritten text segment.

[0081] In a possible implementation, the format inheritance module is further configured to:

[0082] Obtaining the paragraph attributes of the original text segment;

[0083] Apply the paragraph attributes of the original text segment to the target rewritten text segment.

[0084] In a third aspect, the present application provides an electronic device comprising: a processor and a memory, wherein the processor is configured to execute a text rewriting program stored in the memory to implement the text rewriting method described in any one of the first aspects.

[0085] In a fourth aspect, the present application further provides a computer storage medium storing computer-executable instructions, wherein the computer-executable instructions are used to execute the text rewriting method described in any one of the above items of the present application.

[0086] The above technical solution provided by the embodiment of the present application has the following advantages over the prior art: the technical solution provided by the embodiment of the present application generates an initial rewritten text segment by rewriting the content of the original text segment, and matches the original text block with the corresponding rewritten text block. Subsequently, the format attributes of each original text block are applied to the corresponding rewritten text block, thereby obtaining a target rewritten text segment that is optimized and improved in content and retains the style of the original text segment in format. This technical solution effectively solves the technical problem of inconsistent text format before and after text rewriting using a text rewriting tool, which not only significantly improves the practicality and reliability of the text rewriting tool in the field of text processing, but also greatly enhances the overall aesthetics and readability of the text. More importantly, it ensures that the rewritten text is highly consistent with the original text in terms of content, logic and information emphasis, avoiding the troubles caused by inconsistent formatting to text interpretation and understanding. BRIEF DESCRIPTION OF THE DRAWINGS

[0087] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0088] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0089] One or more embodiments are exemplarily illustrated by pictures in the corresponding drawings. These exemplifications do not constitute limitations on the embodiments. Elements with the same reference numerals in the drawings are represented as similar elements. Unless otherwise stated, the figures in the drawings do not constitute proportional limitations.

[0090] Figure 1 A flowchart of an embodiment of a text rewriting method provided in an embodiment of the present application;

[0091] Figure 2 A flowchart of another text rewriting method provided in an embodiment of the present application;

[0092] Figure 3 A schematic diagram of an implementation process for merging consecutive text characters with the same formatting attributes in an original text segment into the same original text block to obtain multiple original text blocks.

[0093] Figure 4 A flowchart of another text rewriting method provided in an embodiment of the present application;

[0094] Figure 5A flowchart of another text rewriting method provided in an embodiment of the present application;

[0095] Figure 6 A block diagram of an embodiment of a text rewriting device provided in an embodiment of the present application;

[0096] Figure 7 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0097] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0098] The disclosure below provides many different embodiments or examples for implementing different structures of the present application. In order to simplify the disclosure of the present application, the components and settings of specific examples are described below. Of course, these are merely examples and are not intended to limit the present application. In addition, the present application may repeat reference numbers and / or letters in different examples. Such repetition is for the purpose of simplicity and clarity and does not in itself indicate the relationship between the various embodiments and / or settings discussed.

[0099] In order to solve the technical problem in the prior art that the text rewritten using text rewriting tools has obvious differences in format from the original text, thereby weakening the overall aesthetics and readability of the text, and is likely to seriously damage the logical structure of the text content and the emphasis effect of information, causing unnecessary troubles in the interpretation and understanding of the text, the present application provides a text rewriting method, device, electronic device and storage medium, which can retain the format attributes of the original text segment after text rewriting, and effectively solve the technical problem of inconsistent text format before and after text rewriting using text rewriting tools.

[0100] Figure 1 This is a flow chart of an embodiment of a text rewriting method provided in an embodiment of the present application.

[0101] like Figure 1 As shown, the method includes the following steps:

[0102] Step 101: rewrite the original text segment to generate an initial rewritten text segment, wherein the original text segment includes at least one original text block.

[0103] In step 101, the original text segment is rewritten to generate an initial rewritten text segment. The original text segment, as a paragraph unit to be rewritten, may include one or more original text blocks. If the original text segment contains only one original text block, the original text segment is the original text block. If the original text segment contains multiple original text blocks, these original text blocks can form a complete original text segment.

[0104] An original text segment can be a paragraph unit in a document, or it can be an independent paragraph with complete meaning. In actual applications, the original text segment may contain redundant expressions, complex sentence structures, or non-standardized technical terms. In such cases, the original text segment can be rewritten. Content rewriting refers to optimizing and adjusting the expression of the text without changing the meaning of the original text. This includes but is not limited to eliminating redundant expressions, splitting long sentences into clearer and more concise short sentences, standardizing terminology, replacing non-standardized technical terms, adding modifiers, etc. The purpose of content rewriting is to make the text easier to understand and improve readability.

[0105] For example, if the original text reads, "Therefore, we attempt to focus our investment on resources, independent innovation, and branded consumption." After the rewriting process, the original rewritten text might become, "Based on this, we are committed to focusing on resources, independent innovation, and branded consumption as our investment strategy." This demonstrates that rewriting preserves the original meaning while making the content clearer and easier to understand.

[0106] In one embodiment, an original text segment is input into a content rewriting model, which then performs content rewriting on the original text segment to generate an initial rewritten text segment. The content rewriting model is an intelligent tool based on natural language processing technology that deeply understands and analyzes the input original text segment and then rewrites it according to preset rules and strategies. The rewritten text not only retains the core meaning of the original content but also optimizes and innovates its expression and language style, making the content more vivid, interesting, and easy to understand. For example, the working principle of the content rewriting model generally includes the following steps: text processing: preprocessing the input original text segment, including word segmentation, removal of invalid vocabulary and stop words, etc., to enable effective text recognition and convert it into a format that can be processed by machines; semantic analysis: in-depth analysis of the preprocessed text to understand its meaning and context, ensuring that the semantic consistency of the original text is maintained during the rewriting process; and rewriting strategy application: rewriting the text according to preset rewriting rules and strategies. This may include synonym replacement, sentence adjustment, paragraph reorganization and other operations to optimize the expression and language style of the text; Quality check: Perform quality check on the rewritten text to ensure the accuracy and fluency of the rewritten results.

[0107] In another embodiment, a non-AI rewriting tool can be used to rewrite the content of the original text segment to generate an initial rewritten text segment. Exemplarily, the non-AI rewriting tool uses the following technical solutions to rewrite the content of the original text segment: First, a mapping table is constructed, which may include synonym / synonymous mapping relationships, such as "fast → rapid", "advantage → advantage" and other mapping relationships; domain-specific term mapping relationships, such as in the medical field: "tumor → neoplasm", and so on. Afterwards, the original text segment is traversed to identify the replaceable words therein, and then the words are replaced according to the mapping table. Afterwards, the sentence structure is adjusted, such as active to passive, conjunction adjustment, word order reorganization and other adjustment operations. Finally, grammatical correction and polishing are performed to obtain the initial rewritten text segment.

[0108] It should be noted that the above is merely an exemplary description. In actual applications, there may be other specific implementation methods for rewriting the content of the original text segment, and the embodiments of the present application do not limit this.

[0109] Step 102: Determine the rewritten text blocks corresponding to each original text block in the initial rewritten text segment.

[0110] In step 102, the purpose is to accurately identify and determine the rewritten text block corresponding to each original text block in the generated initial rewritten text segment. Here, "correspondence" refers to the logical or content relationship between the original text block and the rewritten text block during the rewriting process.

[0111] Among them, since the process of content rewriting may involve multiple operations such as synonym replacement, sentence adjustment, and information reorganization, the original text block and its corresponding rewritten text block may be the same in expression, or there may be significant differences.

[0112] Taking the specific example in step 101 as an example, the original text block "thereby" corresponds to the rewritten text block "based on this" after rewriting, and the vocabulary is replaced here to optimize the expression; the original text block "attempt" corresponds to the rewritten text block "committed to" after rewriting, and the vocabulary is also replaced here to optimize the expression; and the original text block "we" remains "we" after rewriting, indicating that this original text block has not changed during the content rewriting process and remains the same.

[0113] As for how to determine the rewritten text block corresponding to each original text block in the initial rewritten text segment, exemplary explanations are given below through different embodiments, which will not be described in detail here.

[0114] Step 103: Apply the format attributes of each original text block to the corresponding rewritten text block to obtain a target rewritten text segment.

[0115] In step 103, the formatting attributes of the original text block are accurately applied to the corresponding rewritten text block. This is done to ensure that the formatting of the rewritten text content remains consistent, thereby generating a target rewritten text segment that meets the requirements in terms of content and retains the original format.

[0116] Format attributes encompass various parameters of text presentation, including but not limited to font type, size, numbering sequence, indentation level, bolding, underlining, etc. These format attributes are typically used in original text to emphasize information hierarchy, highlight key content, or achieve specific typographical effects.

[0117] For example, the original text segment is:

[0118] “Thus, we attempt The main investment lines are resources, independent innovation and brand consumption.

[0119] Among them, "therefore" is displayed in bold, "main line of investment" is displayed in a larger font size (for example, size 3), and "attempt" is underlined to further emphasize key information.

[0120] After rewriting the content of the original text segment, an initial rewritten text segment with the following content can be obtained:

[0121] "Based on this, we are committed to using resources, independent innovation and brand consumption as our investment strategy."

[0122] Next, by executing steps 102 and 103, the word "based on this" in the initial rewritten text segment can be bolded, "committed to" is underlined, and "investment strategy" is set to size 3 font. The formatting attributes of the corresponding original text blocks are also applied to the other rewritten text blocks in the initial rewritten text segment, and the following target rewritten text segment is finally obtained:

[0123] “Based on this, we Committed to The investment strategy is to use resources, independent innovation and brand consumption."

[0124] This shows that the target rewritten text segment not only meets the content rewriting requirements but also maintains the original formatting. Such a target rewritten text segment is not only easy to read and understand, but also better meets specific text processing needs, such as document editing, information release, data reporting, etc.

[0125] The technical solution provided by the embodiment of the present application generates an initial rewritten text segment by rewriting the content of the original text segment, and matches the original text block with the corresponding rewritten text block. Subsequently, the format attributes of each original text block are applied to the corresponding rewritten text block, thereby obtaining a target rewritten text segment that is optimized and improved in content and retains the style of the original text segment in format. This technical solution effectively solves the technical problem of inconsistent text format before and after text rewriting using a text rewriting tool, which not only significantly improves the practicality and reliability of the text rewriting tool in the field of text processing, but also greatly enhances the overall aesthetics and readability of the text. More importantly, it ensures that the rewritten text is highly consistent with the original text in terms of content, logic and information emphasis, avoiding the troubles caused by inconsistent formatting to text interpretation and understanding.

[0126] Figure 2 This is a flowchart of another text rewriting method provided in an embodiment of the present application. Figure 2 The process shown in Figure 1 Based on the process shown in FIG, a method for determining the rewritten text blocks corresponding to each original text block in the original text segment in the initial rewritten text segment is described in detail. Figure 2 As shown, the following steps are included:

[0127] Step 201: Perform text block processing on the original text segment to obtain multiple original text blocks.

[0128] Step 202: Add an identification tag to each original text block in the original text segment to obtain a labeled text segment.

[0129] Step 203: Input the labeled text segment into the text rewriting model to obtain an initial rewritten text segment containing the identification tag.

[0130] Step 204: For each original text block, determine the text block in the initial rewritten text segment that has the same identification tag as the original text block as the rewritten text block corresponding to the original text block.

[0131] Step 205: Apply the format attributes of each original text block to the corresponding rewritten text block to obtain a target rewritten text segment.

[0132] For ease of understanding, steps 201 to 205 are explained in a unified manner below:

[0133] Figure 2 In the illustrated process, the original text segment is first segmented in step 201 to break the complete original text segment into multiple independent original text blocks. Then, in step 202, identification tags are added to each original text block. This step ensures the identification and tracking of the original text blocks during the rewriting process. Subsequently, in step 203, the tagged text segment, i.e., the labeled text segment, is input into the text rewriting model for content rewriting. The text rewriting model is then constrained to retain the identification tags during content rewriting. The text rewriting model, leveraging its powerful language processing capabilities, can then generate an initial rewritten text segment containing the identification tags of the original text blocks. Then, in step 204, the previously added identification tags are used to accurately identify the rewritten text block corresponding to each original text block by comparing the original text blocks with the text blocks in the initial rewritten text segment. Finally, in step 205, the formatting attributes of each original text block are applied to the corresponding rewritten text block, resulting in a target rewritten text segment that is both optimized in content and strictly maintains the original format.

[0134] Among them, as an optional implementation method, when the text rewriting model rewrites the content of the tagged text segment, it first identifies the identification tags therein, for example, using natural language processing technology, especially the recognition algorithm for specific tags or labels, to locate and understand the identification tags in the tagged text segment, ensuring that they are not accidentally deleted or changed during the rewriting process. Next, when the text rewriting model rewrites the content, it rewrites it at the granularity of the original text block, which involves rewriting the content of the original text block while keeping the identification tags of the original text block unchanged. After the rewriting is completed, the text rewriting model splices each rewritten text block (including the rewritten text content and the unrewritten identification tags) in the original order, thereby obtaining an initial rewritten text segment that contains both the rewritten text content and the complete retention of the identification tags. Here, this implementation method is generally applicable to the case where the original text block is a text unit with independent meaning.

[0135] As another optional implementation method, when rewriting the content of a tagged text segment, the text rewriting model also first identifies the identification tags. It then ignores the identification tags and performs a global content rewrite on the text. After the rewriting is completed, the text rewriting model uses the previously identified tag position and tag content information to perform secondary processing on the rewritten text. This mainly includes: finding the corresponding insertion points in the rewritten text based on the position information of the original identification tags; and accurately inserting the original identification tags into these insertion points to restore the identification tags.

[0136] It can be seen from the above description that steps 201 to 205 together constitute an efficient, accurate and reliable text rewriting process. Applying this process not only ensures the accuracy and readability of the text content, but also greatly maintains the format consistency of the original text, providing users with a more convenient and efficient text processing experience, and can meet the text rewriting needs in various application scenarios.

[0137] exist Figure 2 In the illustrated process, as an embodiment, the specific implementation of text segmentation in step 201 to obtain multiple original text blocks includes: merging consecutive text characters with the same formatting attributes in the original text segment into the same original text block to obtain multiple original text blocks. In other words, when text characters in the original text segment maintain consistent formatting attributes and appear consecutively, these text characters are considered as a whole, i.e., an original text block.

[0138] Among them, as an optional implementation, the specific implementation of merging consecutive text characters with the same format attributes in the original text segment into the same original text block includes: sequentially reading each text character in the original text segment, when reading the first text character, recording a first mark for the first text character; starting from the second text character read, when the format attribute of the currently read text character is different from the format attribute of the previously read text character, recording a first mark for the currently read text character, and recording a second mark for the previously read text character; merging the text characters between adjacent first marks and second marks into the same original text block to obtain multiple original text blocks.

[0139] Taking a specific example, the original text segment is:

[0140] "Thus, we attempt take resources, independent innovation and brand consumption as the main investment line".

[0141] Among them, "Thus" is shown in bold, "main investment line" is shown in a larger font size (for example, size three), and "attempt" is underlined.

[0142] For the above original text segment, starting from the first text character "by", read the text characters sequentially and record a first mark for "by", when reading ",", it can be found that the format attribute of the currently read text character is different from the format attribute of the previously read text character "this", then, record a first mark for ",", and record a second mark for "this". And so on, the final marking result is as Figure 3 shown.

[0143] See Figure 3 , merge the text characters between adjacent first marks and second marks (including the text characters pointed to by the first mark and the second mark respectively) into the same original text block, and merge the text characters after the last first mark (including the text character pointed to by the last first mark) into the same original text block, to obtain the following original text blocks: "Thus", ", we", " attempt ", "take resources, independent innovation and brand consumption as", "main investment line".

[0144] On this basis, perform step 202, add identification labels to each original text block in the original text segment, and the following labeled text segment can be obtained:

[0145] " <r1> thus< / r1> <r2> ,us< / r2> <r3> attempt < / r3> <r4> Based on resources, independent innovation and brand consumption< / r4> <r5> Investment Main Line< / r5> ".

[0146] Furthermore, step 203 is executed to input the labeled text segment into the text rewriting model to obtain the following initial rewritten text segment containing the identification tag:

[0147] “ <r1> Based on this< / r1> <r2> ,us< / r2> <r3> Committed to< / r3> <r4> With resources, independent innovation and brand consumption as< / r4> <r5> Investment Strategy< / r5> ”.

[0148] Then, step 204 is performed. For each original text block, the text blocks in the initial rewritten text segment that have the same identification tag as the original text block are determined as the rewritten text blocks corresponding to the original text block. The corresponding relationship between the original text blocks and the rewritten text blocks can be obtained as shown in Table 1:

[0149] Table 1

[0150]

[0151]

[0152] Finally, step 205 is executed to apply the format attributes of each original text block to the corresponding rewritten text block, thereby obtaining the following target rewritten text segment:

[0153] “Based on this, we Committed to The investment strategy is to use resources, independent innovation and brand consumption."

[0154] The above embodiment has the advantage of directness in maintaining the format consistency when rewriting the original text segment. This directness is reflected in the following two aspects:

[0155] 1. Clarity of Objective: The core objective of the embodiments of this application is to maintain the format consistency of the original text while rewriting the text content. To achieve this goal, the above embodiments directly use format attributes as the primary criterion for text block segmentation. This strategy ensures that during subsequent text processing, the text characters within each segmented original text block maintain a high degree of format consistency.

[0156] 2. Ease of Operation: Compared to some complex text segmentation methods, the method used in the above embodiment is more intuitive and simple. It does not require in-depth grammatical analysis or semantic understanding of the original text, but instead performs a simple merging operation based on the format attributes of the text characters. This operation method not only greatly simplifies the processing flow and reduces processing complexity, but also significantly improves overall processing efficiency. In addition, by avoiding complex grammatical and semantic analysis, this method can also maintain high stability and reliability when processing large-scale text data.

[0157] In addition, in practical applications, a special situation may also be encountered: when a user edits an original text segment, they accidentally perform special processing on the format attributes of non-emphasized text characters, while omitting the text characters that originally needed to be emphasized. For example, assume the original text segment is:

[0158] "Thus, we try The picture is Resources, independent innovation, and brand consumption as the main investment themes".

[0159] Among them, the user originally intended to underline "try", but mistakenly underlined "try to".

[0160] To address this issue, in one embodiment, the following solution is proposed: First, check the original text segment to determine whether the format attributes of the text characters need to be adjusted. If adjustment is required, after the adjustment, rewrite the content of the adjusted original text segment and沿用 the format.

[0161] Specifically, as an optional implementation method, first, divide the original text segment into multiple original text blocks in the manner provided in the above embodiment. Then, identify the first text block with specific format attributes from the multiple original text blocks of the original text segment. Here, the specific format attributes refer to format attributes used for emphasizing and highlighting, such as underline marking, bold, increased font size, etc. For example, in the above example, the first text block with specific format attributes includes but is not limited to "try to". Next, determine whether the semantics of the first text block are complete. If it is determined that the semantics of the first text block are incomplete, determine a second text block with complete semantics from the text character set composed of the first text block and its adjacent text characters. After determining the second text block, apply the specific format attribute to the second text block, and modify the format attributes of the other text characters in the text character set except the second text block to the first format attribute. Among them, the second text block and the first text block have overlapping text characters, and the format attribute refers to other format attributes that appear in the text character set except the specific format attribute.

[0162] For example, the first text block " The picture is " and its adjacent text characters "try", "resources" form a text character set: "try to use resources". In this text character set, the second text block with complete semantics "try" can be determined. This second text block and the first text block have overlapping text characters "try". This is set to avoid mistakenly determining other semantically complete words in the text character set (such as "resources") as the second text block. Subsequently, apply the specific format attribute to this second text block, and modify the format attributes of the other text characters in this text character set (such as "to use resources") to the first format attribute (for example, no underline, size four, Song typeface). After such adjustment, the text segment can be obtained:

[0163] “Thus, we attempt The main investment lines are resources, independent innovation and brand consumption.

[0164] As can be seen from the above description, the technical solutions of the embodiments of this application can effectively correct formatting errors that users may make when editing text, ensuring the correctness and readability of the text at the format level. On this basis, applying the technical solutions provided by the embodiments of this application to rewrite content and retain the format, the target rewritten text segment obtained will also have a high degree of accuracy in terms of content and format.

[0165] exist Figure 2 In the illustrated process, as another embodiment, the specific implementation of performing text segmentation processing on the original text segment in step 201 to obtain multiple original text blocks includes: performing word segmentation processing on the original text segment, determining each obtained word as an original text block, and obtaining multiple original text blocks. Specific word segmentation methods may include rule-based word segmentation, statistical-based word segmentation, and machine learning-based word segmentation, etc., which are not limited in this embodiment of the present application.

[0166] After word segmentation, the original text segment will be cut into multiple original text blocks. Each original text block corresponds to an independent word and can be used as a text unit with independent meaning.

[0167] For example, the original text segment is:

[0168] “Thus, we attempt The main investment lines are resources, independent innovation and brand consumption.

[0169] By performing word segmentation on the original text segment, the following original text block can be obtained (this is just an example):

[0170] Therefore, we try to take resources, independent innovation, and brand consumption as the main investment line.

[0171] On this basis, by executing step 202, an identification tag is added to each original text block in the original text segment, and the following labeled text segment can be obtained:

[0172] “ <r1> thus< / r1> <r2> ,< / r2> <r3> us < / r3> <r4> attempt < / r4> <r5> by< / r5> <r6> resource< / r6> <r7> 、< / r7> <r8> autonomy< / r8> <r9> Innovation< / r9> <r10> and< / r10> <r11> Brand consumption< / r11> <r12> for< / r12> <r13>Investment Main Line

[0173] ”.

[0174] Furthermore, step 203 is executed to input the labeled text segment into the text rewriting model to obtain the following initial rewritten text segment containing the identification tag:

[0175] “ <r1> Based on this< / r1> <r2> ,< / r2> <r3> us < / r3> <r4>Committed to

[0176] <r5> by< / r5> <r6> resource< / r6> <r7> 、< / r7> <r8> autonomy< / r8> <r9> Innovation< / r9> <r10> and< / r10> <r11> Brand consumption< / r11> <r12> As< / r12> <r13> Investment Strategy< / r13> ”.

[0177] Here, the text rewriting model can rewrite the content of the label text segment by partial rewriting. The specific implementation can be found in the above description and will not be repeated here.

[0178] Then, step 204 is performed. For each original text block, the text blocks in the initial rewritten text segment that have the same identification tag as the original text block are determined as the rewritten text blocks corresponding to the original text block. The corresponding relationship between the original text blocks and the rewritten text blocks can be obtained as shown in Table 2:

[0179] Table 2

[0180]

[0181]

[0182] Finally, step 205 is executed to apply the format attributes of each original text block to the corresponding rewritten text block, thereby obtaining the following target rewritten text segment:

[0183] “Based on this, we Committed to The investment strategy is to use resources, independent innovation and brand consumption."

[0184] In the technical solution provided in the above embodiment, the method of segmenting the original text segment and rewriting and formatting each word as an original text block includes but is not limited to the following beneficial effects:

[0185] Accurate formatting: This approach ensures accurate formatting by breaking down the original text into word-based blocks and preserving the original formatting of each rewritten block during the rewriting process. Each word, as the smallest processing unit, has its formatting attributes (such as font, color, and bolding) accurately recorded and applied to the corresponding rewritten block during the rewriting process, preventing formatting information from being lost or mixed up.

[0186] Rewriting flexibility: Dividing text blocks into word-based units makes the rewriting process more flexible. Each word can be rewritten independently, without being restricted by the surrounding text. This flexibility allows for more precise text rewriting.

[0187] In summary, by applying the above embodiments, it is possible to ensure that the format of the entire text remains consistent before and after rewriting, thereby maintaining the readability and coherence of the text.

[0188] In addition, based on this embodiment, a solution is also proposed for the special situation described above.

[0189] Specifically, in one embodiment, the specific implementation of applying the format attributes of each original text block to the corresponding rewritten text block includes: identifying a first text block with specific format attributes from multiple original text blocks in the original text segment; when it is determined that the semantics of the first text block is incomplete, determining a semantically complete second text block from a text character set consisting of the first text block and its adjacent text characters, applying the specific format attributes to the second text block, and modifying the format attributes of other text characters in the text character set except the second text block to the first format attributes; wherein the second text block has overlapping text characters with the first text block; the first format attributes refer to other format attributes appearing in the text character set except the specific format attributes; based on the original text segment after the above processing is performed, executing the step of applying the format attributes of each original text block to the corresponding rewritten text block.

[0190] As for the detailed description of this embodiment, please refer to the above related content and will not be repeated here.

[0191] Figure 4 This is a flow chart of another embodiment of a text rewriting method provided in an embodiment of the present application. Figure 4 The process shown in Figure 1 Based on the process shown in FIG, another implementation method of determining the rewritten text block corresponding to each original text block in the original text segment in the initial rewritten text segment is described in detail. Figure 4 As shown, the following steps are included:

[0192] Step 401: rewrite the original text segment to generate an initial rewritten text segment, wherein the original text segment includes at least one original text block.

[0193] Step 402: perform word segmentation on the initial rewritten text segment, and determine each word obtained by the word segmentation as a rewritten text block, thereby obtaining a plurality of rewritten text blocks.

[0194] Step 403: for each rewritten text block: determine in the original text segment an original text block that satisfies a preset semantic consistency condition with the rewritten text block, and determine the original text block that satisfies the preset semantic consistency condition as the original text block corresponding to the rewritten text block.

[0195] Step 404: Apply the format attributes of each original text block to the corresponding rewritten text block to obtain a target rewritten text segment.

[0196] For ease of understanding, steps 401 to 404 are explained in a unified manner below:

[0197] First of all, unlike Figure 2 In the embodiment, the original text segment is divided into original text blocks in advance, and in the content rewriting process, the original text blocks are used as the granularity to track and locate the rewritten text blocks. Figure 4 In the embodiment, the rewritten content, i.e., the initial rewritten text segment, is first generated, and then the corresponding part in the original text segment is traced back based on the semantics of the rewritten content to determine the format attributes that need to be applied to the rewritten content.

[0198] Specifically, Figure 4 The process of the embodiment includes the following steps:

[0199] First, step 401 is executed to rewrite the original text segment to generate an initial rewritten text segment. The detailed description of step 401 can be found in the relevant description of the above embodiment, and will not be repeated here.

[0200] Subsequently, step 402 is executed to perform word segmentation on the rewritten text segment, i.e., the initial rewritten text segment, so as to divide the initial rewritten text segment into a series of independent words or phrases, each of which is considered as a rewritten text block. Specific word segmentation methods may include rule-based word segmentation, statistical-based word segmentation, and machine learning-based word segmentation, which are not limited in this embodiment of the present application.

[0201] After word segmentation, the initial rewritten text segment is split into multiple rewritten text blocks, each corresponding to an independent word. These rewritten text blocks will be used as the objects for applying formatting properties in subsequent steps.

[0202] Next, step 403 is executed to find an original text block in the original text segment that satisfies a preset semantic consistency condition for each rewritten text block. The preset semantic consistency condition generally means that the two text blocks are semantically consistent or highly similar.

[0203] In this process, some semantic matching or semantic similarity calculation methods can be used to compare the similarities of the two text blocks in terms of vocabulary, syntactic structure, contextual information, etc., so as to determine whether they meet the preset semantic consistency conditions.

[0204] It is worth noting that since the rewriting process may introduce changes such as synonym replacement and sentence structure adjustment, the original text block corresponding to the rewritten text block is not necessarily exactly the same words or phrases, but may also be semantically similar or equivalent expressions.

[0205] Finally, once the original text block corresponding to the rewritten text block is found, step 405 is executed for the rewritten text block to apply the formatting attributes of the original text block to the corresponding rewritten text block. In this way, the rewritten text segment can achieve both content rewriting and optimization while retaining the formatting information of the original text segment.

[0206] Figure 4 In the illustrated process, there's no need for complex segmentation and tagging of the original text segment before rewriting. Instead, the semantic information of the rewritten content is used to trace back to the corresponding parts in the original text segment. This ensures that the rewritten text segment is optimized in content while maintaining the same format as the original, increasing processing flexibility and adaptability.

[0207] Figure 5 This is a flow chart of another embodiment of a text rewriting method provided in an embodiment of the present application. Figure 5 The process shown in Figure 1 Based on the process shown in FIG, another implementation method of determining the rewritten text block corresponding to each original text block in the original text segment in the initial rewritten text segment is described in detail. Figure 5 As shown, the following steps are included:

[0208] Step 501: rewrite the original text segment to generate an initial rewritten text segment, wherein the original text segment includes at least one original text block.

[0209] Step 502: Identify the content differences between the original text segment and the initial rewritten text segment.

[0210] Step 503: Based on the content difference portion, a plurality of original text blocks are defined in the original text segment, and a rewritten text block corresponding to each original text block is defined in the initial rewritten text segment.

[0211] Step 504: Apply the format attributes of each original text block to the corresponding rewritten text block to obtain a target rewritten text segment.

[0212] For ease of understanding, steps 501 to 504 are explained in a unified manner below:

[0213] First of all, Figure 5 Examples and Figure 4 The similarities between the embodiments are that the rewritten content, i.e., the initial rewritten text segment, is first generated, and then the original text block in the original text segment and the rewritten text block corresponding to the original text block in the initial rewritten text segment are determined. The difference is that Figure 4 The embodiment of the invention is to locate the corresponding part in the original text segment based on the semantics of the rewritten content, and then determine the format attributes that need to be applied to the rewritten content. Figure 5 The embodiment starts directly from the content difference level, and determines the correspondence between the original text block and the rewritten text block by identifying the content difference between the original text segment and the initial rewritten text segment.

[0214] Specifically, Figure 5 The process of the embodiment includes the following steps:

[0215] First, step 501 is executed to rewrite the original text segment to generate an initial rewritten text segment. The detailed description of step 501 can be found in the relevant description of the above embodiment, which will not be repeated here.

[0216] Subsequently, step 502 is executed to identify content differences between the original text segment and the initial rewritten text segment. Here, the content differences between the original text segment and the initial rewritten text segment may include word replacement, sentence reorganization, addition or deletion of information, etc. Then, step 503 is executed to define multiple original text blocks in the original text segment based on the content differences, and to define rewritten text blocks corresponding to each original text block in the initial rewritten text segment.

[0217] In one embodiment, the specific implementation of identifying the content difference between the original text segment and the initial rewritten text segment includes: using a preset difference comparison algorithm to identify a set of text positions in the original text segment that have content differences with the initial rewritten text segment. On this basis, based on the content difference, multiple original text blocks are defined in the original text segment, and rewritten text blocks corresponding to each original text block are defined in the initial rewritten text segment. The specific implementation includes: based on the text position set, original text blocks corresponding to text positions with content differences are defined in the original text segment, and rewritten text blocks corresponding to the original text blocks for the parts corresponding to the text position set in the initial rewritten text segment; and continuous text intervals in which the text characters in the original text segment and the initial rewritten text segment are completely consistent are defined as original text blocks and corresponding rewritten text blocks, respectively.

[0218] Among them, the preset difference comparison algorithm is, for example, the diff algorithm (full name Difference algorithm). As an efficient text comparison tool, the diff algorithm can accurately output a set of text positions where there are content differences between the original text segment and the initial rewritten text segment by carefully analyzing the additions, deletions and changes in the text content. The text position set records in detail multiple position points in the original text segment and their corresponding operation instructions (such as deletion or addition).

[0219] For example, suppose the original text is as follows:

[0220] "Therefore, we try to focus our investment on resources, independent innovation and brand consumption."

[0221] The text content in the initial rewritten text segment is:

[0222] "Based on this, we are committed to using resources, independent innovation and brand consumption as our investment strategy."

[0223] After comparing and analyzing the two using the diff algorithm, the text position set obtained is as follows (this is only an example description and is not limited to this embodiment of the present application):

[0224] [{position: 0, operation: "replace", original character: "by", rewritten character: "base"},

[0225] {Position: 1, Operation: "Replace", Original character: "this", Rewritten character: "at"},

[0226] {Position: 2, Operation: "Insert", Original Character: null, Rewrite Character: "this"},

[0227] {Position: 3, Operation: "Retain", Original character: ",", Rewrite character: ","},

[0228] {Position: 4-5, Operation: "Replace", Original character: "Try", Rewritten character: "Committed to"},

[0229] {Position: 10, Operation: "Replace", Original character: "for", Rewritten character: "as"},

[0230] {Position: 15-16, Operation: "Replace", Original Character: "Main Line", Rewritten Character: "Strategy"}]

[0231] Next, based on the above-identified content differences, corresponding text blocks are respectively defined in the original text segment and the initial rewritten text segment. The specific practices include: according to the difference positions marked in the text position set, accurately divide the original text blocks directly related to these difference positions in the original text segment; at the same time, in the initial rewritten text segment, for the parts matching the text position set, correspondingly define the rewritten text blocks that correspond one by one to the original text blocks.

[0232] For example, according to the text position set in the above example, it can be determined that a range of difference positions is 0 - 2, and the text characters within this range include three consecutive text characters, namely "由", "此", and ",". Among them, "由" at position 0 is replaced by "基", "此" at position 1 is replaced by "于", and a new "此" character is inserted at position 2 in the rewritten text segment, resulting in the "," at position 2 in the original text being shifted to position 3 in the rewritten text. Therefore, the original text block needs to cover the complete content from position 0 to 2, which is "由此,", while the rewritten text block needs to be extended to position 0 to 3, including the replaced "基", "于", the inserted "此", and the shifted ",". Finally, the rewritten phrase "基于此," is formed.

[0233] In this division process, the position offset caused by the insertion operation needs to be dynamically processed. For example, "试图" (positions 4 - 5) in the original text segment is replaced by "致力于" (rewritten positions 5 - 7). Since the rewritten phrase has 1 more character than the original text, the position indexes of subsequent difference blocks need to be shifted as a whole. Similarly, "为" at the end of the difference block "为投资主线" (original position 10) is extended to "作为" (rewritten positions 11 - 12), and "主线" (original positions 15 - 16) is entirely replaced by "策略" (rewritten positions 17 - 18). By combining consecutive operations (such as combining the replacement of "由", "此" and the insertion of "此" into the head difference block) and explicitly marking the unchanged common parts (such as "资源、自主创新及品牌消费"), the corresponding blocks between the original text and the rewritten text can be accurately defined.

[0234] In addition, it is worth noting that in addition to the parts clearly marked as differences, there are also text intervals that remain unchanged between the two. Therefore, the intervals in the original text segment and the initial rewritten text segment where the text characters are exactly the same and continuous are also defined as the original text blocks and the corresponding rewritten text blocks to ensure the comprehensiveness and accuracy of the analysis.

[0235] Continuing with the above specific example, the corresponding relationship between the original text blocks and the rewritten text blocks as shown in Table 3 can be obtained:

[0236] Table 3

[0237] Original text block Rewrite a block of text thus, Based on this us us attempt Committed to With resources, independent innovation and brand consumption With resources, independent innovation and brand consumption for As invest invest Main Story Strategy

[0238] Figure 5 The process shown directly starts from the content difference level. By identifying the content difference between the original text segment and the rewritten text segment, determining the correspondence between the original text block and the rewritten text block is easy to implement and more intuitive.

[0239] Furthermore, in one embodiment, before identifying the content differences between the original text segment and the initial rewritten text segment, a rough match can be performed on the original text segment and the initial rewritten text segment. Here, the rough match aims to preliminarily determine the approximate correspondence or mapping range between the two, and then perform a detailed content difference comparison within this preliminarily defined small range. This approach is particularly suitable for processing text segments containing multiple short sentences, as it effectively narrows the focus of the comparison and improves the accuracy and efficiency of identifying content differences.

[0240] For example, the original text segment is:

[0241] The renovation of water-blocking forests within the river management area of ​​XX County is a systematic project that requires the concerted efforts of government departments at all levels, all sectors of society, and the general public. Through the implementation of this plan, we will effectively improve the river's ecological environment, enhance its flood-carrying capacity, and provide a solid ecological foundation for XX County's economic and social development. Let us work together to contribute to building a beautiful Minhou and a harmonious homeland!

[0242] The initial rewritten text segment is:

[0243] The remediation of water-blocking forests within the XX County river management area is a complex and comprehensive systematic project. Its smooth implementation and successful completion urgently require the concerted efforts of all levels of government, various sectors of society, and the general public. The implementation of this plan is expected to significantly improve the river's ecological environment and effectively enhance its flood-carrying capacity, thereby building a solid ecological security barrier for the economic and social development of XX County. We hereby call on all parties to join hands and contribute their wisdom and strength to building a more beautiful XX and a more harmonious home!

[0244] Through rough matching, we can obtain the corresponding relationship shown in Table 4 below:

[0245] Table 4

[0246]

[0247] Accordingly, before identifying the content difference between the original text segment and the initial rewritten text segment, it also includes: identifying the corresponding original text ranges and rewritten text ranges in the original text segment and the initial rewritten text segment; performing the following processing for each group of corresponding original text ranges and rewritten text ranges: taking the original text range therein as the original text segment, and taking the rewritten text range therein as the rewritten text segment, performing the steps of identifying the content difference between the original text segment and the initial rewritten text segment and subsequent steps.

[0248] As an optional implementation, an artificial intelligence model can be used to identify the original text range and the rewritten text range that correspond to the original text segment and the initial rewritten text segment. This artificial intelligence model can be a natural language processing model in the field of deep learning. These models are trained with large amounts of text data and possess rich linguistic knowledge and contextual understanding capabilities, and can accurately capture the semantic correspondence between texts. Then, when processing the task of matching the original text segment with the initial rewritten text segment, they can identify the key information in the text and accurately locate the corresponding part in the rewritten text.

[0249] By utilizing advanced artificial intelligence models, the embodiments of the present application can achieve more efficient and accurate text matching, providing a solid foundation for subsequent text processing and analysis tasks.

[0250] Finally, in one embodiment, the paragraph attributes of the original text segment are obtained, and the paragraph attributes of the original text segment are applied to the target rewritten text segment.

[0251] Paragraph attributes refer to a set of characteristics and settings that define a paragraph's appearance and behavior. In document processing software (such as Microsoft Word) or web design, paragraph attributes may include, but are not limited to, alignment (such as left, center, right, or justified), indentation (first line indent, hanging indent, etc.), line spacing, font style (such as font type, size, color, weight, italicization, etc.), paragraph borders, shading, text direction, and orphan line control. Together, these attributes determine the overall presentation of a paragraph.

[0252] In this embodiment, the paragraph attributes of the original text segment are completely applied to the target rewritten text segment, which means that the target rewritten text segment will inherit the format and layout settings of the original text segment while retaining the rewritten content, thereby maintaining the consistency and professionalism of the document.

[0253] Figure 6 This is a block diagram of an embodiment of a text rewriting device provided in an embodiment of the present application. Figure 6 As shown, the device includes:

[0254] A rewriting module 61 is configured to rewrite the original text segment to generate an initial rewritten text segment, wherein the original text segment includes at least one original text block;

[0255] A mapping module 62, configured to determine, in the initial rewritten text segment, a rewritten text block corresponding to each of the original text blocks;

[0256] The format inheritance module 63 is configured to apply the format attributes of each original text block to the corresponding rewritten text block to obtain a target rewritten text segment.

[0257] In a possible implementation, the rewriting module 61 includes:

[0258] A block division unit is used to perform text block processing on the original text segment to obtain multiple original text blocks;

[0259] a label adding unit, configured to add an identification label to each of the original text blocks in the original text segment to obtain a label text segment;

[0260] The text rewriting unit is used to input the label text segment into a text rewriting model to obtain an initial rewritten text segment containing the identification label.

[0261] In a possible implementation, the mapping module 62 is specifically configured to:

[0262] The following processing is performed for each original text block:

[0263] The text blocks in the initial rewritten text segment that have the same identification tag as the original text block are determined as the rewritten text blocks corresponding to the original text blocks.

[0264] In a possible implementation, the block division unit is specifically configured to:

[0265] Reading each text character in the original text segment in sequence, and when a first text character is read, recording a first tag for the first text character;

[0266] Starting from the second text character read, when the format attribute of the currently read text character is different from the format attribute of the text character read previously, recording the first tag for the currently read text character, and recording the second tag for the text character read previously;

[0267] The text characters between the adjacent first markers and the second markers are merged into the same original text block to obtain multiple original text blocks.

[0268] In a possible implementation, the block division unit is specifically configured to:

[0269] The original text segment is segmented, and each word obtained is determined as an original text block, thereby obtaining a plurality of original text blocks.

[0270] In a possible implementation, the format inheritance module 63 is specifically configured to:

[0271] Identifying a first text block having a specific formatting attribute from a plurality of original text blocks of the original text segment;

[0272] If it is determined that the first text block is semantically incomplete, a semantically complete second text block is determined from a text character set consisting of the first text block and its adjacent text characters, the specific format attributes are applied to the second text block, and the format attributes of the text characters in the text character set other than the second text block are modified to the first format attributes; wherein the second text block and the first text block have overlapping text characters; and the first format attributes are format attributes other than the specific format attributes that appear in the text character set.

[0273] Based on the original text segment after the above processing, the step of applying the format attributes of each original text block to the corresponding rewritten text block is performed.

[0274] In a possible implementation, the mapping module 62 includes:

[0275] A word segmentation unit is used to perform word segmentation on the initial rewritten text segment, and determine each word obtained by word segmentation as a rewritten text block, thereby obtaining a plurality of rewritten text blocks;

[0276] A determination unit, configured to perform the following processing on each of the rewritten text blocks:

[0277] An original text block that satisfies a preset semantic consistency condition with the rewritten text block is determined in the original text segment, and the original text block that satisfies the preset semantic consistency condition is determined as the original text block corresponding to the rewritten text block.

[0278] In a possible implementation, the mapping module 62 includes:

[0279] an identification unit, configured to identify a content difference between the original text segment and the initial rewritten text segment;

[0280] The defining unit is configured to define a plurality of original text blocks in the original text segment based on the content difference portion, and define a rewritten text block corresponding to each of the original text blocks in the initial rewritten text segment.

[0281] In a possible implementation, the identification unit is specifically configured to:

[0282] Using a preset difference comparison algorithm, identifying a set of text locations in the original text segment that differ from the initial rewritten text segment;

[0283] The defining unit is specifically used to:

[0284] Based on the text position set, defining original text blocks corresponding to text positions with different contents in the original text segment, and defining rewritten text blocks corresponding to the original text blocks for portions corresponding to the text position set in the initial rewritten text segment;

[0285] Furthermore, continuous text intervals in which the text characters in the original text segment and the initial rewritten text segment are completely identical are defined as original text blocks and corresponding rewritten text blocks, respectively.

[0286] In a possible implementation, the device further includes:

[0287] a coarse screening module, configured to identify the original text range and the rewritten text range corresponding to each other in the original text segment and the initial rewritten text segment before identifying the content difference between the original text segment and the initial rewritten text segment;

[0288] The recognition unit is used to perform the following processing for each group of corresponding original text ranges and rewritten text ranges: taking the original text range as the original text segment and the rewritten text range as the rewritten text segment, and performing the step of identifying the content difference between the original text segment and the initial rewritten text segment.

[0289] In a possible implementation, the format inheritance module 63 is further configured to:

[0290] Obtaining the paragraph attributes of the original text segment;

[0291] Apply the paragraph attributes of the original text segment to the target rewritten text segment.

[0292] like Figure 7 As shown, an embodiment of the present application provides an electronic device, including a processor 111, a communication interface 112, a memory 113 and a communication bus 114, wherein the processor 111, the communication interface 112, and the memory 113 communicate with each other through the communication bus 114.

[0293] Memory 113, for storing computer programs;

[0294] In one embodiment of the present application, the processor 111 is configured to implement the text rewriting method provided by any of the aforementioned method embodiments when executing a program stored in the memory 113, including:

[0295] Performing content rewriting processing on the original text segment to generate an initial rewritten text segment, wherein the original text segment includes at least one original text block;

[0296] Determining, in the initial rewritten text segment, a rewritten text block corresponding to each of the original text blocks;

[0297] The format attributes of each original text block are applied to the corresponding rewritten text block to obtain a target rewritten text segment.

[0298] An embodiment of the present application further provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of the text rewriting method provided in any of the aforementioned method embodiments are implemented.

[0299] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of this embodiment.

[0300] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a general hardware platform, or of course, by hardware. Based on this understanding, the above technical solution, in essence, or the part that contributes to the relevant technology, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or certain parts of the embodiment.

[0301] It should be understood that the terms used herein are for the purpose of describing specific example embodiments only and are not intended to be limiting. Unless the context clearly indicates otherwise, the singular forms "one", "an" and "said" as used herein may also be meant to include plural forms. The terms "comprise", "include", "contain" and "have" are inclusive and therefore specify the presence of stated features, steps, operations, elements and / or parts, but do not exclude the presence or addition of one or more other features, steps, operations, elements, parts, and / or combinations thereof. The method steps, processes, and operations described herein are not to be construed as necessarily requiring them to be performed in the specific order described or illustrated, unless the order of execution is clearly indicated. It should also be understood that additional or alternative steps may be used.

[0302] The foregoing is merely a list of specific embodiments of the present application, intended to enable those skilled in the art to understand or implement the present application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application is not limited to the embodiments shown herein, but is intended to conform to the broadest scope consistent with the principles and novel features of the present application.

Claims

1. A text rewriting method, characterized in that: The method comprises: Performing content rewriting processing on the original text segment to generate an initial rewritten text segment, wherein the original text segment includes at least one original text block; Determining, in the initial rewritten text segment, a rewritten text block corresponding to each of the original text blocks; The format attributes of each original text block are applied to the corresponding rewritten text block to obtain a target rewritten text segment.

2. The method according to claim 1, characterized in that The step of rewriting the original text segment to generate an initial rewritten text segment includes: Perform text block processing on the original text segment to obtain multiple original text blocks; Adding an identification tag to each of the original text blocks in the original text segment to obtain a labeled text segment; The labeled text segment is input into a text rewriting model to obtain an initial rewritten text segment containing the identification tag.

3. The method according to claim 2, characterized in that Determining the rewritten text blocks corresponding to each of the original text blocks in the initial rewritten text segment includes: The following processing is performed for each original text block: The text blocks in the initial rewritten text segment that have the same identification tag as the original text block are determined as the rewritten text blocks corresponding to the original text blocks.

4. The method according to claim 2, characterized in that The original text segment is subjected to text block processing to obtain multiple original text blocks, including: Reading each text character in the original text segment in sequence, and when a first text character is read, recording a first tag for the first text character; Starting from the second text character read, when the format attribute of the currently read text character is different from the format attribute of the text character read previously, recording the first tag for the currently read text character, and recording the second tag for the text character read previously; The text characters between the adjacent first markers and the second markers are merged into the same original text block to obtain multiple original text blocks.

5. The method according to claim 2, characterized in that The text segment is segmented into blocks to obtain a plurality of original text blocks, including: The original text segment is segmented, and each word obtained is determined as an original text block, thereby obtaining a plurality of original text blocks.

6. The method according to claim 1, characterized in that The step of applying the formatting attributes of each original text block to the corresponding rewritten text block includes: Identifying a first text block having a specific formatting attribute from a plurality of original text blocks of the original text segment; If it is determined that the first text block is semantically incomplete, a semantically complete second text block is determined from a text character set consisting of the first text block and its adjacent text characters, the specific format attributes are applied to the second text block, and the format attributes of the text characters in the text character set other than the second text block are modified to the first format attributes; wherein the second text block and the first text block have overlapping text characters; and the first format attributes are format attributes other than the specific format attributes that appear in the text character set. Based on the original text segment after the above processing, the step of applying the format attributes of each original text block to the corresponding rewritten text block is performed.

7. The method according to claim 1, characterized in that Determining the rewritten text blocks corresponding to each of the original text blocks in the initial rewritten text segment includes: Performing word segmentation on the initial rewritten text segment, and determining each word obtained by the word segmentation as a rewritten text block, thereby obtaining a plurality of rewritten text blocks; The following processing is performed for each rewritten text block: An original text block that satisfies a preset semantic consistency condition with the rewritten text block is determined in the original text segment, and the original text block that satisfies the preset semantic consistency condition is determined as the original text block corresponding to the rewritten text block.

8. The method according to claim 1, characterized in that Determining the rewritten text blocks corresponding to each of the original text blocks in the initial rewritten text segment includes: identifying content differences between the original text segment and the initial rewritten text segment; Based on the content difference portion, a plurality of original text blocks are defined in the original text segment, and a rewritten text block corresponding to each of the original text blocks is defined in the initial rewritten text segment.

9. The method according to claim 8, characterized in that The identifying of the content difference between the original text segment and the initial rewritten text segment includes: Using a preset difference comparison algorithm, identifying a set of text locations in the original text segment that differ from the initial rewritten text segment; The step of defining a plurality of original text blocks in the original text segment based on the content difference portion, and defining a rewritten text block corresponding to each of the original text blocks in the initial rewritten text segment, includes: Based on the text position set, defining original text blocks corresponding to text positions with different contents in the original text segment, and defining rewritten text blocks corresponding to the original text blocks for portions corresponding to the text position set in the initial rewritten text segment; Furthermore, continuous text intervals in which the text characters in the original text segment and the initial rewritten text segment are completely identical are defined as original text blocks and corresponding rewritten text blocks, respectively.

10. The method according to claim 8, characterized in that Before identifying the content difference between the original text segment and the initial rewritten text segment, the method further includes: identifying corresponding original text ranges and rewritten text ranges in the original text segment and the initial rewritten text segment; The following processing is performed for each group of corresponding original text ranges and rewritten text ranges: the original text range is treated as the original text segment, and the rewritten text range is treated as the rewritten text segment, and the steps of identifying the content difference between the original text segment and the initial rewritten text segment and subsequent steps are performed.

11. The method according to any one of claims 1 to 10, characterized in that The method further comprises: Obtaining the paragraph attributes of the original text segment; Apply the paragraph attributes of the original text segment to the target rewritten text segment.

12. A text rewriting device, characterized in that: The device comprises: a rewriting module, configured to perform content rewriting processing on an original text segment to generate an initial rewritten text segment, wherein the original text segment includes at least one original text block; A mapping module, configured to determine, in the initial rewritten text segment, a rewritten text block corresponding to each of the original text blocks; The format inheritance module is used to apply the format attributes of each original text block to the corresponding rewritten text block to obtain a target rewritten text segment.

13. An electronic device, characterized in that: include: A processor and a memory, wherein the processor is configured to execute a text rewriting program stored in the memory to implement the text rewriting method according to any one of claims 1 to 11.