Text rewriting contrast method and device, electronic equipment and storage medium
By dynamically adjusting the text rewriting comparison method of the difference detection granularity, the problem of limited efficiency and accuracy caused by the single granularity in the existing technology is solved, and efficient and accurate text rewriting analysis is achieved in different scenarios.
Patent Information
- Application Number
- CN202510784924.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-09-12
AI Technical Summary
The existing technology uses a single or fixed method to select the granularity of difference detection, which limits the efficiency and accuracy of text rewriting analysis and cannot meet the diverse needs of different users in different scenarios.
A text rewriting comparison method is provided, which performs first granularity difference detection on the original text and the rewritten text, dynamically adjusts the difference detection granularity in response to a granularity conversion trigger condition, forms a second granularity difference unit set, and outputs a text rewriting comparison view with difference markings at the second granularity.
It realizes flexible adjustment of the granularity of difference display, which can meet the difference analysis needs of different users in different scenarios, improves the efficiency and accuracy of text rewriting analysis, and can grasp the overall differences at a macro level and analyze local details at a micro level.
Smart Images

Figure CN120633628A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of text processing, and in particular to a text rewriting and comparison method, device, electronic device, and storage medium. Background Art
[0002] In the field of text comparison technology, difference detection technology is a key means to identify and present the differences between different texts, and is widely used in scenarios such as code comparison and translation comparison.
[0003] Currently, the selection of the granularity for difference detection is mainly based on fixed rules or single strategies. For example, in code comparison scenarios, differences are often compared at the "line" granularity, while in translation comparison scenarios, differences are often compared at the "paragraph" granularity.
[0004] However, this approach of using a single strategy or fixed rules for difference detection has significant limitations. For example, if the difference detection granularity is set too fine, such as down to each character in a text rewrite comparison, while it can accurately identify subtle differences, it can seriously affect text readability and increase the user's information processing burden. If the difference detection granularity is set too coarse, such as using paragraphs as units in a text rewrite comparison, while it can improve macro-comparison efficiency, it can reduce comparison accuracy, leading to the overlooking of important subtle differences, affecting subsequent decision-making and operations. Summary of the Invention
[0005] The present application provides a text rewriting comparison method, device, electronic device and storage medium to solve the technical problem that the existing technology adopts a single or fixed difference detection strategy, which has great limitations and affects the efficiency and accuracy of text rewriting analysis.
[0006] In a first aspect, the present application provides a text rewriting comparison method, the method comprising:
[0007] Performing first-granularity difference detection on the original text and the corresponding rewritten text to obtain a first-granularity difference unit set;
[0008] In response to a granularity conversion trigger condition, determining a second granularity;
[0009] Updating the first granularity difference unit set to a second granularity difference unit set corresponding to the second granularity;
[0010] Based on the second granularity difference unit set, a text rewriting comparison view with difference markings performed at the second granularity is output.
[0011] In a possible implementation, updating the first granularity difference unit set to a second granularity difference unit set corresponding to the second granularity includes:
[0012] The first granularity difference unit in the first granularity difference unit set is adjusted to a second granularity difference unit to form a second granularity difference unit set.
[0013] In a possible implementation, adjusting the first granularity difference unit in the first granularity difference unit set to a second granularity difference unit includes:
[0014] A plurality of first granularity difference units in the first granularity difference unit set that meet the aggregation condition corresponding to the second granularity are aggregated into a second granularity difference unit.
[0015] In a possible implementation, aggregating a plurality of first granularity difference units that meet the aggregation condition corresponding to the second granularity in the first granularity difference unit set into a second granularity difference unit includes:
[0016] Based on the second granularity, dividing the first granularity difference units in the first granularity difference unit set into intervals, wherein the text combination corresponding to the first granularity difference units contained in an interval conforms to the semantic boundary of the second granularity;
[0017] At least one first granularity difference unit in the same interval is merged into a second granularity difference unit, and an operation type identifier is assigned to the second granularity difference unit.
[0018] In a possible implementation, determining the second granularity in response to the granularity conversion triggering condition includes:
[0019] determining a type of rewriting operation from the original text to the rewritten text;
[0020] Obtaining a preset target granularity corresponding to the rewrite operation type;
[0021] In a case where it is determined that the target granularity is inconsistent with the first granularity, determining that a granularity conversion trigger condition is satisfied;
[0022] In response to the granularity conversion triggering condition, the target granularity is determined to be a second granularity.
[0023] In a possible implementation, before performing first-granularity difference detection on the original text and the corresponding rewritten text to obtain a first-granularity difference unit set, the method further includes:
[0024] determining a type of rewriting operation from the original text to the rewritten text;
[0025] Obtaining a preset target granularity corresponding to the rewrite operation type;
[0026] The target particle size is used as the first particle size.
[0027] In a possible implementation, determining the type of rewriting operation from the original text to the rewritten text includes:
[0028] performing intent recognition on a natural language processing instruction input by a user, and determining a type of rewriting operation from the original text to the rewritten text based on the intent recognition result, wherein the natural language processing instruction is used to instruct to rewrite the original text;
[0029] Alternatively, obtain a function entry identifier that triggers the rewriting of the original text; search a preset operation type mapping table based on the function entry identifier to obtain the rewriting operation type from the original text to the rewritten text, wherein the operation type mapping table includes a correspondence between the function entry identifier and the rewriting operation type.
[0030] In a possible implementation, determining the second granularity in response to the granularity conversion triggering condition includes:
[0031] Determining, based on the first granularity difference unit set, a first difference marking density index for marking differences between the original text and the rewritten text at the first granularity;
[0032] When it is determined that the first difference mark density indicator meets a preset condition, determining that a granularity conversion trigger condition is met;
[0033] In response to the granularity conversion triggering condition, a second granularity is determined.
[0034] In a possible implementation manner, before outputting the text rewriting comparison view with difference markings at the second granularity based on the second granularity difference unit set, the method further includes:
[0035] Determining, based on a current set of second-granularity difference units, a second difference marking density index for marking differences between the original text and the rewritten text at the current second granularity;
[0036] When it is determined that the current second difference mark density indicator satisfies the preset condition, determining again that the granularity conversion trigger condition is satisfied;
[0037] In response to the re-determined granularity conversion triggering condition, the second granularity is determined again until the current second difference mark density indicator does not meet the preset condition, or the preset level granularity is determined as the second granularity.
[0038] In a possible implementation, determining the second granularity includes:
[0039] From a preset granularity level sequence, a preset level granularity is selected and determined as the second granularity.
[0040] In a possible implementation, re-determining the second granularity includes:
[0041] From the preset granularity level sequence, a granularity at a level adjacent to the current second granularity is selected and determined as the second granularity again.
[0042] In a possible implementation, determining a first difference marking density index for marking differences between the original text and the rewritten text at the first granularity includes:
[0043] Determine a first granularity difference unit having at least one difference point in the first granularity difference unit set as a target difference unit;
[0044] The proportion of the target difference unit in the first granularity difference unit set is determined, and the proportion is determined as a first difference marking density index for performing difference marking on the original text and the rewritten text at the first granularity.
[0045] In a possible implementation, performing first-granularity difference detection on the original text and its corresponding rewritten text to obtain a first-granularity difference unit set includes:
[0046] Decomposing the original text and the corresponding rewritten text into first granularity units that meet the first granularity respectively;
[0047] Difference detection is performed based on first granularity units corresponding to the original text and its corresponding rewritten text to obtain a first granularity difference unit set.
[0048] In a possible implementation, updating the first granularity difference unit set to a second granularity difference unit set corresponding to the second granularity includes:
[0049] Decomposing the original text and the corresponding rewritten text into second granularity units that conform to the second granularity respectively;
[0050] Difference detection is performed based on the second granularity units corresponding to the original text and the corresponding rewritten text to obtain a set of second granularity difference units.
[0051] In a second aspect, the present application provides a text rewriting comparison device, the device comprising:
[0052] A first difference detection module is configured to perform first-granularity difference detection on the original text and its corresponding rewritten text to obtain a first-granularity difference unit set;
[0053] a granularity determination module, configured to determine a second granularity in response to a granularity conversion trigger condition;
[0054] a granularity conversion module, configured to update the first granularity difference unit set to a second granularity difference unit set corresponding to the second granularity;
[0055] The rewriting comparison module is configured to output a text rewriting comparison view with difference markings performed at the second granularity based on the second granularity difference unit set.
[0056] In a possible implementation, the granularity conversion module is specifically configured to:
[0057] The first granularity difference unit in the first granularity difference unit set is adjusted to a second granularity difference unit to form a second granularity difference unit set.
[0058] In a possible implementation, the granularity conversion module is specifically configured to:
[0059] A plurality of first granularity difference units in the first granularity difference unit set that meet the aggregation condition corresponding to the second granularity are aggregated into a second granularity difference unit.
[0060] In a possible implementation, the granularity conversion module includes:
[0061] a semantic segmentation unit configured to segment the first granularity difference units in the first granularity difference unit set into intervals based on the second granularity, wherein a text combination corresponding to the first granularity difference units contained in an interval conforms to a semantic boundary of the second granularity;
[0062] The merging unit is configured to merge at least one first granularity difference unit in the same interval into a second granularity difference unit, and assign an operation type identifier to the second granularity difference unit.
[0063] In a possible implementation, the granularity determination module includes:
[0064] a scene recognition unit, configured to determine a type of rewriting operation from the original text to the rewritten text;
[0065] A first determining unit, configured to obtain a preset target granularity corresponding to the rewrite operation type;
[0066] a first condition determination unit, configured to determine that a granularity conversion trigger condition is satisfied when it is determined that the target granularity is inconsistent with the first granularity;
[0067] The second determining unit is configured to determine the target granularity as a second granularity in response to the granularity conversion triggering condition.
[0068] In a possible implementation, the device further includes:
[0069] The first granularity determination module is used to determine the type of rewriting operation from the original text to the rewritten text before performing first granularity difference detection on the original text and its corresponding rewritten text to obtain a first granularity difference unit set; obtain a preset target granularity corresponding to the rewriting operation type; and use the target granularity as the first granularity.
[0070] In a possible implementation, the scene recognition unit is specifically configured to:
[0071] performing intent recognition on a natural language processing instruction input by a user, and determining a type of rewriting operation from the original text to the rewritten text based on the intent recognition result, wherein the natural language processing instruction is used to instruct to rewrite the original text;
[0072] Alternatively, obtain a function entry identifier that triggers the rewriting of the original text; search a preset operation type mapping table based on the function entry identifier to obtain the rewriting operation type from the original text to the rewritten text, wherein the operation type mapping table includes a correspondence between the function entry identifier and the rewriting operation type.
[0073] In a possible implementation, the granularity determination module includes:
[0074] a mark density prediction unit, configured to determine, based on the first granularity difference unit set, a first difference mark density index for marking differences between the original text and the rewritten text at the first granularity;
[0075] a second condition determination unit, configured to determine that a granularity conversion trigger condition is satisfied when it is determined that the first difference mark density indicator satisfies a preset condition;
[0076] The third determining unit is configured to determine a second granularity in response to the granularity conversion triggering condition.
[0077] In one possible implementation, the mark density prediction unit is further configured to: before outputting the text rewriting comparison view with difference marks at the second granularity based on the second granularity difference unit set, determine, based on the current second granularity difference unit set, a second difference mark density index for difference marks between the original text and the rewritten text at the current second granularity;
[0078] The second condition determination unit is further configured to: upon determining that the current second difference mark density indicator satisfies the preset condition, determine again that the granularity conversion trigger condition is satisfied;
[0079] The third determining unit is further configured to re-determine the second granularity in response to the re-determined granularity conversion triggering condition, until the current second difference mark density indicator does not meet the preset condition, or the preset level granularity is determined as the second granularity.
[0080] In a possible implementation, the third determining unit determines the second granularity, including:
[0081] From a preset granularity level sequence, a preset level granularity is selected and determined as the second granularity.
[0082] In a possible implementation manner, the third determining unit determining the second granularity again includes:
[0083] From the preset granularity level sequence, a granularity at a level adjacent to the current second granularity is selected and determined as the second granularity again.
[0084] In a possible implementation, the marker density prediction unit is specifically configured to:
[0085] Determine a first granularity difference unit having at least one difference point in the first granularity difference unit set as a target difference unit;
[0086] The proportion of the target difference unit in the first granularity difference unit set is determined, and the proportion is determined as a first difference marking density index for performing difference marking on the original text and the rewritten text at the first granularity.
[0087] In a possible implementation, the first difference detection module includes:
[0088] A first granularity decomposition unit, configured to decompose the original text and its corresponding rewritten text into first granularity units conforming to a first granularity;
[0089] The first difference detection unit is configured to perform difference detection based on first granularity units corresponding to the original text and the corresponding rewritten text, to obtain a first granularity difference unit set.
[0090] In a possible implementation, the granularity conversion module includes:
[0091] A second granularity decomposition unit decomposes the original text and the corresponding rewritten text into second granularity units conforming to a second granularity;
[0092] The second difference detection unit is configured to perform difference detection based on the second granularity units corresponding to the original text and the corresponding rewritten text, to obtain a second granularity difference unit set.
[0093] In a third aspect, the present application provides an electronic device comprising: a processor and a memory, wherein the processor is configured to execute a text rewriting comparison program stored in the memory to implement the text rewriting comparison method described in any one of the first aspects.
[0094] In a fourth aspect, the present application provides a storage medium storing one or more programs, which can be executed by one or more processors to implement the text rewriting comparison method described in any one of the first aspects.
[0095] The technical solution provided by the embodiments of the present application has the following advantages over the prior art: the method provided by the embodiments of the present application performs a first-granularity difference detection between the original text and its corresponding rewritten text to obtain a first-granularity difference unit set; determines a second granularity in response to a granularity conversion trigger condition; updates the first-granularity difference unit set to a second-granularity difference unit set corresponding to the second granularity; and outputs a text rewrite comparison view with difference markings at the second granularity based on the second-granularity difference unit set. This allows for flexible adjustment of the difference display granularity, meeting the diverse needs of different users for difference analysis in different scenarios. Whether a macroscopic grasp of overall differences or a microscopic analysis of local details is required, both can be easily achieved by adjusting the difference detection granularity, greatly improving the efficiency and accuracy of text rewrite analysis. BRIEF DESCRIPTION OF THE DRAWINGS
[0096] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0097] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0098] One or more embodiments are exemplarily illustrated by pictures in the corresponding drawings. These exemplifications do not constitute limitations on the embodiments. Elements with the same reference numerals in the drawings are represented as similar elements. Unless otherwise stated, the figures in the drawings do not constitute proportional limitations.
[0099] Figure 1 A flowchart of an embodiment of a text rewriting comparison method provided in an embodiment of the present application;
[0100] Figure 2 Examples of rewriting contrast views for texts that are differentially marked at different granularities;
[0101] Figure 3 A flowchart of another content rewriting method provided in an embodiment of the present application;
[0102] Figure 4 A flowchart of another content rewriting method provided in an embodiment of the present application;
[0103] Figure 5 A flowchart of another content rewriting method provided in an embodiment of the present application;
[0104] Figure 6 A flowchart of another content rewriting method provided in an embodiment of the present application;
[0105] Figure 7 This is a block diagram of an embodiment of a text rewriting and comparison device provided in an embodiment of the present application;
[0106] Figure 8 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0107] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0108] The disclosure below provides many different embodiments or examples for implementing different structures of the present application. In order to simplify the disclosure of the present application, the components and settings of specific examples are described below. Of course, these are merely examples and are not intended to limit the present application. In addition, the present application may repeat reference numbers and / or letters in different examples. Such repetition is for the purpose of simplicity and clarity and does not in itself indicate the relationship between the various embodiments and / or settings discussed.
[0109] In order to solve the technical problem that the existing technology adopts a single or fixed difference detection strategy, which has great limitations and affects the efficiency and accuracy of text rewriting analysis, the present application provides a text rewriting comparison method, device, electronic device and storage medium, which can realize flexible adjustment of the difference display granularity and meet the diverse needs of different users for difference analysis in different scenarios. Whether it is necessary to grasp the overall difference at a macro level or analyze local details at a micro level, it can be easily achieved by adjusting the difference detection granularity, greatly improving the efficiency and accuracy of text rewriting analysis.
[0110] Figure 1This is a flow chart of an embodiment of a text rewriting comparison method provided in the embodiment of this application. Figure 1 As shown, the method includes the following steps:
[0111] Step 101: Perform first-granularity difference detection on the original text and its corresponding rewritten text to obtain a first-granularity difference unit set.
[0112] Step 102: In response to a granularity conversion triggering condition, determine a second granularity.
[0113] Step 103: Update the first granularity difference unit set to a second granularity difference unit set corresponding to the second granularity.
[0114] Step 104: Based on the second granularity difference unit set, output a text rewriting comparison view with difference markings at the second granularity.
[0115] For ease of understanding, a unified explanation of steps 101 to 104 is provided below:
[0116] First, in general, the description of steps 101 to 104 demonstrates that the core logic of this embodiment lies in dynamically switching the difference detection granularity to flexibly adapt to user needs for text difference display in different scenarios. In one embodiment, three difference detection granularities are designed: word granularity, sentence granularity, and paragraph granularity.
[0117] Word-level difference detection refers to detecting text differences using words as the smallest unit to reflect changes at the lexical level. Specifically, the system can break down the text into word-level units for difference detection to identify changes at the lexical level. Specifically, the system compares the original text with its corresponding rewritten text word by word, identifying and marking the type of difference between each word before and after the rewrite, forming a set of difference units based on words. Unlike languages like English that use spaces to separate words, Chinese uses characters as the basic writing unit, which are independent ideographic symbols. Words are composed of multiple characters, but there are no explicit separators such as spaces between words when written. Therefore, word segmentation can be used to break the text down into word-level units. Alternatively, after performing character-level difference detection on the text, the character-level differences can be aggregated into word-level units. Difference types include deletion, addition, and no difference.
[0118] For example, the original text is: "Mom bought apples, bananas, and oranges", and the rewritten text is: "Mom bought grapes and oranges". By detecting differences at word granularity, we can obtain the difference detection results shown in Table 1 below:
[0119] Table 1
[0120] Original word segmentation Rewrite participle Matching results Difference Type A1: Mom B1: Mom Exact match No difference A2: I bought it B2: Bought Exact match No difference A3: Apple none Unmatched delete A4: Banana none Unmatched delete A5: and none Unmatched delete A6: Orange B4: Orange Exact match No difference none B3: Grapes Unmatched Increase
[0121] In Table 1 above, each row corresponds to a word-granularity difference unit, and these rows together constitute the word-granularity difference unit set. As can be seen from Table 1, each difference unit records a text unit at a certain word-granularity level in the original text, a text unit at the corresponding word-granularity level in the rewritten text, and the type of difference between the two text units.
[0122] As the examples above demonstrate, word-level difference detection is ideal for situations where subtle changes need to be accurately identified and displayed, such as synonym replacement and typo correction. Word-level difference detection allows users to clearly see the additions, deletions, and modifications of each word, helping them gain a deeper understanding of the details of the text changes.
[0123] Sentence-level difference detection involves detecting text differences using sentences as the smallest unit to reflect changes at the sentence level. A sentence is the basic unit of complete semantic expression in a language. It is typically composed of words or phrases combined according to grammatical rules and possesses independent intonation, a logical subject, and a predicate (some omitted sentences require contextual interpretation). In written language, punctuation marks such as periods (.), question marks (?), exclamation points (!), and semicolons (;) are often used as end-of-sentence markers, which are important for segmenting text. Specifically, the system can break down the text into sentence-level units for difference detection to identify changes at the sentence level. Specifically, the system compares the original text with its corresponding rewritten text sentence by sentence, identifying and marking the difference type between each sentence before and after the rewrite, forming a set of difference units based on sentences. Alternatively, after performing character-level (or word-level) difference detection on the text, the differences at the character-level (or word-level) are aggregated into sentence-level differences. Here, difference types also include deletions, additions, and no differences.
[0124] For example, the original text is: "Mom bought apples and bananas today", and the rewritten text is: "Mom bought apples and bananas today". By detecting differences at the sentence granularity, the difference detection results shown in Table 2 below can be obtained:
[0125] Table 2
[0126]
[0127] In Table 2 above, each row corresponds to a sentence-level difference unit, and these rows together constitute the set of sentence-level difference units. As can be seen from Table 2, each difference unit records a text unit at a certain sentence-level granularity in the original text, a corresponding text unit at a certain sentence-level granularity in the rewritten text, and the type of difference between the two text units.
[0128] As the examples above demonstrate, sentence-level difference detection is suitable for scenarios where you need to display sentence-level changes while maintaining a certain level of text readability. For example, when reorganizing or optimizing sentences within a paragraph, sentence-level difference detection can help users quickly grasp the changes at the sentence level without having to delve into the details of each word.
[0129] Paragraph-level difference detection further expands the scope of difference comparison. It uses paragraphs as the smallest unit for detecting text differences, reflecting changes at the paragraph level. A paragraph is a relatively independent semantic unit in a text, consisting of a group of sentences centered around a common theme, visually separated by line breaks or blank lines. Its core function is to break down long texts into logical hierarchies to improve readability. Paragraphs can be segmented from the original text based on line breaks, blank lines, etc. Paragraphs can also be segmented from the original text based on logical similarity. Specifically, the system can perform difference detection on the paragraph level to identify changes at the paragraph level. Specifically, the system compares the original text with its corresponding rewritten text paragraph by paragraph, identifying and marking the overall differences within each paragraph, forming a set of difference units based on the paragraph as the basic unit. Alternatively, after performing difference detection on the text at the character (or word, or sentence) level, these differences can be aggregated to the paragraph level. Here, the types of differences still include deletion, addition, and no difference, but the focus is on the overall changes at the paragraph level rather than changes to individual sentences or words.
[0130] For example, if the original text contains a paragraph describing a shopping list, and the rewritten text completely deletes this paragraph and adds a paragraph describing cooking steps, the difference detection at the paragraph granularity can produce the difference detection results shown in Table 3 below:
[0131] Table 3
[0132]
[0133] As can be seen from the above examples, paragraph-level difference detection is suitable for scenarios where overall structural changes or major modifications need to be displayed, such as paragraph order adjustments, adding or deleting paragraphs, etc., helping users grasp the overall changes of the text from a macro perspective.
[0134] It can be seen from the above description that difference detection of different granularities has different characteristics and applicable scenarios in text comparison. In order to further meet the diverse needs, the embodiment of the present application provides a dynamic conversion mechanism for difference detection granularity. Exemplarily, the dynamic conversion of difference detection granularity includes the conversion from word granularity to sentence granularity, the conversion from sentence granularity to paragraph granularity, and the cross-level conversion from word granularity to paragraph granularity. At the same time, it can also include reverse conversion, such as the conversion from sentence granularity to word granularity, the conversion from paragraph granularity to sentence granularity, and even the cross-level conversion from paragraph granularity directly to word granularity, which ensures that a suitable text difference display method can be provided in different demand scenarios.
[0135] Specifically, first, in step 101 , a first-granularity difference detection is performed on the original text and the corresponding rewritten text to obtain a first-granularity difference unit set.
[0136] In one embodiment, the first granularity is a preset or user-selected initial granularity, and its specific form is not limited, such as word granularity, sentence granularity, or paragraph granularity. For example, the user can select from the granularity options provided by the system based on their needs, and the system determines the first granularity based on the user's selection.
[0137] In another embodiment, the first granularity is determined by: determining the type of rewriting operation from the original text to the rewritten text; obtaining a preset target granularity corresponding to the rewriting operation type; and using the target granularity as the first granularity.
[0138] Among them, the rewriting operation type refers to the specific method or purpose of modifying the original text, which can be divided into three categories: first, rewriting the overall content, second, rewriting the local content, and third, adjusting the text structure. Overall rewriting covers operations such as polishing and expansion, aiming to comprehensively optimize the overall content of the text. Local rewriting focuses on specific parts of the text, such as correcting typos, beautifying wording, and other minor adjustments. Structural adjustments involve significant changes to the text structure, such as converting the text into a table or list format, or changing the narrative order from chronological to reverse.
[0139] In one embodiment, an exemplary implementation of determining the type of rewriting operation from the original text to the rewritten text includes: performing intent recognition on a natural language processing instruction input by the user, and determining the type of rewriting operation from the original text to the rewritten text based on the intent recognition result, wherein the natural language processing instruction is used to instruct the rewriting of the original text.
[0140] When users need to rewrite the original text, they can enter natural language processing instructions, such as "Please polish this text to make it more vivid" or "Please expand this paragraph and add some details." After receiving these instructions, the system uses natural language processing technology to identify the instructions' intent. Specifically, natural language processing technology analyzes the keywords, grammatical structure, and semantic information in the instructions to understand the user's desired rewriting goals. For example, if the instructions include words such as "polish" and "vivid," the system will recognize that the user wants to optimize the overall content of the text; if the instructions include words such as "expand" and "add details," the system will determine that the user needs to enrich the content locally. Based on the intent recognition results, the system can accurately determine the type of rewriting operation from the original text to the rewritten text, such as global content rewriting (polishing, expansion, etc.), local content rewriting (correcting typos, replacing words, etc.), or text structure adjustment (converting to a table, adjusting paragraph order, etc.).
[0141] This method directly utilizes the user's natural language input and conforms to the user's operating habits, allowing the user to express rewriting requirements in a natural and convenient way, and the system can also quickly and accurately understand and determine the type of rewriting operation.
[0142] In another embodiment, an exemplary implementation of determining the type of rewriting operation from the original text to the rewritten text includes: obtaining a function entry identifier that triggers the rewriting process of the original text; searching a preset operation type mapping table according to the function entry identifier to obtain the type of rewriting operation from the original text to the rewritten text, wherein the operation type mapping table includes a correspondence between the function entry identifier and the rewriting operation type.
[0143] In software or systems, different rewriting functions often have different entry points, such as buttons or options in the menu bar for "polishing," "expanding," and "structuring." Each entry point has a corresponding identifier. When a user clicks a function entry point to trigger the rewriting of the original text, the system retrieves the identifier for that function entry point.
[0144] The system has a pre-set operation type mapping table that details the correspondence between function entry identifiers and rewrite operation types. For example, the "Embellishment" entry identifier corresponds to the "Embellishment" operation type in the overall content rewrite, while the "Expand" entry identifier corresponds to the "Expand" operation type in the overall content rewrite. Based on the function entry identifiers obtained, the system searches the operation type mapping table to determine the rewrite operation type from the original text to the rewritten text.
[0145] This approach makes the determination of rewrite operation types more standardized and accurate through clear function entrances and preset mapping relationships, reduces recognition errors caused by inaccurate or unclear user input, and also facilitates system management and maintenance.
[0146] Furthermore, different types of rewriting operations have different requirements for difference detection. Therefore, in scenarios with different types of rewriting operations, the applicable difference detection granularity is also different. For example, in the scenario of overall content rewriting, in order to facilitate users to fully understand the changes, the granularity of difference detection can be flexibly adjusted according to the degree of change. In the scenario of partial rewriting, since users usually expect to observe the rewriting effect at a small granularity level, the rewriting results are displayed at a small granularity (such as word granularity). In the scenario of structural adjustment, since users are more concerned about the overall structure after rewriting, the rewriting results are presented at a larger granularity (such as paragraph granularity).
[0147] Accordingly, in this embodiment, a corresponding difference detection granularity can be pre-set for each rewriting operation type and stored in a preset mapping table. Then, when displaying the original text and its corresponding rewritten text for comparison, the target granularity corresponding to the rewriting operation type from the original text to the rewritten text can be retrieved from the aforementioned mapping table based on the actual rewriting operation type, and then this target granularity can be used as the first granularity for difference detection. This process can significantly improve the accuracy and efficiency of difference detection, ensuring that users can obtain the most suitable difference display effect in different rewriting scenarios.
[0148] Exemplarily, the scenario corresponding to this embodiment is: the user has turned on the function of intelligently setting the difference detection granularity according to the rewriting scenario. In this scenario, the embodiment of the present application can automatically adapt the difference detection granularity that matches it based on the accurate identification of the rewriting operation type. For example, when encountering a local rewrite, the difference is accurately located at a small granularity, allowing the user to focus on the specific details of the text change, providing strong support for in-depth analysis of the rationality and accuracy of the local rewrite. When the text structure is adjusted, the difference is displayed at a larger granularity, allowing the user to quickly grasp the macro-change trend of the text structure, which helps the user to evaluate the rewriting effect as a whole. This intelligent difference detection granularity setting mechanism effectively improves the pertinence and effectiveness of text difference detection, and significantly improves the efficiency and quality of users in handling text rewriting tasks.
[0149] In one embodiment, performing first-granularity difference detection on the original text and its corresponding rewritten text to obtain a first-granularity difference unit set is an exemplary implementation including: decomposing the original text and its corresponding rewritten text into first-granularity units that conform to the first granularity respectively; performing difference detection based on the first-granularity units corresponding to the original text and its corresponding rewritten text to obtain a first-granularity difference unit set.
[0150] Taking the first granularity as word granularity as an example, the process of performing first granularity difference detection on the original text and its corresponding rewritten text to obtain a first granularity difference unit set includes:
[0151] S1. Perform word segmentation on the original text and the rewritten text to obtain the original word segmentation sequence and the rewritten word segmentation sequence. Word segmentation can be performed using existing word segmentation tools (such as Jieba word segmentation and HanLP word segmentation), and the appropriate word segmentation granularity (such as word level or phrase level) can be selected according to specific needs.
[0152] S2. Construct a matrix for comparing word segmentation differences. Create a two-dimensional matrix with the length of the original word segmentation sequence as the row and the length of the rewritten word segmentation sequence as the column. This matrix records the matching of word segmentations at different positions in the two sequences.
[0153] S3. Calculate the similarity or matching between the segmented words. For each segmented word in the original segmented word sequence and each segmented word in the rewritten segmented word sequence, calculate the similarity between them (using methods such as edit distance and cosine similarity), and fill the similarity value into the matrix constructed in step S2.
[0154] S4. Find the optimal matching path. Use a dynamic programming algorithm (e.g., a variant of the longest common subsequence algorithm) to find an optimal path from the upper left corner to the lower right corner of the similarity matrix obtained in step S3. The optimal path conditions are: the path must meet the conditions of maximum sum of similarities or minimum difference and be the shortest path.
[0155] S5. Determine the difference in word segmentation. Based on the optimal matching path found in step S4, compare the unmatched word segmentation in the original word segmentation sequence and the rewritten word segmentation sequence to determine differences such as addition, deletion, and modification. For example, a word segmentation that exists in the original word segmentation sequence but does not exist in the rewritten word segmentation sequence is a deleted word segmentation; a word segmentation that exists in the rewritten word segmentation sequence but does not exist in the original word segmentation sequence is an added word segmentation.
[0156] Here is an additional explanation: in the process of determining word segmentation differences, there is a special case, that is, the original text deletes a word, and the rewritten text inserts the word, but in the end it can be determined that the word is a complete match. This situation usually occurs in the following scenario: Based on the problem of the dynamic programming algorithm used in step S4, the found path may identify a word that is a complete match between the original text and the rewritten text as unmatched, and first delete it and then insert it. However, from the perspective of overall semantics and text coherence, the deletion and insertion of the word did not cause a substantial change to the core meaning of the text.
[0157] For example, the original text reads: "Mom bought apples, bananas, and oranges," and the rewritten text reads: "Mom bought grapes and oranges." The "oranges" in the original text are first matched with the "grapes" in the rewritten text. Both are shown as unmatched, so the difference is displayed: "oranges" in the original text are deleted and "grapes" in the rewritten text are inserted. Then, when the "oranges" in the rewritten text are matched, since all the words in the original text have been deleted and an unmatched state is displayed, the "oranges" in the rewritten text are inserted. Therefore, the "oranges" in the original text are first deleted and then the "oranges" in the rewritten text are inserted. However, since the two words are consistent and the insertion positions are consistent, it can be determined that the "oranges" in the original text and the "oranges" in the rewritten text are actually matched and can be marked as an exact match.
[0158] Therefore, in one embodiment, for an insert operation and a delete operation at the same position, if the unit corresponding to the insert operation and the unit corresponding to the delete operation are the same, the insert operation and the delete operation are combined into an unmodified operation, and the unit is deemed not to be modified.
[0159] S6. Output the word segmentation difference detection results. Output the word segmentation differences determined in step S5 in a clear manner, such as by marking the word segmentation differences between the original text and the rewritten text with different marks (such as "_" to indicate added word segments and "-" to indicate deleted word segments).
[0160] For example:
[0161] Original text: "Mom bought apples, bananas and oranges" (sequence after word segmentation: [A1: Mom, A2: bought, A3: apples, A4: bananas, A5: and, A6: oranges]).
[0162] Rewrite the text: "Mom bought grapes and oranges" (sequence after word segmentation: B = [B1: Mom, B2: Bought, B3: Grapes, B4: Oranges])
[0163] S2: Constructing a word segmentation difference comparison matrix
[0164] The length of the original word segmentation sequence is 6 (rows), and the length of the rewritten word segmentation sequence is 4 (columns). A 6×4 two-dimensional matrix (initial value is 0) is constructed as shown in Table 4 below:
[0165] Table 4
[0166] B1: Mom B2: Bought B3: Grapes B4: Orange A1: Mom 0 0 0 0 A2: I bought it 0 0 0 0 A3: Apple 0 0 0 0 A4: Banana 0 0 0 0 A5: and 0 0 0 0 A6: Orange 0 0 0 0
[0167] S3: Calculate word similarity and fill the matrix
[0168] The similarity is calculated using the edit distance (the minimum number of operations required to convert word a into word b, including insertion, deletion, and substitution). The specific formula is as follows:
[0169] Similarity = 1 - (Edit Distance) / max(len(a), len(b));
[0170] Where, (max(len(a), len(b)) represents the larger value of the lengths of the two word segmentations for normalization), max() is the function to take the maximum value, and len(b) is the function to calculate the length of the word segmentation of b.
[0171] The specific calculation results are as follows:
[0172] A1: Mom vs B1: Mom: Edit Distance = 0; Similarity = 1 - 0 / 2 = 1;
[0173] A2: Bought vs B2: Bought: Edit Distance = 0; Similarity = 1 - 0 / 2 = 1;
[0174] A3: Apple vs B3: Grape: Need to replace: "Apple" with "Grape", "Fruit" with "Peel" (2 replacements); Therefore, Edit Distance = 2; Similarity = 1 - 2 / 2 = 0;
[0175] A3: Apple vs B4: Orange: Need to replace "Apple" with "Orange", "Fruit" with "Seed" (2 replacements); Therefore, Edit Distance = 2; Similarity = 1 - 2 / 2 = 0;
[0176] A4: Banana vs B3: Grape: Need to replace "Banana" with "Grape", "Fruit" with "Peel" (2 replacements); Therefore, Edit Distance = 2; Similarity = 1 - 2 / 2 = 0;
[0177] A4: Banana vs B4: Orange: Need to replace "Banana" with "Orange", "Fruit" with "Seed" (2 replacements); Therefore, Edit Distance = 2; Similarity = 1 - 2 / 2 = 0;
[0178] A5: And vs B3: Grape: Need to delete "And" (1 deletion), insert "Grape", "Peel" (2 insertions); Therefore, Edit Distance = 3; Similarity = 1 - 3 / 2 = -0.5 (negative number indicates complete mismatch, take 0)
[0179] A5: And vs B4: Orange: Need to delete "And" (1 deletion), insert "Orange", "Seed" (2 insertions); Therefore, Edit Distance = 3; Similarity = 1 - 3 / 2 = -0.5 (take 0);
[0180] A6: Orange vs B4: Orange: Edit Distance = 0; Similarity = 1 - 0 / 2 = 1;
[0181] The matrix filled with similarities is as shown in Table 5 below (where only the similarity values at key positions are shown):
[0182] Table 5
[0183] B1: Mom B2: Bought B3: Grapes B4: Orange A1: Mom 1 0 0 0 A2: I bought it 0 1 0 0 A3: Apple 0 0 0 0 A4: Banana 0 0 0 0 A5: and 0 0 0 0 A6: Orange 0 0 0 1
[0184] S4: Find the optimal matching path (minimum difference)
[0185] The goal is to find a path from the upper left corner (A1, B1) to the lower right corner (A6, B4) in the matrix such that the sum of similarities along the path is maximized (meaning the difference is minimized). If multiple paths have the same difference, the priority is as follows: "match exactly the same segmentation first > process deletions / additions later" > process deletions and additions simultaneously.
[0186] Among them, the core logic of the path rule:
[0187] In the word segmentation difference detection algorithm, a path is a moving track from the starting point (upper left corner) to the end point (lower right corner) of the original word segmentation sequence (row) and the rewritten word segmentation sequence (column). Each step of the path corresponds to a matching relationship between the word segments in the two sequences, and its rules must simultaneously meet the following conditions:
[0188] 1. Movement direction restriction: Each step can only choose one of the following three movement methods:
[0189] Move right (column +1, row unchanged): indicates that the rewritten text has an additional participle at the current position (the original text has no corresponding participle);
[0190] Move down (row +1, column unchanged): indicates that a segmentation has been deleted from the original text at the current position (the rewritten text has no corresponding segmentation);
[0191] Move to the lower right (row +1, column +1): Indicates that the original text and the rewritten text try to match the word segmentation at the current position (which may be a complete match, a partial match, or no match).
[0192] It should be noted that:
[0193] 1. Exact match (similarity = 1)
[0194] Condition: The original participle and the rewritten participle are exactly the same (e.g., "mother" vs. "mother", "bought" vs. "bought").
[0195] Action: No changes are required to the original or rewritten text; simply mark it as "No Difference".
[0196] Example:
[0197] Original word segmentation sequence: [Mom, bought, apples, bananas, and, oranges]
[0198] Rewrite the word sequence: [mom, bought, grapes, oranges]
[0199] Among them, "Mom" and "Mom", "bought" and "bought", "orange" and "orange" are completely matched, and no operation is required.
[0200] 2. Partial match (0 < similarity < 1)
[0201] Condition: The original word and the rewritten word have some similarity (such as "today" vs. "yesterday"), but are not completely consistent (edit distance > 0 but not to the extent of complete mismatch).
[0202] Operation: The operation marked as "modify" means rewriting the text to change the original word segmentation to the current word segmentation.
[0203] Example:
[0204] Original word segmentation sequence: [Xiao Ming, today, went to the supermarket, bought apples]
[0205] Rewrite the word sequence: [Xiao Ming went to the fruit store yesterday and bought bananas]
[0206] Replacing "today" with "yesterday" is a partial match and needs to be marked as modified.
[0207] 3. No match (similarity = 0)
[0208] Conditions: The original word and the rewritten word have no obvious correlation (such as "apple" vs. "grape", "banana" vs. "orange"), and the edit distance is equal to or exceeds the maximum possible distance (for example, the two words are the same length but completely different).
[0209] In the event of a mismatch, you can take one of the following actions:
[0210] If the path selection moves to the lower right (i.e., forced matching): it is still marked as a "modify" operation, but the degree of difference is the greatest (such as replacing "apple" with "grape").
[0211] If the path selection does not move to the lower right (i.e., it is processed by moving downward / right first): it may be considered that the original text deletes the participle (moves downward) or the text is rewritten to add the participle (moves right).
[0212] Example:
[0213] Original word segmentation sequence: [Mom, bought, apples, bananas, and, oranges]
[0214] Rewrite the word sequence: [mom, bought, grapes, oranges]
[0215] If the path forcefully matches "apple" and "grape" (moving diagonally down to the right), it is marked as "modified" (replacing apple with grape); if the path chooses to move down (deleting "apple"), then "apple" is marked for deletion, and "grape" is marked for addition (moving to the right).
[0216] Goal: The path needs to maximize the total similarity (or minimize the differences), that is, select the path with the highest similarity to reflect the most reasonable matching relationship between word segmentations.
[0217] Example path analysis:
[0218] 1. Use example path analysis (original sequence A = [mom, bought, apple, banana, and, orange]; rewritten sequence B = [mom, bought, grape, orange])
[0219] 2. Background of path selection
[0220] The length of the original word segmentation sequence is 6 (rows), and the length of the rewritten word segmentation sequence is 4 (columns). Therefore, the end point is the lower right corner of the matrix (A6, B4). The path needs to start from (A1, B1) and finally reach (A6, B4).
[0221] 3. Specific path examples
[0222] Path 1 (sub-optimal path):
[0223] Moving steps: (A1, B1) → diagonally down to the right → (A2, B2) → down → (A3, B2) → down → (A4, B2) → down → (A5, B2) → diagonally down to the right → (A6, B3) → right → (A6, B4) Corresponding operations and similarity calculations:
[0224] (A1, B1): The original word segmentation "mom" and the rewritten word segmentation "mom" are completely matched (moving diagonally down to the right), similarity = 1, indicating no modification.
[0225] (A2, B2): The original word segmentation "bought" and the rewritten word segmentation "bought" are completely matched (moving diagonally down to the right), similarity = 1.
[0226] (A3, B2): There is no rewritten word segmentation matching the original word segmentation "apple" (moving down, deleting "apple"), similarity = 0.
[0227] (A4, B2): There is no rewritten word segmentation matching the original word segmentation "banana" (moving down, deleting "banana"), similarity = 0.
[0228] (A5, B2): There is no rewritten word segmentation matching the original word segmentation "and" (moving down, deleting "and"), similarity = 0.
[0229] (A6,B3): The original segmented word "orange" does not fully match the rewritten segmented word "grape" (move diagonally down to the right, delete "orange", add "grape", that is, modify "orange" to "grape"), similarity = 0.
[0230] (A6,B4): There is no original segmented word that matches the rewritten segmented word "orange" (move to the right), similarity = 1, add "orange".
[0231] Total similarity: 1 + 1 + 0 + 0 + 0 + 0 + 1 = 3.
[0232] Path 2 (optimal path):
[0233] Moving steps: (A1,B1) → diagonally down to the right → (A2,B2) → down → (A3,B2) → down → (A4,B2) → down → (A5,B2) → right → (A5,B3) → diagonally down to the right → (A6,B4). Corresponding operations and similarity calculations:
[0234] (A1,B1): The original segmented word "mom" fully matches the rewritten segmented word "mom" (move diagonally down to the right), similarity = 1, indicating no modification.
[0235] (A2,B2): The original segmented word "bought" fully matches the rewritten segmented word "bought" (move diagonally down to the right), similarity = 1.
[0236] (A3,B2): There is no rewritten segmented word that matches the original segmented word "apple" (move down, delete "apple"), similarity = 0.
[0237] (A4,B2): There is no rewritten segmented word that matches the original segmented word "banana" (move down, delete "banana"), similarity = 0.
[0238] (A5,B2): There is no rewritten segmented word that matches the original segmented word "and" (move down, delete "and"), similarity = 0.
[0239] (A5,B3): There is no original segmented word that matches the rewritten segmented word "grape" (move to the right, add "grape"), similarity = 0.
[0240] (A6,B4): The original segmented word "orange" matches the rewritten segmented word "orange" (move diagonally down to the right), similarity = 1, add "orange".
[0241] Total similarity: 1 + 1 + 0 + 0 + 0 + 0 + 1 = 3, the same as path 2.
[0242] When the total similarities of multiple paths are the same, determine the path that conforms to the priority logic as the optimal path.
[0243] For example, although the above paths 1 and 2 have the same total similarity, path 2 conforms to the priority logic and is more concise overall. Specifically, path 1 attempts to match the irrelevant segmentation words "orange" and "grape" at (A6, B3), resulting in the need to delete "orange" and add "grape". Path 2, on the other hand, directly skips the invalid match. If an invalid match occurs when moving to the lower right, it switches to moving to the right, only adding "grape", avoiding the deletion of "orange" and the subsequent reinsertion of "orange". Therefore, the priority logic can be "match exactly the same segmentation first > process deletion / addition" > process deletion and addition simultaneously.
[0244] Therefore, the final conclusion is: the optimal path is path 2 (the total similarity is the highest, and it conforms to the logic of "matching exactly the same word segmentation first > then processing deletion / addition" > processing deletion and addition at the same time).
[0245] S5: Determine word segmentation differences
[0246] According to the optimal path (path 2), the matching of the original word segmentation sequence and the rewritten word segmentation sequence is compared, as shown in Table 1 above. The final difference summary is:
[0247] Deletion operation: The original text deletes the word segmentation "apple", "banana", and "and".
[0248] Additional operation: The rewritten text adds the participle "grape".
[0249] Exactly matched participles: "mom", "bought", "orange".
[0250] The above is an exemplary implementation process for detecting the difference in word granularity between the original text and its corresponding rewritten text to obtain a first granularity difference unit set, which is not limited in the embodiments of the present application.
[0251] In another embodiment, performing a difference detection of a first granularity on the original text and its corresponding rewritten text to obtain a first granularity difference unit set may include: performing a difference detection of a preset granularity on the original text and its corresponding rewritten text to obtain a preset granularity difference unit set; adjusting the difference units in the preset granularity difference unit set to the first granularity difference unit to form the first granularity difference unit set. The preset granularity may be a granularity smaller than the first granularity. For example, when the first granularity is sentence granularity, the preset granularity may be word granularity; for example, when the first granularity is word granularity, the preset granularity may be character granularity.
[0252] Subsequently, in step 102, in response to the granularity conversion trigger condition, a second granularity is determined. It is noteworthy that the second granularity is different from the first granularity. That is, in the technical solution provided in the embodiment of the present application, once the granularity conversion trigger condition is met, the second granularity to be converted to (e.g., from word granularity to sentence granularity) is determined according to a preset rule.
[0253] In one embodiment, the granularity conversion trigger condition is actively selected by the user. For example, when a user is performing a text comparison and wants to switch from a more detailed word-level granularity display to a sentence-level granularity display to more quickly grasp the overall changes to the text, the user can actively trigger the granularity conversion operation. For another example, when a user is performing a text comparison and wants to switch from a macro sentence-level granularity display to a more detailed word-level granularity display to more deeply grasp the detailed changes to the text, the user can also actively trigger the granularity conversion operation.
[0254] In this embodiment, as an optional implementation, the second granularity is manually selected by the user. For example, the user can select from the granularity options provided by the system according to their own needs, and the system determines the second granularity based on the user's selection.
[0255] In another embodiment, the granularity conversion trigger condition is intelligently triggered by the system based on actual conditions. Specifically, in response to the granularity conversion trigger condition, determining the second granularity in an exemplary implementation includes: determining the type of rewriting operation from the original text to the rewritten text; obtaining a preset target granularity corresponding to the rewriting operation type; if it is determined that the target granularity is inconsistent with the first granularity, determining that the granularity conversion trigger condition is met; and in response to the granularity conversion trigger condition, determining the target granularity to be the second granularity.
[0256] The scenario corresponding to this embodiment is, for example, that the system first performs difference detection at a first granularity (such as a word granularity manually selected by the user, or a default character granularity) and displays the corresponding difference detection results. Subsequently, the user turns on the function of intelligently adjusting the difference detection granularity according to the rewriting scenario. At this time, the system will determine the type of rewriting operation from the original text to the rewritten text, and obtain a preset target granularity corresponding to the rewriting operation type. As for how to obtain the preset target granularity corresponding to the rewriting operation type, please refer to the relevant description above and will not go into details here.
[0257] The acquired target granularity is then compared with the first granularity. If the target granularity is determined to be inconsistent with the first granularity, the granularity conversion trigger condition is determined to be met. In response to this trigger condition, the target granularity is determined to be the second granularity, allowing for subsequent conversion to the difference detection granularity. If the target granularity is determined to be consistent with the first granularity, granularity conversion is determined to be unnecessary, and difference detection and display continue at the current first granularity.
[0258] For example, if the rewriting operation type is determined to be text structure adjustment, and the preset target granularity is paragraph granularity, which is inconsistent with the current first granularity (word granularity), the granularity conversion is triggered, and the paragraph granularity is determined to be the second granularity, so that the difference detection granularity can be subsequently converted from word granularity to paragraph granularity. This conversion allows the difference detection results to be presented in paragraph units again, which helps users to quickly grasp the changes in text structure from a macro perspective, such as the adjustment of paragraph order, the merging or splitting of paragraph content, etc., thereby improving the user's efficiency in grasping the overall rewriting situation. At the same time, it avoids the tediousness and inefficiency of analyzing structural adjustments word by word at the word granularity level.
[0259] If the rewrite operation type is determined to be a local content rewrite, and the preset target granularity is word granularity, which is consistent with the current first granularity (word granularity), there is no need to trigger granularity conversion, and difference detection and display will continue at word granularity. This ensures that users can accurately locate the locally modified vocabulary and meet their needs for analyzing local subtle changes.
[0260] If the rewrite operation type is determined to be a partial content rewrite, and the preset target granularity is word granularity, which is inconsistent with the current first granularity (paragraph granularity), a granularity conversion is triggered, and the word granularity is determined to be the second granularity, so that the difference detection granularity can be subsequently converted from paragraph granularity to word granularity. This conversion allows the difference detection results to be presented in units of words again, allowing users to intuitively see the differences between each word in the original text and the rewritten text, and to deeply analyze the rationality and accuracy of the partial rewrite. At the same time, it avoids ignoring local details due to excessive focus on the overall structure at the paragraph granularity level.
[0261] It can be seen that this mechanism of intelligently adjusting the granularity of difference detection according to the rewriting scenario can flexibly adapt to different text rewriting needs. While ensuring the accuracy of difference detection, it improves the efficiency and pertinence of difference detection, providing users with a more intelligent and efficient text difference detection experience.
[0262] Then, in step 103, the first granularity difference unit set is updated to a second granularity difference unit set corresponding to the second granularity. This process is intended to achieve the conversion of the difference detection granularity.
[0263] In one embodiment, an exemplary implementation of updating the first granularity difference unit set to the second granularity difference unit set corresponding to the second granularity includes: decomposing the original text and its corresponding rewritten text into second granularity units that conform to the second granularity respectively; performing difference detection based on the second granularity units corresponding to the original text and its corresponding rewritten text to obtain the second granularity difference unit set.
[0264] This embodiment is equivalent to re-performing difference detection at the second granularity. In practical applications, the applicable scenario of this embodiment is mainly the case where large granularity is converted to small granularity. For example, when the user initially performs difference detection at the paragraph granularity, but then finds that a more detailed analysis of the vocabulary or phrase changes in the text is required, the process of this embodiment can be triggered. The system will re-decompose the original text and the rewritten text into units of word granularity, and perform difference detection based on these units, thereby generating a set of difference units of word granularity. This processing method enables users to grasp subtle changes in the text more accurately, meeting the needs of scenarios with high requirements for text analysis accuracy.
[0265] Of course, this embodiment is not only applicable to the conversion of large granularity to small granularity, but also has a wide range of applicable scenarios. For example, if the user initially performs difference detection at the sentence granularity, but later finds that the adjustment of paragraph structure is crucial to understanding the text changes, the second granularity can be set to paragraph granularity. The system will then re-decompose the original text and the rewritten text into paragraph units and generate a set of difference units at the paragraph granularity, helping the user to grasp the changes in text structure from a macro perspective.
[0266] This embodiment is applicable to a variety of granularity conversion scenarios. Any user needing to adjust the granularity of difference detection can be achieved through this embodiment. This flexibility enables this embodiment to adapt to the text difference analysis needs of different users and scenarios, providing users with more accurate and efficient text difference detection services.
[0267] In another embodiment, an exemplary implementation of updating the first granularity difference unit set to a second granularity difference unit set corresponding to the second granularity includes: adjusting the first granularity difference unit in the first granularity difference unit set to a second granularity difference unit to form the second granularity difference unit set.
[0268] In this embodiment, rather than directly re-performing difference detection at the second granularity, the first granularity difference unit set already generated is adjusted to obtain the second granularity difference unit, thereby forming the second granularity difference unit set. This approach avoids the tedious process of repeated difference detection and can improve the processing efficiency of difference detection.
[0269] Specifically, adjusting the first granularity difference unit in the first granularity difference unit set to the second granularity difference unit includes: aggregating multiple first granularity difference units in the first granularity difference unit set that meet the aggregation condition corresponding to the second granularity into the second granularity difference unit.
[0270] Similarly, this embodiment is also applicable to various granularity conversion scenarios. As long as the user has the need to adjust the difference detection granularity, it can be achieved through this embodiment. As for the specific implementation of this embodiment, please refer to the detailed embodiment provided below, which will not be described in detail here.
[0271] Finally, in step 104, based on the second granularity difference unit set, a text rewriting comparison view with difference markings at the second granularity is output. This view will clearly show the differences between the original text and the rewritten text at the second granularity level, helping users to more intuitively understand the specific circumstances of the text changes. Figure 2 , which is an example of a text rewrite comparison view with difference marking at different granularities.
[0272] The technical solution provided by the embodiments of the present application performs difference detection at a first granularity between the original text and its corresponding rewritten text to obtain a first granularity difference unit set; determines a second granularity in response to a granularity conversion trigger condition; updates the first granularity difference unit set to a second granularity difference unit set corresponding to the second granularity; and, based on the second granularity difference unit set, outputs a text rewrite comparison view with difference markings at the second granularity. This allows for flexible adjustment of the granularity of text difference display, meeting the diverse needs of different users for text difference analysis in different scenarios. Whether a macroscopic grasp of overall differences or a microscopic analysis of local details is required, these can be easily achieved by adjusting the difference detection granularity, significantly improving the efficiency and accuracy of text rewrite analysis.
[0273] In one embodiment, multiple first granularity difference units in the first granularity difference unit set that meet the aggregation conditions corresponding to the second granularity are aggregated into a second granularity difference unit, including: dividing the first granularity difference units in the first granularity difference unit set into intervals based on the second granularity, wherein the text combination corresponding to the first granularity difference units contained in an interval conforms to the semantic boundary of the second granularity; merging at least one first granularity difference unit in the same interval into a second granularity difference unit, and assigning an operation type identifier to the second granularity difference unit.
[0274] In this embodiment, in order to adapt to the display requirements of different granularities, the first granularity difference unit set is reprocessed to obtain the second granularity difference unit set. Specifically, first, based on the second granularity, the first granularity difference unit in the first granularity difference unit set is divided into intervals. The basis for division is to ensure that the text combination corresponding to the first granularity difference unit contained in an interval conforms to the semantic boundary of the second granularity, which means that the interval should be divided according to the range of semantic units represented by the second granularity (such as words, sentences, paragraphs, etc.). When performing interval division, a mapping relationship from the first granularity to the second granularity is actually constructed. The specific method can be to traverse the first granularity difference unit set and divide the first granularity difference units whose texts are continuous and semantically related into the same interval according to the semantic boundary of the second granularity. For example, if the second granularity is the word level, then the character-level difference units whose texts can constitute a complete word will be divided into the same interval here. Afterwards, at least one first granularity difference unit in the same interval is merged into a second granularity difference unit, and an operation type identifier is assigned to the newly generated second granularity difference unit so that the corresponding rewrite operation can be clearly identified later.
[0275] The core of the above process lies in the aggregation operation. It should be noted that the aggregation here is not a simple splicing from small-granularity units to large-granularity units, but covers more complex logic, which may involve the re-splitting of parts that were originally considered unmodified, and flexible integration of the split parts. Specifically, when adjusting the granularity, the text parts that were originally judged to be unmodified at a small granularity may need to be re-examined from a large-granularity perspective. The system will first split these unmodified parts according to the semantic boundaries and display requirements of the large granularity, and then integrate the split parts with other related parts based on the range of semantic units represented by the second granularity, thereby forming a difference unit that meets the requirements of the large granularity.
[0276] For example, the original text is "Yesterday it rained," and the rewritten text is "Tomorrow it rains." When performing difference detection at the character level (small granularity), the first-granularity difference unit set is {delete(yesterday), insert(tomorrow), equal(tianxiayu)}. Here, delete(yesterday) indicates that the character "zuo" in the original text has been deleted, insert(tomorrow) indicates that the character "tomorrow" has been inserted into the rewritten text, and equal(tianxiayu) indicates that the part "tianxiayu" ("tianxiayu") is unchanged in both the original and rewritten texts.
[0277] When adjusting the difference detection granularity to the word level (coarse granularity), corresponding processing is required. First is the conversion from the character level to the word level, which involves re-splitting the originally regarded unmodified part "天下雨". According to the semantic boundaries of words, "天下雨" should be split into two parts: "天" and "下雨". Accordingly, the first granularity difference unit equal(天下雨) can be split into two first granularity difference units: equal(天), equal(下雨).
[0278] Then, integrate the split first granularity difference units with other first granularity difference units that are not involved in splitting (such as delete(昨), insert(明)). Among them, from the word level, "昨" and "天" together form the complete word "昨天". Therefore, divide delete(昨) and equal(天) into one interval according to the semantic boundaries of words. And because "昨" is deleted, the operation type assigned to the second granularity difference unit formed by this interval is "delete", obtaining the second granularity difference unit delete(昨天).
[0279] Similarly, "明" and "天" together form the complete word "明天". Therefore, divide insert(明) and equal(天) into one interval according to the semantic boundaries of words. And because "明" is inserted, the operation type assigned to the second granularity difference unit formed by this interval is "insert", obtaining the second granularity difference unit insert(明天).
[0280] For equal(下雨), it is not modified at both the character level and the word level. Therefore, in the set of difference units at the word level, it is retained as an independent and unchanged interval.
[0281] Based on the above integration logic, the final set of second granularity difference units is {delete(昨天), insert(明天), equal(下雨)}.
[0282] From this example, it can be seen that when adjusting the difference detection granularity, through splitting and integration operations, the set of second granularity difference units can more accurately display the text differences at different granularities, meeting the needs of users to analyze text rewriting situations at different granularities.
[0283] Figure 3 It is a flowchart of an embodiment of another content rewriting method provided for the embodiments of this application. Figure 3 The shown process is Figure 1 based on the shown process and includes the following steps:
[0284] Step 301, determine the rewriting operation type from the original text to the rewritten text.
[0285] Step 302: Obtain a preset target granularity corresponding to the rewrite operation type; and use the target granularity as the first granularity.
[0286] Figure 3 The embodiment shown is in Figure 1 Based on the illustrated embodiment, the focus is on implementation in a specific scenario, that is, the function of intelligently setting the difference detection granularity based on the rewriting scenario is initially enabled. This function can be enabled by the user or by default. The embodiments of this application do not limit this.
[0287] As for the specific implementation of step 301 and step 302, please refer to the relevant description in the above embodiment, which will not be repeated here.
[0288] Step 303: Perform first-granularity difference detection on the original text and its corresponding rewritten text to obtain a first-granularity difference unit set.
[0289] Step 304: Based on the first granularity difference unit set, output a text rewriting comparison view with difference markings at the first granularity.
[0290] Step 305: In response to the granularity conversion triggering condition, determine the second granularity.
[0291] Step 306: Update the first granularity difference unit set to a second granularity difference unit set corresponding to the second granularity.
[0292] Step 307: Based on the second granularity difference unit set, output a text rewriting comparison view with difference markings at the second granularity.
[0293] It can be seen from the description of step 303 to step 307 that Figure 3 In the application scenario focused on in the illustrated embodiment, the original and rewritten texts are first displayed with differences marked at a first granularity, allowing users to gain a preliminary understanding of the text changes from the perspective of the first granularity. Subsequently, if the granularity conversion trigger condition is met, the system re-determines the second granularity and re-marks the differences at the second granularity, providing users with difference analysis results at different granularity levels to meet their diverse analysis needs.
[0294] In one embodiment, the granularity conversion trigger condition is manually triggered by the user. Specifically, the system interface can provide the user with a function entry for manually switching the difference detection granularity, such as setting a drop-down menu or button, so that the user can select different granularity levels according to their needs, such as switching from paragraph granularity to word granularity, or from word granularity to character granularity. When the user makes a selection, it is considered that the user has manually triggered the granularity conversion trigger condition.
[0295] This method allows users to flexibly select the appropriate granularity for text difference analysis based on their own needs, thereby improving the user experience.
[0296] In another embodiment, the granularity conversion trigger condition is automatically triggered. As an optional implementation method, a first difference mark density index for performing difference markup on the original text and the rewritten text at a first granularity is determined; if it is determined that the first difference mark density index meets a preset condition, the granularity conversion trigger condition is determined to be met; if it is determined that the first difference mark density index does not meet the preset condition, the granularity conversion trigger condition is determined to be not met.
[0297] This automatic triggering method intelligently adjusts the granularity of difference detection based on the actual differences in the text, eliminating the need for manual user intervention and improving the system's intelligence and automation. It also ensures that users always receive the most appropriate difference display, helping them conduct text rewriting analysis more efficiently.
[0298] The first difference marker density indicator reflects the degree of difference between the original text and the rewritten text at the first granularity. This indicator quantifies the density of differences, providing an appropriate analytical perspective for determining whether the difference display at the current granularity meets user needs and whether a granularity shift should be triggered.
[0299] Specifically, determining a first difference marking density index for marking the difference between the original text and the rewritten text at a first granularity includes: determining a first granularity difference unit with at least one difference point in the first granularity difference unit set as a target difference unit. For example, in the word granularity difference unit set, if a word is different between the original text and the rewritten text (such as being deleted or inserted), the difference unit where the word is located is the target difference unit. Then, determining the proportion of the target difference unit in the first granularity difference unit set, and determining the proportion as the first difference marking density index for marking the difference between the original text and the rewritten text at the first granularity.
[0300] Among them, when determining the proportion of the target difference unit in the first granularity difference unit set, as an optional implementation method, the total number of characters contained in the target difference unit and the total number of characters contained in all difference units in the first granularity difference unit set are counted respectively. Finally, the ratio of the total number of characters of the target difference unit to the total number of characters of all difference units is calculated, and the ratio is determined as the first difference marking density index for difference marking the original text and the rewritten text at the first granularity. Here, referring to the example of Table 1 above, all difference units in the first granularity difference unit set refer to: text units (such as word granularity units) at all corresponding positions in the original text and the rewritten text, the matching status between the two (complete match / no match), and the difference type (deletion / addition / no difference) determined by the status. For example, the first granularity difference unit set corresponding to Table 1 above includes 7 difference units, including 3 target difference units (respectively, the difference units shown in the 1st row, the 2nd row and the 6th row).
[0301] As another optional implementation, the number of target difference units and the number of all difference units in the first-granularity difference unit set are counted. Finally, the ratio of the number of target difference units to the number of all difference units is calculated, and this ratio is determined as a first difference marking density indicator for difference marking the original text and the rewritten text at the first granularity.
[0302] The preset conditions can be set according to the actual application scenario and user needs, and can be expressed in the form of thresholds. The following explains different situations:
[0303] Case 1: Difference Marker Density is Too High. When the Difference Marker Density exceeds a certain threshold (set as the first threshold), the preset condition is met and this indicates that the difference markers at the current granularity are too dense, making it difficult for users to quickly grasp key changes and significantly affecting their reading experience. At this point, a granularity switch is automatically triggered, adjusting the difference detection granularity to a more macroscopic one, allowing users to grasp text differences at a macro level.
[0304] For example, assuming that the original and rewritten texts are differentially marked at the word granularity (first granularity), and the calculated first differential marking density index is 0.6 (i.e., the number of differentially marked words accounts for 60% of the total number of words). If the preset condition is that the differential marking density index is greater than 0.5 to trigger granularity conversion, then the granularity conversion trigger condition will be determined to be met, and the second granularity (such as sentence granularity) will be re-determined, and the differential marking will be displayed again at the sentence granularity.
[0305] Case 2: Difference Marker Density is Too Low. When the Difference Marker Density falls below a certain threshold (set as the second threshold), the pre-set condition is met and indicates that the difference markers at the current granularity are too sparse and may not fully reflect the text changes. In this case, a granularity conversion is automatically triggered, adjusting the difference detection granularity to a finer granularity to provide a more detailed difference view.
[0306] For example, suppose that after the original and rewritten texts are differentially marked at the paragraph granularity (first granularity), the calculated first differential mark density index is 0.1 (i.e., the number of differentially marked paragraphs accounts for 10% of the total number of paragraphs). If the preset condition is that the granularity conversion is triggered when the differential mark density index is less than 0.5, then the granularity conversion trigger condition will be determined to be met, and the second granularity (such as sentence granularity) will be re-determined, and the differential marks will be displayed again at the sentence granularity.
[0307] Case 3: Set two thresholds for finer control. Two thresholds can be set: a first threshold and a second threshold. A difference marker density index above the first threshold or below the second threshold indicates that the preset conditions are met, and different granularity conversion strategies are adopted accordingly. If the difference marker density index falls between the two thresholds, the preset conditions are not met. In this case, the difference marker display at the current granularity can be retained without granularity conversion.
[0308] Figure 3 The process shown in Figure 1 Based on the process shown, the system initially enables intelligent setting of the difference detection granularity based on the rewriting scenario. This automatically adapts the appropriate granularity to different rewriting scenarios, meeting user needs and reducing the need for manual setup. Subsequently, it also supports dynamic conversion of the difference detection granularity, intelligently adjusting the difference detection granularity based on the actual differences in the text without manual user intervention, thus improving the system's intelligence and automation. At the same time, it ensures that users always receive the most appropriate difference display effect, helping them conduct text rewriting analysis more efficiently.
[0309] Figure 4 This is a flowchart of another embodiment of a content rewriting method provided in an embodiment of the present application. Figure 4 The process shown in Figure 1 Based on the process shown, the following steps are included:
[0310] Step 401: Perform first-granularity difference detection on the original text and its corresponding rewritten text to obtain a first-granularity difference unit set.
[0311] Step 402: determine the type of rewriting operation from the original text to the rewritten text; obtain a preset target granularity corresponding to the rewriting operation type; if it is determined that the target granularity is inconsistent with the first granularity, execute step 403; if it is determined that the target granularity is consistent with the first granularity, execute step 405.
[0312] Step 403: Determine whether a granularity conversion triggering condition is satisfied; in response to the granularity conversion triggering condition, determine the target granularity as the second granularity.
[0313] Step 404: Update the first granularity difference unit set to a second granularity difference unit set corresponding to the second granularity; and output a text rewriting comparison view with difference markings at the second granularity based on the second granularity difference unit set.
[0314] Step 405: Determine that the granularity conversion trigger condition is not satisfied; based on the first granularity difference unit set, output a text rewriting comparison view with difference markings at the first granularity.
[0315] Figure 4 The embodiment shown is in Figure 1 On the basis of the illustrated embodiment, the focus is on the implementation in another specific scenario, that is, firstly perform an initial first-granularity difference detection on the original text and the rewritten text. The first granularity can be actively selected by the user or by the system default. Then, the function of intelligently adjusting the difference detection granularity according to the rewriting scenario is turned on. This function can be actively turned on by the user to meet the immediate need for granularity adjustment in a specific analysis stage; it can also be turned on by the system at a scheduled time, for example, it is automatically triggered when it is detected that the rewritten text reaches a certain length or complexity to ensure the accuracy of the difference analysis. The embodiments of the present application do not limit this.
[0316] Figure 4 The illustrated embodiment allows for dynamic adjustment of the difference detection granularity based on the actual rewriting operation type, ensuring that the difference display always matches the rewriting scenario. When the target granularity is inconsistent with the first granularity, it can quickly respond and complete the granularity conversion, promptly outputting a difference markup view that better meets the current analysis needs; when the two are consistent, the differences are directly displayed at the first granularity, avoiding unnecessary conversion operations. This intelligent granularity adjustment mechanism not only improves the efficiency and accuracy of text rewriting analysis, but also provides users with a more convenient and personalized user experience, helping users complete text rewriting and analysis tasks more efficiently.
[0317] As for Figure 4 For a detailed description of each step in the process shown, please refer to the relevant description in the above embodiment, which will not be repeated here.
[0318] Figure 5 This is a flowchart of another embodiment of a content rewriting method provided in an embodiment of the present application. Figure 5 The process shown in Figure 1 Based on the process shown, the following steps are included:
[0319] Step 501: Perform first-granularity difference detection on the original text and its corresponding rewritten text to obtain a first-granularity difference unit set.
[0320] Step 502: Based on the first granularity difference unit set, determine the first difference mark density index for difference marking the original text and the rewritten text at the first granularity; if it is determined that the first difference mark density index meets the preset conditions, execute step 503; if it is determined that the first difference mark density index does not meet the preset conditions, execute step 508.
[0321] Step 503: Determine whether a granularity conversion triggering condition is satisfied; and determine a second granularity in response to the granularity conversion triggering condition.
[0322] Step 504: Update the first granularity difference unit set to a second granularity difference unit set corresponding to the second granularity.
[0323] Step 505: Based on the current second-granularity difference unit set, determine the second difference mark density index for difference marking the original text and the rewritten text at the current second granularity; if it is determined that the current second difference mark density index meets the preset conditions, execute step 506; if it is determined that the current second difference mark density index does not meet the preset conditions, or the preset level granularity is determined as the second granularity, execute step 507.
[0324] Step 506 : Determine again whether the granularity conversion triggering condition is satisfied; in response to the re-determined granularity conversion triggering condition, determine the second granularity again; and return to step 504 .
[0325] Step 507: Based on the second granularity difference unit set, output a text rewriting comparison view with difference markings at the second granularity.
[0326] Step 508: Based on the first granularity difference unit set, output a text rewriting comparison view with difference markings at the first granularity.
[0327] Figure 5 The embodiment shown is in Figure 1On the basis of the illustrated embodiment, the focus is on the implementation in another specific scenario, that is, first perform an initial first-granularity difference detection on the original text and the rewritten text. The first granularity can be actively selected by the user or by the system default. Then, based on the density of the difference marking at the first granularity, that is, whether the first difference marking density index meets the preset conditions, it is intelligently determined whether granularity conversion is required. If the conditions are met, a second granularity different from the first granularity is selected, and the operations of updating the difference unit set, re-judging the difference marking density index at the new granularity, etc. are completed in sequence until the granularity that meets the preset conditions is found, and finally the text rewriting comparison view of the difference marking at the granularity is output; if the conditions are not met, the difference marking comparison view at the first granularity is directly output.
[0328] This method of dynamically adjusting the granularity based on the density of difference markers ensures that users always obtain the most appropriate difference display effect, greatly improving the efficiency and accuracy of text rewriting analysis.
[0329] In step 503, a granularity at a level adjacent to the first granularity is selected from the preset granularity level sequence and determined as the second granularity. Similarly, in step 506, a granularity at a level adjacent to the current second granularity is selected from the preset granularity level sequence and determined as the second granularity again.
[0330] The preset granularity hierarchy sequence includes multiple different levels of granularity, and these granularities are arranged in a hierarchical order. The arrangement can be from low to high, such as word granularity → sentence granularity → paragraph granularity, or from high to low, such as paragraph granularity → sentence granularity → word granularity.
[0331] When determining the second granularity for the first time and again, the selection targets are the adjacent levels of the current granularity. This design is intended to avoid repeated adjustments that may be caused by cross-level selection, ensuring a more efficient and stable granularity conversion process.
[0332] As an optional implementation method, in the selection of adjacent level granularity, whether to select the adjacent high level or the adjacent low level depends on the comparison result of the difference mark density index with a specific threshold.
[0333] Taking the setting of the first threshold and the second threshold as an example, according to the above rules, when the difference mark density index is higher than the first threshold or lower than the second threshold, the preset conditions are met. However, in different situations, the granularity conversion strategies triggered are different. Specifically, when the difference mark density index is higher than the first threshold, it means that the difference marks at the current granularity are too dense, and it is difficult for users to quickly grasp the key changes. At this time, the adjacent high-level granularity is selected as the second granularity so that users can grasp the text differences from a more macro level; when the difference mark density index is lower than the second threshold, it indicates that the difference marks at the current granularity are too sparse and may not be able to fully display the changes in the text. At this time, the adjacent low-level granularity is selected as the second granularity to provide a more detailed difference view.
[0334] In addition, the preset level granularity mentioned in the above step 505 may refer to the highest or lowest level granularity in the granularity level sequence.
[0335] Figure 5 In the illustrated embodiment, when the first difference mark density index meets the preset conditions, that is, when it indicates that the difference mark display at the first granularity is not ideal, it can respond quickly, select the appropriate adjacent level granularity from the preset granularity hierarchy sequence for conversion, and continuously monitor the difference mark density index at the new granularity. If the conditions are still not met, continue to adjust until the most appropriate granularity is found, and timely output the difference mark view that best reflects the current text rewriting situation; and when the first difference mark density index does not meet the preset conditions, the difference is directly displayed at the first granularity, avoiding unnecessary granularity conversion and saving system resources and time. This intelligent, multi-level granularity adjustment mechanism based on the difference mark density index greatly improves the accuracy and flexibility of text rewriting analysis. It can automatically select the most appropriate granularity for display based on the actual situation of the text difference, thereby providing users with a more accurate, convenient and personalized user experience.
[0336] Figure 6 This is a flowchart of another embodiment of a content rewriting method provided in an embodiment of the present application. Figure 6 The process shown in Figure 1 Based on the process shown, the following steps are included:
[0337] Step 601: Determine the type of rewriting operation from the original text to the rewritten text. If the rewriting operation type is the first type, execute step 602; if the rewriting operation type is other types, execute step 603.
[0338] Step 602 : Select a preset level granularity from a preset granularity level sequence and determine it as the first granularity; and execute step 604 .
[0339] Step 603: Obtain a preset target granularity corresponding to the rewrite operation type, and use the target granularity as the first granularity.
[0340] Step 604: Perform first-granularity difference detection on the original text and its corresponding rewritten text to obtain a first-granularity difference unit set; if the rewriting operation type is the first type, execute step 605; if the rewriting operation type is other types, execute step 610.
[0341] Step 605: Based on the first granularity difference unit set, determine the first difference mark density index for difference marking the original text and the rewritten text at the first granularity; if it is determined that the first difference mark density index meets the preset conditions, execute step 606; if it is determined that the first difference mark density index does not meet the preset conditions, execute step 610.
[0342] Step 606: Determine whether a granularity conversion triggering condition is satisfied; in response to the granularity conversion triggering condition, select a granularity at a level adjacent to the current granularity from a preset granularity level sequence and determine it as a second granularity.
[0343] Step 607: Update the first granularity difference unit set to a second granularity difference unit set corresponding to the second granularity.
[0344] Step 608: Based on the current second-granularity difference unit set, determine the second difference mark density index for difference marking the original text and the rewritten text at the current second granularity; if it is determined that the current second difference mark density index meets the preset conditions, return to execute step 606; if it is determined that the current second difference mark density index does not meet the preset conditions, or the preset level granularity is determined as the second granularity, execute step 609.
[0345] Step 609: Based on the second granularity difference unit set, output a text rewriting comparison view with difference markings at the second granularity.
[0346] Step 610: Based on the first granularity difference unit set, output a text rewriting comparison view with difference markings at the first granularity.
[0347] Figure 6 The embodiment shown is in Figure 1 Based on the embodiment shown, we focus on the implementation in another specific scenario, that is, the function of intelligently setting the difference detection granularity according to the rewriting scenario is initially enabled. This function can be enabled by the user actively or by the system default. This embodiment of the application does not limit this. Figure 3 The difference between the illustrated embodiments is that, depending on the type of rewriting operation, the subsequent process will lead to different processing paths to adapt to diverse rewriting scenario requirements.
[0348] Specifically, referring to the description of steps 601, 603, 604, and 610, when the rewrite operation type is other than the first type, a preset target granularity corresponding to the rewrite operation type is obtained, and the target granularity is used as the first granularity. In other words, for types other than the first type, a clear corresponding difference detection granularity is pre-set, and difference detection is then performed based on the clear first granularity. This differentiated granularity setting strategy fully considers the characteristics of different rewrite scenarios and can ensure that text changes are captured more accurately in the subsequent difference detection process.
[0349] Referring to the description of step 601, step 602, and step 604 to step 609, when the rewrite operation type is the first type, first, from the preset granularity level sequence, the preset level granularity is selected to be determined as the first granularity, and the first granularity is used as the initial difference detection granularity for processing the first type of rewrite operation. Subsequently, based on the density of difference marking at the first granularity, that is, whether the first difference marking density index meets the preset conditions, it is intelligently determined whether granularity conversion is required. If the conditions are met, a second granularity different from the first granularity is selected, and the operations of updating the difference unit set, re-judging the difference marking density index at the new granularity, etc. are completed in sequence until the granularity that meets the preset conditions is found, and finally the text rewrite comparison view of the difference marking at the granularity is output; if the conditions are not met, the difference marking comparison view at the first granularity is directly output. This intelligent, multi-level granularity adjustment mechanism based on rewriting operation type and difference mark density index greatly improves the accuracy and flexibility of text rewriting analysis. It can automatically select the most appropriate granularity for display based on the actual situation of text differences, providing users with a more accurate, convenient and personalized user experience.
[0350] Referring to the description of step 606, when the granularity conversion trigger condition is met, a granularity at a level adjacent to the current granularity is selected from the preset granularity hierarchy sequence and determined as the second granularity. Initially, the current granularity is the preset level granularity selected from the preset granularity hierarchy sequence. During subsequent updates, the current granularity is the currently determined second granularity. The adjacent level granularity selection rule here is intended to avoid repeated adjustments that may be caused by cross-level selection, ensuring a more efficient and stable granularity conversion process.
[0351] For example, the first type is overall content rewriting, and other types include partial content rewriting and text structure adjustment.
[0352] In this example, according to Figure 6In the process shown, if the rewrite operation type is determined to be text structure adjustment, the paragraph granularity is selected as the first granularity from the preset granularity hierarchy sequence, and difference detection is performed at the paragraph granularity. The difference detection results are presented in paragraph units, which helps users quickly grasp the changes in text structure from a macro perspective, such as the adjustment of paragraph order, the merging or splitting of paragraph content, etc., thereby improving the user's grasp of the overall rewrite situation. At the same time, it avoids the tedious and inefficient analysis of structural adjustments at the word granularity level.
[0353] If the rewrite operation type is determined to be a partial content rewrite, word granularity is selected as the first granularity from the preset granularity hierarchy sequence, and difference detection is performed at the word granularity. The difference detection results are presented in word units. This ensures that users can accurately locate the locally modified words and meet their needs for analyzing local subtle changes.
[0354] If the rewrite operation type is determined to be a full content rewrite, the lowest level of granularity is selected from the granularity hierarchy sequence, such as word granularity as the initial first granularity, and then difference detection is performed at word granularity. Afterwards, assuming that the original text and the rewritten text are differentially marked at word granularity, the calculated first difference mark density index is 0.6. If the preset condition is that the granularity conversion is triggered when the difference mark density index is greater than 0.5, then the granularity conversion trigger condition will be determined to be met at this time, and the second granularity (such as sentence granularity) will be re-determined, and the difference marks will be displayed again at sentence granularity.
[0355] If the original and rewritten texts are differentially marked at the sentence granularity (the current second granularity), the calculated second differential mark density index is 0.7. If the preset condition is that the differential mark density index is greater than 0.5 to trigger granularity conversion, the granularity conversion trigger condition is determined to be met, and the second granularity (such as paragraph granularity) is re-determined, and the differential marks are displayed again at the paragraph granularity.
[0356] This method of gradually transitioning from small granularity to large granularity can dynamically adjust the granularity of difference display according to the actual situation of difference marking, ensuring that users can obtain clear and accurate difference information at different granularities, thereby gaining a more comprehensive and in-depth understanding of the content and extent of text rewriting.
[0357] Figure 6The process shown, by enabling the function of intelligently setting the difference detection granularity according to the rewriting scenario at the initial stage, can flexibly adapt to different needs, whether it is actively enabled by the user or enabled by the system by default. Different processing paths are directed according to the type of rewriting operation, fully considering the characteristics of diverse rewriting scenarios. Among them, for types other than the first type, the corresponding target granularity is pre-set to ensure that the difference detection accurately captures the text changes; for the first type, the initial granularity is first determined, and then the granularity conversion is intelligently determined based on the difference mark density index. This intelligent, multi-level granularity adjustment mechanism greatly improves the accuracy and flexibility of text rewriting analysis, and can automatically select the most appropriate granularity for display, providing users with an accurate, convenient and personalized user experience. At the same time, the adjacent level granularity selection rules avoid the tedious cross-level adjustments, making the granularity conversion process efficient and stable.
[0358] Figure 7 This is a block diagram of an embodiment of a text rewriting comparison device provided in an embodiment of the present application. Figure 7 As shown, the device includes:
[0359] A first difference detection module 71 is configured to perform a first-granularity difference detection on the original text and its corresponding rewritten text to obtain a first-granularity difference unit set;
[0360] a granularity determination module 72 for determining a second granularity in response to a granularity conversion trigger condition;
[0361] A granularity conversion module 73 is configured to update the first granularity difference unit set to a second granularity difference unit set corresponding to the second granularity;
[0362] The rewriting comparison module 74 is configured to output a text rewriting comparison view with difference markings performed at the second granularity based on the second granularity difference unit set.
[0363] In a possible implementation, the granularity conversion module 73 is specifically configured to:
[0364] The first granularity difference unit in the first granularity difference unit set is adjusted to a second granularity difference unit to form a second granularity difference unit set.
[0365] In a possible implementation, the granularity conversion module 73 is specifically configured to:
[0366] A plurality of first granularity difference units in the first granularity difference unit set that meet the aggregation condition corresponding to the second granularity are aggregated into a second granularity difference unit.
[0367] In a possible implementation, the granularity conversion module 73 includes:
[0368] a semantic segmentation unit configured to segment the first granularity difference units in the first granularity difference unit set into intervals based on the second granularity, wherein a text combination corresponding to the first granularity difference units contained in an interval conforms to a semantic boundary of the second granularity;
[0369] The merging unit is configured to merge at least one first granularity difference unit in the same interval into a second granularity difference unit, and assign an operation type identifier to the second granularity difference unit.
[0370] In a possible implementation, the granularity determination module 72 includes:
[0371] a scene recognition unit, configured to determine a type of rewriting operation from the original text to the rewritten text;
[0372] A first determining unit, configured to obtain a preset target granularity corresponding to the rewrite operation type;
[0373] a first condition determination unit, configured to determine that a granularity conversion trigger condition is satisfied when it is determined that the target granularity is inconsistent with the first granularity;
[0374] The second determining unit is configured to determine the target granularity as a second granularity in response to the granularity conversion triggering condition.
[0375] In a possible implementation, the device further includes:
[0376] The first granularity determination module is used to determine the type of rewriting operation from the original text to the rewritten text before performing first granularity difference detection on the original text and its corresponding rewritten text to obtain a first granularity difference unit set; obtain a preset target granularity corresponding to the rewriting operation type; and use the target granularity as the first granularity.
[0377] In a possible implementation, the scene recognition unit is specifically configured to:
[0378] performing intent recognition on a natural language processing instruction input by a user, and determining a type of rewriting operation from the original text to the rewritten text based on the intent recognition result, wherein the natural language processing instruction is used to instruct to rewrite the original text;
[0379] Alternatively, obtain a function entry identifier that triggers the rewriting of the original text; search a preset operation type mapping table based on the function entry identifier to obtain the rewriting operation type from the original text to the rewritten text, wherein the operation type mapping table includes a correspondence between the function entry identifier and the rewriting operation type.
[0380] In a possible implementation, the granularity determination module 72 includes:
[0381] a mark density prediction unit, configured to determine, based on the first granularity difference unit set, a first difference mark density index for marking differences between the original text and the rewritten text at the first granularity;
[0382] a second condition determination unit, configured to determine that a granularity conversion trigger condition is satisfied when it is determined that the first difference mark density indicator satisfies a preset condition;
[0383] The third determining unit is configured to determine a second granularity in response to the granularity conversion triggering condition.
[0384] In one possible implementation, the mark density prediction unit is further configured to: before outputting the text rewriting comparison view with difference marks at the second granularity based on the second granularity difference unit set, determine, based on the current second granularity difference unit set, a second difference mark density index for difference marks between the original text and the rewritten text at the current second granularity;
[0385] The second condition determination unit is further configured to: upon determining that the current second difference mark density indicator satisfies the preset condition, determine again that the granularity conversion trigger condition is satisfied;
[0386] The third determining unit is further configured to re-determine the second granularity in response to the re-determined granularity conversion triggering condition, until the current second difference mark density indicator does not meet the preset condition, or the preset level granularity is determined as the second granularity.
[0387] In a possible implementation, the third determining unit determines the second granularity, including:
[0388] From a preset granularity level sequence, a preset level granularity is selected and determined as the second granularity.
[0389] In a possible implementation manner, the third determining unit determining the second granularity again includes:
[0390] From the preset granularity level sequence, a granularity at a level adjacent to the current second granularity is selected and determined as the second granularity again.
[0391] In a possible implementation, the marker density prediction unit is specifically configured to:
[0392] Determine a first granularity difference unit having at least one difference point in the first granularity difference unit set as a target difference unit;
[0393] The proportion of the target difference unit in the first granularity difference unit set is determined, and the proportion is determined as a first difference marking density index for performing difference marking on the original text and the rewritten text at the first granularity.
[0394] In a possible implementation, the first difference detection module includes:
[0395] A first granularity decomposition unit, configured to decompose the original text and its corresponding rewritten text into first granularity units conforming to a first granularity;
[0396] The first difference detection unit is configured to perform difference detection based on first granularity units corresponding to the original text and the corresponding rewritten text, to obtain a first granularity difference unit set.
[0397] In a possible implementation, the granularity conversion module 73 includes:
[0398] A second granularity decomposition unit decomposes the original text and the corresponding rewritten text into second granularity units conforming to a second granularity;
[0399] The second difference detection unit is configured to perform difference detection based on the second granularity units corresponding to the original text and the corresponding rewritten text, to obtain a second granularity difference unit set.
[0400] like Figure 8 As shown, an embodiment of the present application provides an electronic device, including a processor 111, a communication interface 112, a memory 113 and a communication bus 114, wherein the processor 111, the communication interface 112, and the memory 113 communicate with each other through the communication bus 114.
[0401] Memory 113, for storing computer programs;
[0402] In one embodiment of the present application, the processor 111 is configured to implement the text rewriting comparison method provided by any one of the aforementioned method embodiments when executing a program stored in the memory 113, including:
[0403] Performing first-granularity difference detection on the original text and the corresponding rewritten text to obtain a first-granularity difference unit set;
[0404] In response to a granularity conversion trigger condition, determining a second granularity;
[0405] Updating the first granularity difference unit set to a second granularity difference unit set corresponding to the second granularity;
[0406] Based on the second granularity difference unit set, a text rewriting comparison view with difference markings performed at the second granularity is output.
[0407] An embodiment of the present application further provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of the text rewriting comparison method provided in any of the aforementioned method embodiments are implemented.
[0408] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of this embodiment.
[0409] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a general hardware platform, or of course, by hardware. Based on this understanding, the above technical solution, in essence, or the part that contributes to the relevant technology, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or certain parts of the embodiment.
[0410] It should be understood that the terms used herein are for the purpose of describing specific example embodiments only and are not intended to be limiting. Unless the context clearly indicates otherwise, the singular forms "one", "an" and "said" as used herein may also be meant to include plural forms. The terms "comprise", "include", "contain" and "have" are inclusive and therefore specify the presence of stated features, steps, operations, elements and / or parts, but do not exclude the presence or addition of one or more other features, steps, operations, elements, parts, and / or combinations thereof. The method steps, processes, and operations described herein are not to be construed as necessarily requiring them to be performed in the specific order described or illustrated, unless the order of execution is clearly indicated. It should also be understood that additional or alternative steps may be used.
[0411] The foregoing is merely a list of specific embodiments of the present application, intended to enable those skilled in the art to understand or implement the present application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application is not limited to the embodiments shown herein, but is intended to conform to the broadest scope consistent with the principles and novel features of the present application.
Claims
1. A text rewriting comparison method, characterized in that: The method comprises: Performing first-granularity difference detection on the original text and the corresponding rewritten text to obtain a first-granularity difference unit set; In response to a granularity conversion trigger condition, determining a second granularity; Updating the first granularity difference unit set to a second granularity difference unit set corresponding to the second granularity; Based on the second granularity difference unit set, a text rewriting comparison view with difference markings performed at the second granularity is output.
2. The method according to claim 1, characterized in that The updating of the first granularity difference unit set to a second granularity difference unit set corresponding to the second granularity includes: The first granularity difference unit in the first granularity difference unit set is adjusted to a second granularity difference unit to form a second granularity difference unit set.
3. The method according to claim 2, characterized in that The adjusting the first granularity difference unit in the first granularity difference unit set to a second granularity difference unit includes: A plurality of first granularity difference units in the first granularity difference unit set that meet the aggregation condition corresponding to the second granularity are aggregated into a second granularity difference unit.
4. The method according to claim 3, characterized in that The step of aggregating a plurality of first granularity difference units that meet the aggregation condition corresponding to the second granularity in the first granularity difference unit set into a second granularity difference unit includes: Based on the second granularity, dividing the first granularity difference unit in the first granularity difference unit set into intervals, wherein the text combination corresponding to the first granularity difference unit contained in an interval conforms to the semantic boundary of the second granularity; At least one first granularity difference unit in the same interval is merged into a second granularity difference unit, and an operation type identifier is assigned to the second granularity difference unit.
5. The method according to claim 1, characterized in that In response to a granularity conversion triggering condition, determining a second granularity includes: Determine the type of rewriting operation from the original text to the rewritten text; Obtaining a preset target granularity corresponding to the rewrite operation type; In a case where it is determined that the target granularity is inconsistent with the first granularity, determining that a granularity conversion trigger condition is satisfied; In response to a granularity conversion triggering condition, the target granularity is determined to be a second granularity.
6. The method according to claim 1, characterized in that Before performing first-granularity difference detection on the original text and the corresponding rewritten text to obtain a first-granularity difference unit set, the method further includes: determining a type of rewriting operation from the original text to the rewritten text; Obtaining a preset target granularity corresponding to the rewrite operation type; The target particle size is used as the first particle size.
7. The method according to claim 5 or 6, characterized in that Determine the type of rewriting operation from the original text to the rewritten text, including: Performing intent recognition on a natural language processing instruction input by a user, and determining a type of rewriting operation from the original text to the rewritten text based on the intent recognition result, wherein the natural language processing instruction is used to instruct to rewrite the original text; Alternatively, obtain a function entry identifier that triggers the rewriting of the original text; search a preset operation type mapping table based on the function entry identifier to obtain the rewriting operation type from the original text to the rewritten text, wherein the operation type mapping table includes a correspondence between the function entry identifier and the rewriting operation type.
8. The method according to claim 1, characterized in that In response to a granularity conversion triggering condition, determining a second granularity includes: Determining a first difference marking density index for marking differences between the original text and the rewritten text at a first granularity based on the first granularity difference unit set; When it is determined that the first difference mark density indicator meets a preset condition, determining that a granularity conversion trigger condition is met; In response to a granularity conversion triggering condition, a second granularity is determined.
9. The method according to claim 8, characterized in that Before outputting the text rewriting comparison view with difference marks at the second granularity based on the second granularity difference unit set, the method further includes: Determining, based on the current second granularity difference unit set, a second difference marking density index for marking differences between the original text and the rewritten text at the current second granularity; When it is determined that the current second differential mark density indicator satisfies the preset condition, determining again that a granularity conversion trigger condition is satisfied; In response to the re-determined granularity conversion triggering condition, the second granularity is determined again until the current second difference mark density indicator does not meet the preset condition, or the preset level granularity is determined as the second granularity.
10. The method according to claim 8, characterized in that Determine the second granularity, including: From a preset granularity level sequence, a preset level granularity is selected and determined as the second granularity.
11. The method according to claim 9, characterized in that The re-determining of the second granularity includes: From the preset granularity level sequence, a granularity at a level adjacent to the current second granularity is selected and determined as the second granularity again.
12. The method according to claim 8, characterized in that The determining of a first difference marking density index for marking differences between the original text and the rewritten text at a first granularity includes: Determine a first granularity difference unit having at least one difference point in the first granularity difference unit set as a target difference unit; The proportion of the target difference unit in the first granularity difference unit set is determined, and the proportion is determined as a first difference marking density index for performing difference marking on the original text and the rewritten text at the first granularity.
13. The method according to claim 1, characterized in that The first-granularity difference detection is performed on the original text and the corresponding rewritten text to obtain a first-granularity difference unit set, including: Decomposing the original text and its corresponding rewritten text into first granularity units that conform to the first granularity respectively; Difference detection is performed based on first granularity units corresponding to the original text and the corresponding rewritten text to obtain a first granularity difference unit set.
14. The method according to claim 1, wherein The updating of the first granularity difference unit set to a second granularity difference unit set corresponding to the second granularity includes: Decomposing the original text and the corresponding rewritten text into second granularity units that conform to the second granularity respectively; Difference detection is performed based on the second granularity units corresponding to the original text and the corresponding rewritten text to obtain a set of second granularity difference units.
15. A text rewriting comparison device, characterized in that: The device comprises: A first difference detection module is configured to perform first-granularity difference detection on the original text and its corresponding rewritten text to obtain a first-granularity difference unit set; a granularity determination module, configured to determine a second granularity in response to a granularity conversion trigger condition; a granularity conversion module, configured to update the first granularity difference unit set to a second granularity difference unit set corresponding to the second granularity; The rewriting comparison module is configured to output a text rewriting comparison view with difference markings performed at the second granularity based on the second granularity difference unit set.
16. An electronic device, characterized in that: include: A processor and a memory, wherein the processor is configured to execute a text rewriting comparison program stored in the memory to implement the text rewriting comparison method according to any one of claims 1 to 14.
17. A storage medium, characterized in that: The storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the text rewriting comparison method according to any one of claims 1 to 14.