Editing management method based on recorded text interaction recognition

The recording device enters speech and uses the natural language understanding model for text editing, and automatically performs text conversion, semantic analysis and paragraph division, solving the problem of inefficiency of existing tools and achieving efficient and accurate text editing.

CN120146007APending Publication Date: 2025-06-13GUANGZHOU CHINA MOBILE SOFTWARE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510203642.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

Existing text editing tools based on speech recognition and natural language understanding are difficult to automatically perform text conversion, semantic analysis and paragraph division, resulting in users still needing to manually perform a large amount of semantic associations and paragraph division during the editing process, which is inefficient.

Method used

Voice is entered through recording equipment, speech is converted into text using language recognition technology, and sentence segmentation, semantic meaning and keyword extraction are performed through natural language understanding model, edit correlation index and basic editing index, automatically divide editing related paragraphs and determine the sentences to be edited.

Benefits of technology

It realizes the automation and intelligence of text editing, improves the efficiency and accuracy of text editing, and reduces the workload of manual editing by users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120146007A_ABST
    Figure CN120146007A_ABST
Patent Text Reader

Abstract

The invention discloses an editing management method based on recorded text interaction recognition, and relates to the technical field of text editing, and the method comprises the following steps: step 3, obtaining an editing priority index of each editing associated paragraph, and marking the editing associated paragraph with the maximum editing priority index value as a primary editing paragraph; 4, acquiring a basic editing index of each sentence in the paragraph to be edited, setting a basic editing standard index, marking the sentence as a sentence to be edited when the basic editing index of the sentence is greater than the basic editing standard index, firstly selecting the sentence to be edited with the maximum basic editing index, and then selecting the sentence to be edited with the maximum basic editing index; according to the method, the user is supported to execute editing operation on the to-be-edited sentence, the text paragraphs needing to be preferentially edited and the text sentences needing to be preferentially edited are rapidly determined through multi-dimensional analysis by performing word vector disassembling analysis and historical comparison analysis on the sentence in each edited paragraph, and the text editing pertinence, efficiency and precision of the user are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of text editing, and more specifically, it relates to an editing management method based on interactive recognition of recorded text. Background Art

[0002] In the current field of text editing technology, when users perform text editing, they usually rely on manual input or copy-and-paste methods to enter content into the editing system, and then edit and modify the text by means of manual reading and understanding. This traditional text editing method is not only inefficient but also error-prone. Especially when dealing with long texts or complex content, users often need to spend a lot of time and effort to locate and correct errors or inappropriate parts.

[0003] With the continuous development of speech recognition technology and natural language understanding technology, more and more text editing tools have begun to try to apply these technologies to the text editing process to achieve a more intelligent and automated editing method. However, existing text editing tools based on speech recognition and natural language understanding still have some problems. For example, these tools can often only perform simple speech recognition and text conversion, lacking in-depth understanding and analysis of text semantics, resulting in users still having to manually perform a large amount of semantic association and paragraph division work during the editing process.

[0004] In addition, existing text editing tools also have deficiencies in providing editing suggestions. Therefore, there is an urgent need for a text editing management method that can automatically perform text conversion, semantic analysis, and paragraph division, while accurately determining the editing location. Summary of the Invention

[0005] Aiming at the deficiencies of the existing technology, the purpose of the present invention is to provide an editing management method based on interactive recognition of recorded text.

[0006] To achieve the above purpose, the present invention provides the following technical solutions:

[0007] An editing management method based on interactive recognition of recorded text, comprising the following steps:

[0008] Step 1: The user records speech through a recording device, converts the speech into text through speech recognition technology, and performs sentence segmentation on the text through a natural language understanding model to obtain multiple sentences;

[0009] Step 2: Through the natural language understanding model, extract the semantic intention and keywords of each sentence, extract the keywords and semantic intentions in each sentence, obtain the editing association index of adjacent sentences, and based on the comparison result between the editing association index and the editing association standard index, divide to obtain multiple editing association paragraphs;

[0010] Step 3: Obtain the editing priority index of each edit-related paragraph, and mark the edit-related paragraph with the largest editing priority index value as the primary editing paragraph;

[0011] Step 4: Obtain the basic editing index of each sentence in the primary editing paragraph, set the basic editing standard index. When the basic editing index of a sentence is greater than the basic editing standard index, mark the sentence as a to-be-edited sentence. First, select the to-be-edited sentence with the largest basic editing index, support the user to perform an editing operation on the to-be-edited sentence. After the user finishes editing, the to-be-edited sentence is completed. Then, first select the to-be-edited sentence with the second largest basic editing index.

[0012] Further, the editing association index of adjacent sentences is obtained as follows: Denote the adjacent sentences as sentence A and sentence B respectively. Obtain the keyword vector and semantic association value of sentence A, and mark them as [A1, A2], where A1 is the keyword vector of sentence A and A2 is the semantic association value of sentence A. Obtain the keyword vector and semantic association value of sentence B, and mark them as [B1, B2], where B1 is the keyword vector of sentence B and B2 is the semantic association value of sentence B. Use the cosine similarity algorithm to calculate the editing association index of adjacent sentences.

[0013] Further, the keyword vector of a sentence is obtained as follows: Obtain the word vector of each keyword in the sentence through the Word2Vec model, and combine the word vectors of all keywords in the sentence to obtain the keyword vector of the sentence.

[0014] Further, the semantic association value of a sentence is obtained as follows: Denote the adjacent sentences as sentence A and sentence B respectively. Extract all semantic intents of sentence A, compare all semantic intents in sentence A with all semantic intents in sentence B one by one to obtain the semantic intent knowledge graph. When there is a connection between the two compared semantic intents in the semantic intent knowledge graph, add one to the number of associated intents. When there is no connection between the two compared semantic intents in the semantic intent knowledge graph, add one to the number of unrelated intents. Mark the number of associated intents as Num(asso) and the number of unrelated intents as Num(nor). Calculate the semantic association value sen(tn) of the sentence through sen(tn) = Num(asso) / (Num(asso) + Num(nor)).

[0015] Further, the editing priority index of the edit-related paragraph is obtained as follows: Obtain the basic editing index Bma of each sentence in the edit-related paragraph, where m = 1, 2,..., M, m is the sentence number, and M is the total number of sentences. Set the basic editing index coefficient as va, and use the formula to calculate the editing priority index index(pe) of the edit-related paragraph.

[0016] Further, the basic editing index of a sentence is obtained as follows: For all semantic intents included in the sentence, obtain the single-sentence vector deviation value of each semantic intent, sum and average all the single-sentence vector deviation values of the semantic intents to calculate the basic editing index of the sentence.

[0017] Further, the single-sentence vector deviation value of a semantic intent is obtained as follows: Select a sentence in the editing-related paragraph, mark this sentence as the preselected sentence, obtain the single-sentence vector value of the preselected sentence, obtain all the sentences that the user has edited before the current time of the system and contain this semantic intent, obtain the single-sentence vector value of each sentence, sum and average all the single-sentence vector values to calculate the historical single-sentence vector value, and calculate the absolute difference between the single-sentence vector value of the preselected sentence and the historical single-sentence vector value to calculate the single-sentence vector deviation value of this semantic intent.

[0018] Further, the single-sentence vector value is obtained as follows: Obtain the word vectors of each word in the sentence through the Word2Vec model, sum the word vectors of adjacent words to calculate the fused value of adjacent word vectors, sum and average all the fused values of adjacent word vectors to obtain the average fused value of adjacent word vectors Price(df), calculate the absolute difference between the word vectors of adjacent words to obtain the fluctuation value of adjacent word vectors, sum and average all the fluctuation values of adjacent word vectors to obtain the average fluctuation value of adjacent word vectors Price(ac), and calculate the single-sentence vector value BBs of this sentence through calculating the single-sentence vector value BBs of this sentence.

[0019] Compared with the prior art, the present invention has the following beneficial effects:

[0020] The method of the present invention, through steps one to four, performs text conversion on the user's recording through language recognition technology and natural language understanding model, and through in-depth analysis of the semantic intents and keywords of each sentence in the text, quickly divides the text into editing paragraphs, facilitating subsequent rapid positioning of the sentences to be edited in the semantically related paragraphs, performing disassembly analysis and historical comparison analysis of the word vectors of the sentences in each editing paragraph, and quickly determining the text paragraphs and text sentences that need to be edited preferentially through multi-dimensional analysis, improving the pertinence, efficiency and accuracy of the user's text editing. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 is a flowchart of an editing management method based on interactive recognition of recorded text;

[0022] Figure 2 is a flowchart for obtaining the single-sentence vector deviation value of a semantic intent. DETAILED DESCRIPTION OF THE INVENTION

[0023] Refer to Figures 1 to 2 , an editing management method based on interactive recognition of recorded text, includes the following steps:

[0024] Step 1: The user inputs speech through a recording device, converts the speech into text through speech recognition technology, and performs sentence segmentation on the text through a natural language understanding model to obtain multiple sentences;

[0025] Step 2: Through the natural language understanding model, extract the semantic intention and keywords for each sentence, extract the keywords and semantic intention in each sentence, obtain the editing correlation index of adjacent sentences, and set an editing correlation standard index (the editing correlation standard index is a preset index used for comparison with the editing correlation index). When the editing correlation index of adjacent sentences is less than the editing correlation standard index, perform paragraph division between the two sentences (when the editing correlation index of adjacent sentences is greater than or equal to the editing correlation standard index, there is no subsequent processing), and then divide to obtain multiple editing correlation paragraphs;

[0026] The natural language understanding model includes but is not limited to the NLP model. The keyword can be one or more, and the semantic intention can also be one or more. In the sentence "I like to take a walk and have a picnic in the park on weekends", the keywords include "weekends", "park", "take a walk", and "have a picnic", and the semantic intentions include "personal preference", "time selection", "activity location", "activity content";

[0027] The specific method for obtaining the editing correlation index of adjacent sentences is as follows: Denote the adjacent sentences as sentence A and sentence B respectively, obtain the keyword vector and semantic correlation value of sentence A, and mark them as [A1, A2], where A1 is the keyword vector of sentence A and A2 is the semantic correlation value of sentence A. Obtain the keyword vector and semantic correlation value of sentence B, and mark them as [B1, B2], where B1 is the keyword vector of sentence B and B2 is the semantic correlation value of sentence B. Use the cosine similarity algorithm to calculate the editing correlation index of adjacent sentences;

[0028]

[0029] The specific method for obtaining the keyword vector of a sentence is as follows: Obtain the word vector of each keyword in the sentence through the Word2Vec model, and combine the word vectors of all keywords in the sentence to obtain the keyword vector of the sentence;

[0030] The semantic association value of a sentence can be obtained as follows: Denote adjacent sentences as sentence A and sentence B respectively. Extract all semantic intents of sentence A, and compare each semantic intent in sentence A with all semantic intents in sentence B one by one (for example, if the semantic intents in sentence A include "personal preference" and "activity location", and the semantic intents in sentence B include "time selection" and "activity content", then compare "personal preference" in sentence A with "time selection" in sentence B, compare "personal preference" in sentence A with "activity content" in sentence B, compare "activity location" in sentence A with "time selection" in sentence B, and compare "activity location" in sentence A with "activity content" in sentence B) to obtain a semantic intent knowledge graph. When there is a connection between two compared semantic intents in the semantic intent knowledge graph, increment the number of associated intents by one. When there is no connection between two compared semantic intents in the semantic intent knowledge graph, increment the number of unrelated intents by one. Denote the number of associated intents as Num(asso) and the number of unrelated intents as Num(nor). Calculate the semantic association value sen(tn) of this sentence through sen(tn) = Num(asso) / (Num(asso) + Num(nor));

[0031] The semantic intent knowledge graph contains all semantic intents and displays all semantic intents in the form of a knowledge tree. Each node represents a semantic intent, and the connections between nodes represent the association or dependency relationship between them;

[0032] Step 3: Obtain the editing priority index of each editing-related paragraph, and mark the editing-related paragraph with the largest editing priority index value as the primary editing paragraph;

[0033] The editing priority index of an editing-related paragraph can be obtained as follows: Obtain the basic editing index Bma of each sentence in the editing-related paragraph, where m = 1, 2,..., M, m is the sentence number, and M is the total number of sentences. Set the basic editing index coefficient as va, where a = 1, 2, 3,..., a, and v1 < v2 <... < va - 1 < va. Each basic editing index coefficient corresponds to a range of basic editing indices. The ranges of basic editing indices include (0, Bm1], (Bm1, Bm2],..., (Bma - 1, Bma]. When Bma ∈ (0, Bm1], the basic editing index coefficient is v1. Use the formula to calculate the editing priority index index(pe) of this editing-related paragraph;

[0034] The basic editing index of a sentence can be obtained as follows: For all semantic intents included in the sentence, obtain the single-sentence vector deviation value of each semantic intent, and perform a sum and average calculation on the single-sentence vector deviation values of all semantic intents to calculate the basic editing index of this sentence;

[0035] The deviation value of the single-sentence vector of the semantic intention is obtained as follows: Select a sentence from the edited related paragraph, mark the sentence as the preselected sentence, obtain the single-sentence vector value of the preselected sentence, obtain all the sentences that the user has edited before the current time of the system and contain this semantic intention, obtain the single-sentence vector value of each sentence, calculate the sum mean of the single-sentence vector values of all sentences to obtain the historical single-sentence vector value, and calculate the absolute difference between the single-sentence vector value of the preselected sentence and the historical single-sentence vector value to obtain the deviation value of the single-sentence vector of this semantic intention;

[0036] The single-sentence vector value is obtained as follows: Obtain the word vectors of each word in the sentence through the Word2Vec model (including the word vectors of keywords and non-keywords), calculate the sum of the word vectors of adjacent words to obtain the fused value of adjacent word vectors, calculate the sum mean of all the fused values of adjacent word vectors to obtain the average fused value of adjacent word vectors Price(df), calculate the absolute difference between the word vectors of adjacent words to obtain the fluctuation value of adjacent word vectors, calculate the sum mean of all the fluctuation values of adjacent word vectors to obtain the average fluctuation value of adjacent word vectors Price(ac), and calculate the single-sentence vector value BBs of this sentence;

[0037] Step 4: Obtain the basic editing index of each sentence in the primary edited paragraph, set the basic editing standard index (the basic editing standard index is a preset index used for comparison with the basic editing index). When the basic editing index of a sentence is greater than the basic editing standard index, mark the sentence as a sentence to be edited (when the basic editing index of a sentence is less than or equal to the basic editing standard index, there is no further processing). First, select the sentence to be edited with the largest basic editing index, support the user to perform editing operations (such as insertion, deletion, replacement, etc.) on this sentence to be edited. After the user finishes editing, this sentence to be edited is completed. Then, first select the sentence to be edited with the second largest basic editing index (and so on. When all the sentences to be edited in the primary edited paragraph are completed, mark the edited related paragraph with the second largest editing priority index value as the primary edited paragraph).

[0038] Through Steps 1 to 4, the recorded voice of the user is textually converted through language recognition technology and natural language understanding models, and through in-depth analysis of the semantic intention and keywords of each sentence in the text, the edited paragraphs are quickly divided, which is convenient for quickly locating the sentences to be edited in the semantically related paragraphs later. The word vectors of the sentences in each edited paragraph are disassembled and analyzed and compared with the history. Through multi-dimensional analysis, the text paragraphs that need to be edited first and the text sentences that need to be edited first are quickly determined, improving the pertinence, efficiency, and accuracy of the user's text editing.

[0039] The above formulas are all dimensionless and take their numerical values for calculation. The formulas are obtained by collecting a large amount of data and performing software simulations to get a formula closest to the actual situation. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.

[0040] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired or wireless (such as infrared, wireless, microwave, etc.) manner. The computer-readable storage medium can be any available medium that the computer can access or a data storage device such as a server or data center that contains one or more collections of available media. The available media can be magnetic media (such as floppy disks, hard disks, magnetic tapes), optical media (such as DVDs), or semiconductor media. The semiconductor media can be a solid-state drive.

[0041] It should be understood that in various embodiments of the present application, the magnitudes of the serial numbers of the above processes do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0042] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0043] Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be repeated here.

[0044] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the devices or units can be in electrical, mechanical, or other forms.

[0045] If the described function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0046] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claimed rights.

Claims

1. An editing management method based on interactive recognition of recorded text, characterized in that: The steps include: Step 1: The user records voice through a recording device, the voice is converted into text through language recognition technology, and the text is segmented into sentences through a natural language understanding model to obtain multiple sentences; Step 2: Use the natural language understanding model to extract the semantic intent and keywords of each sentence, extract the keywords and semantic intent in each sentence, obtain the editing relevance index of adjacent sentences, and divide them into multiple editing relevance paragraphs based on the comparison results of the editing relevance index and the editing relevance standard index; Step 3: Obtain the editing priority index of each editing-related paragraph, and mark the editing-related paragraph with the largest editing priority index value as the primary editing paragraph; Step 4: Obtain the basic editing index of each sentence in the primary editing paragraph, set the basic editing standard index, and when the basic editing index of a sentence is greater than the basic editing standard index, mark the sentence as a sentence to be edited. First, select the sentence to be edited with the largest basic editing index, and support the user to perform editing operations on the sentence to be edited. After the user finishes editing, the sentence to be edited completes editing, and then select the sentence to be edited with the second largest basic editing index first.

2. The editing management method based on interactive recognition of audio and text according to claim 1 is characterized in that: The editing association index of adjacent sentences is specifically obtained as follows: the adjacent sentences are respectively recorded as sentence A and sentence B, the keyword vector and semantic association value of sentence A are obtained, and marked as [A1, A2], where A1 is the keyword vector of sentence A and A2 is the semantic association value of sentence A, the keyword vector and semantic association value of sentence B are obtained, and marked as [B1, B2], where B1 is the keyword vector of sentence B and B2 is the semantic association value of sentence B, and the editing association index of adjacent sentences is calculated using the cosine similarity algorithm.

3. The editing management method based on interactive recognition of audio and text according to claim 2 is characterized in that: The keyword vector of a sentence is obtained as follows: the word vector of each keyword in the sentence is obtained through the Word2Vec model, and the word vectors of all keywords in the sentence are combined to obtain the keyword vector of the sentence.

4. The editing management method based on interactive recognition of audio and text according to claim 2 is characterized in that: The semantic association value of a sentence is specifically obtained as follows: adjacent sentences are respectively recorded as sentence A and sentence B, all semantic intentions of sentence A are extracted, all semantic intentions in sentence A are compared one by one with all semantic intentions in sentence B, and the semantic intention knowledge graph is obtained. When the two compared semantic intentions are connected in the semantic intention knowledge graph, the number of associated intentions is increased by one; when the two compared semantic intentions are not connected in the semantic intention knowledge graph, the number of irrelevant intentions is increased by one; the number of associated intentions is marked as Num(asso), and the number of irrelevant intentions is marked as Num(nor); the semantic association value sen(tn) of the sentence is calculated by sen(tn)=Num(asso) / (Num(asso)+Num(nor)).

5. The editing management method based on interactive recognition of audio and text according to claim 1 is characterized in that: The editing priority index of the editing-related paragraph is obtained as follows: obtain the basic editing index Bma of each sentence in the editing-related paragraph, m = 1, 2, ..., M, m is the sequence number of the sentence, M is the total number of sentences, set the basic editing index coefficient to va, and use the formula Calculate the editing priority index index (pe) of the editing-related paragraph.

6. The editing management method based on interactive recognition of audio and text according to claim 5 is characterized in that: The basic editing index of a sentence is obtained by the following method: for all the semantic intents contained in the sentence, obtain the single-sentence vector deviation value of each semantic intent, calculate the sum and mean of the single-sentence vector deviation values ​​of all semantic intents, and calculate the basic editing index of the sentence.

7. The editing management method based on interactive recognition of audio and text according to claim 6 is characterized in that: The single-sentence vector deviation value of the semantic intent is specifically obtained as follows: select a sentence in the edit-related paragraph, mark the sentence as a pre-selected sentence, obtain the single-sentence vector value of the pre-selected sentence, obtain all sentences that the user has completed editing before the current time of the system and contain the semantic intent, obtain the single-sentence vector value of each sentence, calculate the sum and mean of the single-sentence vector values ​​of all sentences, calculate the historical single-sentence vector value, calculate the absolute difference between the single-sentence vector value of the pre-selected sentence and the historical single-sentence vector value, and calculate the single-sentence vector deviation value of the semantic intent.

8. The editing management method based on interactive recognition of audio and text according to claim 7 is characterized in that: The specific method of obtaining the sentence vector value is as follows: obtain the word vector of each word in the sentence through the Word2Vec model, sum the word vectors of adjacent words to obtain the adjacent word vector fusion value, sum and average all adjacent word vector fusion values ​​to obtain the average adjacent word vector fusion value Price(df), calculate the absolute difference of the word vectors of adjacent words to obtain the adjacent word vector fluctuation value, sum and average all adjacent word vector fluctuation values ​​to obtain the average adjacent word vector fluctuation value Price(ac), and use Calculate the sentence vector value BBs of the sentence.