Text editing, display method, device, electronic device and storage medium
By determining the operation intention based on quantity difference and pronunciation similarity in text editing operations, and maintaining the consistency of voice timestamps when replacing text, the problem of text editing destroying the correspondence between voice and transcribed text is solved, and the synchronous display and comparative editing of voice and text are achieved.
Patent Information
- Application Number
- CN202211723650.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-30
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2042-12-30
AI Technical Summary
In the prior art, text editing operations destroy the correspondence between speech and transcribed text, resulting in an inability to effectively compare, view, and edit.
By determining the operation intention based on the quantitative difference and pronunciation similarity between the deletion operation and the addition operation in the text editing operation, and maintaining the consistency of the voice timestamp when replacing the text, the correspondence between the voice and text is ensured not to be destroyed.
After text editing, the voice and text are displayed synchronously, providing a better comparison experience and ensuring that users can effectively compare, view and edit.
Smart Images

Figure CN115935918B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a text editing and display method, device, electronic device and storage medium. Background Art
[0002] A typical speech transcription solution needs to have an online editing function for the transcribed text, and be able to simultaneously display the transcribed text corresponding to the currently displayed content while satisfying the speech display, so that users can check whether the corresponding transcribed text needs to be edited and corrected by comparing it with the currently displayed content.
[0003] However, editing operations on the transcribed text often destroy the correspondence between the speech and the transcribed text, causing the correspondence between the edited transcribed text and the speech to become confusing, making it impossible to perform comparative viewing and comparative editing. Summary of the Invention
[0004] The present invention provides a text editing and display method, device, electronic device and storage medium, which are used to solve the defect in the prior art that editing and modifying transcribed text will cause the comparison display of speech and transcribed text to fail.
[0005] The present invention provides a text editing method, comprising:
[0006] Receive text editing operations;
[0007] In a case where the text editing operation includes a adding operation and a deleting operation at the same position in the corresponding text, determining the operation intention of the text editing operation based on a difference in the number of deleted objects of the deleting operation and the added objects of the adding operation, and / or a pronunciation similarity between the added objects and the deleted objects;
[0008] In the case where the operation intention is text replacement, the adding operation and the deleting operation are performed in the text, and based on the voice timestamp of the deletion object corresponding to the deletion operation in the text, the voice timestamp of the adding object corresponding to the adding operation in the text is determined.
[0009] According to a text editing method provided by the present invention, determining the operation intention based on the difference in quantity between the deleted objects and the newly added objects, and / or the pronunciation similarity between the newly added objects and the deleted objects, includes:
[0010] When the difference in quantity between the deleted objects and the newly added objects is less than a first threshold, determining that the operation intention is text replacement;
[0011] When the difference in number between the deleted objects and the newly added objects is greater than or equal to the first threshold and less than a second threshold, determining the operation intention based on the pronunciation similarity between the newly added objects and the deleted objects;
[0012] When the difference in quantity between the deleted objects and the newly added objects is greater than or equal to the second threshold, it is determined that the operation intention is independent editing.
[0013] According to a text editing method provided by the present invention, determining the voice timestamp of the object added in the text corresponding to the deletion operation based on the voice timestamp of the deleted object in the text, includes:
[0014] When the number of the deleted objects and the number of the newly added objects are equal, determining a correspondence between the deleted objects and the newly added objects based on an arrangement order of the deleted objects and an arrangement order of the newly added objects, and using the voice timestamps of the deleted objects as the voice timestamps of the corresponding newly added objects;
[0015] When the number of the deleted objects and the number of the newly added objects are not equal, the voice time interval of the deleted object segment is determined based on the voice timestamp of the deleted object, and the voice timestamp of the newly added object is determined based on the voice time interval and the number of the newly added objects.
[0016] According to a text editing method provided by the present invention, after receiving the text editing operation, the method further includes:
[0017] When the text editing operation includes adding and deleting operations corresponding to different positions in the text, or only includes adding or deleting operations, or the operation is intended to be an independent edit, the deleting operation is performed in the text, and / or the adding operation is performed in the text and the voice timestamp of the newly added object is set to empty.
[0018] The present invention also provides a display method, comprising:
[0019] Receive display control operations;
[0020] Performing voice and text display based on the display position corresponding to the display control operation and the voice timestamps of each object in the text corresponding to the voice;
[0021] The text is obtained through text editing. When the operation intention of the text editing operation is text replacement, the voice timestamp of the newly added object in the text is determined based on the voice timestamp of the deleted object. The operation intention is, when the text editing operation includes a new operation and a deletion operation at the same position in the corresponding text, based on the difference in quantity between the deleted objects of the deletion operation and the new objects of the new operation, and / or the pronunciation similarity between the new objects and the deleted objects.
[0022] According to a display method provided by the present invention, performing voice and text display based on the display position corresponding to the display control operation and the voice timestamps of each object in the text corresponding to the voice includes:
[0023] determining a presentation start object from the text based on the presentation position;
[0024] When the voice timestamp of the display starting object is not empty, performing voice and text display based on the voice timestamp;
[0025] When the voice timestamp of the display starting object is empty, the start time of the voice display is determined based on the voice timestamp of the first object with a non-empty voice timestamp after the display starting object in the text, voice display is performed based on the start time, and text display is performed based on the display starting object.
[0026] The present invention also provides a text editing device, comprising:
[0027] An editing receiving unit, used for receiving text editing operations;
[0028] an intention determining unit for determining, when the text editing operation includes a adding operation and a deleting operation at the same position in the corresponding text, the operation intention of the text editing operation based on a difference in the number of deleted objects of the deleting operation and the added objects of the adding operation, and / or a pronunciation similarity between the added objects and the deleted objects;
[0029] An execution unit is used to execute the adding operation and the deleting operation in the text when the operation intention is text replacement, and determine the voice timestamp of the adding object corresponding to the adding operation in the text based on the voice timestamp of the deletion object corresponding to the deletion operation in the text.
[0030] The present invention also provides a display device, comprising:
[0031] A presentation receiving unit, configured to receive a presentation control operation;
[0032] A presentation unit, configured to present the voice and text based on a presentation position corresponding to the presentation control operation and a speech timestamp of each object in the text corresponding to the voice;
[0033] The text is obtained through text editing. When the operation intention of the text editing operation is text replacement, the voice timestamp of the newly added object in the text is determined based on the voice timestamp of the deleted object. The operation intention is, when the text editing operation includes a new operation and a deletion operation at the same position in the corresponding text, based on the difference in quantity between the deleted objects of the deletion operation and the new objects of the new operation, and / or the pronunciation similarity between the new objects and the deleted objects.
[0034] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the text editing method or display method described above is implemented.
[0035] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the above-mentioned text editing methods or display methods.
[0036] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements any of the above-mentioned text editing methods or display methods.
[0037] The text editing and display method, device, electronic device and storage medium provided by the present invention determine the operational intention of the text editing operation based on the difference in quantity between the deleted objects of the deletion operation and the added objects of the addition operation and / or the pronunciation similarity between the added objects and the deleted objects when there are addition operations and deletion operations at the same position in the corresponding text, thereby ensuring that when the operational intention is text replacement, the addition operation and the deletion operation will not damage the correspondence between the voice and text. The text after text editing can still be used for the synchronous display of voice and text, providing users with a better comparison experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0039] Figure 1 It is a flowchart of the text editing method provided by the present invention;
[0040] Figure 2 It is a schematic flow chart of the display method provided by the present invention;
[0041] Figure 3 It is a structural diagram of the text editing device provided by the present invention;
[0042] Figure 4 It is a structural schematic diagram of the display device provided by the present invention;
[0043] Figure 5 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0044] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0045] In smart office or other voice recognition scenarios, you can record and transcribe your voice. After the transcription is complete, you can drag the audio progress bar to a specific position to view and modify the transcript. The system will automatically jump to the transcript at that point in time. Similarly, you can click on the desired transcription in the transcript to automatically jump to the corresponding audio location.
[0046] In order to achieve such an effect, during the actual voice recording and transcription process, it is necessary to match each word in the transcribed text with the voice according to the timestamp information. However, due to the differences in the actual use environment, the accuracy of voice transcription varies, so after the voice recording and transcription are completed, the user usually needs to edit the transcribed text. In order to avoid the editing operation of the transcribed text destroying the corresponding relationship between the voice and the transcribed text, an embodiment of the present invention provides a text editing method.
[0047] Figure 1 It is a flowchart of the text editing method provided by the present invention, such as Figure 1 As shown, the method includes:
[0048] Step 110: Receive a text editing operation.
[0049] Here, a text editing operation is a user-initiated operation to edit text associated with speech. A text editing operation may include one or more operations, and include the location of each operation in the text. It is understood that the speech-related text referred to herein may be text obtained by recognizing or transcribing speech, or may be text that is temporally associated with the speech input, and this is not specifically limited in the present embodiment.
[0050] For example, for the text "This is a paragraph that has been deleted only. Each time it is displayed, the sound is displayed starting from the character after the punctuation mark", the text editing operation can be a deletion operation, specifically deleting "each time it is displayed"; the text editing operation can also be a adding operation, specifically adding "This is a paragraph after compilation" before "each time it is displayed"; the text editing operation can also include deletion and adding operations at the same time, and the adding and deletion operations can correspond to the same execution position or different execution positions, specifically deleting "only" and adding "without".
[0051] Step 120, when the text editing operation includes a adding operation and a deleting operation at the same position in the corresponding text, the operational intention of the text editing operation is determined based on the difference in quantity between the deleted objects of the deletion operation and the added objects of the adding operation, and / or the pronunciation similarity between the added objects and the deleted objects.
[0052] Specifically, after the text editing operation is acquired, corresponding operation actions need to be performed for different text editing operations to ensure that the text editing does not affect the correspondence between each word in the text and the voice.
[0053] For the case where the text editing operation includes a single operation, the user's intention to edit the text is the text editing operation itself, such as deletion or addition; for the case where the text editing operation includes multiple operations, the user's intention to edit the text may be to regard each operation as an independent operation, and delete or add separately, or to regard multiple operations as a combined operation, and realize the intention of text replacement through the combination of operations.
[0054] Therefore, in the case where the text editing operation includes a new operation and a deletion operation at the same position in the corresponding text, it is necessary to further determine whether the operational intention of the text editing operation is text replacement or independent editing. Among them, text replacement, that is, through the new operation and the deletion operation at the same position, the operation of replacing the deleted object of the deletion operation with the new object of the new operation is realized; independent editing, that is, the new operation and the deletion operation are independently performed at the same position, and there is no replacement relationship between the new object of the new operation and the deleted object of the deletion operation. It can be understood that the new objects and deleted objects referred to here are units relative to the text. For example, the object can be a character, a word, a character in a word, a phrase, a short sentence, etc., and the embodiment of the present invention does not make specific limitations on this.
[0055] In order to obtain the operation intention, the number of deleted objects of the deletion operation and the number of newly added objects of the newly added operation can be obtained, and the difference between the two can be calculated. It is understandable that a common error in speech transcription is a wrong character, and it is rare for a word in the speech to be transcribed into a text of multiple words. Therefore, if text replacement is performed to correct the speech transcription error, the number of deleted objects and newly added objects may be equal, or the difference in quantity between the two may be small. Thus, it is possible to determine whether the operation intention is to replace text based on the size of the difference in quantity between the deleted objects and the newly added objects.
[0056] For example, a difference threshold can be set in advance. If the difference in quantity between the deleted objects and the newly added objects is greater than the difference threshold, the operation intention is determined to be independent editing. If the difference in quantity between the deleted objects and the newly added objects is less than or equal to the difference threshold, the operation intention is determined to be text replacement. The difference threshold here can be set to 3, or to 4, 5, etc., and the embodiment of the present invention does not make specific limitations on this.
[0057] In addition, to obtain the operation intention, the pronunciation similarity between the deleted object of the delete operation and the newly added object of the add operation can also be referred to. Here, the pronunciation similarity between the newly added object and the deleted object can be calculated based on the encoding similarity of the pronunciation sequence encoding of the newly added object and the pronunciation sequence encoding of the deleted object. It is understandable that the higher the pronunciation similarity between the two, the more likely the text editing operation is a text replacement operation performed to correct errors in speech transcription.
[0058] Therefore, the operation intention can be distinguished based on the pronunciation similarity between the newly added object and the deleted object. For example, if the pronunciation similarity between the newly added object and the deleted object is greater than the similarity threshold, the operation intention can be determined to be text replacement, otherwise the operation intention can be determined to be independent editing.
[0059] In addition, although most of the text replacements performed to correct speech transcription errors have a small difference in the number of deleted objects and newly added objects, it cannot be ruled out that the text replacement situation in which the difference in the number of deleted objects and newly added objects is greater than a pre-set difference threshold cannot be ruled out. In order to avoid confusion in the correspondence between speech and text due to incorrect judgment of the operation intention, the operation intention can be judged by combining the difference in the number of deleted objects and newly added objects, as well as the pronunciation similarity between the newly added objects and the deleted objects. For example, the first score of the operation intention for text replacement can be evaluated by the size of the difference, and the second score of the operation intention for text replacement can be evaluated by the pronunciation similarity between the newly added objects and the deleted objects. Then, based on the first score and the second score, the operation intention can be distinguished.
[0060] Step 130, when the operation intention is text replacement, perform the adding operation and the deleting operation in the text, and determine the voice timestamp of the adding object corresponding to the adding operation in the text based on the voice timestamp of the deleting object corresponding to the deleting operation in the text.
[0061] Specifically, for the case where the operation intention is to replace text, the text before and after the text replacement should correspond to the same timestamp, that is, the correspondence between the text and the voice before and after the text replacement should remain unchanged. Thus, after completing the add and delete operations, the newly added object in the text based on the add operation can inherit the voice timestamp of the object deleted in the text by the delete operation, that is, the voice timestamp of the deleted object is used as the voice timestamp of the newly added object. It can be understood that the voice timestamp here is based on the object, and each object corresponds to a voice timestamp. The voice timestamp of an object is used to reflect the corresponding time position of the object in the voice, which can specifically be the start time and end time in the voice.
[0062] In addition, for the addition operation performed by independent editing, the newly added object has nothing to do with the voice itself. The voice timestamp of the newly added object can be directly set to empty, or the voice timestamp of the original object in the text arranged after the newly added object can be used as the voice timestamp of the newly added object; for the deletion operation performed by independent editing, directly delete the object to be deleted in the text, and delete the voice timestamp of the object at the same time.
[0063] The method provided by an embodiment of the present invention determines the operational intention of a text editing operation based on the quantitative difference between the deleted objects of the deletion operation and the added objects of the addition operation, and / or the pronunciation similarity between the added objects and the deleted objects when there are addition operations and deletion operations at the same position in the corresponding text. This ensures that when the operational intention is to replace text, the addition operations and deletion operations will not damage the correspondence between voice and text. The text after text editing can still be used for the synchronous display of voice and text, providing users with a better comparison experience.
[0064] Based on any of the above embodiments, in step 120, determining the operation intention based on the difference in number between the deleted objects and the newly added objects, and / or the pronunciation similarity between the newly added objects and the deleted objects, includes:
[0065] When the difference in quantity between the deleted objects and the newly added objects is less than a first threshold, determining that the operation intention is text replacement;
[0066] When the difference in number between the deleted objects and the newly added objects is greater than or equal to the first threshold and less than a second threshold, determining the operation intention based on the pronunciation similarity between the newly added objects and the deleted objects;
[0067] When the difference in quantity between the deleted objects and the newly added objects is greater than or equal to the second threshold, it is determined that the operation intention is independent editing.
[0068] Specifically, in the process of distinguishing the operation intention by combining the difference in the number of deleted objects and the newly added objects, as well as the pronunciation similarity between the newly added objects and the deleted objects, the difference in the number of deleted objects and the newly added objects can be used as a prerequisite, and two thresholds for measuring the size of the difference in the number of objects, namely a first threshold and a second threshold, can be set, thereby specifically dividing three situations to facilitate the distinction of the operation intention. It is understandable that the first threshold here is smaller than the second threshold. For example, the first threshold can be set to 1 or 0, and the second threshold can be set to 4 or 5.
[0069] When the quantity difference between the deleted objects and the newly added objects is less than the first threshold, the quantity difference between the deleted objects and the newly added objects is extremely small or even close to 0. It can be directly determined that the newly added object is a replacement for the deleted object, that is, it can be directly determined that the operation intention is text replacement.
[0070] When the difference in number between the deleted objects and the newly added objects is greater than or equal to the second threshold, the difference in number between the deleted objects and the newly added objects is large, and it can be directly determined that the newly added object is not a replacement for the deleted object, that is, it can be directly determined that the operation intention is independent editing.
[0071] If the difference in number between the deleted and newly added objects is greater than or equal to a first threshold and less than a second threshold, it is difficult to directly determine whether the newly added object is a replacement for the deleted object based solely on the difference in number between the deleted and newly added objects. Therefore, it is necessary to further calculate the pronunciation similarity between the deleted and newly added objects, that is, the pronunciation similarity between the two. It is understandable that in this case, if the pronunciation similarity is greater than the pre-set similarity threshold, the operation intention is determined to be text replacement; otherwise, the operation intention is determined to be independent editing.
[0072] The method provided by an embodiment of the present invention divides the operation intention into three situations based on the difference in the number of deleted objects and newly added objects. When the difference is less than a first threshold or greater than or equal to a second threshold, the operation intention can be determined without calculating the pronunciation similarity. While ensuring the reliability of the determination of the operation intention, the amount of calculation required for text editing is further reduced.
[0073] Based on any of the above embodiments, in step 130, determining the voice timestamp of the object corresponding to the newly added object in the text by the adding operation based on the voice timestamp of the deleted object in the text by the deleting operation includes:
[0074] When the number of the deleted objects and the number of the newly added objects are equal, determining a correspondence between the deleted objects and the newly added objects based on an arrangement order of the deleted objects and an arrangement order of the newly added objects, and using the voice timestamps of the deleted objects as the voice timestamps of the corresponding newly added objects;
[0075] When the number of the deleted objects and the number of the newly added objects are not equal, the voice time interval of the deleted object segment is determined based on the voice timestamp of the deleted object, and the voice timestamp of the newly added object is determined based on the voice time interval and the number of the newly added objects.
[0076] Specifically, in the case of text replacement, in order to allow the newly added objects after replacement to postpone the voice timestamps of the deleted objects and ensure that the correspondence between the text and voice after editing is still reliable, it is necessary to set the voice timestamps of the newly added objects based on whether the number of newly added objects is equal to the number of deleted objects.
[0077] Regarding the case where the number of deleted objects and newly added objects is equal, it is understandable that in this case, the deleted objects and the newly added objects are usually one-to-one corresponding, and the correspondence between the deleted objects and the newly added objects can be directly determined based on the arrangement order of the deleted objects and the arrangement order of the newly added objects. For example, if the deleted objects include "for only" and the newly added objects include "without", "for" is the first deleted object, "without" is the first newly added object, "only" is the second deleted object, and "within" is the second newly added object, then based on the arrangement order, "for" can be determined to correspond to "without", and "only" can be determined to correspond to "within".
[0078] After the corresponding relationship is determined, the voice timestamp of the deleted object corresponding to the newly added object can be used as the voice timestamp of the newly added object. For example, the voice timestamp of "for" can be assigned to "not yet", and the voice timestamp of "only" can be assigned to "through".
[0079] Regarding the situation where the number of deleted objects and newly added objects is not equal, it is understandable that in this case, the deleted objects and the newly added objects cannot correspond one to one, and therefore the voice timestamp of the deleted objects cannot be directly used as the voice timestamp of the newly added objects. However, considering that the deleted objects as a whole of the text replacement correspond to the newly added objects as a whole, the voice time interval of the deletion field composed of the deleted objects can be determined based on the voice timestamps of all deleted objects. It is understandable that the voice time interval here is the interval with the start time of the voice timestamp of the first deleted object in the deletion field and the end time of the voice timestamp of the last deleted object in the deletion field as the two time endpoints. After obtaining the voice time interval, the voice time interval can be evenly divided by the number of newly added objects, so as to obtain voice timestamps that can correspond one to one with the newly added objects.
[0080] The method provided by the embodiment of the present invention determines the speech timestamp of the newly added object by combining whether the number of deleted objects and newly added objects is equal, thereby ensuring the reliability of the correspondence between text and speech during the text replacement operation.
[0081] Based on any of the above embodiments, after step 110, the method further includes:
[0082] When the text editing operation includes adding and deleting operations corresponding to different positions in the text, or only includes adding or deleting operations, or the operation is intended to be an independent edit, the deleting operation is performed in the text, and / or the adding operation is performed in the text and the voice timestamp of the newly added object is set to empty.
[0083] Specifically, the text editing operation received in step 110 may also include only one independent operation, such as only a new operation, or only a delete operation; or, although the text editing operation includes multiple operations, they are independent operations at different positions in the corresponding text, or although the text editing operation includes multiple operations at the same position in the corresponding text, it is determined that the operation intention is independent editing, that is, it is clear that the multiple operations here are independent of each other.
[0084] For the above situation, that is, the operation included in the text editing operation is an independent operation, for the newly added operation that is executed independently, a new object corresponding to the newly added operation can be added to the text, and considering that the newly added object is not related to the voice, the voice timestamp of the newly added object can be set to empty to avoid the newly added object destroying the correspondence between voice and text.
[0085] For independently executed deletion operations, the deletion object corresponding to the deletion operation can be deleted in the text, and the voice timestamp of the deleted object can also be deleted together. The deletion operation on the voice timestamp of the deleted object does not affect the voice timestamps of other objects in the text.
[0086] Based on any of the above embodiments, Figure 2 It is a flow chart of the display method provided by the present invention, such as Figure 2 As shown, the display method is used to realize the comparative display of voice and text based on the above text editing method, and the method includes:
[0087] Step 210: receiving a display control operation;
[0088] Step 220: Perform voice and text display based on the display position corresponding to the display control operation and the voice timestamps of each object in the text corresponding to the voice;
[0089] The text is obtained through text editing. When the operation intention of the text editing operation is text replacement, the voice timestamp of the newly added object in the text is determined based on the voice timestamp of the deleted object. The operation intention is, when the text editing operation includes a new operation and a deletion operation at the same position in the corresponding text, based on the difference in quantity between the deleted objects of the deletion operation and the new objects of the new operation, and / or the pronunciation similarity between the new objects and the deleted objects.
[0090] Specifically, the display method described here implements a comparative display of speech and the text associated with the speech. This comparative display is possible only if there is a corresponding relationship between the speech and the text. That is, each object in the text has a corresponding speech timestamp, which reflects the corresponding position of each object in the speech.
[0091] Considering that text editing operations are likely to destroy the correspondence between speech and text, when editing text, it is necessary to determine whether the multiple operations on the same position contained in the text editing operation are for text replacement or are executed independently. This ensures that under the intention of text replacement operations, the newly added objects after replacement can inherit the speech timestamps of the deleted objects before replacement, thereby ensuring that text replacement will not adversely affect the correspondence between speech and text.
[0092] By applying the voice timestamps of each object in the text maintained based on the above text editing method, the comparison display of voice and text can be achieved.
[0093] Specifically, during the execution of the comparison display, the user may initiate a control operation for the comparison display, i.e., a display control operation. The display control operation may indicate the start time of the voice display, or indicate the starting sentence or starting object of the text display, which is not specifically limited in the embodiment of the present invention.
[0094] After receiving a display control operation, the corresponding display position can be determined. The display position can be the start time of the voice display, or indicate the starting sentence or starting object of the text display. Subsequently, the display position and the voice timestamps of each object in the text can be combined to achieve synchronous comparison of voice and text.
[0095] During this process, in the case where the display position is the starting time of the voice display, voice playback can be performed based on the starting time, and based on the voice timestamps of each object in the text, the object corresponding to the voice timestamp corresponding to the starting time is found as the starting object of the text display, and text display is performed based on the starting object, thereby maintaining a corresponding relationship between the voice playback and the text display, which is convenient for comparison and viewing.
[0096] In the case where the display position is the starting sentence or starting object of text display, the starting time of the voice display can be determined based on the voice timestamp of the first object in the starting sentence, or based on the voice timestamp of the starting object, so that voice playback is performed based on the starting time, and text display is performed based on the starting sentence or starting object. This keeps the voice playback and text display in correspondence, which is convenient for comparison and viewing.
[0097] The method provided by the embodiment of the present invention applies the operation intention under text editing to update the voice timestamp of each object in the text, thereby maintaining the correspondence between voice and text, and based on this correspondence, realizes the comparative display of voice and text under display control operation, providing convenience for comparative viewing and comparative editing of the voice transcription system.
[0098] Based on any of the above embodiments, step 220 includes:
[0099] determining a presentation start object from the text based on the presentation position;
[0100] When the voice timestamp of the display starting object is not empty, performing voice and text display based on the voice timestamp;
[0101] When the voice timestamp of the display starting object is empty, the start time of the voice display is determined based on the voice timestamp of the first object with a non-empty voice timestamp after the display starting object in the text, voice display is performed based on the start time, and text display is performed based on the display starting object.
[0102] Specifically, in the case where the display position indicated by the display control operation is the display starting sentence of the text, the first object in the display starting sentence can be used as the display starting object; or, in the case where the display position indicated by the display control operation is the display starting object of the text, the object at the indicated display position can be directly determined as the display starting object.
[0103] After obtaining the display start object, the corresponding voice timestamp of the display start object can be obtained. Since the text has been edited, if the display start object is an existing object in the text, or if the display start object is a newly added object obtained by text replacement, the voice timestamp of the display start object contains a specific time. The start time of the voice display can be directly determined based on the voice timestamp, and then the voice display is based on the display time, and the text display is based on the display start object.
[0104] However, if the display starting object is a new object added to the text through an independent addition operation, and the new object itself is not related to the voice, the voice timestamp of the display starting object is empty. At this time, you can traverse the objects after the display starting object in the text until you get an object with a non-empty voice timestamp, and use the start time in the voice timestamp of the object as the start time of the voice display, and perform voice display based on the start time, and also perform text display based on the display starting object.
[0105] Based on any of the above embodiments, an embodiment of the present invention provides a text editing method and a display method based on the above text editing method, which may include the following steps:
[0106] First receive the text editing operation.
[0107] For situations where a text editing operation includes both a new operation and a deletion operation at the same position in the corresponding text, the operational intent of the text editing operation can be determined based on the difference in the number of objects deleted by the deletion operation and the objects added by the addition operation. For example, when the difference in the number of deleted objects and newly added objects is ≤5, the default operational intent of this text editing operation is text replacement, and the voice timestamp of the deleted object in the text is assigned to the voice timestamp of the newly added object in the text, that is, the correspondence between the text and the voice at the text replacement location is retained, and the voice corresponding to the newly added object after the replacement can be played completely during subsequent display without being blocked; when the difference in the number of deleted objects and newly added objects is >5, it is defaulted that the content of the text does not meet expectations, and the operational intent is independent editing, that is, the addition and deletion operations in the text editing operation are independent of each other.
[0108] In the case where the operations included in the text editing operation are independent operations, for the newly added operations that are executed independently, a new object corresponding to the newly added operation can be added to the text. In addition, considering that the newly added object is not related to the voice, the voice timestamp of the newly added object can be set to empty to avoid the newly added object destroying the correspondence between the voice and the text. For example, the text is "This is a paragraph that has not been deleted. Each time it is played, the sound starts from the character after the punctuation mark." Assuming that the new object "This is a compiled paragraph" is added before "Each time it is played, the sound starts from the character after the punctuation mark," the voice timestamps corresponding to the newly added object "This is a compiled paragraph" are all empty. If the subsequent display control operation is to click "This is a compiled paragraph. Each time it is played, the sound starts from the character after the punctuation mark," the voice will still start playing from "Each time it is played, the sound starts from the character after the punctuation mark."
[0109] For independently executed deletion operations, the deletion object corresponding to the deletion operation can be deleted in the text, and the voice timestamp of the deleted object can also be deleted together. The deletion operation on the voice timestamp of the deleted object does not affect the voice timestamps of other objects in the text. For example, the text is "This is a paragraph that has not been deleted. Each time it is played, the sound will be played from the characters after the punctuation mark." Suppose that "each time it is played" is deleted through the deletion operation, then the voice timestamp corresponding to "each time it is played" is also deleted. If the subsequent display control operation is to click "Play the sound from the characters after the punctuation mark", the voice will also be played from "Play the sound from the characters after the punctuation mark".
[0110] The method provided by the embodiment of the present invention determines the operational intention of the text editing operation based on the difference in the number of deleted objects of the deletion operation and the added objects of the new operation when there are new operations and deletion operations at the same position in the corresponding text, thereby ensuring that when the operational intention is to replace the text, the new operations and deletion operations will not damage the correspondence between the voice and the text. The text after text editing can still be used for the synchronous display of voice and text, providing users with a better comparison experience.
[0111] Based on any of the above embodiments, Figure 3 It is a structural diagram of the text editing device provided by the present invention. Figure 3 As shown, the device includes:
[0112] The editing receiving unit 310 is used to receive a text editing operation;
[0113] an intention determination unit 320 for determining, when the text editing operation includes a adding operation and a deleting operation at the same position in the corresponding text, the operation intention of the text editing operation based on a difference in the number of deleted objects of the deleting operation and the number of added objects of the adding operation, and / or a pronunciation similarity between the added objects and the deleted objects;
[0114] The execution unit 330 is used to perform the adding operation and the deleting operation in the text when the operation intention is text replacement, and determine the voice timestamp of the adding object corresponding to the adding operation in the text based on the voice timestamp of the deletion object corresponding to the deletion operation in the text.
[0115] The device provided by an embodiment of the present invention determines the operational intention of a text editing operation based on the difference in quantity between the deleted objects of the deletion operation and the added objects of the addition operation, and / or the pronunciation similarity between the added objects and the deleted objects when there are addition operations and deletion operations at the same position in the corresponding text. This ensures that when the operational intention is to replace text, the addition operation and the deletion operation will not damage the correspondence between voice and text. The text after text editing can still be used for the synchronous display of voice and text, providing users with a better comparison experience.
[0116] Based on any of the above embodiments, the intention determination unit 320 is specifically configured to:
[0117] The operation intention is determined based on the difference in quantity between the deleted objects and the newly added objects, and the pronunciation similarity between the newly added objects and the deleted objects.
[0118] Based on any of the above embodiments, the intention determination unit 320 is specifically configured to:
[0119] When the difference in quantity between the deleted objects and the newly added objects is less than a first threshold, determining that the operation intention is text replacement;
[0120] When the difference in number between the deleted objects and the newly added objects is greater than or equal to the first threshold and less than a second threshold, determining the operation intention based on the pronunciation similarity between the newly added objects and the deleted objects;
[0121] When the difference in quantity between the deleted objects and the newly added objects is greater than or equal to the second threshold, it is determined that the operation intention is independent editing.
[0122] Based on any of the above embodiments, the execution unit 330 is specifically configured to:
[0123] When the number of the deleted objects and the number of the newly added objects are equal, determining a correspondence between the deleted objects and the newly added objects based on an arrangement order of the deleted objects and an arrangement order of the newly added objects, and using the voice timestamps of the deleted objects as the voice timestamps of the corresponding newly added objects;
[0124] When the number of the deleted objects and the number of the newly added objects are not equal, the voice time interval of the deleted object segment is determined based on the voice timestamp of the deleted object, and the voice timestamp of the newly added object is determined based on the voice time interval and the number of the newly added objects.
[0125] Based on any of the above embodiments, the execution unit 330 is further configured to:
[0126] When the text editing operation includes adding and deleting operations corresponding to different positions in the text, or only includes adding or deleting operations, or the operation is intended to be an independent edit, the deleting operation is performed in the text, and / or the adding operation is performed in the text and the voice timestamp of the newly added object is set to empty.
[0127] Based on any of the above embodiments, Figure 4 It is a structural diagram of the display device provided by the present invention, such as Figure 4 As shown, the device includes:
[0128] A presentation receiving unit 410, configured to receive a presentation control operation;
[0129] A display unit 420 for displaying the voice and text based on the display position corresponding to the display control operation and the voice timestamps of each object in the text corresponding to the voice;
[0130] The text is obtained through text editing. When the operation intention of the text editing operation is text replacement, the voice timestamp of the newly added object in the text is determined based on the voice timestamp of the deleted object. The operation intention is, when the text editing operation includes a new operation and a deletion operation at the same position in the corresponding text, based on the difference in quantity between the deleted objects of the deletion operation and the new objects of the new operation, and / or the pronunciation similarity between the new objects and the deleted objects.
[0131] The device provided by the embodiment of the present invention applies the operation intention under text editing to update the voice timestamp of each object in the text, thereby maintaining the correspondence between voice and text, and based on this correspondence, realizes the comparative display of voice and text under display control operation, providing convenience for comparative viewing and comparative editing of the voice system.
[0132] Based on any of the above embodiments, the display unit 420 is specifically configured to:
[0133] determining a presentation start object from the text based on the presentation position;
[0134] When the voice timestamp of the display starting object is not empty, performing voice and text display based on the voice timestamp;
[0135] When the voice timestamp of the display starting object is empty, the start time of the voice display is determined based on the voice timestamp of the first object with a non-empty voice timestamp after the display starting object in the text, voice display is performed based on the start time, and text display is performed based on the display starting object.
[0136] Figure 5 An example of a physical structure diagram of an electronic device is shown below. Figure 5As shown, the electronic device may include: a processor 510, a communication interface 520, a memory 530, and a communication bus 540, wherein the processor 510, the communication interface 520, and the memory 530 communicate with each other via the communication bus 540. The processor 510 may call logic instructions in the memory 530 to execute a text editing method, which includes: receiving a text editing operation; when the text editing operation includes a new operation and a deletion operation corresponding to the same position in the text, determining the operation intention of the text editing operation based on the difference in quantity between the deleted objects of the deletion operation and the added objects of the new operation, and / or the pronunciation similarity between the added objects and the deleted objects; when the operation intention is text replacement, performing the new operation and the deletion operation in the text, and determining the voice timestamp of the new object corresponding to the new operation in the text based on the voice timestamp of the deleted object corresponding to the deletion operation in the text.
[0137] The processor 510 can call the logic instructions in the memory 530 to execute the display method, which includes: receiving a display control operation; performing voice and text display based on the display position corresponding to the display control operation and the voice timestamps of each object in the text corresponding to the voice; the text is obtained through text editing, and when the operation intention of the text editing operation is text replacement, the voice timestamp of the newly added object in the text is determined based on the voice timestamp of the deleted object, and the operation intention is when the text editing operation includes a new operation and a deletion operation at the same position in the corresponding text, based on the difference in quantity between the deleted objects of the deletion operation and the new objects of the new operation, and / or the pronunciation similarity between the new objects and the deleted objects.
[0138] In addition, the logic instructions in the above-mentioned memory 530 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0139] On the other hand, the present invention also provides a computer program product, which includes a computer program, which can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the text editing method provided by the above methods, the method including: receiving a text editing operation; when the text editing operation includes a new operation and a deletion operation at the same position in the corresponding text, determining the operation intention of the text editing operation based on the difference in quantity between the deleted objects of the deletion operation and the new objects of the new operation, and / or the pronunciation similarity between the new objects and the deleted objects; when the operation intention is text replacement, performing the new operation and the deletion operation in the text, and determining the voice timestamp of the new object corresponding to the new operation in the text based on the voice timestamp of the deletion object corresponding to the deletion operation in the text.
[0140] The present invention also provides a computer program product, which includes a computer program, which can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can also execute the display method provided by the above-mentioned methods, the method including: receiving a display control operation; performing voice and text display based on the display position corresponding to the display control operation and the voice timestamps of each object in the text corresponding to the voice; the text is obtained through text editing, and when the operation intention of the text editing operation is text replacement, the voice timestamp of the newly added object in the text is determined based on the voice timestamp of the deleted object, and when the text editing operation includes a new operation and a deletion operation at the same position in the corresponding text, the operation intention is determined based on the difference in quantity between the deleted objects of the deletion operation and the new objects of the new operation, and / or the pronunciation similarity between the new objects and the deleted objects.
[0141] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the text editing method provided by the above-mentioned methods, the method comprising: receiving a text editing operation; in a case where the text editing operation includes a new addition operation and a deletion operation at the same position in the corresponding text, determining the operation intention of the text editing operation based on the difference in quantity between the deleted objects of the deletion operation and the new objects of the new addition operation, and / or the pronunciation similarity between the new objects and the deleted objects; in a case where the operation intention is text replacement, executing the new addition operation and the deletion operation in the text, and determining the voice timestamp of the new object corresponding to the new addition operation in the text based on the voice timestamp of the deletion object corresponding to the deletion operation in the text.
[0142] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the display methods provided by the above-mentioned methods, the method comprising: receiving a display control operation; performing voice and text display based on the display position corresponding to the display control operation and the voice timestamps of each object in the text corresponding to the voice; the text is obtained through text editing, and when the operation intention of the text editing operation is text replacement, the voice timestamp of the newly added object in the text is determined based on the voice timestamp of the deleted object, and when the text editing operation includes a new operation and a deletion operation at the same position in the corresponding text, the operation intention is determined based on the difference in quantity between the deleted objects of the deletion operation and the new objects of the new operation, and / or the pronunciation similarity between the new objects and the deleted objects.
[0143] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.
[0144] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, or of course, by hardware. Based on this understanding, the essence of the above technical solution or the part that contributes to the existing technology can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or certain parts of the embodiments.
[0145] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A text editing method, characterized in that: include: Receive text editing operations; In a case where the text editing operation includes a adding operation and a deleting operation at the same position in the corresponding text, determining the operation intention of the text editing operation based on a difference in the number of deleted objects of the deleting operation and the added objects of the adding operation, and / or a pronunciation similarity between the added objects and the deleted objects; In the case where the operation intention is text replacement, the adding operation and the deleting operation are performed in the text, and based on the voice timestamp of the deletion object corresponding to the deletion operation in the text, the voice timestamp of the adding object corresponding to the adding operation in the text is determined.
2. The text editing method according to claim 1, characterized in that: The determining of the operation intention based on the difference in quantity between the deleted objects and the newly added objects, and / or the pronunciation similarity between the newly added objects and the deleted objects, includes: When the difference in quantity between the deleted objects and the newly added objects is less than a first threshold, determining that the operation intention is text replacement; When the difference in number between the deleted objects and the newly added objects is greater than or equal to the first threshold and less than a second threshold, determining the operation intention based on the pronunciation similarity between the newly added objects and the deleted objects; When the difference in quantity between the deleted objects and the newly added objects is greater than or equal to the second threshold, it is determined that the operation intention is independent editing.
3. The text editing method according to claim 1 or 2, characterized in that: The determining, based on the voice timestamp of the deleted object corresponding to the deletion operation in the text, the voice timestamp of the added object corresponding to the addition operation in the text comprises: When the number of the deleted objects and the number of the newly added objects are equal, determining a correspondence between the deleted objects and the newly added objects based on an arrangement order of the deleted objects and an arrangement order of the newly added objects, and using the voice timestamps of the deleted objects as the voice timestamps of the corresponding newly added objects; When the number of the deleted objects and the number of the newly added objects are not equal, the voice time interval of the deleted object segment is determined based on the voice timestamp of the deleted object, and the voice timestamp of the newly added object is determined based on the voice time interval and the number of the newly added objects.
4. The text editing method according to claim 1 or 2, characterized in that: After receiving the text editing operation, the method further includes: When the text editing operation includes adding and deleting operations corresponding to different positions in the text, or only includes adding or deleting operations, or the operation is intended to be an independent edit, the deleting operation is performed in the text, and / or the adding operation is performed in the text and the voice timestamp of the newly added object is set to empty.
5. A display method, characterized in that: include: Receive display control operations; Performing voice and text display based on the display position corresponding to the display control operation and the voice timestamps of each object in the text corresponding to the voice; The text is obtained through text editing. When the operation intention of the text editing operation is text replacement, the voice timestamp of the newly added object in the text is determined based on the voice timestamp of the deleted object. The operation intention is, when the text editing operation includes a new operation and a deletion operation at the same position in the corresponding text, based on the difference in quantity between the deleted objects of the deletion operation and the new objects of the new operation, and / or the pronunciation similarity between the new objects and the deleted objects.
6. The display method according to claim 5, characterized in that: The performing of voice and text display based on the display position corresponding to the display control operation and the voice timestamp of each object in the text corresponding to the voice includes: determining a presentation start object from the text based on the presentation position; When the voice timestamp of the display starting object is not empty, performing voice and text display based on the voice timestamp; When the voice timestamp of the display starting object is empty, the start time of the voice display is determined based on the voice timestamp of the first object with a non-empty voice timestamp after the display starting object in the text, voice display is performed based on the start time, and text display is performed based on the display starting object.
7. A text editing device, characterized in that: include: An editing receiving unit, used for receiving text editing operations; an intention determining unit for determining, when the text editing operation includes a adding operation and a deleting operation at the same position in the corresponding text, the operation intention of the text editing operation based on a difference in the number of deleted objects of the deleting operation and the added objects of the adding operation, and / or a pronunciation similarity between the added objects and the deleted objects; An execution unit is used to execute the adding operation and the deleting operation in the text when the operation intention is text replacement, and determine the voice timestamp of the adding object corresponding to the adding operation in the text based on the voice timestamp of the deletion object corresponding to the deletion operation in the text.
8. A display device, characterized in that: include: A presentation receiving unit, configured to receive a presentation control operation; A presentation unit, configured to present the voice and text based on a presentation position corresponding to the presentation control operation and a speech timestamp of each object in the text corresponding to the voice; The text is obtained through text editing. When the operation intention of the text editing operation is text replacement, the voice timestamp of the newly added object in the text is determined based on the voice timestamp of the deleted object. The operation intention is, when the text editing operation includes a new operation and a deletion operation at the same position in the corresponding text, based on the difference in quantity between the deleted objects of the deletion operation and the new objects of the new operation, and / or the pronunciation similarity between the new objects and the deleted objects.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the text editing method according to any one of claims 1 to 4 or the display method according to claim 5 or 6 is implemented.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the text editing method according to any one of claims 1 to 4 or the display method according to claim 5 or 6 is implemented.
Citation Information
Patent Citations
Method and device for searching contract modification part, computer equipment and storage medium
CN109933754A
Text editing apparatus and text editing method based on speech signal
US20180018308A1