Recording fusion method, related device, equipment and storage medium

By detecting touch tracks and performing text modification during the speech transcription process, the problems of low information recording efficiency and insufficient customized recording in the existing technology are solved, and efficient customized information recording is achieved.

CN120723136APending Publication Date: 2025-09-30CHENGDU READING & WRITING TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510643147.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-19
Publication Date
2025-09-30

AI Technical Summary

Technical Problem

Existing handwriting input and audio-to-text input methods have shortcomings in information recording efficiency and customized recording needs, and cannot meet the needs of efficient recording and customized modification at the same time.

Method used

During the real-time output of speech transcription, the system detects whether the recording interface of the transcribed text triggers the touch track, and responds to the touch track to modify the text in the editable state, supporting the collaborative work of handwriting input and audio-to-text input.

Benefits of technology

It improves the efficiency of information recording, especially when recording large amounts of content. It supports customized information recording and realizes customized modification of transcribed text through the integration of handwriting input and audio-to-text input.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120723136A_ABST
    Figure CN120723136A_ABST
Patent Text Reader

Abstract

The invention discloses a fusion recording method, a related device, equipment and a storage medium, and the method comprises the steps: detecting whether a recording interface of a transliteration text is triggered to form a touch track or not in a voice transliteration real-time output process, and detecting whether the transliteration text is in a modifiable state or not; and in response to the fact that the transferred text is in the modifiable state and a touch track is detected in the recording interface, performing text modification on the transferred text based on a modification operation matched with the touch track. According to the scheme, the information recording efficiency can be improved, and self-defined information recording is supported.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of human-computer interaction technology, and in particular to a fusion recording method and related devices, equipment and storage media. Background Art

[0002] In modern digital office, learning and other scenarios, people have increasing demands for the efficiency and convenience of information recording.

[0003] Currently, electronic devices such as smartphones and tablets mainly have two recording methods: handwriting input and audio-to-text input. For example, when using the handwriting input recording method, you can use a stylus to write on the electronic screen, but when you need to record a large amount of content, this recording method is relatively inefficient; or, for example, when using the audio-to-text recording method, you can call out the voice transcription function to automatically transcribe the recording into text. Although this is more efficient than handwriting input, it cannot meet the recorder's customized recording needs, and the recorder can only passively accept the recording content. In view of this, how to improve the efficiency of information recording and support customized information recording has become an urgent problem to be solved. Summary of the Invention

[0004] The main technical problem solved by this application is to provide a fusion recording method and related devices, equipment and storage media, which can improve information recording efficiency and support customized information recording.

[0005] In order to solve the above technical problems, the first aspect of the present application provides a fusion recording method, including: in the process of real-time output of speech transcription, detecting whether a touch track is triggered in the recording interface of the transcribed text, and detecting whether the transcribed text is in a modifiable state; in response to detecting a touch track in the recording interface when the transcribed text is in a modifiable state, performing text modification on the transcribed text based on the modification operation matching the touch track.

[0006] In order to solve the above technical problems, the second aspect of the present application provides a fusion recording device, including: a detection module and a modification module, the detection module is used to detect whether a touch track is triggered in the recording interface of the transcribed text during the real-time output of speech transcription, and detect whether the transcribed text is in a modifiable state; the modification module is used to perform text modification on the transcribed text based on a modification operation matching the touch track in response to the transcribed text being in a modifiable state and a touch track being detected in the recording interface.

[0007] In order to solve the above technical problems, the third aspect of the present application provides an electronic device, which includes at least a touch screen, a memory and a processor. The touch screen and the memory are respectively coupled to the processor. The memory stores at least program instructions, and the processor is used to execute the program instructions to implement the fusion recording method in the above first aspect.

[0008] In order to solve the above technical problems, the fourth aspect of the present application provides a computer-readable storage medium storing program instructions that can be executed by a processor, and the program instructions are used to implement the fusion recording method of the first aspect.

[0009] The above scheme, during the process of real-time output of speech transcription, detects whether a touch track is triggered in the recording interface of the transcribed text, and detects whether the transcribed text is in a modifiable state. In response to the transcribed text being in a modifiable state and the touch track being detected in the recording interface, the transcribed text is modified based on the modification operation matching the touch track. Therefore, on the one hand, compared with simply using handwriting input, integrating speech transcription while supporting handwriting input can improve information recording efficiency, especially when a large amount of content needs to be recorded. On the other hand, compared with simply using audio-to-text input, integrating handwriting input while supporting audio-to-text input, and further responding to the touch track in the modifiable state to modify the transcribed text using the modification operation matching it, can support custom information recording. Therefore, by the collaborative work and mutual integration of handwriting input and audio-to-text input, it is possible to support custom modification of the transcribed text by screen touch during the real-time output of speech transcription, which helps to improve information recording efficiency and support custom information recording. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] Figure 1 This is a flow chart of an embodiment of the fusion recording method of the present application; Figure 2a This is a schematic diagram showing the effect of an embodiment of the fusion recording method of the present application; Figure 2b This is a schematic diagram showing the effect of an embodiment of the deletion operation of the present application; Figure 2c This is a schematic diagram showing the effect of an embodiment of the exchange operation of the present application; Figure 2d This is a schematic diagram showing the effect of an embodiment of the highlight operation of the present application; Figure 2e This is a schematic diagram showing the effect of adding an indentation operation in an embodiment of the present application; Figure 2f This is a schematic diagram of the effect of an embodiment of the indentation reduction operation of the present application; Figure 2g This is a schematic diagram showing the effect of an embodiment of the line break operation of the present application; Figure 2h This is a schematic diagram showing the effect of an embodiment of canceling the line feed operation of the present application; Figure 2i This is a schematic diagram showing the effect of an embodiment of the text insertion operation of this application; Figure 2j This is a schematic diagram of the effect of a touch track embodiment when the transcribed text of the present application is in an unmodifiable state; Figure 2k This is a schematic diagram of the effect of an embodiment of transcribing historical sentences and latest sentences in the text of this application; Figure 3 This is a schematic diagram of the framework of an embodiment of the fusion recording device of the present application; Figure 4 This is a schematic diagram of the framework of an embodiment of the electronic device of the present application; Figure 5 It is a schematic diagram of a framework of an embodiment of a computer-readable storage medium of the present application. DETAILED DESCRIPTION

[0011] The following describes the embodiments of the present application in detail with reference to the accompanying drawings.

[0012] In the following description, for the purpose of explanation rather than limitation, specific details such as specific system structures, interfaces, and technologies are provided to facilitate a thorough understanding of the present application.

[0013] The terms "system" and "network" are often used interchangeably in this document. The term "and / or" is simply a description of an association between related objects, indicating that three possible relationships exist. For example, "A and / or B" can mean: A exists alone, A and B exist simultaneously, or B exists alone. Furthermore, the fragment " / " generally indicates that the related objects are in an "or" relationship. Furthermore, "multiple" in this document refers to two or more than two.

[0014] See also Figure 1 , Figure 1 This is a flow chart of an embodiment of the fusion recording method of the present application. Specifically, it may include the following steps: Step S11: During the real-time output of the voice transcription, it is detected whether a touch track is triggered in the recording interface of the transcribed text, and whether the transcribed text is in a modifiable state.

[0015] In one implementation scenario, as a possible implementation method, a preset icon can be set at a fixed position on the recording interface. The preset icon is used to switch between starting and stopping speech transcription. For example, the preset icon can be fixedly set at the top or bottom of the recording interface. When the preset icon is clicked and triggered, speech transcription can be started. At this time, the transcribed text output by the speech transcription can be updated in real time on the recording interface. When the preset icon is clicked and triggered again, speech transcription can be turned off. At this time, the transcribed text can be stopped from being updated on the recording interface.

[0016] In another implementation scenario, as another possible implementation method, a preset icon can also be displayed on the recording interface along with the text line updated in real time by the voice transcription, and the preset icon is used to switch between starting and closing the voice transcription. It should be noted that the specific process of using the preset icon to switch the start and close of the voice transcription can be referred to the above-mentioned related description, which will not be repeated here. For ease of understanding, please refer to Figure 2a , Figure 2a This is a schematic diagram of the effect of an embodiment of the fusion recording method of the present application. Figure 2a The text line is marked in the dotted box, such as Figure 2a As shown in the middle left figure, the record interface can contain the text content "Today, we gather together to discuss artificial intelligence technology." Since the preset icon follows the latest text line, the preset icon is now located in the text line "Artificial Intelligence Technology." Figure 2a As shown in the middle right image, when the speech transcription produces a new line of text, "Artificial intelligence technology has developed rapidly.", the preset icon now appears on the text line "Artificial intelligence technology has developed rapidly." Of course, the above example is only one possible example in actual application, and other possible scenarios will not be given one by one here. In addition, the preset icon can be a static icon, or the preset icon can also be a dynamic icon. For example, when speech transcription is turned off, the preset icon following the text line can be a variation of the dynamic icon, and when speech transcription is turned on, the preset icon following the text line can be another variation of the dynamic icon.

[0017] It should be noted that the embodiment of the present disclosure is mainly used to support the integration of handwritten records in the process of voice transcription, so as to solve the problem that it is impossible to take into account both recording efficiency and record customization during the voice transcription process. However, this does not mean that the transcribed text can no longer be customized (such as custom modifications, etc.) after the voice transcription is completed. For example, after the voice transcription is completely completed (such as, in the case of application to in-class recording, since the teacher's lecture usually ends at the end of class, it can be regarded as the complete completion of the voice transcription), it is still possible to continue to detect whether the touch track is triggered in the recording interface of the transcribed text, and detect whether the transcribed text is in a modifiable state, as well as the subsequent steps in the embodiment of the present disclosure.

[0018] In an implementation scenario, please refer to Figure 2a As a possible implementation, the recording interface may contain only one operation window, which may be used to display the transcribed text, and the touch track may also be triggered to form in the operation window. In other words, the recording interface may not need to be divided into multiple operation windows, so that different operation windows can each perform different functions, such as one operation window is used to display the transcribed text, and another operation window is used to sense handwritten text. Of course, the above example is only a possible example of the recording interface in actual application, and other possible situations will not be given one by one here.

[0019] In one implementation scenario, the touch track can be formed by a stylus touching the screen, and the stylus can be connected to the electronic device that performs voice transcription. It should be noted that the writing pen may include but is not limited to: resistive pressure type, electromagnetic pressure sensing type, capacitive touch type, etc., and the specific type of the stylus is not limited here. In addition, the stylus can be connected to the electronic device through Bluetooth or 2.4G protocol to achieve wireless communication. Of course, the stylus can also be connected to the electronic device through wired communication such as USB. The communication method between the stylus and the electronic device is not limited here. Of course, the touch track can be triggered by other tools besides the stylus, such as the touch track can be triggered by tools such as fingers and ordinary pens, which are not limited here.

[0020] In one implementation scenario, when the stylus triggers the generation of a modification instruction, the electronic device can enter a modifiable state in response to the modification instruction; conversely, when the stylus does not trigger the generation of a modification instruction (or cancels the modification instruction), the electronic device can enter an unmodifiable state in response thereto. It should be noted that when the electronic device enters the modifiable state, the electronic device allows text modification of the transcribed text through the touch track, and when the electronic device enters the unmodifiable state, the electronic device does not allow text modification of the transcribed text through the touch track. In addition, text modification means modifying the transcribed text itself, such as including but not limited to: modifying the text content such as adding or deleting text, modifying the text layout such as line breaks and indents, and modifying the text style such as font size and color. The specific situations of text modification will not be given one by one here. For example, in the case where the touch track is formed by the stylus touching the screen, the stylus can generate a modification instruction when a key trigger action is detected. For example, the stylus pen may be provided with a trigger switch, such as a button. When the trigger switch is triggered (e.g., the button is pressed), it can be determined that a modification instruction has been generated. Conversely, when the trigger switch is de-activated (e.g., the button is pressed again), it can be determined that the modification instruction has not been generated or that the modification instruction has been canceled. Of course, the above example is only one possible example of triggering or canceling the generation of a modification instruction, and other methods are not limited here, and will not be given one by one.

[0021] Step S12: In response to the transcribed text being in a modifiable state and a touch trajectory being detected in the recording interface, perform text modification on the transcribed text based on a modification operation that matches the touch trajectory.

[0022] Specifically, when the transcribed text is in a modifiable state, if a touch trajectory is detected in the recording interface, it indicates that the user currently expects to trigger the modification of the transcribed text. At this time, text modification can be performed on the transcribed text based on a modification operation that matches the touch trajectory. On the contrary, although the transcribed text is in a modifiable state, if no touch trajectory is detected in the recording interface, it means that the user currently does not expect to trigger the modification of the transcribed text, and thus the transcribed text can be left unmodified. It should be noted that when using the touch trajectory to perform text modification on the transcribed text, the speech transcription still continues without pausing. On the one hand, a modification operation that matches the touch trajectory can be determined as the target operation, and on the other hand, according to the trigger position of the touch trajectory in the recording interface, the text range targeted by the target operation in the transcribed text can be determined. Then, text modification can be performed on the determined text range according to the target operation. Of course, different touch trajectories can match different modification operations. Correspondingly, different trigger positions can also determine different text ranges. As long as different touch trajectories do not cause confusion, the specific matching between the touch trajectory and the modification operation and the specific correspondence between the trigger position and the text range are not specifically limited in the embodiments of the present disclosure. For example, the specific matching between the touch trajectory and the modification operation and the specific correspondence between the trigger position and the text range can be designed for the convenience of user operation and memory. For the sake of understanding, several different modification operations will be exemplified below, but this does not mean being limited to the following several situations, and custom adjustments can also be made by referring to the following examples.

[0023] In an implementation scenario, as a possible example, in response to the touch trajectory being a horizontal line continuously passing through several characters in the transcribed text, a modification operation that matches the touch trajectory can be determined as a target operation of the deletion operation type, and the characters continuously passed by the touch trajectory can be selected as the text range targeted by the target operation. At this time, the text range targeted by the target operation includes the range to be deleted targeted by the deletion operation, and then text modification is performed on the range to be deleted according to the deletion operation. Please refer to Figure 2b , Figure 2b is a schematic diagram of the effect of an embodiment of the deletion operation of this application. As shown in the left figure of Figure 2b As shown in the left figure, the red dotted line represents the touch trajectory, which continuously passes through several characters "hot" in the transcribed text. Then, a modification operation that matches the touch trajectory can be determined as the deletion operation, and the characters "hot" continuously passed by the touch trajectory can be selected as the range to be deleted targeted by the deletion operation. Finally, text modification can be performed on the range to be deleted "hot" according to the deletion operation, resulting in Figure 2b The middle right image shows the transcribed text after text modification: "Artificial intelligence technology has developed." It should be noted that the "horizontal line" described in the embodiments of this disclosure does not require the touch track to be completely horizontal; it can also be tilted at least partially, as long as it extends horizontally overall. Of course, the above example is only one possible example of a deletion operation in actual application, and other possible scenarios will not be given one by one here.

[0024] In one implementation scenario, as another possible example, in response to the touch track being a step-type line, the modification operation that matches the touch track can be determined to be a target operation of the swap type, and the step-type line includes at least two horizontal lines that form a step, and the text covered by the two horizontal lines in the touch track is selected as the text range targeted by the target operation, and at this time the text range targeted by the target operation includes the two to-be-exchanged ranges targeted by the swap operation, and then the text modification is performed on the two to-be-exchanged ranges according to the swap operation. It should be noted that the specific meaning of "horizontal line" can be found in the aforementioned related description and will not be repeated here. In addition, the vertical line connecting the two horizontal lines in the step-type line is not required to be completely vertical, and may also be tilted in at least part of the position, as long as it presents a longitudinal extension as a whole. Please refer to Figure 2c , Figure 2c This is a schematic diagram showing the effect of an embodiment of the switching operation of this application. Figure 2c As shown in the middle left figure, the red dotted line shows the touch trajectory, which is at least two horizontal lines that form a step. One horizontal line covers the word "hot" as one of the ranges to be exchanged, and the other horizontal line covers the word "development" as the other range to be exchanged. Finally, the text of the exchange range "hot" and the range to be exchanged "development" can be modified according to the exchange operation, and the result is as follows Figure 2c The middle-right image shows the resulting text after text modification: "Artificial intelligence technology has developed rapidly." Of course, the above example is only one possible example of a swap operation in actual application, and other possible scenarios will not be listed here.

[0025] In an implementation scenario, as another possible example, in response to the touch trajectory being a closed line that encloses several characters in the transcribed text, it is possible to determine the modification operation that matches the touch trajectory as the target operation of the type of highlighting operation, and select the characters enclosed by the touch trajectory as the text range targeted by the target operation. At this time, the text range targeted by the target operation includes the range to be highlighted targeted by the highlighting operation. Then, perform text modification on the range to be highlighted according to the highlighting operation. It should be noted that the specific shape of the "closed line" is not limited here. For example, the "closed line" can be a regular shape such as a rounded rectangle, a right-angled rectangle, an oval, etc., or it can also be an irregular shape. No further examples of the specific shape of the "closed line" will be given here. Please refer to Figure 2d , Figure 2d is a schematic diagram of the effect of an embodiment of the highlighting operation of this application. As shown in the left figure of Figure 2d , the red dotted line shown is the touch trajectory, which is a closed line that encloses several characters "hot" in the transcribed text. Then, it can be determined that the modification operation that matches the touch trajectory is the highlighting operation, and select the characters "hot" it encloses as the range to be highlighted. Finally, perform text modification on the range to be highlighted "hot" according to the highlighting operation, and obtain the transcribed text "Artificial intelligence technology has developed rapidly." after text modification as shown in the right figure of Figure 2d (where the character "hot" has a highlighted background). Of course, the above example is only one possible example of the highlighting operation in the actual application process, and no further examples of other possible situations will be given here.

[0026] In an implementation scenario, as another possible example, in response to the touch trajectory being a vertical line located at the beginning of a text line in the transcribed text and in the first direction, it is possible to determine the modification operation that matches the touch trajectory as the target operation of the type of increasing indentation operation, and select the text line where the touch trajectory is located as the text range targeted by the target operation. At this time, the text range targeted by the target operation includes the range to be indented targeted by the increasing indentation operation. Then, perform text modification on the range to be indented according to the increasing indentation operation. It should be noted that the first direction can be set according to actual application needs. For example, it can be set according to the user's operation habit, and the first direction can be the upward direction. That is to say, when the touch trajectory is located at the beginning of a text line in the transcribed text and is a vertical line extending upward, it can be determined that the modification operation that matches the touch trajectory is the increasing indentation operation. In addition, the specific meaning of the "vertical line" is similar to that of the aforementioned "horizontal line". In the actual application process, it is not required to be completely vertical, and there can also be an inclination at least in some positions, as long as it shows a vertical extension as a whole. Please refer to Figure 2e , Figure 2e is a schematic diagram of the effect of an embodiment of the increasing indentation operation of this application. Taking the first direction as the "upward direction" as an example, as shown in Figure 2eAs shown in the middle left picture, the red dotted line shows the touch track, and the arrow indicates the extension direction of the touch track. Figure 2e The touch track shown in the middle left image is a vertical line at the beginning of the text line "Artificial intelligence technology has developed rapidly" and pointing in the first direction. Therefore, it can be determined that the modification operation matching the touch track is an increase indentation operation. Then, the text line where the touch track is located, namely "Artificial intelligence technology has developed rapidly", can be selected as the range to be increased or decreased for the increase indentation operation. Finally, the text modification can be performed on the increase or decrease range "Artificial intelligence technology has developed rapidly" according to the increase indentation operation, and the result is as follows: Figure 2e The middle right image shows "artificial intelligence technology has developed rapidly" with a 2-character indent at the beginning of a line. It should be noted that the number of indent characters for the "increase indent operation" can be unlimited, such as "2 characters" in the above example, or it can be set to another number of characters, which is not limited here. In addition, after the increase indent operation is performed on a text line, other text lines belonging to the same paragraph can automatically adjust their positions to avoid text overlap. Of course, the above example is only one possible example of the increase indent operation in actual application, and other possible situations will not be given examples one by one here.

[0027] In one implementation scenario, as another possible example, in response to the touch track being a vertical line located at the beginning of the first text line in the transcribed text and facing the second direction, the modification operation that matches the touch track can be determined to be a target operation of the type of reduce indentation operation. It should be noted that the first text line itself has an indented text line, and the specific number of characters indented in the first text line is not limited here, such as 2 characters. In addition, the second direction can be opposite to the aforementioned first direction. For example, when the aforementioned first direction is an upward direction, the second direction can be a downward direction. On this basis, the text line where the touch track is located can be selected as the text range targeted by the target operation, and at this time the text range targeted by the target operation includes the range to be reduced targeted by the reduce indentation operation, and then the text modification is performed on the range to be reduced according to the reduce indentation operation. Please refer to Figure 2f , Figure 2f This is a schematic diagram of the effect of reducing the indentation operation of an embodiment of the present application. Taking the second direction as the "downward direction" as an example, Figure 2f As shown in the middle left picture, the red dotted line shows the touch track, and the arrow indicates the extension direction of the touch track. Figure 2f The touch track shown in the middle left image is a vertical line located at the beginning of the text line "Artificial intelligence technology has developed rapidly" and pointing in the second direction. Therefore, it can be determined that the modification operation matching the touch track is a reduce indentation operation. Then, the text line where the touch track is located, namely "Artificial intelligence technology has developed rapidly", can be selected as the range to be reduced for the reduce indentation operation. Finally, the text modification can be performed on the reduced range "Artificial intelligence technology has developed rapidly" according to the reduce indentation operation, and the result is as follows: Figure 2f The middle-right image shows the "artificial intelligence technology that has been rapidly developing" for reducing the indentation of a line by 2 characters. It should be noted that the number of characters reduced in the "reduce indentation operation" can be unlimited, such as "2 characters" in the above example, or it can be set to another number of characters, which is not limited here. In addition, after the text line performs the reduce indentation operation, other text lines belonging to the same paragraph can automatically adjust their positions to avoid blank areas between the characters. Of course, the above example is only one possible example of the reduce indentation operation in actual application, and other possible situations will not be given examples one by one here.

[0028] In one implementation scenario, as another possible example, in response to the touch track being located on a vertical line in a transcribed text line, the modification operation that matches the touch track can be determined to be a target operation of the type of line break operation. It should be noted that the touch track that matches the line break operation may not require its extension direction to be a specific direction in the actual application process, such as it can be either "upward" or "downward". As a special example, in order to be unified with the operational logic of the aforementioned "increase indentation operation", the extension direction of the touch track that matches the line break operation may be required to be consistent with the trajectory direction of the aforementioned "increase indentation operation", that is, the first direction. On this basis, the text line where the touch track is located can be selected as the text range targeted by the target operation, and at this time, the text range targeted by the target operation includes the range to be broken targeted by the line break operation, and then at the trigger position of the touch track, the text modification is performed on the range to be broken according to the line break operation. Please refer to Figure 2g , Figure 2g This is a schematic diagram showing the effect of an embodiment of the line break operation of this application. Figure 2g As shown in the middle left picture, the red dotted line shows the touch track. Figure 2g The touch track shown in the middle left image is a vertical line in the text line "Artificial intelligence technology has developed rapidly." The modification operation matching the touch track can be determined to be a line break operation, and the text line "Artificial intelligence technology has developed rapidly" where the touch track is located is selected as the line break range for the line break operation. At the trigger position of the touch track, that is, between "Artificial intelligence technology" and "has developed rapidly," the text modification is performed according to the line break operation, resulting in the following result: Figure 2g The middle-right image shows the transcribed text after text modification: "Artificial intelligence technology (line break here) has developed rapidly." Of course, the above example is only one possible example of line break operation in actual application, and other possible scenarios will not be listed here.

[0029] In an implementation scenario, as another possible example, in response to the touch trajectory being a vertical line at the beginning of the second text line in the transcribed text and in the second direction, it can be determined that the modification operation matching the touch trajectory is a target operation of the type of line break cancellation operation. It should be noted that the second text line is a text line without indentation. In addition, for the specific meaning of the "second direction", reference can be made to the relevant description of the "indent reduction operation" mentioned above, which will not be elaborated here. On this basis, the text line where the touch trajectory is located can be selected as the text range targeted by the target operation, and at this time, the text range targeted by the target operation includes the range to be cancelled targeted by the line break cancellation operation, and then the text modification is performed on the range to be cancelled according to the line break cancellation operation. Please refer to Figure 2h , Figure 2h is a schematic diagram of the effect of an embodiment of the line break cancellation operation in this application. As Figure 2h shown in the left figure in, taking the "second direction" as the "upward direction" as an example, the red dotted line represents the touch trajectory. Obviously Figure 2h the touch trajectory shown in the left figure in is a vertical line at the beginning of a text line without indentation and in the upward direction. Therefore, it can be determined that the modification operation matching the touch trajectory is a line break cancellation operation. Then, the text line "has developed vigorously." where the touch trajectory is located can be selected as the range to be cancelled targeted by the line break cancellation operation. Finally, the text modification is performed on the range to be cancelled "has developed vigorously." according to the line break cancellation operation, and the transcribed text "Artificial intelligence technology has developed vigorously." after the text modification as shown in Figure 2h the right figure in is obtained. Of course, the above example is only one possible example of the line break cancellation operation in the actual application process, and other possible situations will not be exemplified one by one here.

[0030] In an implementation scenario, as another possible example, in response to the touch trajectory being a trajectory point at any position in the transcribed text, it can be determined that the modification operation matching the touch trajectory is a target operation of the type of text insertion operation, and the trigger position of the touch trajectory is selected as the text range targeted by the target operation. At this time, the text range targeted by the target operation includes the range to be inserted targeted by the text insertion operation, and then the text modification is performed on the range to be inserted according to the text insertion operation. Please refer to Figure 2i ,​​​​​The touch trajectory shown in the left middle figure is the trajectory point between "already" and "hot" in the transcribed text "Artificial intelligence technology has been booming". Therefore, it can be determined that the modification operation matching the touch trajectory is a text insertion operation. Then, the trigger position of the touch trajectory, that is, between "already" and "hot", can be selected as the range to be inserted for the text insertion operation. Finally, the text modification is performed on the range to be inserted according to the text insertion operation. It should be noted that the text insertion operation can be performed by voice input in the range to be inserted, that is, at this time, the user's voice can be collected, transcribed and inserted, and the input method of the inserted text is not limited here. For example, the inserted text can also be input by handwriting recognition input, soft keyboard input, etc., and no further examples will be given here. Please continue to refer to Figure 2i , such as Figure 2i As shown in the right middle figure, the inserted text "more and more" can be input between "already" and "hot". Finally, the transcribed text after performing the text insertion operation can be obtained as "Artificial intelligence technology has become more and more booming." In addition, the display style of the inserted text in the recording interface can also be distinguished from the transcribed text. For example, the transcribed text can be in black, and the inserted text can be in blue, or the transcribed text can be in regular script, and the inserted text can be in boldface, or the transcribed text can be in normal point size, and the inserted text can be in bold, or the transcribed text can be in roman, and the inserted text can be in italic. Of course, the above examples are only one possible example of the text insertion operation in the actual application process, and no further examples of other possible situations will be given here.

[0031] It should be noted that the above examples are only several possible examples of the modification operation, and the modification operation is not limited to the several examples shown above. For example, other types of modification operations can also be added according to the actual application needs, or the modification operations shown above can be adjusted adaptively. Here, no further examples of the modification operation will be given.

[0032] In an implementation scenario, the transcribed text and the touch trajectory can be set in layers. For example, the transcribed text can be located in the text layer, and a writing layer can be set above the text layer. Then, when in the modifiable state, the touch trajectory can be non-persistently displayed on the writing layer. For example, Figures 2b to 2i As shown, when in the modifiable state, during the formation of the touch trajectory triggered, the touch trajectory can be displayed in a preset style (for example, the color is specified as "red", the line type is specified as "dashed line", etc.), and after the text modification is performed on the transcribed text, the touch trajectory automatically disappears. It should be noted that the transparency of the writing layer can be higher than the transparency of the text layer, so that the transcribed text on the text layer below the writing layer can be observed by the user through the writing layer.

[0033] In one implementation scenario, in actual application, the transcribed text may be in an unmodifiable state in addition to being in an editable state. It should be noted that the switching method of the transcribed text between the "editable state" and the "unmodifiable state" can refer to the aforementioned related descriptions and will not be repeated here. On this basis, in response to the touch track being detected on the recording interface when the transcribed text is in an unmodifiable state, the touch track can be persistently displayed on the writing layer, so that graffiti annotations can be made to the transcribed text when the transcribed text is in an unmodifiable state. Please refer to Figure 2j , Figure 2j This is a schematic diagram of the effect of the touch track embodiment when the transcribed text of this application is in an unmodifiable state. Figure 2j As shown, when the transcribed text is in an unmodifiable state, when a touch track is detected on the recording interface (such as Figure 2j In other words, when the touch track is detected in the unmodifiable state, the transcribed text in the text layer will not be modified accordingly, but the touch track will only be persistently displayed in the writing layer above the text layer. Of course, "persistent display" in the unmodifiable state is only relative to "non-persistent display" in the modifiable state. For example, when an operation such as erasing is performed on the persistently displayed touch track, the touch track can also disappear in the writing layer accordingly. It should be noted that after the touch track is triggered in the writing layer in the unmodifiable state, if it is switched to the modifiable state and a modification operation (such as deletion, increase indentation, decrease indentation, line break, cancel line break, text insertion, etc.) is triggered, the transcribed text will change accordingly, and the touch track originally triggered in the writing layer in the unmodifiable state can be adjusted accordingly.

[0034] In one implementation scenario, the latest sentence in the transcribed text (such as the sentence currently being transcribed) can be distinguished from the historical sentences before the latest sentence in the transcribed text and displayed in a different style. For example, the historical sentences can be displayed in black, while the latest sentence can be displayed in blue; or the historical sentences can be displayed in italics, while the latest sentence can be displayed in bold; or the historical sentences can be displayed in regular weight, while the latest sentence can be displayed in bold; or the historical sentences can be displayed in regular weight, while the latest sentence can be displayed in italics. For easier understanding, please refer to Figure 2k , Figure 2k This is a schematic diagram showing the effect of an embodiment of transcribing historical sentences and the latest sentences in the text of this application. Figure 2k As shown, historical sentences can be expressed in black italics, while the latest sentences can be expressed in blue italics. Figure 2kThe example shown is only one possible example of distinguishing historical statements from latest statements in actual application. Other possible methods are not limited here and will not be given examples one by one.

[0035] The above scheme, during the process of real-time output of speech transcription, detects whether a touch track is triggered in the recording interface of the transcribed text, and detects whether the transcribed text is in a modifiable state. In response to the transcribed text being in a modifiable state and the touch track being detected in the recording interface, the transcribed text is modified based on the modification operation matching the touch track. Therefore, on the one hand, compared with simply using handwriting input, integrating speech transcription while supporting handwriting input can improve information recording efficiency, especially when a large amount of content needs to be recorded. On the other hand, compared with simply using audio-to-text input, integrating handwriting input while supporting audio-to-text input, and further responding to the touch track in the modifiable state to modify the transcribed text using the modification operation matching it, can support custom information recording. Therefore, by the collaborative work and mutual integration of handwriting input and audio-to-text input, it is possible to support custom modification of the transcribed text by screen touch during the real-time output of speech transcription, which helps to improve information recording efficiency and support custom information recording.

[0036] See also Figure 3 , Figure 3 The figure is a schematic diagram of a framework of an embodiment of a fusion recording device of the present application. The fusion recording device 30 includes: a detection module 31 and a modification module 32. The detection module 31 is used to detect whether a touch track is triggered in the recording interface of the transcribed text during the real-time output of speech transcription, and to detect whether the transcribed text is in a modifiable state; the modification module 32 is used to perform text modification on the transcribed text based on a modification operation matching the touch track in response to the transcribed text being in a modifiable state and the touch track being detected in the recording interface.

[0037] In the above scheme, the fusion recording device 30 detects whether a touch track is triggered in the recording interface of the transcribed text during the real-time output of the voice transcription, and detects whether the transcribed text is in a modifiable state. In response to the transcribed text being in a modifiable state and the touch track being detected in the recording interface, the transcribed text is modified based on the modification operation matching the touch track. Therefore, on the one hand, compared with simply using handwriting input, integrating voice transcription while supporting handwriting input can improve information recording efficiency, especially when a large amount of content needs to be recorded. On the other hand, compared with simply using audio-to-text input, integrating handwriting input while supporting audio-to-text input, and further responding to the touch track in the modifiable state, so as to modify the transcribed text using the modification operation matching it, can support custom information recording. Therefore, by working together and integrating handwriting input and audio-to-text input, it is possible to support custom modification of the transcribed text by screen touch during the real-time output of voice transcription, which helps to improve information recording efficiency and support custom information recording.

[0038] In some disclosed embodiments, the modification module 32 includes an operation determination submodule for determining a modification operation that matches the touch trajectory as a target operation; the modification module 32 includes a range determination submodule for determining the text range targeted by the target operation in the transcribed text based on the trigger position of the touch trajectory in the recording interface; the modification module 32 includes a text modification submodule for performing text modification on the text range according to the target operation.

[0039] In some disclosed embodiments, the modification module 32 includes a first response submodule for determining, in response to the touch trajectory being a horizontal line that continuously passes through several characters in the transcribed text, a modification operation that matches the touch trajectory, which is a target operation of the deletion type; the modification module 32 includes a first determination submodule for selecting the characters that the touch trajectory continuously passes through as the text range targeted by the target operation, and the text range targeted by the target operation includes the range to be deleted targeted by the deletion operation; the modification module 32 includes a first modification submodule for performing text modification on the range to be deleted according to the deletion operation.

[0040] In some disclosed embodiments, the modification module 32 includes a second response submodule for determining, in response to the touch trajectory being a step-type line, a modification operation that matches the touch trajectory, as a target operation of the swap type; wherein the step-type line includes at least two horizontal lines that form a step; the modification module 32 includes a second determination submodule for selecting the text respectively covered by the two horizontal lines in the touch trajectory as the text range targeted by the target operation, the text range targeted by the target operation including the two ranges to be swapped targeted by the swap operation; the modification module 32 includes a second modification submodule for performing text modification on the two ranges to be swapped according to the swap operation.

[0041] In some disclosed embodiments, the modification module 32 includes a third response submodule for determining, in response to the touch trajectory being a closed line surrounding several characters in the transcribed text, a modification operation that matches the touch trajectory, as a target operation of the highlight operation type; the modification module 32 includes a third determination submodule for selecting the characters surrounded by the touch trajectory as the text range targeted by the target operation, the text range targeted by the target operation including the range to be highlighted targeted by the highlight operation; the modification module 32 includes a third modification submodule for performing text modification on the highlighted range according to the highlight operation.

[0042] In some disclosed embodiments, the modification module 32 includes a fourth response submodule for determining, in response to the touch track being a vertical line located at the beginning of a text line in the transcribed text and facing a first direction, a modification operation that matches the touch track, which is a target operation of the type of increase indentation operation; the modification module 32 includes a fourth determination submodule for selecting the text line where the touch track is located as the text range targeted by the target operation, the text range targeted by the target operation including the range to be increased or decreased targeted by the increase indentation operation; the modification module 32 includes a fourth modification submodule for performing text modification on the range to be increased or decreased according to the increase indentation operation.

[0043] In some disclosed embodiments, the modification module 32 includes a fifth response submodule for determining, in response to the touch trajectory being a vertical line located at the beginning of the first text line in the transcribed text and facing the second direction, a modification operation that matches the touch trajectory, which is a target operation of the type of reduce indentation operation; wherein, the first text line has an indentation; the modification module 32 includes a fifth determination submodule for selecting the text line where the touch trajectory is located as the text range targeted by the target operation, and the text range targeted by the target operation includes the range to be reduced targeted by the reduce indentation operation; the modification module 32 includes a fifth modification submodule for performing text modification on the range to be reduced according to the reduce indentation operation.

[0044] In some disclosed embodiments, the modification module 32 includes a sixth response submodule for determining, in response to the touch trajectory being a vertical line located in a text line in the transcribed text, a modification operation that matches the touch trajectory, as a target operation of the type of line break operation; the modification module 32 includes a sixth determination submodule for selecting the text line where the touch trajectory is located as the text range targeted by the target operation, the text range targeted by the target operation including the range to be broken for the line break operation; the modification module 32 includes a sixth modification submodule for performing text modification on the range to be broken according to the line break operation at the trigger position of the touch trajectory.

[0045] In some disclosed embodiments, the modification module 32 includes a seventh response submodule for determining, in response to the touch track being a vertical line located at the beginning of the second text line in the transcribed text and facing the second direction, a modification operation that matches the touch track, which is a target operation of the type of cancel line break operation; wherein, the second text line is not indented; the modification module 32 includes a seventh determination submodule for selecting the text line where the touch track is located as the text range targeted by the target operation, and the text range targeted by the target operation includes the range to be canceled targeted by the cancel line break operation; the modification module 32 includes a seventh modification submodule for performing text modification on the range to be canceled according to the cancel line break operation.

[0046] In some disclosed embodiments, the modification module 32 includes an eighth response submodule for determining, in response to the touch trajectory being a trajectory point located at any position in the transcribed text, a modification operation that matches the touch trajectory, which is a target operation of the type of text insertion operation; the modification module 32 includes an eighth determination submodule for selecting the trigger position of the touch trajectory as the text range targeted by the target operation, the text range targeted by the target operation including the range to be inserted targeted by the text insertion operation; the modification module 32 includes an eighth modification submodule for performing text modification in accordance with the text insertion operation in the range to be inserted.

[0047] In some disclosed embodiments, the transcribed text is located in a text layer, a writing layer is provided above the text layer, the touch track is non-persistently displayed on the writing layer when in a modifiable state, and the transparency of the writing layer is higher than that of the text layer.

[0048] In some disclosed embodiments, the fusion recording device 30 includes a graffiti writing module for persistently displaying a touch track on the writing layer in response to the transcribed text being in an unmodifiable state and a touch track being detected in the recording interface.

[0049] In some disclosed embodiments, when in a modifiable state, during the process of triggering the formation of a touch track, the touch track is displayed in a preset style, and after the text modification is performed on the transcribed text, the touch track automatically disappears.

[0050] In some disclosed embodiments, the fusion recording device 30 includes an icon following module for displaying a preset icon on the recording interface following the text line updated in real time by the voice transcription; wherein the preset icon is used to switch between starting and stopping the voice transcription.

[0051] In some disclosed embodiments, the touch trajectory is formed by a stylus touching the screen, the stylus is in communication with an electronic device that performs voice transcription, and when the stylus triggers the generation of a modification instruction, the electronic device enters a modifiable state in response to the modification instruction; and / or, the recording interface contains only one operation window, the operation window is used to display the transcribed text, and the touch trajectory is triggered to be formed in the operation window.

[0052] See also Figure 4 , Figure 4 It is a schematic diagram of the framework of an embodiment of an electronic device of the present application. The electronic device 40 at least includes a touch screen 41, a memory 42 and a processor 43. The touch screen 41 and the memory 42 are respectively coupled to the processor 43. The memory 42 stores at least program instructions. The processor 43 is used to execute the program instructions to implement the steps in any of the above-mentioned fusion recording method embodiments, and the graph data is displayed on the touch screen 41. For details, please refer to the aforementioned disclosed embodiments, which will not be repeated here. As a possible example, the touch screen 41 can be an ink screen. In this case, the electronic device 40 can include but is not limited to e-books, ink screen office books, ink screen displays, ink screen mobile phones and other ink screen devices; or, as a possible example, the touch screen 41 can include but is not limited to non-ink screens such as LCD and OLED. In this case, the electronic device 40 can include but is not limited to displays, tablet computers, learning machines, smart large screens and other devices. The specific types of the touch screen 41 and the electronic device 40 are not limited here.

[0053] Specifically, the processor 43 is used to control itself, the touch screen 41, and the memory 42 to implement the steps in any of the above-mentioned fusion recording method embodiments. The processor 43 can also be called a CPU (Central Processing Unit). The processor 43 may be an integrated circuit chip with signal processing capabilities. The processor 43 can also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The general-purpose processor can be a microprocessor or any conventional processor. In addition, the processor 43 can be implemented by an integrated circuit chip.

[0054] In the above scheme, the electronic device 40 detects whether a touch track is triggered in the recording interface of the transcribed text during the real-time output of voice transcription, and detects whether the transcribed text is in a modifiable state. In response to the transcribed text being in a modifiable state and the touch track being detected in the recording interface, the transcribed text is modified based on the modification operation matching the touch track. Therefore, on the one hand, compared with simply using handwriting input, integrating voice transcription while supporting handwriting input can improve information recording efficiency, especially when a large amount of content needs to be recorded. On the other hand, compared with simply using audio-to-text input, integrating handwriting input while supporting audio-to-text input, and further responding to the touch track in the modifiable state, so as to modify the transcribed text using the modification operation matching it, can support custom information recording. Therefore, by working together and integrating handwriting input and audio-to-text input, it is possible to support custom modification of the transcribed text by screen touch during the real-time output of voice transcription, which helps to improve information recording efficiency and support custom information recording.

[0055] See also Figure 5 , Figure 5 The computer-readable storage medium 50 stores program instructions 51 that can be executed by a processor, and the program instructions 51 are used to implement the steps of any of the above-mentioned fusion recording method embodiments.

[0056] In the above scheme, the computer-readable storage medium 50 detects whether a touch track is triggered in the recording interface of the transcribed text during the real-time output of speech transcription, and detects whether the transcribed text is in a modifiable state. In response to the transcribed text being in a modifiable state and the touch track being detected in the recording interface, the transcribed text is modified based on the modification operation matching the touch track. Therefore, on the one hand, compared with simply using handwriting input, while supporting handwriting input, speech transcription is integrated, which can improve the efficiency of information recording, especially when a large amount of content needs to be recorded. On the other hand, compared with simply using audio-to-text input, while supporting audio-to-text input, handwriting input is integrated, and further in the modifiable state, a touch track is responded to, so that the transcribed text is modified using a modification operation matching the touch track, which can support custom information recording. Therefore, by working together and integrating handwriting input and audio-to-text input, it is possible to support custom modification of the transcribed text by screen touch during the real-time output of speech transcription, which helps to improve information recording efficiency and support custom information recording.

[0057] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the method described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.

[0058] The above description of the various embodiments tends to emphasize the differences between the various embodiments. The same or similar aspects can be referenced with each other and will not be repeated herein for the sake of brevity.

[0059] In the several embodiments provided in this application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device implementation methods described above are only schematic. For example, the division of modules or units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, and the indirect coupling or communication connection of devices or units can be electrical, mechanical or other forms.

[0060] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0061] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0062] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the existing technology, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of the various implementation methods of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, and other media that can store program code.

[0063] If the technical solution of this application involves personal information, the product that applies the technical solution of this application has clearly informed the personal information processing rules and obtained the individual's voluntary consent before processing personal information. If the technical solution of this application involves sensitive personal information, the product that applies the technical solution of this application has obtained the individual's separate consent before processing sensitive personal information, and at the same time meets the "explicit consent" requirement. For example, on personal information collection devices such as cameras, a clear and prominent sign is set to inform that the personal information collection scope has been entered and personal information will be collected. If the individual voluntarily enters the collection scope, it is deemed that they agree to the collection of their personal information; or on the personal information processing device, when the personal information processing rules are notified by obvious signs / information, the individual's authorization is obtained through pop-up information or by asking the individual to upload their personal information; among which, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the type of personal information processed.

Claims

1. A fusion recording method, characterized in that: include: During the real-time output of the speech transcription, detecting whether a touch track is triggered in the recording interface of the transcribed text, and detecting whether the transcribed text is in a modifiable state; In response to the transcribed text being in the modifiable state and the touch track being detected in the recording interface, text modification is performed on the transcribed text based on a modification operation matching the touch track.

2. The method according to claim 1, characterized in that The performing text modification on the transcribed text based on the modification operation matching the touch track includes: determining a modification operation matching the touch trajectory as a target operation; Determining, based on a trigger position of the touch track in the recording interface, a text range targeted by the target operation in the transcribed text; Perform text modification on the text range according to the target operation.

3. The method according to claim 1, characterized in that When the transcribed text is in the modifiable state, in the case where the touch track is detected on the recording interface, the method further includes: In response to the touch track being a horizontal line that continuously passes through a plurality of characters in the transcribed text, determining that the modification operation matching the touch track is a target operation of a delete operation type; Selecting the text that the touch track passes through continuously as the text range targeted by the target operation; wherein the text range targeted by the target operation includes the to-be-deleted range targeted by the delete operation; Perform text modification on the to-be-deleted range according to the deletion operation.

4. The method according to claim 1, wherein When the transcribed text is in the modifiable state, in the case where the touch track is detected on the recording interface, the method further includes: In response to the touch track being a step-type line, determining that a modification operation matching the touch track is a target operation of a swap type; wherein the step-type line includes at least two horizontal lines forming a step; Selecting the text respectively covered by the two horizontal lines in the touch track as the text range targeted by the target operation; wherein the text range targeted by the target operation includes the two to-be-exchanged ranges targeted by the swap operation; Perform text modification on the two to-be-exchanged ranges according to the exchange operation.

5. The method according to claim 1, wherein When the transcribed text is in the modifiable state, in the case where the touch track is detected on the recording interface, the method further includes: In response to the touch track being a closed line surrounding a plurality of characters in the transcribed text, determining that a modification operation matching the touch track is a target operation of a highlight operation type; Selecting the text surrounded by the touch track as the text range targeted by the target operation; wherein the text range targeted by the target operation includes the to-be-highlighted range targeted by the highlight operation; Perform text modification on the to-be-highlighted range according to the highlighting operation.

6. The method according to claim 1, characterized in that When the transcribed text is in the modifiable state, in the case where the touch track is detected on the recording interface, the method further includes: In response to the touch track being a vertical line located at the beginning of a text line in the transcribed text and extending in a first direction, determining that a modification operation matching the touch track is a target operation of an increase indentation type; selecting the text line where the touch track is located as a text range targeted by the target operation; wherein the text range targeted by the target operation includes the range to be increased or decreased targeted by the increase indentation operation; and performing text modification on the range to be increased or decreased according to the increase indentation operation; Alternatively, in response to the touch trajectory being a vertical line located at the beginning of the first text line in the transcribed text and facing in the second direction, the modification operation matching the touch trajectory is determined to be a target operation of the reduce indentation type; wherein, the first text line is indented; the text line where the touch trajectory is located is selected as the text range targeted by the target operation; wherein, the text range targeted by the target operation includes the range to be reduced targeted by the reduce indentation operation; and text modification is performed on the range to be reduced according to the reduce indentation operation.

7. The method according to claim 1, characterized in that When the transcribed text is in the modifiable state, in the case where the touch track is detected on the recording interface, the method further includes: In response to the touch track being a vertical line within a text line in the transcribed text, determining that a modification operation matching the touch track is a target operation of a line break type; selecting the text line within which the touch track is located as a text range targeted by the target operation; wherein the text range targeted by the target operation includes a range to be broken targeted by the line break operation; and performing text modification on the range to be broken according to the line break operation at a trigger position of the touch track; Alternatively, in response to the touch track being a vertical line located at the beginning of the second text line in the transcribed text and facing in the second direction, the modification operation matching the touch track is determined to be a target operation of the type of cancel line break operation; wherein, the second text line is not indented; the text line where the touch track is located is selected as the text range targeted by the target operation; wherein, the text range targeted by the target operation includes the range to be canceled targeted by the cancel line break operation; and text modification is performed on the range to be canceled according to the cancel line break operation.

8. The method according to claim 1, characterized in that When the transcribed text is in the modifiable state, in the case where the touch track is detected on the recording interface, the method further includes: In response to the touch track being a track point located at any position in the transcribed text, determining that a modification operation matching the touch track is a target operation of a text insertion operation type; Selecting a trigger position of the touch track as a text range targeted by the target operation; wherein the text range targeted by the target operation includes a to-be-inserted range targeted by the text insertion operation; Perform text modification in the to-be-inserted range according to the text insertion operation.

9. The method according to claim 1, characterized in that The transcribed text is located in a text layer, a writing layer is provided above the text layer, the touch track is non-persistently displayed on the writing layer when in the modifiable state, and the transparency of the writing layer is higher than that of the text layer.

10. The method according to claim 9, characterized in that The method further comprises: In response to the transcribed text being in an unmodifiable state and the touch track being detected in the recording interface, the touch track is persistently displayed on the writing layer.

11. The method according to claim 9, characterized in that When in the modifiable state, during the process of triggering the formation of the touch track, the touch track is displayed in a preset style, and after the text modification is performed on the transcribed text, the touch track automatically disappears.

12. The method according to claim 1, characterized in that The method further comprises: A preset icon is displayed on the text line that is updated in real time with the voice transcription on the recording interface; wherein the preset icon is used to switch between starting and stopping the voice transcription.

13. The method according to claim 1, wherein The touch track is formed by a stylus touching the screen, the stylus is in communication with an electronic device that performs speech transcription, and when the stylus triggers the generation of a modification instruction, the electronic device enters the modifiable state in response to the modification instruction; And / or, the recording interface only includes one operation window, the operation window is used to display the transcribed text, and the touch track is triggered to be formed in the operation window.

14. The method according to claim 13, characterized in that In the case that the touch track is formed by a stylus touching the screen, the stylus generates the modification instruction when detecting a key triggering action.

15. A fusion recording device, characterized in that: include: A detection module is used to detect whether a touch track is triggered in the recording interface of the transcribed text during the real-time output of the voice transcription, and to detect whether the transcribed text is in a modifiable state; The modification module is configured to, in response to the transcribed text being in the modifiable state and the touch track being detected on the recording interface, perform text modification on the transcribed text based on a modification operation matching the touch track.

16. An electronic device, characterized in that: The device comprises at least a touch screen, a memory and a processor, wherein the touch screen and the memory are respectively coupled to the processor, the memory stores at least program instructions, and the processor is used to execute the program instructions to implement the fusion recording method according to any one of claims 1 to 14.

17. The device according to claim 16, characterized in that The touch screen includes an ink screen.

18. A computer-readable storage medium, characterized in that Program instructions that can be executed by a processor are stored, and the program instructions are used to implement the fusion recording method according to any one of claims 1 to 14.