Display method and related device

By receiving and storing user editing operations during the recording process and processing the voice recognition data when appropriate, the problem of user editing content being lost during the voice recognition process is solved, and a seamless combination of user editing and voice recognition is achieved.

WO2025194303A1PCT designated stage Publication Date: 2025-09-25HONOR DEVICE CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/082177
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-03-18
Publication Date
2025-09-25

AI Technical Summary

Technical Problem

During the voice recognition process, the electronic device cannot respond to the user's editing operations, resulting in the loss of user editing content or affecting the processing of voice recognition data.

Method used

During the recording process, the electronic device receives the user's editing operations on the recognized text and puts the newly acquired voice recognition data into a queue, and processes it after the user finishes editing to avoid the newly recognized text overwriting the user's edited content.

Benefits of technology

This ensures that user editing and speech recognition do not affect each other during the speech recognition process, avoids the loss of user edited content, and ensures the integrity of speech recognition data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024082177_25092025_PF_FP_ABST
    Figure CN2024082177_25092025_PF_FP_ABST
Patent Text Reader

Abstract

The present application discloses a display method and a related device. According to the method, in a recording process, an electronic device can put newly acquired ASR data into a queue when a user can trigger to edit a recognized text, and processes the ASR data in the queue when the user triggers to end editing. It can be understood that when the user triggers to end editing, the electronic device still can continuously acquire ASR data, and under this condition, the electronic device can first process ASR data in the queue, and when the ASR data in the queue is completely processed, processes the ASR data newly acquired by the electronic device when the user triggers to end editing. According to the method, the electronic device can edit the recognized text in the ASR process, and does not loss the content edited by the user and speed recognition content acquired in the editing process.
Need to check novelty before this filing date? Find Prior Art

Description

A display method and related equipment Technical Field

[0001] The present application relates to the field of terminal technology, and in particular to a display method and related equipment. Background Art

[0002] Speech recognition technology, also known as automatic speech recognition (ASR), aims to convert the lexical content of human speech into computer-readable input (such as keystrokes, binary codes, or character sequences). ASR is currently widely used in various scenarios, including office work, entertainment, and daily life, providing great convenience. For example, electronic devices can use ASR to convert audio data collected during recording into text. However, during this process, the electronic device cannot respond to user editing operations.

[0003] Summary of the Invention

[0004] This application provides a display method and related devices. According to this method, a user can trigger editing during the recording process of an electronic device. After the user triggers editing, the electronic device can put the newly acquired ASR data into a queue and then process the ASR data in the queue after the user finishes editing. Using this method, the electronic device can edit recognized text during the ASR process without losing the user's edited content or the speech recognition content acquired during the editing process.

[0005] In the first aspect, the present application provides a display method. The method can be applied to a first device (i.e., the electronic device involved in the present application). The method may include: the electronic device may display a first interface, the first interface including a first control; in response to a first operation on the first control, the electronic device displays a second interface, and after the first operation, the electronic device may start recording. In a first time period, the electronic device may obtain first audio data obtained through recording, and the first text content corresponding to the first audio data may be displayed in a first display area of ​​the second interface. During the first time period, the first display area does not include a cursor. After the first time period, in response to a second operation on the first display area, a cursor may be displayed in the first display area; the method may also include: after the second operation, the electronic device may receive an editing operation, and the first text content modified by the editing operation may be displayed in the first display area; after the second operation, the electronic device may obtain second audio data obtained through recording, and the second text content corresponding to the second audio data is not displayed in the first display area.

[0006] The electronic device can record audio and identify the audio data obtained from the recording to obtain corresponding text content, which can be displayed in the recognized text display area (such as the first display area). It is understandable that during the recording process, the electronic device can update the text content displayed in the recognized text display area. In the solution provided in the present application, the electronic device can receive the user's editing operation on the recognized text display area while recording and turning on the voice-to-text function. In the process of receiving the user's editing operation, the electronic device can modify the text content in the recognized text display area based on the editing operation, and the newly identified text content in the process of receiving the user's editing operation is not displayed in the recognized text display area. In this way, the loss of user-edited content due to the addition of newly identified text content can be avoided. In this way, the electronic device can achieve ASR and user editing without affecting each other, and the ASR data obtained by the electronic device will not be lost due to user editing, nor will the user-edited content be lost due to processing of ASR data.

[0007] In some embodiments of the present application, the first interface may be a new note interface (for example, the new note interface 1d shown in FIG6D ), the first control may be a recording addition control (for example, the recording addition control 13 shown in FIG6D ), and the first operation may be a user operation of clicking the recording addition control.

[0008] In some embodiments of the present application, the first interface may be a note interface to which a recording has been added, for example, the user interface shown in (2) of FIG. 1A . The first control may be a continue recording control, for example, the continue recording control 106 shown in (2) of FIG. 1A and (1) and (2) of FIG. 1B . The first operation may be a user operation of clicking the continue recording control.

[0009] Optionally, after receiving the first operation on the first control, the electronic device may display the second interface. Optionally, the electronic device may detect the first operation on the first control, and in response to the first operation, the electronic device may display the second interface.

[0010] In some embodiments of the present application, the electronic device may start recording after displaying the second interface. In some other embodiments of the present application, the electronic device may start recording while displaying the second interface. In some other embodiments of the present application, the electronic device may display the second interface after starting recording.

[0011] It can be understood that the second interface involved in this application can be understood as a user interface with changing content, that is, the content included in the second interface can change.

[0012] In some embodiments of the present application, the electronic device can start recording after displaying the second interface. In this case, in response to the first operation on the first control, the second interface initially displayed by the electronic device does not include the text content corresponding to the audio data obtained by the recording (or the text content obtained by recognition). Instead, after the recording starts, the text content corresponding to the audio data obtained by the recording will be displayed in the corresponding display area of ​​the second interface (for example, the first display area).

[0013] In some embodiments of the present application, after the electronic device starts recording, it can display a second interface. In this case, in response to the first operation on the first control, the second interface displayed earliest by the electronic device may include text content corresponding to the audio data obtained from the recording, and as the recording duration increases, the text content corresponding to the audio data obtained from the recording displayed in the corresponding display area of ​​the second interface (for example, the first display area) may change.

[0014] In some embodiments of the present application, the second interface may be a user interface as shown in FIG. 6E to FIG. 6J .

[0015] In some embodiments of the present application, the first display area may be a recognized text display area (e.g., recognized text display area 130), which may be understood as a display area corresponding to the final result obtained by performing speech recognition within a period of time (e.g., a first time period). It is understood that the final result can be specifically referred to the relevant description below, which is not elaborated in this application. In one possible implementation, the first display area may specifically be a display area corresponding to a certain speaker in the recognized text display area (e.g., the second speaker text display area 1031).

[0016] In some embodiments of the present application, as shown in (1) in FIG6E , the first time period may be calculated from the start of recording to the 29th second of recording. In this case, the first audio data may be the audio data obtained from the start of recording to the 29th second of recording. Accordingly, the first text content may include "111222333444555666" in the first speaker text display area 1032 and "abcde3336666999" in the second speaker text display area 1031.

[0017] In some embodiments of the present application, as shown in (1) in FIG6E , the first time period may be from the 22nd second after the start of recording to the 29th second of recording. In this case, the first audio data may be the audio data obtained from the recording from the 22nd second to the 29th second of recording. Accordingly, the first text content may include "abcde3336666999" in the second speaker text display area 1031.

[0018] In some embodiments of the present application, as shown in (1) in FIG. 6E , the second operation may be a user operation of clicking on the second speaker text display area 1031 .

[0019] In some embodiments of the present application, the second operation may be a user operation of clicking on the entire speaker text display area 1034 .

[0020] Optionally, after the electronic device receives the second operation on the first display area, the cursor may be displayed in the first display area. Optionally, the electronic device may detect the second operation on the first display area, and in response to the second operation, the cursor may be displayed in the first display area.

[0021] In some embodiments of the present application, as shown in ( 1 ) of FIG. 1B , the cursor displayed in the first display area may be cursor 104 .

[0022] In some embodiments of the present application, as shown in FIG. 6F , the cursor displayed in the first display area may be a cursor 15 .

[0023] In some embodiments of the present application, the cursor may be displayed in the first display area, which may be understood as: the electronic device may display the cursor in the first display area, or display the cursor in the first display area.

[0024] It can be understood that the specific position where the cursor is displayed in the first display area can be determined based on the specific position of the second operation.

[0025] In some embodiments of the present application, receiving an edit operation by an electronic device may specifically include: the electronic device receiving a user operation of clicking a keyboard key. In some embodiments of the present application, receiving an edit operation by an electronic device may specifically include: the electronic device receiving a user operation of writing with a finger or a stylus in an input method handwriting area. It is understood that receiving an edit operation may also include other implementations, which are not specifically limited by the present application.

[0026] In some embodiments of the present application, the editing operation may specifically include operations such as adding text and deleting text, and the present application does not impose specific limitations on this. It is understood that the editing operation may involve the target editing content, editing position, editing method, and editing content mentioned in step S307 below.

[0027] In some embodiments of the present application, as shown in (1) in FIG6E , the first text content may include “abcde3336666999” in the second speaker text display area 1031 . The editing operation may include deleting the text “6” to the left of the cursor 15 as shown in FIG6F . The editing operation may also include adding the text “888” to the left of the cursor 15 as shown in FIG6G . In this case, the modified first text content (or the text content obtained after modifying the first text content) may be “abcde333666888999”.

[0028] It is understandable that during the process of receiving the editing operation, the text content in the first display area may change accordingly. In some embodiments of the present application, as shown in Figures 6G-6H, during the process of receiving the editing operation, the text content in the first display area may change from "abcde3336666999" to "abcde333666999", and then from "abcde333666999" to "abcde333666888999".

[0029] In some embodiments of the present application, as shown in Figures 6F-6H, the second time period may be from the 31st second after the start of recording to the 42nd second of recording. In this case, the second audio data may be the audio data obtained from the recording from the 31st second to the 40th second of recording. Accordingly, the second text content may include "1234567890", which is not displayed in the recognized text display area 130.

[0030] In some embodiments of the present application, the first text content corresponding to the first audio data can be displayed in the first display area of ​​the second interface, which can be understood as: the electronic device can display the first text content corresponding to the first audio data in the first display area of ​​the second interface, or display the first text content corresponding to the first audio data in the first display area of ​​the second interface.

[0031] In some embodiments of the present application, the first text content modified by the editing operation can be displayed in the first display area, which can be understood as: the electronic device can display the first text content modified by the editing operation in the first display area, or display the first text content modified by the editing operation in the first display area.

[0032] In some embodiments of the present application, the text, text content and text data involved in the present application have the same meaning.

[0033] In conjunction with the first aspect, in a possible implementation, the second interface may include a second display area. After the second operation, the second text content may be displayed in the second display area.

[0034] In the solution provided in the present application, the electronic device can receive the user's editing operation on the recognized text display area when recording and turning on the voice-to-text function. In the process of receiving the user's editing operation, the electronic device may not add the newly recognized text content to the recognized text display area. That is, the newly recognized text content may not be displayed in the recognized text display area, but the newly recognized text can be displayed in the recognized text display area. In this way, during the user's editing process, the electronic device can promptly display the text content corresponding to the audio data obtained by the recording during the editing process in other display areas, so that the user can also promptly view the newly recognized text during the editing process.

[0035] In some embodiments of the present application, the second display area may be a recognized text display area (e.g., recognized text display area 1033), which may be understood as a display area corresponding to the intermediate results obtained by voice recognition over a period of time. It is understood that the intermediate results may be specifically described below, and this application does not elaborate on them here.

[0036] In some embodiments of the present application, as shown in Figures 6F-6H, the second time period can be from the 31st second after the start of recording to the 42nd second of the recording. In this case, the second audio data can be: audio data obtained by recording from the 31st second to the 40th second of the recording, and the second text content can include "1234567890". It is worth noting that, as shown in Figure 6H, after the second operation, "1234567890" is not displayed in the recognized text display area 130, but "1234567890" is displayed in the text display area 1033 being recognized.

[0037] In combination with the first aspect, in a possible implementation, the method may further include: in response to an event of exiting editing, the second text content and the modified first text content may be displayed in the first display area, and after the event of exiting editing, the cursor is not displayed in the first display area.

[0038] In the solution provided in the present application, the electronic device can receive the user's editing operations on the recognized text display area when recording and turning on the voice-to-text function. In the process of receiving the user's editing operations, the electronic device may not add the newly recognized text content to the recognized text display area, but wait until the user's editing is completed before adding the newly recognized text content to the recognized text display area, thereby avoiding the loss of user edited content due to adding newly recognized text content during the process of receiving user editing.

[0039] In some embodiments of the present application, the event of exiting editing may specifically include any of the following: receiving an operation in a non-editing area; or not receiving an editing operation on the first display area for a period longer than T1. Specifically, not receiving an editing operation on the first display area for a period longer than T1 may specifically include: not receiving a user operation for a period longer than T1 (for details, see step S308).

[0040] In some embodiments of the present application, according to the following, the non-editing area may include an area other than the recognized text display area and the area where the keyboard is located.

[0041] In some embodiments of the present application, the non-editing area may also include areas other than the recognized text display area and the handwriting area.

[0042] In some embodiments of the present application, the time to exit editing may specifically include: a user operation of clicking a blank area as shown in FIG6H . In response to the user operation of clicking a blank area as shown in FIG6H , the second text content "1234567890" and the modified first text content "abcde333666888999" are displayed in the second speaker text display area 1031 as shown in FIG6I .

[0043] In combination with the first aspect, in a possible implementation, in response to the first operation, the electronic device may start recording and enable a speech-to-text function.

[0044] In the solution provided by the present application, the electronic device can create a new note or open a note that does not contain a recording, and in response to a user operation of clicking a recording add control, the electronic device can start recording, and when the recording starts, the speech-to-text function is already turned on. That is, after the electronic device starts recording, it can obtain the text content corresponding to the audio data obtained by the recording and display it in the corresponding display area (for example, the first display area and the second display area), and the user does not need to trigger the electronic device to turn on the speech-to-text function by clicking the corresponding control (for example, the speech-to-text control 141).

[0045] In conjunction with the first aspect, in one possible implementation, the second interface may include a second control. After the first operation, the electronic device begins recording, which may specifically include: in response to the first operation, the electronic device may begin recording. During the first time period, the electronic device obtains first audio data obtained through recording, which may specifically include: in response to a third operation on the second control, the electronic device may obtain the first audio data and first text content during the first time period.

[0046] In the solution provided in the present application, the electronic device can display a note interface containing a recording and continue to start recording, and when the recording is continued, the speech-to-text function is not turned on. During the process of continuing recording, the user can turn on the speech-to-text function. After turning on the speech-to-text function, the electronic device can obtain the text content corresponding to the audio data obtained from the recording and display the text content on the display screen. In addition, after turning on the speech-to-text function, the electronic device can also receive the user's editing operation for the recognized text display area. In the process of receiving the user's editing operation, the electronic device can modify the text content in the recognized text display area based on the editing operation, and the newly recognized text content in the process of receiving the user's editing operation is not displayed in the recognized text display area. In this way, the loss of the user's edited content due to the addition of the newly recognized text content can be avoided.

[0047] According to the above, in some embodiments of the present application, the first interface may be a note interface to which a recording has been added, for example, the user interface shown in (2) of FIG1A . The first control may be a continue recording control (for example, continue recording control 106 ), and the first operation may be a user operation of clicking the continue recording control. Furthermore, in some embodiments of the present application, the second control may be a speech-to-text control (for example, speech-to-text control 141 ).

[0048] In the above case, in one possible implementation, in response to the first operation, the electronic device can display the second interface, and the speech-to-text control in the second interface that is first displayed by the electronic device is in an unselected state, and the first displayed second interface does not include the text content corresponding to the audio data obtained by the recording. That is, the third operation on the second control can be understood as: a user operation (for example, a click operation) on the speech-to-text control that is in an unselected state. Furthermore, in response to the user operation of clicking on the speech-to-text control that is in an unselected state, the speech-to-text control that is in a selected state is displayed on the second interface, that is, the speech-to-text function is turned on. In this case, the electronic device can not only obtain the audio data obtained by the recording in the first time period (that is, the first audio data), but also obtain the text content corresponding to the audio data (that is, the first text content).

[0049] In combination with the first aspect, in a possible implementation, after the electronic device obtains the first audio data obtained by recording, it can send the first audio data to the second device (i.e., the speech recognition server involved in this application). After the electronic device sends the first audio data to the second device, the electronic device can receive multiple audio recognition results sent by the second device. Wherein, the text in the first category of audio recognition results included in the multiple audio recognition results can constitute the first text content, and the multiple audio recognition results include the first recognition result. After the electronic device receives the first recognition result, the electronic device can process the first recognition result to obtain the third text content. In the case where the first recognition result is the first category of audio recognition result, the third text content can be displayed in the first display area, and the first text content can include the third text content. However, in the case where the first recognition result is the second category of audio recognition result, the third text content can replace the content originally displayed in the second display area, that is, the third text content can be displayed in the second display area, and it replaces the content originally displayed in the second display area. It should be noted that in the process of processing the first recognition result, the first marker is the first content, and the first queue does not include the audio recognition result.

[0050] In the solution provided by this application, an electronic device can obtain the audio recognition results corresponding to the audio data obtained from the recording when the voice-to-text function is enabled, and determine whether the text content is displayed in the first display area or the second display area based on the type of the audio recognition result. In this way, for different types of audio recognition results, the electronic device can display the text content contained therein in different display areas for user viewing.

[0051] In some embodiments of the present application, the audio recognition result may be the speech-to-text result mentioned below.

[0052] In some embodiments of the present application, the first type of audio recognition result can be understood as the final result of speech recognition within a period of time, and the second type of audio recognition result can be understood as the intermediate result of speech recognition within a period of time.

[0053] In some embodiments of the present application, the first type of audio recognition result may be the final type of speech-to-text result mentioned below (or referred to as the speech-to-text result of type final), and the second type of audio recognition result may be the part type of speech-to-text result mentioned below (or referred to as the part type of speech-to-text result).

[0054] It is understandable that after the electronic device sends the first audio data to the speech recognition server, the speech recognition server can recognize the first audio data, obtain multiple audio recognition results, and send the multiple audio recognition results to the electronic device. It is understandable that the multiple audio recognition results can be understood as the audio recognition results corresponding to the first audio data. According to the above, the electronic device can receive the audio recognition results corresponding to the first audio data. It should be noted that the electronic device can process and display the audio recognition results in the order in which they are received.

[0055] It is understandable that the audio recognition result corresponding to the first audio data may include one or more first-category audio recognition results. In some embodiments of the present application, the first-category audio recognition results in the audio recognition results corresponding to the first audio data may be composed of the first text content in the order of generation (or the order in which the electronic device receives them) and displayed in the first display area. In other words, the text content displayed in the first display area is spliced ​​together from the first-category audio recognition results.

[0056] In some embodiments of the present application, the first queue may be the queue mentioned below, the first flag bit may be the flag bit mentioned below, and the first content may be false mentioned below. Of course, the first content may also be represented in other ways, which is not limited by the present application.

[0057] It is understood that after receiving the first recognition result, the electronic device can determine whether the first mark bit is the first content and whether the queue is empty. If the first mark bit is the first content and the queue is empty, the electronic device can process the first recognition result and display the processed text on the display screen. The specific process of the electronic device processing and displaying the first recognition result can be referred to steps S203-S212, which will not be described in detail in this application.

[0058] It is understandable that the third text content is displayed in the first display area, which can specifically include: the third text content can be spliced ​​behind the text content originally displayed in the first display area to obtain the spliced ​​text content, and the spliced ​​text content can be displayed in the first display area. It is understandable that the first text content can include the spliced ​​text content, that is, the first text content can include the third text content and the text content originally displayed in the first display area. It is understandable that the above-mentioned splicing and display operations can also be performed on other first-category audio recognition results in the audio recognition results corresponding to the first audio data. Please refer to the following for details.

[0059] Exemplarily, the audio recognition result corresponding to the first audio data includes multiple first-category audio recognition results, and the first recognition result is a first-category audio recognition result. The text content originally displayed in the second speaker text display area 1031 is "abcde3336666", and the third text content may be "999". The electronic device may add "999" to the end of "abcde3336666" to obtain "abcde3336666999", and "abcde3336666999" may be displayed in the second speaker text display area 1031, as specifically shown in (1) in FIG6E.

[0060] It should be noted that, in the process of receiving multiple audio recognition results corresponding to the first audio data, the electronic device can determine whether to display the text content contained therein in the first display area or the text content contained therein in the second display area based on the type of each audio recognition result after receiving the multiple audio recognition results, instead of waiting until all the multiple audio recognition results are received and then determining their types one by one in the order of receipt and whether to display the text content contained therein in the first display area or the second display area.

[0061] In combination with the first aspect, in a possible implementation, the method may further include: in response to the first operation, creating a first queue, and setting the first mark bit to the first content.

[0062] In the solution provided in the present application, the electronic device can achieve the goal of ASR and user editing not affecting each other by setting a flag bit and a queue. Specifically, the electronic device can indicate whether the user is editing the recognized text by setting a flag bit, and create a queue to temporarily store the audio recognition results after the user triggers the editing of the recognized text, thereby avoiding the text content in the audio recognition results from affecting the user's editing content.

[0063] According to the above, in some embodiments of the present application, the first interface can be a new note interface (for example, the new note interface 1d shown in FIG6D ), the first control can be a recording addition control (for example, the recording addition control 13 shown in FIG6D ), and the first operation can be a user operation of clicking the recording addition control. In this case, in response to the operation of clicking the recording addition control, the electronic device can create a first queue and set the first flag bit to the first content.

[0064] In conjunction with the first aspect, in one possible implementation, the first interface may be a note interface to which a recording has been added. In this case, before displaying the first interface, the electronic device may display a note creation interface (note creation interface 1c as shown in FIG6C ). The note creation interface may include an icon corresponding to a note containing a recording (the icon corresponding to the “recording note” as shown in FIG6C ). The electronic device may create a first queue and set the first marker bit to the first content in response to a user operation of clicking the icon corresponding to the note containing the recording.

[0065] In the solution provided in the present application, for notes to which recordings have been added, after the electronic device receives a user operation that triggers the display of the user interface corresponding to the note to which recordings have been added, the electronic device can create a queue and set a mark bit, without having to wait until the recording starts to create a queue and set a mark bit, thus saving time.

[0066] In combination with the first aspect, in a possible implementation, after the electronic device obtains the second audio data obtained by recording, the electronic device may send the second audio data to the second device. After the electronic device sends the second audio data to the second device, the electronic device may receive multiple audio recognition results sent by the second device, and process the multiple audio recognition results to obtain the second text content. Each of the multiple audio recognition results includes part or all of the second text content. In the process of receiving the multiple audio recognition results sent by the second device and processing the multiple audio recognition results, the first marker is the second content. After sending the second audio data to the second device and receiving the multiple audio recognition results sent by the second device, the first queue includes the multiple audio recognition results received after sending the second audio data to the second device, or the first queue includes the first type of audio recognition results among the multiple audio recognition results received after sending the second audio data to the second device.

[0067] In the solution provided by this application, the electronic device can receive user editing operations on the display area of ​​recognized text when recording and the voice-to-text function is turned on. In the process of receiving the user's editing operations, on the one hand, the electronic device can modify the recognized text, and on the other hand, the electronic device can still obtain the audio recognition results corresponding to the audio data obtained from the recording and put the audio recognition results into the queue. In this way, the electronic device does not need to add newly recognized text when modifying the recognized text, thereby avoiding the loss of user edited content due to the addition of newly recognized text.

[0068] In some embodiments of the present application, the second content may be the true mentioned below. Of course, the second content may also be expressed in other ways, which are not limited by the present application.

[0069] It is understandable that after the electronic device sends the second audio data to the speech recognition server, the speech recognition server can recognize the second audio data, obtain multiple audio recognition results, and send the multiple audio recognition results to the electronic device. It is understandable that the multiple audio recognition results can be understood as the audio recognition results corresponding to the second audio data. According to the above, the electronic device can receive the audio recognition result corresponding to the second audio data. It should be noted that in the process of receiving user editing operations, on the one hand, the electronic device can process and display all or part of the audio recognition results in the order of receiving them, and on the other hand, the electronic device can put all the audio recognition results into a queue, or put the first category of audio recognition results in the queue.

[0070] It is understandable that after the electronic device receives the audio recognition result corresponding to the second audio data, it can determine whether the first mark bit is the first content. In the case where the first mark bit is the second content, the electronic device can process the audio recognition result corresponding to the second audio data and display the processed text on the display screen. Among them, the electronic device's processing and display process of the audio recognition result corresponding to the second audio data can be specifically referred to steps S214-step S217, and the relevant description of Table 1, which will not be elaborated in this application.

[0071] In combination with the first aspect, in a possible implementation, the multiple audio recognition results received by the electronic device after sending the second audio data to the second device include a second recognition result and a third recognition result. The second recognition result and the third recognition result are first-category audio recognition results, and the second recognition result is received earlier than the third recognition result. The second recognition result includes a fourth text content, and the third recognition result includes a fifth text content. The electronic device receives the multiple audio recognition results sent by the second device and processes the multiple audio recognition results to obtain the second text content. Specifically, it may include: the electronic device may receive the second recognition result, and after receiving the second recognition result, it may process the second recognition result to obtain the fourth text content, and the fourth text content may be displayed in the second display area; the electronic device may receive the third recognition result, and after receiving the third recognition result, it may process the third recognition result to obtain the fifth text content; the fifth text content and the fourth text content form the second text content and are displayed in the second display area.

[0072] In the solution provided by this application, the electronic device can receive user editing operations on the recognized text display area when recording and the voice-to-text function is turned on. In the process of receiving the user editing operation, the electronic device can receive multiple audio recognition results corresponding to the second audio data, and process the first type of audio recognition results in the multiple audio recognition results in the order of receipt, and display the processed text content in the second display area. In this way, during the user editing process, the user can also view the final result obtained by relatively accurate recognition over a period of time in the second display area.

[0073] It should be noted that, in the process of receiving user editing operations, the electronic device can receive multiple audio recognition results corresponding to the second audio data, and after receiving each of the multiple audio recognition results, determine whether to display the text content contained therein in the second display area based on the type of the audio recognition result, instead of waiting until all the multiple audio recognition results are received, and then determining their types one by one in the order of receipt, and whether to display the text content contained therein in the second display area.

[0074] In some embodiments of the present application, the fifth text content and the fourth text content constitute the second text content, which may specifically include: the fifth text content is spliced ​​after the fourth text content to constitute the second text content. As shown in Figure 6F, the second display area can be the text display area 1033 in the recognition process, and the fourth text content can be "123456", and "123456" is displayed in the text display area 1033 in the recognition process. As shown in Figure 6G, the fifth text content can be "789", and "789" can be spliced ​​after "123456" to form "123456789", and "123456789" can be displayed in the text display area 1033 in the recognition process.

[0075] In combination with the first aspect, in a possible implementation, the multiple audio recognition results received by the electronic device after sending the second audio data to the second device may include a fourth recognition result and a fifth recognition result. The fourth recognition result is received earlier than the fifth recognition result. The fourth recognition result may include a sixth text content, and the fifth recognition result may include a second text content. The electronic device receives multiple audio recognition results sent by the second device and processes the multiple audio recognition results to obtain the second text content. Specifically, it may include: the electronic device may receive the fourth recognition result, and after receiving the fourth recognition result, it may process the fourth recognition result to obtain the sixth text content, and the sixth text content is displayed in the second display area; the electronic device may receive the fifth recognition result, and after receiving the fifth recognition result, it may process the fifth recognition result to obtain the second text content, and the second text content replaces the sixth text content and is displayed in the second display area.

[0076] In the solution provided in the present application, the electronic device can receive the user's editing operation on the recognized text display area when recording and turning on the voice-to-text function. In the process of receiving the user's editing operation, the electronic device can receive multiple audio recognition results corresponding to the second audio data, and process some of the multiple audio recognition results (the first type of audio recognition results or the second type of audio recognition results) or all of the audio recognition results in the order of reception, and overlay the processed text content on the second display area. In this way, during the user editing process, the user can also view the recognition results obtained over a period of time in the second display area.

[0077] In some embodiments of the present application, as shown in FIG6F , the second display area may be the recognized text display area 1033, and the sixth text content may be "123456", which is displayed in the recognized text display area 1033. As shown in FIG6G , the second text content may be "123456789", with "123456789" replacing "123456" and displayed in the recognized text display area 1033.

[0078] In combination with the first aspect, in one possible implementation, the fourth recognition result and the fifth recognition result may be first-category audio recognition results; or, the fourth recognition result and the fifth recognition result may be second-category audio recognition results; or, the fourth recognition result may be first-category audio recognition result, and the fifth recognition result may be second-category audio recognition result; or, the fourth recognition result may be second-category audio recognition result, and the fifth recognition result may be first-category audio recognition result.

[0079] In the solution provided by this application, regardless of whether the electronic device processes only the first type of audio recognition results received, only the second type of audio recognition results received, or all the audio recognition results received during the user editing operation, the electronic device can overlay the processed text content on the second display area. In this way, the user can also see the newly recognized text content during the editing process. Moreover, this method is simple in logic, easy to implement, and convenient to maintain.

[0080] In conjunction with the first aspect, in one possible implementation, after the first time period, the method may further include: in response to a second operation, the electronic device may set the first marker to the second content. After the electronic device obtains the second text content, the method may further include: in response to an event of exiting editing, the electronic device may set the first marker to the first content.

[0081] In the solution provided in the present application, the electronic device can indicate that the user is not currently editing the recognized text by setting the first mark bit to the first content, and can indicate that the user is currently editing the recognized text by setting the first mark bit to the second content. In this way, the electronic device can determine whether the user is currently editing the recognized text based solely on the specific content of the first mark bit, and thus adopt different strategies to process the received audio recognition results accordingly. This method is concise and effective.

[0082] In a second aspect, the present application provides an electronic device, comprising one or more memories and one or more processors; the one or more memories are coupled to the one or more processors, the memory being used to store computer program code, the computer program code comprising computer instructions, the one or more processors calling the computer instructions to enable the electronic device to execute the method described in the first aspect or any one of the implementations of the first aspect.

[0083] In a third aspect, the present application provides a computer storage medium comprising computer instructions, which, when executed on an electronic device, causes the electronic device to execute the method described in the first aspect or any one of the implementations of the first aspect.

[0084] In a fourth aspect, an embodiment of the present application provides a chip. The chip can be applied to an electronic device, and the chip includes one or more processors configured to invoke computer instructions to cause the electronic device to execute the method described in the first aspect or any one of the implementations of the first aspect.

[0085] In some embodiments of the present application, the chip system may be an application processor (AP) or a system on chip (SoC) including an AP. The method described in the first aspect or any one of the implementations of the first aspect may be implemented by an AP, and the method described in the second aspect or any one of the implementations of the second aspect may be implemented by an AP.

[0086] In some other embodiments of the present application, the chip system may include an AP and other modules, wherein the other modules may be a modem processor (also referred to as a baseband processor).

[0087] In a fifth aspect, an embodiment of the present application provides a computer program product comprising instructions. When the computer program product is run on an electronic device, the electronic device executes the method described in the first aspect or any one of the implementations of the first aspect.

[0088] It is understood that the electronic device provided in the second aspect, the computer storage medium provided in the third aspect, the chip provided in the fourth aspect, and the computer program product provided in the fifth aspect are all used to perform the method described in the first aspect or any one of the implementations of the first aspect. Therefore, the beneficial effects that can be achieved can be referenced to the beneficial effects of any possible implementation of the first aspect, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0089] FIG1A-FIG1B are schematic diagrams of a set of user interfaces provided in an embodiment of the present application;

[0090] 2A-2B are schematic diagrams of a method for implementing editing during speech recognition according to an embodiment of the present application;

[0091] FIG3 is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application;

[0092] FIG4 is a schematic diagram of a software structure of an electronic device provided in an embodiment of the present application;

[0093] FIG5 is a diagram illustrating an architecture of a speech recognition system provided in an embodiment of the present application;

[0094] 6A-6J are schematic diagrams of another set of user interfaces provided in an embodiment of the present application;

[0095] 7A-7D are flowcharts of a display method provided in an embodiment of the present application;

[0096] FIG8 is a flowchart of another display method provided in an embodiment of the present application;

[0097] 9A-9H are schematic diagrams of enqueuing and dequeuing speech-to-text results provided by an embodiment of the present application;

[0098] FIG10 is another schematic diagram of a method for implementing editing during speech recognition according to an embodiment of the present application;

[0099] FIG11 is a flowchart of two threads operating on one object together according to an embodiment of the present application. DETAILED DESCRIPTION

[0100] The following will be combined with the accompanying drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. In the description of the embodiments of the present application, unless otherwise specified, " / " means or, for example, A / B can mean A or B; "and / or" in the text is only a description of the association relationship of associated objects, indicating that there can be three relationships, for example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, in the description of the embodiments of the present application, "multiple" means two or more than two.

[0101] It should be understood that the terms "first," "second," and the like in the specification, claims, and drawings of this application are used to distinguish between different objects, rather than to describe a particular order. Furthermore, the terms "including," "having," and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or elements is not limited to the listed steps or elements, but may optionally include steps or elements not listed, or may optionally include other steps or elements inherent to the process, method, product, or apparatus.

[0102] It should be understood that the term "user interface" in the specification, claims, and drawings of this application refers to the media interface for interaction and information exchange between an application or operating system and a user. A common form of user interface is a graphical user interface (GUI), which refers to a user interface related to computer operations that is displayed graphically. It can be an interface element such as an icon, window, or control displayed on the display screen of an electronic device. Controls can include visual interface elements such as icons, buttons, menus, tabs, text boxes, dialog boxes, status bars, navigation bars, and widgets.

[0103] References to "embodiments" in this application mean that a particular feature, structure, or characteristic described in connection with the embodiment may be included in at least one embodiment of the application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described in this application may be combined with other embodiments.

[0104] According to the above, the electronic device can convert the collected audio data (also understood as voice data) into text through ASR. However, during this process, the electronic device cannot respond to the user's editing operation.

[0105] The following uses the audio data conversion (also known as voice-to-text) and editing in the note-taking app as an example to illustrate:

[0106] After the electronic device opens the note application, it can respond to user operations, create notes, and record audio in the notes (i.e., collect audio data), and obtain text converted (or recognized) from the collected audio data. At the 29th second of the recording, the electronic device can display the user interface shown in (1) in Figure 1A. The user interface shown in (1) in Figure 1A can include a stop recording control 101 and a voiceprint waveform display area 102. Among them, the stop recording control 101 can be used to trigger the electronic device to stop recording, and the voiceprint waveform display area 102 can include the voiceprint waveform corresponding to the audio data collected by the electronic device. It can be understood that if the electronic device displays both the stop recording control 101 and the voiceprint waveform display area 102, it can indicate that the electronic device is recording. The user interface shown in (1) in Figure 1A can also include a speech recognition text display area 103. The speech recognition text display area 103 is used to display the text obtained by recognizing the collected audio data. The speech recognition text display area 103 can include a recognized text display area 130 and a recognition-in-progress text display area 1033. Among them, the recognized text display area 130 may include recognized text, and the identifying text display area 1033 includes the text being recognized. In some embodiments of the present application, the recognized text display area 130 may include text display areas corresponding to different speakers. For example, the first speaker text display area 1032 and the second speaker text display area 1031 as shown in (1) in Figure 1A. In some embodiments of the present application, the recognized text display area 130 may include all speaker text display areas. For example, the all speaker text display area 1034 as shown in (2) in Figure 6E. The relevant description of all speaker text display areas can be specifically referred to below, and this application will not elaborate on it here.

[0107] It is understandable that the recognized text involved in this application can also be referred to as the text that has been recognized. It represents the corresponding text finally obtained after the audio data is recognized, which can be understood as the final result obtained by recognition (or the final result of voice recognition). It is understandable that the final result is displayed in the recognized text display area (for example, the recognized text display area 130). The text being recognized involved in this application represents the corresponding text obtained in the process of recognizing the audio data, which can be understood as the intermediate result obtained by recognition (or the intermediate result of voice recognition), rather than the final result obtained by recognition. It is understandable that the intermediate result is displayed in the text display area (for example, the text display area 1033) in recognition.

[0108] It should be noted that the final result will not change with subsequent speech recognition, while the intermediate result may be revised with subsequent speech recognition to obtain a new intermediate result or final result. For example, as shown in (1) in Figure 1A, the final result can be displayed in the second speaker text display area 1031 and the first speaker text display area 1032, and the intermediate result can be displayed in the recognition text display area 1033.

[0109] It should also be noted that, since audio data is collected continuously, its recognition is carried out continuously, that is, each time a part of audio data is collected, this part of audio data will be recognized. Then, the final result refers to the final result obtained by identifying this part of audio data after each part of audio data is collected during the continuous recognition process, rather than the final result obtained by identifying all audio data collected during the entire continuous recognition process. Similarly, the above-mentioned intermediate result refers to the intermediate result obtained by identifying this part of audio data after each part of audio data is collected during the continuous recognition process, rather than the intermediate result obtained by identifying all audio data collected during the entire continuous recognition process.

[0110] The user can click on the second speaker text display area 1031 (including the recognized text) to trigger editing of the text included in the second speaker text display area 1031. However, in the related art, when the electronic device is recording and converting speech to text, the user cannot edit the speech recognition content, so the electronic device will not respond to the user operation of clicking on the second speaker text display area 1031. In other words, the user cannot edit the text included in the second speaker text display area 1031 (i.e., the recognized text).

[0111] After the user clicks the stop recording control 101, the electronic device can stop recording and display the user interface shown in (2) in Figure 1A. Optionally, in response to the user operation of clicking the stop recording control 101, the electronic device can stop recording and display the user interface shown in (2) in Figure 1A. Optionally, the electronic device can detect the user operation acting on the stop recording control 101, and in response to the user operation, the electronic device can stop recording and display the user interface shown in (2) in Figure 1A. Similar to the user interface shown in (1) in Figure 1A, the user interface shown in (2) in Figure 1A can also include a second speaker text display area 1031.

[0112] After the user clicks on the second speaker text display area 1031 as shown in (2) in FIG1A , the cursor is displayed in the second speaker text display area 1031 , and the electronic device may display the user interface as shown in (1) in FIG1B . Optionally, in response to the user operation of clicking on the second speaker text display area 1031 as shown in (2) in FIG1A , the cursor is displayed in the second speaker text display area 1031 , and the electronic device may display the user interface as shown in (1) in FIG1B . Optionally, the electronic device may detect the user operation acting on the second speaker text display area 1031 as shown in (2) in FIG1A , and in response to the user operation, the cursor is displayed in the second speaker text display area 1031 , and the electronic device may display the user interface as shown in (1) in FIG1B . The user interface shown in (1) in FIG1B may include a cursor 104 and a keyboard 105. The keyboard 105 may include a key 1051. The key 1051 is a delete key. In some embodiments of the present application, a key may also be referred to as a button, which may be understood as a control.

[0113] The user can click button 1051 to delete part or all of the text included in the second speaker text display area 1031. As shown in (1) in Figure 1B, after the user clicks button 1051, the electronic device can delete the text located before the cursor 104 in the second speaker text display area 1031 and display the user interface shown in (2) in Figure 1B. Optionally, in response to the user operation of clicking button 1051 as shown in (1) in Figure 1B, the electronic device can delete the text located before or after the cursor 104 in the second speaker text display area 1031 and display the user interface shown in (2) in Figure 1B. Optionally, the electronic device can detect the user operation acting on button 1051 as shown in (1) in Figure 1B, and in response to the user operation, the electronic device can delete the text located before the cursor 104 in the second speaker text display area 1031 and display the user interface shown in (2) in Figure 1B.

[0114] It is understood that the text before or after the cursor refers to the text before or after the cursor in the writing order. For example, if the writing order is from left to right, the text before the cursor refers to the text to the left of the cursor, and the text after the cursor refers to the text to the right of the cursor. For another example, if the writing order is from right to left, the text before the cursor refers to the text to the right of the cursor, and the text after the cursor refers to the text to the left of the cursor.

[0115] In some embodiments of the present application, the electronic device may be set with a default writing order. For example, the default writing order set in the electronic device may be from left to back.

[0116] In some embodiments of the present application, a control refers to an interactive element in a user interface that is used to receive user input, display information, or perform a specific function. In other embodiments of the present application, a control refers to a graphical user interface element. For example, a control can be an icon.

[0117] As shown in Figures 1A and 1B, currently, when the electronic device converts the collected audio data into text, the converted text can no longer be edited. However, when the electronic device stops recording and no longer recognizes the collected audio data, the electronic device can edit the converted text.

[0118] If you edit the recognized text during the speech-to-text conversion process, the edited content and the recognized content will overwrite each other, resulting in the loss of the edited content or the recognized content. The following is a detailed explanation:

[0119] During the speech-to-text conversion process, the electronic device can continuously write the text converted from the audio data into the corresponding speech recognition object. The speech recognition object can include the text converted from the audio data. In some embodiments of the present application, the speech recognition object can include recognized text. Exemplarily, the speech recognition object can be SpeakerDataList. As shown in Figure 2A, the text originally stored in the speech recognition object is 123456, that is, the recognized text (the text converted from the collected audio data) is 123456. The user can click the corresponding area on the display screen to edit the recognized text (i.e., 123456). In response to this click operation, at time t1, the electronic device reads the recognized text, i.e., 123456, from the speech recognition object. The user can edit the recognized text, such as adding the number 0 in front of the recognized text. Accordingly, the electronic device can obtain the edited text, i.e., 0123456. Furthermore, at time t3, the electronic device can write the edited text into the speech recognition object, so that the text in the speech recognition object changes from 123456 to 0123456. It is understandable that t1 is earlier than t3.

[0120] Since the process of collecting audio data and converting the audio data into text (i.e., recognizing the audio data to obtain text) is ongoing, the electronic device is also constantly acquiring ASR data during the user's editing process, and the ASR data may include text converted from the audio data. As shown in Figure 2A, during the user's editing of the read recognized text, the electronic device can acquire new ASR data (e.g., 7 and 8 shown in Figure 2A), and then add the text in the ASR data (which can be understood as new recognized text) to the read recognized text. It can be understood that the text in the speech recognition object is still 123456 before the above addition is successfully performed. Specifically, at time t2, the electronic device can read the recognized text from the speech recognition object. Since the user has not completed the above editing at this time, the recognized text read by the electronic device is still 123456. Further, the electronic device adds 7 and 8 included in the acquired new ASR data to the read recognized text to obtain the updated text, i.e., 12345678. At time t4, the electronic device can write the updated text into the speech recognition object, so that the text in the speech recognition object changes from 0123456 to 12345678. It can be understood that t2 is earlier than t4, and t1 is earlier than t2, t2 is earlier than t3, and t3 is earlier than t4.

[0121] As shown in FIG2A , the user can edit the recognized text, and during this editing process, the electronic device can update the recognized text based on the acquired new ASR data. However, since the user has not yet completed editing the recognized text when the update begins, the electronic device updates the original recognized text, rather than the user-edited text. This results in the electronic device being able to display the user-edited text but unable to retain the edited text, resulting in the loss of the user's edited content.

[0122] It is understood that the electronic device can update the recognized text based on the new ASR data it obtains. During this update process, the user can trigger the editing of the recognized text. Similar to the situation described above, because the electronic device has not yet completed the update of the recognized text when the user triggers the start of editing, the user edits the original recognized text instead of the updated recognized text. This results in the electronic device being able to retain the edited content but unable to retain the updated recognized text, resulting in the loss of the recognized content.

[0123] Specifically, as shown in Figure 2B, the electronic device can obtain new ASR data (including text 7 and 8 converted from audio data). At time t2', the electronic device can read the recognized text from the speech recognition object. Furthermore, the electronic device can add the text 7 and 8 in the obtained new ASR data to the read recognized text to obtain an updated text, namely 12345678. At time t4', the electronic device can write the updated text into the speech recognition object, so that the text in the speech recognition object changes from 123456 to 12345678. It can be understood that t2' is earlier than t4'.

[0124] During the above-mentioned updating process, the user can click on the corresponding area in the display screen to edit the recognized text. In response to the click operation, the electronic device can read the recognized text, i.e., 123456, from the speech recognition object at time t1'. As shown in Figure 2B, the user can edit the recognized text, such as adding the number 0 in front of the recognized text. Accordingly, the electronic device can obtain the edited text, i.e., 0123456. At time t3', the electronic device can write the edited text into the speech recognition object, so that the text in the speech recognition object will change from 12345678 to 0123456. It can be understood that t1' is earlier than t3', and t2' is earlier than t1', t1' is earlier than t4', and t4' is earlier than t3'.

[0125] In summary, users can edit the recognized text obtained by the electronic device, and the electronic device can also update the recognized text based on the new ASR data obtained, but the two do not affect each other and are relatively independent, which may cause the edited content and the recognized content to overwrite each other, resulting in the loss of the edited content or the recognized content.

[0126] Based on the above content, an embodiment of the present application provides a display method and related equipment. According to this method, a user can trigger editing during the recording process of an electronic device. After the user triggers editing, the electronic device can put the newly acquired ASR data into a queue, and then process the ASR data in the queue after the user finishes editing. It can be understood that after the user finishes editing, the electronic device can still continue to acquire ASR data. In this case, the electronic device can first process the ASR data in the queue, and after the ASR data in the queue is processed, it can process the newly acquired ASR data of the electronic device after the user finishes editing. Through this method, the electronic device can edit the recognized text during the ASR process without losing the user's edited content and the voice recognition content acquired during the editing process.

[0127] It is understood that the electronic device may be a mobile phone, a tablet computer, a personal computer (PC), an in-vehicle device, an augmented reality (AR) / virtual reality (VR) device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA). It is understood that this application does not limit the specific type of electronic device.

[0128] The following first introduces the device involved in the embodiments of the present application.

[0129] 1. Electronic devices

[0130] Please refer to FIG3 , which is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application.

[0131] As shown in Figure 3, the electronic device may include: a processor, an external memory interface, an internal memory, a Universal Serial Bus (USB) interface, a charging management module, a power management module, a battery, antenna 1, antenna 2, a mobile communication module, a wireless communication module, a sensor module, buttons, a motor, an indicator, a camera, a display, and a Subscriber Identity Module (SIM) card slot, etc. The audio module may include a speaker, a receiver, a microphone, a headphone jack, etc., and the sensor module may include a pressure sensor, a gyroscope sensor, an air pressure sensor, a magnetic sensor, an acceleration sensor, a distance sensor, a proximity light sensor, a fingerprint sensor, a temperature sensor, a touch sensor, an ambient light sensor, a bone conduction sensor, etc.

[0132] It is understandable that the structure illustrated in the embodiments of the present application does not constitute a specific limitation on the electronic device. It is understandable that the illustrated components can be implemented in hardware, software, or a combination of software and hardware. In some embodiments of the present application, the electronic device may include more components than illustrated. For example, the electronic device may include other types of sensors. In some other embodiments of the present application, the electronic device may include fewer components than illustrated, or combine certain components, or split certain components, or arrange the components differently. The interface connection relationship between the modules illustrated in the embodiments of the present application is only a schematic illustration and does not constitute a structural limitation on the electronic device.

[0133] A processor may include one or more processing units, such as an application processor (AP), a modem (also known as a baseband processor), a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), an audio digital signal processor (ADSP), a sensor hub, and / or a neural-network processing unit (NPU). The AP is the processor responsible for running the operating system and applications, while the modem is the processor responsible for processing various communication protocols.

[0134] The wireless communication function of the electronic device can be implemented through antenna 1, antenna 2, a mobile communication module, a wireless communication module, and a modem. The modem can interact with the base station through antennas (e.g., antenna 1, antenna 2, etc.). In some embodiments, antenna 1 of the electronic device is coupled to the mobile communication module, and antenna 2 is coupled to the wireless communication module, allowing the electronic device to communicate with the network and other devices using wireless communication technologies.

[0135] Electronic devices can achieve display functions through GPU, display, and application processor.

[0136] A GPU is a microprocessor for image processing that connects the display screen to the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. A processor may include one or more GPUs, which execute program instructions to generate or modify display information. A display screen is used to display images, videos, and the like. In some embodiments, an electronic device may include one or more display screens.

[0137] A camera is used to capture still images or video. An ISP processes the data from the camera. Light passes through the lens and is transmitted to the camera's photosensitive element, where it is converted into an electrical signal. The camera's photosensitive element then transmits this electrical signal to the ISP for processing, transforming it into a visible image. An electronic device may include one or more cameras.

[0138] Internal memory can include one or more RAMs and one or more non-volatile memories (NVMs). RAM can be directly read and written by the processor and can be used to store executable programs (e.g., machine instructions) for the operating system or other running programs, as well as user and application data. NVM can also store executable programs and user and application data, and can be pre-loaded into RAM for direct reading and writing by the processor.

[0139] In the embodiment of the present application, the code for implementing the method described in the embodiment of the present application may be stored in a non-volatile memory. When running the note application, the electronic device may load the executable code stored in the non-volatile memory into the random access memory.

[0140] The external memory interface can be used to connect to an external non-volatile memory to expand the storage capacity of the electronic device.

[0141] Electronic devices can implement audio functions through audio modules, speakers, receivers, microphones, headphone jacks, and application processors.

[0142] In some embodiments of the present application, when the electronic device turns on the voice-to-text function (i.e., the audio data-to-text function mentioned above), it can enable the microphone to collect the sound signal and convert the sound signal into a corresponding electrical signal. In some embodiments of the present application, the audio data involved in the present application can be understood as the electrical signal corresponding to the sound signal.

[0143] The operating system of the electronic device can adopt a layered architecture, an event-driven architecture, a micro-kernel architecture, a microservice architecture, or a cloud architecture. The embodiment of the present application takes the Android operating system with a layered architecture as an example to illustrate the software structure of the electronic device. It should be noted that although the embodiment of the present application is described using the Android operating system (which may be referred to as the Android system) as an example, its basic principles are also applicable to electronic devices based on operating systems such as iOS or Windows.

[0144] FIG4 is a schematic diagram of a software structure of an electronic device provided in an embodiment of the present application.

[0145] The software structure of an electronic device adopts a layered architecture, which divides the software into several layers, each of which has a clear role and division of labor. The layers communicate with each other through software interfaces. Taking the Android system running on an AP as an example, in some embodiments of the present application, the software structure of the Android system is divided into five layers, from top to bottom: the application layer, the application framework layer (Framework), the Android runtime (Android runtime) and system library, the hardware abstraction layer (HAL), and the system kernel layer (Kernel).

[0146] Among them, the application layer may include a series of application packages. The application package may include applications such as camera, gallery, calendar, call, map, WLAN, Bluetooth, music, video, short message, etc. The application layer may also include a note application. The note application can be used to record and store various types of notes, such as multimedia content such as text, pictures and recordings. In some embodiments of the present application, the note application may be a system application. In some embodiments of the present application, the note application may be a third-party application. It is understandable that the name of the note application is only an example given in this application, and this application does not limit the name of the application. The application layer may also include a system UI (System User Interface, system UI). The system UI is used to display the interface of the electronic device, such as displaying the user interface of the note application, displaying the signal icon corresponding to the SIM card, displaying the call interface, etc.

[0147] The application framework layer provides an application programming interface (API) and programming framework for the applications in the application layer. The application framework layer may include some predefined functions. For example, the application framework layer may include a window manager, a content provider, a view system, a phone manager, a resource manager, a notification manager, etc. The phone manager is used to provide call functions for electronic devices, such as the management of call status (including answering, hanging up, etc.). The application framework layer may also include an audio framework. The audio framework can be understood as providing a bridge for applications to access system libraries (e.g., media libraries).

[0148] The runtime is responsible for system scheduling and management. It consists of core libraries and a virtual machine. The core libraries consist of two parts: one containing the functions called by programming languages ​​(e.g., Java) and the other containing the system's core libraries. The application layer and application framework layer run in the virtual machine. The virtual machine executes the application layer and application framework layer programming files (e.g., Java files) as binary files. The virtual machine manages object lifecycles, stacks, threads, security, exceptions, and garbage collection.

[0149] The system library can include multiple functional modules, such as the Surface Manager, Media Libraries, 3D graphics processing libraries (e.g., OpenGL ES), and 2D graphics engines (e.g., SGL). The specific meanings and functions of these functional modules can be found in the relevant technical documentation and will not be explained in detail here.

[0150] The Hardware Abstraction Layer (HAL) is an interface layer between the operating system kernel and upper-level software. Its purpose is to abstract the hardware. The HAL is an abstract interface driven by the device kernel, providing application programming interfaces (APIs) that access the underlying device to higher-level Java API frameworks. The HAL provides a standard interface that exposes device hardware capabilities to higher-level Java API frameworks. The HAL consists of multiple library modules (for example, the camera HAL, audio HAL, etc.). When the system framework API requests access to the portable device's hardware, the operating system loads the library module for that hardware component.

[0151] As you can understand, the audio HAL provides an interface between applications and audio hardware (e.g., microphones, speakers, etc.) to manage the input, output, configuration, and control of audio streams. The main purpose of the audio HAL is to provide a consistent audio interface for upper-layer applications and hide the details of the underlying audio hardware. By using the audio HAL, applications can access and control the functions of the audio hardware, such as audio input, output, encoding, decoding, and mixing.

[0152] The kernel layer is the foundation of the Android system. It is responsible for hardware drivers, networking, power supply, system security, and memory management. The kernel layer acts as an intermediate layer between hardware and software, relaying application requests to the hardware. The kernel layer includes audio drivers, display drivers, camera drivers, and sensor drivers. The audio driver, an intermediate layer between audio software and audio hardware, relays access requests from upper-layer software modules to the audio hardware.

[0153] It should be noted that the software structure diagram of the electronic device shown in FIG4 provided in this application is merely an example and does not limit the specific module divisions within the different layers of the Android system. For details, please refer to the introduction of the Android system software structure in conventional technology. In addition, the method provided in this application can also be implemented based on other operating systems, and this application will not cite examples one by one.

[0154] 2. ASR systems including electronic devices

[0155] As shown in Figure 5, the ASR system may include an electronic device and a speech recognition server. The electronic device and the speech recognition server may establish a communication connection. The communication connection may be a wireless fidelity (Wi-Fi) or an access network such as 3G / 4G / 5G. The network mentioned here may be an access wide area network (e.g., the Internet). In this case, the electronic device may communicate with the speech recognition server via the wide area network (e.g., the Internet). The speech recognition server may be used to recognize audio data and obtain corresponding text.

[0156] As shown in Figure 5, the note-taking application, audio framework, audio HAL, and audio driver in the electronic device all run on the device's access point. The electronic device can detect a user action triggering the start of recording in the note-taking application. In response to this user action, the note-taking application in the electronic device can request the audio framework in the electronic device to collect audio data. The audio framework in the electronic device can then receive the request and, through the audio HAL and audio driver, control the microphone in the electronic device to collect the sound signal. After the microphone in the electronic device collects the sound signal, it can convert the sound signal into an electrical signal, i.e., audio data, and then send the audio data to the access point. The audio driver running on the access point can then obtain the audio data and transmit it upstream to the note-taking application via the audio HAL and audio framework. If the speech-to-text function (i.e., the audio data-to-text conversion mentioned above) is enabled, the note-taking application in the electronic device can transmit the obtained audio data to the modem in the electronic device, which is then transmitted via the antenna and transmitted over the network to the speech recognition server. After receiving the audio data sent by the electronic device, the speech recognition server can recognize the audio data, obtain a speech-to-text result (i.e., the ASR data mentioned above), and then send the speech-to-text result to the electronic device. Accordingly, the electronic device can receive the speech-to-text result via the antenna and modem and transmit the speech-to-text result to the AP. The note-taking application running on the AP can receive the speech-to-text result, process the speech-to-text result to obtain text data, and then send the text data to the display screen for display.

[0157] It is understandable that the text data mentioned in this application can be a binary character code that can be recognized by electronic devices, and different binary character codes can specifically correspond to different texts (for example, Chinese characters, Latin letters, etc.).

[0158] The following describes an editing scenario provided by an embodiment of the present application.

[0159] 1. Record audio data in notes

[0160] The user can trigger the electronic device to start the note application by clicking the application icon of the note application, and can also trigger the electronic device to record in the note application after the electronic device starts the note application. It is understood that the icons involved in this application can refer to files.

[0161] Exemplarily, the desktop 1a shown in FIG6A may include multiple application icons (for example, a weather application icon, a calendar application icon, an email application icon, a settings application icon, an application store application icon, a note application icon 10, an album application icon, etc.). These application icons can be used to trigger the electronic device to start the corresponding application. Among them, the note application icon 10 is an icon of the note application. Users can record text information, picture information, etc. through the note application. After the electronic device receives the user operation of clicking the note application icon 10 as shown in FIG6A, it can display the note application main interface 1b as shown in FIG6B. Optionally, in response to the user operation of clicking the note application icon 10 as shown in FIG6A, the electronic device can display the note application main interface 1b as shown in FIG6B. In some embodiments of the present application, after the electronic device receives the user operation of clicking the note application icon 10 as shown in FIG6A (or in response to the user operation of clicking the note application icon 10 as shown in FIG6A), it can first start the note application and then display the user interface as shown in FIG6B. Optionally, the electronic device may detect a user operation on the note application icon 10 shown in FIG6A . In response to the user operation, the electronic device may launch the note application and display the note application main interface 1b shown in FIG6B . The note application main interface 1b may include icons corresponding to created notes (such as the icon corresponding to "Untitled Notebook" and the icon corresponding to "Audio Notes" as shown in FIG6B ). The note application main interface 1b may also include a new note control 11. The new note control 11 can be used to create a new note.

[0162] Exemplarily, after the electronic device receives the user operation of clicking the new note control 11 as shown in FIG6B , it may display the note creation interface 1c as shown in FIG6C . Optionally, in response to the user operation of clicking the new note control 11 as shown in FIG6B , the electronic device may display the note creation interface 1c as shown in FIG6C . Optionally, the electronic device may detect the user operation acting on the new note control 11 as shown in FIG6B , and in response to the user operation, the electronic device may display the note creation interface 1c as shown in FIG6C . The note creation interface 1c may include a text note creation control 121 and a handwritten note creation control 122. The user may trigger the electronic device to create a text note by clicking the text note creation control 121 or other methods. Similarly, the user may trigger the electronic device to create a handwritten note by clicking the handwritten note creation control 122 or other methods.

[0163] Exemplarily, after the electronic device receives the user operation of clicking the text note create control 121 as shown in Figure 6C, it can display the new note interface 1d as shown in Figure 6D. Optionally, in response to the user operation of clicking the text note create control 121 as shown in Figure 6C, the electronic device can display the new note interface 1d as shown in Figure 6D. In some embodiments of the present application, after the electronic device receives the user operation of clicking the text note create control 121 as shown in Figure 6C (or in response to the user operation of clicking the text note create control 121 as shown in Figure 6C), it can first create a text note and then display the new note interface 1d as shown in Figure 6C. Optionally, the electronic device can detect the user operation acting on the text note create control 121 as shown in Figure 6C, and in response to the user operation, the electronic device can display the new note interface 1d as shown in Figure 6D. The new note interface 1d may include a recording add control 13. The recording add control 13 can be used to trigger the recording of audio data.

[0164] Exemplarily, after the electronic device receives the user operation of clicking the recording addition control 13 as shown in Figure 6D, it can start recording and display the corresponding interface (the user interface as shown in Figure 6E). Optionally, in response to the user operation of clicking the recording addition control 13 as shown in Figure 6D, the electronic device can start recording and display the corresponding interface (the user interface as shown in Figure 6E). Optionally, the electronic device can detect the user operation acting on the recording addition control 13 as shown in Figure 6D, and in response to the user operation, the electronic device can start recording and display the corresponding interface (the user interface as shown in Figure 6E).

[0165] It should be noted that the above content is only an example provided by this application. Users can also create new notes in other ways (for example, in response to the user's operation of clicking the new note control 11, the electronic device displays the new note interface 1d), and this application does not limit this.

[0166] 2. Enable speech-to-text conversion

[0167] In some embodiments of the present application, if an electronic device creates a new note and then records audio in the new note, or records audio in a previously created note that does not include an audio recording, the electronic device may enable the speech-to-text function by default, as well as the display speaker function by default.

[0168] Exemplarily, after the electronic device creates a new note, it may display a new note interface 1d as shown in FIG6D. Further, in response to a user operation of clicking the recording add control 13 included in the new note interface 1d, the electronic device may start recording and enable the voice-to-text function and the speaker display function by default. At the 29th second of the recording, the electronic device may display a recording interface 1e that displays the speaker and the recognized text as shown in (1) in FIG6E. The recording interface 1e that displays the speaker and the recognized text may include a recording-related information display area 14. The recording-related information display area 14 may include relevant information about the electronic device recording, such as the recording time, voiceprint waveform, text recognized based on the recorded audio data, and the corresponding speaker. The recording-related information display area 14 may include a voice-to-text control 141 and a speaker control 142. The voice-to-text control 141 may be used to control the opening and closing of the voice-to-text function, and the speaker control 142 may be used to control the opening and closing of the speaker display function (displaying the speaker corresponding to the recognized text on the display screen). The voice-to-text control 141 in the recording interface 1e that displays the speaker and the recognized text is filled with black, indicating that the speaker control 142 is in the selected state. When the speech-to-text control 141 is selected, the speech-to-text function is enabled. Similarly, the speaker control 142 in the recording interface 1e that displays the speaker and the recognized text is filled with black, indicating that the speaker control 142 is selected. When the speaker control 142 is selected, it indicates that the speaker function is enabled.

[0169] It should be noted that, since the electronic device is recording with the voice-to-text function turned on, the electronic device can recognize the collected audio data during the recording process to obtain the corresponding text. Similar to the user interface shown in (1) in FIG1A , the recording-related information display area 14 in the recording interface 1e that displays the speaker and the recognized text can also include a stop recording control 101, a voiceprint waveform display area 102, and a voice recognition text display area 103. Among them, the voice recognition text display area 103 can display the text recognized by the electronic device based on the audio data collected by it, and the recognized text can specifically include the recognized text (for example, the text shown in the second speaker text display area 1031 and the first speaker text display area 1032) and the text being recognized (for example, the text shown in the recognition text display area 1033).

[0170] It is understandable that since the electronic device not only turns on the speech-to-text function but also turns on the speaker display function, the electronic device can display the speaker corresponding to the text obtained by recognizing the collected audio data on the display screen. As shown in (1) in FIG6E , “00:22” and “Speaker 2” are displayed above the second speaker text display area 1031, indicating that the speaker corresponding to the text in the second speaker text display area 1031 is Speaker 2, and the text is the text obtained by recognizing the words spoken by Speaker 2 starting from the 22nd second of the recording. Similarly, “00:19” and “Speaker 1” are displayed above the first speaker text display area 1032, indicating that the speaker corresponding to the text in the first speaker text display area 1032 is Speaker 1, and the text is the text obtained by recognizing the words spoken by Speaker 1 starting from the 19th second of the recording.

[0171] It should be noted that the electronic device can also indicate that the speech-to-text control 141 and the speaker control 142 are in the selected state in other ways, and this application does not limit this. For example, if a bold line is displayed below the speech-to-text control 141, it indicates that the speech-to-text control 141 is in the selected state. Similarly, if a bold line is displayed below the speaker control 142, it indicates that the speaker control 142 is in the selected state. For another example, if a black box is displayed around the speech-to-text control 141, it indicates that the speech-to-text control 141 is in the selected state. Similarly, if a black box is displayed around the speaker control 142, it indicates that the speaker control 142 is in the selected state.

[0172] It should also be noted that in some scenarios, the speech-to-text control 141 is in an unselectable state (or unavailable state). For example, as shown in Figure 1B, after the electronic device stops recording in the note application, the speech-to-text control 141 only has a pattern outline, and the pattern outline is gray, indicating that the speech-to-text control 141 is in an unselectable state. It can be understood that when the speech-to-text control 141 is in an unselectable state, the electronic device does not respond to the user's operation of clicking on the speech-to-text control 141. This means that when the speech-to-text control 141 is unselectable, the electronic device cannot turn on the speech-to-text function.

[0173] Similarly, in some embodiments of the present application, in some scenarios, the speaker control 142 and other controls in the user interface of the note application may also be in an unselectable state. For details, please refer to the above description of the speech-to-text control 141, and this application will not go into details here.

[0174] In some embodiments of the present application, if an electronic device creates a new note and then records audio in the new note, or records audio in a previously created note that does not include an audio recording, the electronic device may enable the speech-to-text function by default, but disable the display speaker function by default.

[0175] Exemplarily, in response to a user operation on clicking the recording add control 13 included in the new note interface 1d, the electronic device can start recording, and the speech-to-text function is turned on by default, and the display speaker function is turned off by default. At the 29th second of the recording, the electronic device can display a recording interface 2e that displays the recognition text but does not display the speaker as shown in (2) in Figure 6E. Similar to the recording interface 1e that displays the speaker and the recognition text, the recording interface 2e that displays the recognition text but does not display the speaker may include a recording-related information display area 14. The speech-to-text control 141 in the recording interface 2e that displays the recognition text but does not display the speaker is filled with black, which indicates that the speech-to-text control 141 is in the selected state, but the speaker control 142 only has a pattern outline and is not filled with black, which indicates that the speaker control 142 is in the unselected state. In this case, the speech-to-text function is turned on, and the display speaker function is not turned on.

[0176] It should be noted that since the electronic device is recording with the speech-to-text function turned on but the speaker display function not turned on, the electronic device can recognize the collected audio data during the recording process to obtain the corresponding text, but the speaker corresponding to the text is not displayed. Similar to the recording interface 1e that displays the speaker and the recognized text, the recording interface 2e that displays the recognized text but does not display the speaker can include the text being recognized. As shown in Figure 6E, the speech recognition text display area 103 in the recording interface 2e that displays the recognized text but does not display the speaker can include a recognition text display area 1033, and the recognition text display area 1033 includes the text being recognized. However, unlike the recording interface 1e that displays the speaker and the recognized text, in the recording interface 2e that displays the recognized text but does not display the speaker, all the recognized text is directly displayed in one area, rather than being displayed separately in the areas corresponding to the corresponding speakers. As shown in FIG6E , the recording interface 2e that displays the recognized text but does not display the speaker includes the text display area 1034 for all speakers, but does not include the second speaker text display area 1031 and the first speaker text display area 1032, and the text in the text display area 1034 for all speakers includes the text in the second speaker text display area 1031 and the first speaker text display area 1032. In this case, the recognized text display area 130 includes the text display area 1034 for all speakers, but does not include the display areas corresponding to the speakers (e.g., the second speaker text display area 1031 and the first speaker text display area 1032).

[0177] It is understandable that this application does not restrict whether the voice-to-text function and the speaker display function are turned on by default when the electronic device is recording.

[0178] It is understandable that the user can choose whether to turn on the speech-to-text function and the display speaker function. Exemplarily, after the electronic device receives the user operation of clicking the speaker control 142 as shown in (2) in Figure 6E, it can display the recording interface 1e of displaying the speaker and recognizing the text as shown in (1) in Figure 6E. Optionally, in response to the user operation of clicking the speaker control 142 as shown in (2) in Figure 6E, the electronic device can display the recording interface 1e of displaying the speaker and recognizing the text as shown in (1) in Figure 6E. In some embodiments of the present application, optionally, after the electronic device receives the user operation of clicking the speaker control 142 as shown in (2) in Figure 6E (or in response to the user operation of clicking the speaker control 142), the electronic device can turn on the display speaker function and then display the recording interface 1e of displaying the speaker and recognizing the text as shown in (1) in Figure 6E. The electronic device can detect the user operation acting on the speaker control 142 as shown in (1) in Figure 6E, and in response to the user operation, the electronic device can turn on the display speaker function and display the recording interface 1e of displaying the speaker and recognizing the text as shown in (1) in Figure 6E.

[0179] In some embodiments of the present application, if the electronic device continues recording in a note that has been created and includes a recording, the electronic device can determine whether to turn on the speech-to-text function and the display speaker function during the continued recording process based on whether the speech-to-text function and the display speaker function were turned on when the recording was stopped last time.

[0180] In some embodiments of the present application, if the electronic device is recording in a note for the first time, the electronic device may turn on the speech-to-text function by default and turn off the display speaker function by default. However, if this is not the first time that the electronic device is recording in a note, the electronic device may determine whether the speech-to-text function and the display speaker function need to be turned on for this recording based on the user's historical selections.

[0181] For example, if the user triggered the electronic device to turn on the speech-to-text function during the last recording in the notes, then the electronic device can turn on the speech-to-text function by default when the user triggers the electronic device to record in the notes this time. Similarly, if the user triggered the electronic device to turn off the speech-to-text function during the last recording in the notes, then the electronic device can turn off the speech-to-text function by default when the user triggers the electronic device to record in the notes this time. Similarly, the on / off mechanism of the display speaker function can refer to the on / off mechanism of the above-mentioned speech-to-text function, and this application will not repeat it here.

[0182] 3. Edit the recognized text during speech-to-text conversion

[0183] If the speech-to-text function is turned on, after the electronic device starts recording, it can obtain the text obtained by recognition based on the collected audio data (including the recognized text and the text being recognized) and display the text on the display. During the process of the electronic device obtaining and displaying the text recognized by the collected audio data, the user can edit the recognized text. After the editing is completed, the electronic device can continue to display the recognized text.

[0184] Exemplarily, at the 29th second of the recording, the electronic device may display a recording interface 1e displaying the speaker and recognition text as shown in (1) in FIG6E . The relevant description of the recording interface 1e displaying the speaker and recognition text can be referred to above. After the electronic device receives the user operation of clicking the second speaker text display area 1031 as shown in (1) in FIG6E , it may display the initial editing interface 1f as shown in FIG6F . Optionally, in response to the user operation of clicking the second speaker text display area 1031 as shown in (1) in FIG6E , the electronic device may display the initial editing interface 1f as shown in FIG6F . In some embodiments of the present application, after the electronic device receives the user operation of clicking the second speaker text display area 1031 as shown in (1) in FIG6E (or in response to the user operation of clicking the second speaker text display area 1031), the electronic device may determine the cursor position, continue recording, and then display the initial editing interface 1f as shown in FIG6F . Alternatively, the electronic device may detect a user operation on the second speaker text display area 1031 as shown in (1) of FIG6E . In response to the user operation, at the 31st second of the recording, the electronic device may display an initial editing interface 1f as shown in FIG6F . The initial editing interface 1f may include a cursor 15 and a keyboard 16. The cursor 15 is displayed in the second speaker text display area 1031. The keyboard 16 may include a delete button 161. Delete button 161.

[0185] Exemplarily, after receiving a user operation of clicking the delete button 161 in the initial editing interface 1f, the electronic device may display the editing interface 1g shown in FIG6G . Optionally, in response to the user operation of clicking the delete button 161 in the initial editing interface 1f, the electronic device may display the editing interface 1g shown in FIG6G . In some embodiments of the present application, after receiving a user operation of clicking the delete button 161 in the initial editing interface 1f (or in response to the user operation of clicking the delete button 161 in the initial editing interface 1f), the electronic device may delete the text before the cursor 15 in the second speaker text display area 1031 (e.g., the number "6" to the left of the cursor), continue recording, and then display the editing interface 1g again. Optionally, the electronic device may detect the user operation of clicking the delete button 161 in the initial editing interface 1f. In response to the user operation, the electronic device may delete the text before the cursor 15 (e.g., to the left of the cursor 15) in the second speaker text display area 1031, and during the text deletion process, the electronic device may continue recording. As shown in FIG6G , the electronic device may delete the number “6” located to the left of the cursor 15 in the text included in the second speaker text display area 1031 and display the editing interface 1g at the 35th second of the recording.

[0186] For example, after receiving a user operation of clicking the keyboard 16 in the editing interface 1g, the electronic device may display the editing completion interface 1h shown in FIG6H . Optionally, in response to the user operation of clicking the keyboard 16 in the editing interface 1g, the electronic device may display the editing completion interface 1h shown in FIG6H . In some embodiments of the present application, after receiving a user operation of clicking the keyboard 16 in the editing interface 1g (or in response to the user operation of clicking the keyboard 16 in the editing interface 1g), the electronic device may add characters (e.g., the numbers "888") between the text before the cursor 15 and the text after the cursor 15 included in the second speaker text display area 1031, continue recording, and then display the editing completion interface 1h shown in FIG6H . Optionally, the electronic device may detect a user operation on the keyboard 16 in the editing interface 1g. In response to the user operation, the electronic device may add characters, such as the three numbers "888", between the text before the cursor 15 and the text after the cursor 15 included in the second speaker text display area 1031, and may continue recording. As shown in FIG6H , at the 40th second of the recording, the electronic device may display the editing completion interface for 1 hour.

[0187] Exemplarily, after the electronic device receives a user operation of clicking a blank area in the editing completion interface 1h, it can display the recording interface 1i for exiting editing as shown in FIG6I . Optionally, in response to the user operation of clicking a blank area in the editing completion interface 1h, the electronic device can display the recording interface 1i for exiting editing as shown in FIG6I . In some embodiments of the present application, after the electronic device receives a user operation of clicking a blank area in the editing completion interface 1h (or in response to the user operation of clicking a blank area in the editing completion interface 1h), it can exit the editing state and add the recognized text obtained during the editing process to the second speaker text display area 1031, and then display the recording interface 1i for exiting editing as shown in FIG6I . Optionally, the electronic device can detect the user operation acting on the blank area in the editing completion interface 1h. In response to the user operation, the electronic device can exit the editing state and add the recognized text obtained during the editing process to the corresponding display area corresponding to the recognized text. It can be understood that since the speaker corresponding to the recognized text obtained during the editing process is speaker 2, the electronic device can add the text to the text shown in the second speaker text display area 1031, and the electronic device can continue recording and display the recording interface 1i for exiting editing as shown in Figure 6I at the 42nd second of the recording.

[0188] It is understandable that the user can also edit the text in the first speaker text display area 1032 / all speakers text display area 1034. The specific process can be referred to above and will not be described in detail in this application.

[0189] It should be noted that if the speaker corresponding to the recognized text obtained during the editing process is not Speaker 2, but other speakers (for example, Speaker 1, or speakers other than Speaker 1 and Speaker 2), the electronic device can display the recognized text obtained during the editing process in a new display area. It can be understood that the new display area is different from the display areas corresponding to the existing Speaker 1 and Speaker 2, that is, different from the first speaker text display area 1032 and the second speaker text display area 1031. In some embodiments of the present application, after the user edits the text in the all-speaker text display area 1034, the electronic device can add the recognized text obtained during the editing process to the all-speaker text display area 1034.

[0190] For example, if the speaker corresponding to the recognized text obtained during the editing process is Speaker 1, the recognized text obtained during the editing process can be displayed in the third speaker text display area 1035 (as shown in FIG6J ). The third speaker text display area 1035 is a display area different from the second speaker text display area 1031 , the first speaker text display area 1032 , and the identifying text display area 1033 .

[0191] It can be understood that when the user is editing the recognized text, the electronic device no longer adds the recognized text obtained in the editing process to the corresponding display area corresponding to the recognized text, but still displays the newly recognized text in the editing process in the display area corresponding to the text being recognized.

[0192] For example, as shown in Figures 6F, 6G and 6H, when the user deletes and adds text in the second speaker text display area 1031, the newly recognized text is displayed in the recognition text display area 1033, while the text in the second speaker text display area 1031 and the first speaker text display area 1032 does not change. Specifically, as shown in Figure 6E, at the 29th second of the recording, the text in the text display area 1033 in the recognition is "123456", and after the user triggers editing, as shown in Figure 6F, at the 31st second of the recording, the text in the text display area 1033 in the recognition is still "123456", and the user deletes the text in the second speaker text display area 1031 (for example, deletes the number "6"), as shown in Figure 6G, at the 35th second of the recording, the text in the text display area 1033 in the recognition becomes "123456789", and the user further adds text in the second speaker text display area 1031 (for example, adds the number "888"), as shown in Figure 6H, at the 40th second of the recording, the text in the text display area 1033 in the recognition becomes "1234567890".

[0193] In some embodiments of the present application, when the electronic device is in an editing state, the text recognized during the editing process is continuously accumulated in the recognized text display area 1033, that is, the newly recognized text in the recognized text display area 1033 can be continuously accumulated.

[0194] For example, when the electronic device is in editing mode, at the 31st second of the recording, the text in the recognized text display area 1033 may be "123456". From the 31st second to the 35th second of the recording, the newly recognized text "789" is displayed after the original text in the recognized text display area 1033, that is, "789" is displayed after "123456". Therefore, at the 35th second of the recording, the text in the recognized text display area 1033 is no longer "123456", but "123456789". Similarly, from the 35th second to the 40th second of the recording, the newly recognized text "0" is displayed after the original text in the display area, that is, "0" is displayed after "123456789". Therefore, at the 40th second of the recording, the text in the recognized text display area 1033 is no longer "123456789", but "1234567890".

[0195] In some embodiments of the present application, when the electronic device is in an editing state, the currently recognized text in the recognition text display area 1033 is continuously updated, that is, the newly recognized text in the recognition text display area 1033 can be continuously replaced.

[0196] For example, when the electronic device is in the editing state, at the 31st second of the recording, the text in the text display area 1033 in the recognition process may be "123456". From the 31st second to the 35th second of the recording, the original text in the text display area 1033 in the recognition process is replaced by the newly added recognition text "123456789", that is, the "123456" in the text display area 1033 in the recognition process is replaced by "123456789". Therefore, at the 35th second of the recording, the text in the text display area 1033 in the recognition process is no longer It is not "123456" but "123456789". Similarly, from the 35th to the 40th second of the recording, the original text in the recognized text display area 1033 is replaced by the newly recognized text "1234567890", that is, the "123456789" in the recognized text display area 1033 is replaced by "1234567890". Therefore, at the 40th second of the recording, the text in the recognized text display area 1033 is no longer "123456789", but "1234567890".

[0197] It should be noted that the user interfaces shown in Figures 6A-6J are only examples provided by this application. The user interface displayed by the electronic device during the speech-to-text and editing process may also include other elements and layouts, and this application does not impose any restrictions on this.

[0198] A display method provided in an embodiment of the present application is introduced below based on Figures 7A-7D.

[0199] Please refer to Figures 7A-7D, which are a set of flow charts of a display method provided in an embodiment of the present application.

[0200] The method may include but is not limited to the following steps:

[0201] 1. Start recording, create a queue, and obtain the speech-to-text results (as shown in Figure 7A)

[0202] S101: In response to a user operation triggering the start of recording, the note-taking application creates a queue and sets a flag bit to false.

[0203] The user can trigger the electronic device to start recording in the note application. In some embodiments of the present application, when there is no recording in the current note, the note application can create a new recording task and start recording after receiving the user operation that triggers the start of recording (for example, clicking the recording add control 13). Exemplarily, as shown in Figure 6D, the user can trigger the start of recording by clicking the recording add control 13, and accordingly, the note application can receive the click operation on the recording add control 13. In some other embodiments of the present application, when there is a recording in the current note, the note application can continue recording after receiving the user operation that triggers the start of recording. Exemplarily, as shown in (1) in Figure 1B, the user can trigger the start of recording by clicking the continue recording control 106, and accordingly, the note application can receive the click operation on the continue recording control 106.

[0204] Of course, the user can also trigger the electronic device to start recording in the note application through other methods (for example, gestures, voice control, etc.), and this application does not limit this.

[0205] In some embodiments of the present application, once the user triggers the electronic device to start recording, the electronic device may enable the speech-to-text function. In this case, the electronic device may execute steps S101 to S109, and the steps shown in Figures 7B to 7D. In some embodiments of the present application, although the user triggers the electronic device to start recording, the electronic device does not enable the speech-to-text function. In this case, the electronic device may execute steps S101 to S106. Once the user triggers the electronic device to enable the speech-to-text function, the electronic device may continue to execute steps S107 to S109, and the steps shown in Figures 7B to 7D.

[0206] In some embodiments of the present application, if there is no recording in the current note and the note application receives a user operation that triggers the start of recording, then in response to the user operation that triggers the start of recording, the note application can create a queue and set a mark bit.

[0207] In some other embodiments of the present application, the presence of an audio recording in the current note indicates that the note application has created a queue and set a flag. If, in this case, the note application receives a user operation that triggers the start of audio recording, the note application can create a queue and set a flag after the user triggers the display of the current note, without having to wait until the user triggers the start of audio recording before creating a queue and setting a flag. For example, in the case where the user triggers the electronic device to display the user interface shown in FIG1B , the note application can create a queue and set a flag without having to wait until the user clicks the continue recording control 106 to trigger the electronic device to continue recording before creating a queue and setting a flag.

[0208] It is understandable that the flag bit is the flag bit corresponding to the queue. In some embodiments of the present application, each time the user triggers to exit the note, the note application can clear the flag bit and the queue. When the user triggers to open the note next time, the note application can set the flag bit to false again and create a queue.

[0209] In some embodiments of the present application, the flag bit can be set to false by default. When the flag bit is false, it indicates that the electronic device is not in the editing state, that is, the user has not triggered the editing of the recognized text. In this case, the text converted from the audio data will not be placed in the queue, but will be processed directly. When the flag bit is true, it indicates that the electronic device is in the editing state, that is, the user has triggered the editing of the recognized text. In this case, the text converted from the audio data will be placed in the queue instead of being processed directly.

[0210] It is understood that this application does not limit the representation of the flag bit. For example, the flag bit can be represented as isEditByUser. In this case, in response to the user operation that triggers the start of recording, the note-taking application can set isEditByUser=false.

[0211] It is understood that after the note-taking application detects the user operation that triggers the start of recording, it can also set the flag bit to other content (e.g., a number, other string, etc.), and this application does not limit this. For example, in response to the user operation that triggers the start of recording, the note-taking application can set the flag bit to 0.

[0212] S102: In response to a user operation triggering the start of recording, the note-taking application sends an audio data request to the audio framework.

[0213] After the note-taking application receives the user operation that triggers the start of recording, in response to the user operation that triggers the start of recording, the note-taking application may request the audio framework to collect audio data.

[0214] In some embodiments of the present application, in response to a user operation triggering the start of recording, the note-taking application may send information requesting audio data to the audio framework.

[0215] In some embodiments of the present application, in response to the user operation that triggers the start of recording, the note-taking application can call the startRecording() method to record. In the startRecording() method, an object can be created, the audio source can be set to a microphone, and the recorded audio data can be written to a file in Pulse Code Modulation (PCM) format. It can be understood that the steps of recording using the startRecording() method can specifically include steps S102 to S106.

[0216] It is understandable that this application does not limit the order in which the electronic device executes step S101 and step S102. The note-taking application in the electronic device can execute step S101 first and then step S102, or can execute step S102 first and then step S101, or can execute step S101 and step S102 at the same time.

[0217] S103: The audio framework configures recording parameters.

[0218] After the audio framework receives the request from the note application to collect audio data, it can create an object (for example, a MediaRecorder instance, an AudioRecord instance) and configure the recording parameters. It is understandable that the recording parameters may include one or more of the following: audio source, format of the output file (i.e., the output audio file), save path of the output file, sampling rate, audio channel, and audio encoder. In some embodiments of the present application, the audio framework can set the audio source to a microphone audio source. In this way, the audio data recorded by the note application is the audio data collected by the microphone.

[0219] It is understood that the audio source may include a default audio source, a microphone audio source, a camera audio source, a voice recognition audio source, and a call audio source. The format of the output file may include a default output format, an MPEG-4 file format, and may also include an 8-bit PCM encoding format, a 16-bit PCM encoding format, and a floating-point PCM encoding format. The sampling rate may include 44.1 kilohertz (kHz), 22.05kHz, 16kHz, and 8kHz. The audio channel may include mono and stereo. The audio encoder may include a default audio encoder, an ACC encoder, and a high-efficiency AAC (HE-AAC) audio encoder.

[0220] MPEG, short for Moving Picture Experts Group, is an organization dedicated to developing international standards in the multimedia field and the creator of the MPG / MPEG format. Currently, the MPEG compression standards include MPEG-1, MPEG-2, and MPEG-4. PCM is a coding method for digital communications. Its primary encoding process involves sampling analog signals such as voice and images at regular intervals to discretize them, rounding the sampled values ​​to integers based on hierarchical units, and representing the amplitude of the sampling pulses using a set of binary codes.

[0221] It is understandable that the full English name of ACC is Advanced Audio Coding, which means advanced audio coding in Chinese. It is an audio coding technology based on MPEG-2.

[0222] S104: The audio framework notifies the microphone to collect audio data through the audio HAL and audio driver.

[0223] After the audio framework configures the recording parameters, you can start recording. Since the audio source has been set to the microphone, the audio framework can notify the microphone to collect audio data through the audio HAL and audio driver.

[0224] Accordingly, the microphone can receive notifications of collected audio data transmitted by the audio framework through the audio HAL and audio driver.

[0225] S105: The microphone collects audio data.

[0226] After receiving the notification of collecting audio data transmitted by the audio framework through the audio HAL and audio driver, the microphone can collect sound signals and generate audio data.

[0227] S106: The microphone sends audio data to the note-taking application through the audio driver, audio HAL, and audio framework.

[0228] After the microphone collects audio data, it can be sent to the note-taking application through the audio driver, audio HAL, and audio framework.

[0229] It is understandable that during the recording process, step S105 and step S106 will be continuously executed, that is, the microphone will continuously collect audio data and continuously send the collected audio data to the note-taking application.

[0230] S107: The note-taking application sends the audio data to the speech recognition server via the modem.

[0231] After receiving the audio data collected by the microphone, the note-taking application may send the audio data to the modem, and then send the audio data to the speech recognition server through the modem.

[0232] Correspondingly, the speech recognition server can receive the audio data sent by the Modem.

[0233] S108: The speech recognition server recognizes the audio data and obtains a speech-to-text conversion result.

[0234] After receiving the audio data sent by the modem, the speech recognition server can recognize the audio data and obtain the corresponding speech-to-text result. It should be noted that the speech-to-text result is marked (or determined / defined) by the speech recognition server.

[0235] It is understandable that the speech-to-text result may include the text data (which can also be understood as the corresponding text) and type obtained after the audio data is recognized. Among them, the type can be used to indicate whether the text included in the speech-to-text result is recognized text or text being recognized. In some embodiments of the present application, the speech-to-text result may also include a speaker identification, etc. The speaker identification is used to indicate the speaker corresponding to the text data included in the speech-to-text result. It is understandable that the speech-to-text result may also include other parameters related to audio and text (for example, offset), which is not limited by the present application.

[0236] In some embodiments of the present application, the type of the speech-to-text result can be represented by part or final. If the type of the speech-to-text result is part (for example, the speech-to-text result includes part, or the field indicating the type in the speech-to-text result is part, etc.), it indicates that the text included in the speech-to-text result is the text being recognized. If the type of the speech-to-text result is final (for example, the speech-to-text result includes final, or the field indicating the type in the speech-to-text result is final, etc.), it indicates that the text included in the speech-to-text result is recognized text.

[0237] It is understood that the type of the speech-to-text result can be represented by text, character strings, etc., and this application does not impose specific restrictions on this. For example, the type of the speech-to-text result can be represented by 0 and 1. If the type of the speech-to-text result is 0, it indicates that the text included in the speech-to-text result is the text being recognized. If the type of the speech-to-text result is 1, it indicates that the text included in the speech-to-text result is the recognized text.

[0238] In some embodiments of the present application, a speech-to-text result of type part represents the intermediate result mentioned above. The text obtained by processing the speech-to-text result of type part can be displayed in the display area corresponding to the text being recognized. For example, as shown in Figures 6E-6J, the text obtained by processing the speech-to-text result of type part can be displayed in the recognition text display area 1033.

[0239] In some embodiments of the present application, a speech-to-text result of type final represents the final result mentioned above. The text obtained by processing the speech-to-text result of type final can be displayed in the display area corresponding to the recognized text. For example, as shown in Figures 6E-6J, the text obtained by processing the speech-to-text result of type final can be displayed in the second speaker text display area 1031, the first speaker text display area 1032, and the third speaker text display area 1035.

[0240] It is understandable that in some embodiments of the present application, the speech-to-text result may be a custom class composed of a variety of data structures.

[0241] S109: The speech recognition server sends the speech-to-text result to the note-taking application via the modem.

[0242] After the speech recognition server obtains the speech-to-text result, it can send the speech-to-text result to the modem of the electronic device, and then the modem sends the speech-to-text result to the note-taking application. Correspondingly, the note-taking application can receive the speech-to-text result sent by the speech recognition server.

[0243] It is understood that while the speech-to-text function is enabled during recording, steps S107-S109 will continue to execute. That is, the note-taking application can continuously send audio data to the speech recognition server, and the speech recognition server can continuously recognize the audio data and obtain corresponding speech-to-text results, which the speech recognition server can also continuously send the obtained speech-to-text results to the note-taking application.

[0244] In some embodiments of the present application, among the multiple speech-to-text results sent by the speech recognition server to the note application within a period of time (for example, 5 seconds), the number of speech-to-text results of type part is greater than the number of speech-to-text results of type final. In one possible implementation, after the speech recognition server sends k speech-to-text results of type part to the note application, it sends 1 speech-to-text result of type final to the note application, and then repeats the action. Wherein, k is a positive integer greater than 1. It should be noted that k can change each time the speech recognition server performs this action. For example, within 30 seconds, the speech recognition server first continuously sends 5 speech-to-text results of type part to the note application, then sends 1 speech-to-text result of type final to the note application, then continuously sends 3 speech-to-text results of type part to the note application, and then sends 1 speech-to-text result of type final to the note application.

[0245] Optionally, after executing step S109, the electronic device may further execute at least one of steps S201 to S217.

[0246] Optionally, after executing step S109 , the electronic device may further execute at least one of steps S301 to S308 .

[0247] 2. Process the speech-to-text results (as shown in Figure 7B)

[0248] S201: The note application determines whether the flag bit is true.

[0249] After receiving the speech-to-text result from the speech recognition server, the note-taking application can determine whether the flag bit is true. If the flag bit is true, the electronic device can execute steps S214-S217; if the flag bit is not true, the electronic device can execute steps S202-S213.

[0250] S202: The note application determines whether the queue is empty.

[0251] If the flag is false instead of true, it indicates that the user has not triggered editing at this time. In this case, the note-taking application can determine whether the queue is empty. If the queue is empty, it indicates that there are no speech-to-text results in the queue, and the electronic device can continue to perform steps S204-S212. If the queue is not empty, it indicates that there are no speech-to-text results in the queue, and the electronic device can proceed to step S213.

[0252] S203: The note-taking application processes the speech-to-text conversion result to obtain text data.

[0253] If the flag is not true and the queue is empty, the note-taking application can process the speech-to-text result to obtain the corresponding text data.

[0254] S204: When the type of the speech-to-text result is part, the note application writes the processed text data into the first speech recognition object, and overwrites the original text data in the first speech recognition object.

[0255] If the flag bit is not true and the queue is empty, the note application can determine the type of the speech-to-text result. If the type of the speech-to-text result is part, it indicates that the speech-to-text result received by the note application includes the text being recognized, that is, the text data obtained after the note application processes the speech-to-text result is the text being recognized. In this case, the note application can write the text data obtained by processing the speech-to-text result (which can be abbreviated as the processed text data) into the first speech recognition object, and overwrite the original text data in the first speech recognition object (which can be understood as replacing the original text data in the first speech recognition object).

[0256] It is understood that the first speech recognition object may be the speech recognition object mentioned above, which stores the text being converted from the audio data, ie the text being recognized. In some embodiments of the present application, the first speech recognition object may be a string object.

[0257] For example, the original text data in the first speech recognition object may be "01234556", and the text data obtained by the note-taking application from the speech-to-text conversion result may be "01234567". The note-taking application can then write "01234567" into the first speech recognition object, overwriting "01234556". In this case, the text data in the first speech recognition object changes from "01234556" to "01234567".

[0258] In some embodiments of the present application, the first speech recognition object may be a string object corresponding to the recognized text display area 1033. The text stored in the first speech recognition object is displayed in the recognized text display area 1033.

[0259] For example, the original text data in the first speech recognition object may be "123456" in the recognition text display area 1033 shown in Figure 6E. If the user does not trigger editing, the note-taking application can process the speech-to-text result and obtain the text data "1234567". The note-taking application can write "1234567" into the first speech recognition object, overwriting the original "123456". In this case, the text data in the first speech recognition object changes from "123456" to "1234567".

[0260] For example, the original text data in the first speech recognition object may be "987654" in the recognition text display area 1033 shown in Figures 6I and 6J. If the user does not trigger editing, the note-taking application can process the speech-to-text result and obtain the text data "987654321". The note-taking application can write "987654321" into the first speech recognition object, overwriting the original "987654". In this case, the text data in the first speech recognition object changes from "987654" to "987654321".

[0261] S205: The note application transmits the text data in the first voice recognition object to the display screen.

[0262] The note application can transmit the text data in the first voice recognition object to the display screen. Correspondingly, the display screen can receive the text data in the first voice recognition object transmitted by the note application, that is, the text data obtained by processing the speech-to-text result mentioned in step S203.

[0263] S206: The display screen displays the text data in the first voice recognition object.

[0264] After the display screen receives the text data (i.e., text data) in the first voice recognition object transmitted by the note application, the text data can be displayed in the display area corresponding to the first voice recognition object (for example, the recognized text display area 1033 shown in Figures 6E, 6I, and 6J).

[0265] For example, the text data obtained by the note-taking application when processing the speech-to-text result may be "123456", then the text data in the first voice recognition object received by the display screen is "123456", and the display screen can display the text data in the display area corresponding to the first voice recognition object, such as displaying "123456" in the text display area 1033 in the recognition as shown in Figure 6E.

[0266] For example, the text data obtained by the note-taking application when processing the speech-to-text result may be "987654", then the text data in the first voice recognition object received by the display screen is "987654", and the display screen can display the text data in the display area corresponding to the first voice recognition object, such as displaying "987654" in the recognition text display area 1033 as shown in Figures 6I and 6J.

[0267] For example, the speech recognition server may first return a speech-to-text result of the "part" type to the electronic device. This speech-to-text result includes the text "01." After receiving the speech-to-text result, the electronic device may process it to obtain the text "01." "01" may be displayed in the recognition text display area 1033. The speech recognition server may then return a speech-to-text result of the "part" type to the electronic device. This speech-to-text result includes the text "0134." After receiving the speech-to-text result, the electronic device may process it to obtain the text "0134." "0134" may replace "01" and be displayed in the recognition text display area 1033. Similarly, the speech recognition server may return a speech-to-text result of the "part" type to the electronic device. This speech-to-text result includes the text "013456." After receiving the speech-to-text result, the electronic device may process it to obtain the text "013456." "013456" may replace "0134" and be displayed in the recognition text display area 1033.

[0268] In some embodiments of the present application, when the type of the speech-to-text result is part, the note-taking application writes the processed text data into the first speech recognition object and splices it behind the original text data in the first speech recognition object.

[0269] For example, the speech recognition server may first return a speech-to-text result of type part to the electronic device, the speech-to-text result including the text "01". After receiving the speech-to-text result, the electronic device may process it to obtain the text "01". "01" may be displayed in the recognition text display area 1033. The speech recognition server may then return a speech-to-text result of type part to the electronic device, the speech-to-text result including the text "34". After receiving the speech-to-text result, the electronic device may process it to obtain the text "34". "34" may be displayed in the recognition text display area 1033 immediately following "01", that is, the text in the recognition text display area 1033 includes "0134". Similarly, the speech recognition server may then return a speech-to-text result of type part to the electronic device, the speech-to-text result including the text "56". After receiving the speech-to-text result, the electronic device may process it to obtain the text "56". “56” may be displayed immediately following “0134” in the recognized text display area 1033 , that is, the text in the recognized text display area 1033 includes “013456”.

[0270] S207: If the type of the speech-to-text result is final and the speaker corresponding to the speech-to-text result is the same as the speaker corresponding to the second speech recognition object, the note-taking application adds the processed text data to the second speech recognition object, which is the most recently created object among the currently created objects.

[0271] If the flag bit is not true and the queue is empty, the note application can determine the type of the speech-to-text result and the speaker corresponding to the speech-to-text result. If the type of the speech-to-text result is final, it means that the speech-to-text result received by the note application includes recognized text, which means that the text data obtained after the note application processes the speech-to-text result is recognized text. In this case, if the speaker corresponding to the speech-to-text result is the same as the speaker corresponding to the second speech recognition object of the note application, the note application can add the processed text data to the latest object created by the note application, specifically, the processed text data can be spliced ​​after the original text data included in the latest created object. In other words, the text data in the second speech recognition object = the original text data in the second speech recognition object + the text data obtained by processing the speech-to-text result.

[0272] In some embodiments of the present application, if the queue is empty, the note application may first determine the type of speech-to-text result, and then determine whether the speaker corresponding to the speech-to-text result is the same as the speaker corresponding to the second speech recognition object. The note application determines whether the speaker corresponding to the speech-to-text result is the same as the speaker corresponding to the most recently created object, which may specifically include: the note application may determine whether the speaker identifier included in the speech-to-text result is the same as the speaker identifier corresponding to the second speech recognition object. If the speaker identifier included in the speech-to-text result is the same as the speaker identifier corresponding to the second speech recognition object, it indicates that the speaker corresponding to the speech-to-text result is the same as the speaker corresponding to the second speech recognition object.

[0273] In some embodiments of the present application, the second speech recognition object may be a character string object.

[0274] In some embodiments of the present application, the second voice recognition object may be a string object corresponding to the second speaker text display area 1031 . The text stored in the second voice recognition object is displayed in the second speaker text display area 1031 .

[0275] For example, as shown in FIG6H , the speaker corresponding to the second speech recognition object may be speaker 2, and the corresponding speaker identifier is 2. The original text data in the second speech recognition object may be "abcde333666888999". After receiving the speech-to-text result, the note-taking application may determine that the type of the speech-to-text result is final, the speaker identifier included in the speech-to-text result is 2, and the text included in the speech-to-text result is "1234567890", that is, the processed text data is "1234567890". The note-taking application may then concatenate the processed text data after the original text data in the second speech recognition object. In this way, the text data in the second speech recognition object becomes "abcde3336668889991234567890".

[0276] In some embodiments of the present application, the note-taking application can call the getAsrContent() method to obtain the original text data in the second voice recognition object, and then add the text data obtained by processing the speech-to-text result to the second voice recognition object through the updateToAsrContent() method. Specifically, the processed text data is spliced ​​behind the original text data in the second voice recognition object to obtain the spliced ​​text data (that is, the original text data in the second voice recognition object + the processed text data), and finally the spliced ​​text data is written into the second voice recognition object through the setAsrContent() method.

[0277] S208: The note application transmits the text data in the second voice recognition object to the display screen.

[0278] The note application can transmit the text data in the second voice recognition object to the display screen. Correspondingly, the display screen can receive the text data in the second voice recognition object transmitted by the note application, including the original text data and text data of the second voice recognition object.

[0279] S209: The display screen displays the text data in the second voice recognition object.

[0280] After the display screen receives the text data in the second voice recognition object transmitted by the note application, the text data can be displayed in the display area corresponding to the second voice recognition object (for example, the second speaker text display area 1031 shown in Figures 6E-6I).

[0281] For example, the original text data in the second speech recognition object may be "abcde333666888999", and the text data obtained by processing the speech-to-text result may be "1234567890". The text data in the second speech recognition object received by the display screen is obtained by concatenating the original text data in the second speech recognition object and the processed text data, that is, "abcde3336668889991234567890". The display screen can display the concatenated text data in the display area corresponding to the second speech recognition object, for example, displaying "abcde3336668889991234567890" in the second speaker text display area 1031 as shown in FIG6I.

[0282] S210: When the type of the speech-to-text result is final and the speaker corresponding to it is different from the speaker corresponding to the second speech recognition object, the note application creates a third speech recognition object and adds the processed text data to the third speech recognition object.

[0283] If the flag is not true and the queue is empty, the note application can determine the type of the speech-to-text result and the speaker corresponding to the speech-to-text result. If the type of the speech-to-text result is final, it means that the speech-to-text result received by the note application includes recognized text, which means that the text data obtained after the note application processes the speech-to-text result is recognized text. In this case, if the speaker corresponding to the speech-to-text result is different from the speaker corresponding to the second speech recognition object, the note application can create a new object (i.e., a third speech recognition object) and add the text data obtained by processing the speech-to-text result to the newly created object.

[0284] In some embodiments of the present application, if the speaker identifier included in the speech-to-text result is different from the speaker identifier corresponding to the second speech recognition object, it indicates that the speaker corresponding to the speech-to-text result is different from the speaker corresponding to the object most recently created by the note application.

[0285] In some embodiments of the present application, the third speech recognition object may be a character string object.

[0286] In some embodiments of the present application, the third voice recognition object may be a string object corresponding to the third speaker text display area 1035 . The text stored in the third voice recognition object is displayed in the third speaker text display area 1035 .

[0287] For example, as shown in FIG6E , the second speech recognition object may be a string object corresponding to the second speaker text display area 1031 , the speaker corresponding to the second speech recognition object is speaker 2, and the corresponding speaker identifier is 2. If the note-taking application receives a speech-to-text result and determines that the type of the speech-to-text result is final and the speaker identifier included in the speech-to-text result is 1, the note-taking application may create a new string object and add the recognized text included in the speech-to-text result to the string object.

[0288] It should be noted that after the note application determines that the flag bit is not true and the queue is empty, it can execute steps S204, S207 and S209 based on Figure 8. As shown in Figure 8, after the note application determines that the flag bit is not true and the queue is empty, it can execute the following steps:

[0289] S2011: The note application determines whether the type of the speech-to-text result is final.

[0290] If the flag is not true and the queue is empty, the note-taking application may first determine whether the speech-to-text result type is final or part. If the speech-to-text result type is final, the note-taking application may proceed to step S2013. If the speech-to-text result type is part, rather than final, the note-taking application may proceed to step S2012.

[0291] S2012: The note application writes the processed text data into the first voice recognition object, overwriting the original text data in the first voice recognition object.

[0292] If the note-taking application determines that the type of the speech-to-text result is not final but part, the note-taking application can write the text data obtained by processing the speech-to-text result into the first speech recognition object and overwrite the original text data in the first speech recognition object. The specific implementation method can refer to the relevant description of step S204, and this application will not go into details about it.

[0293] S2013: The note-taking application determines whether the speaker corresponding to the speech-to-text result is the same as the speaker corresponding to the second speech recognition object.

[0294] If the speaker corresponding to the speech-to-text result is the same as the speaker corresponding to the second speech recognition object, the note application may execute step S2014; if the speaker corresponding to the speech-to-text result is different from the speaker corresponding to the second speech recognition object, the note application may execute step S2015.

[0295] It is understandable that the specific implementation method of step S2031 can be referred to above, and this application will not go into details.

[0296] S2014: The note application adds the processed text data to the second speech recognition object.

[0297] If the note-taking application determines that the speaker corresponding to the speech-to-text result is the same as the speaker corresponding to the second speech recognition object, the note-taking application can add the text data obtained by processing the speech-to-text result to the second speech recognition object. The specific implementation method can refer to the relevant description of step S207, and this application will not go into details about it.

[0298] S2015: The note application creates another new object and adds the processed text data to the newly created object.

[0299] If the note-taking application determines that the speaker corresponding to the speech-to-text result is different from the speaker corresponding to the second speech recognition object, the note-taking application can create another new object (for example, a third speech recognition object) and add the text data obtained by processing the speech-to-text result to the newly created object.

[0300] S211: The note application transmits the text data in the third voice recognition object to the display screen.

[0301] The note application can transmit the text data in the third voice recognition object to the display screen. Correspondingly, the display screen can receive the text data in the third voice recognition object transmitted by the note application, including the text data.

[0302] S212: The display screen displays the text data in the third voice recognition object.

[0303] After the display screen receives the text data in the third voice recognition object transmitted by the note application, it can determine the display area corresponding to the third voice recognition object and display the text data in the display area corresponding to the third voice recognition object.

[0304] For example, the text data in the third voice recognition object of the note application can be "1234567890". Correspondingly, the text data in the third voice recognition object received by the display screen is "1234567890". According to the above, the speaker identifier corresponding to the third voice recognition object can be 1, that is, the speaker corresponding to the third voice recognition object is speaker 1, then the display screen can divide another display area different from the second speaker text display area 1031, the first speaker text display area 1032 and the identifying text display area 1033 to display the text data in the third voice recognition object. As shown in Figure 6J, the display screen can display "1234567890" in the third speaker text display area 1035.

[0305] S213: The note-taking application processes the speech-to-text results in the queue in the order in which they were queued, and then processes the newly received speech-to-text results.

[0306] If the flag bit is not true and the queue is not empty, the note-taking application can process the speech-to-text results in the queue in the order in which they entered the queue, and then process the newly received speech-to-text results. It is understood that the specific implementation of the above-mentioned speech-to-text result processing can be referred to steps S204-S212, and this application will not be repeated here.

[0307] S214: The note application puts the speech-to-text result into a queue.

[0308] If the flag is true, it indicates that the user has triggered editing, that is, the note application enters the editing state (or the electronic device enters the editing state). In this case, the note application can put the speech-to-text result into the queue.

[0309] In some embodiments of the present application, no matter whether the type of the speech-to-text result received by the note-taking application is part or final, the note-taking application may put the speech-to-text result into a queue.

[0310] In some embodiments of the present application, when the type of the speech-to-text result received by the note application is final, the note application may put the speech-to-text result into a queue, and when the type of the speech-to-text result received by the note application is part, the note application may not put the speech-to-text result into a queue, but directly process it according to step S204.

[0311] S215: When the type of the speech-to-text result is final, the note application processes the speech-to-text result to obtain text data. If the original text data in the first speech recognition object is obtained based on the speech-to-text result of type part, the text data obtained by processing the speech-to-text result is used to overwrite the original text data in the first speech recognition object. If the original text data in the first speech recognition object is obtained based on the speech-to-text result of type final, the text data obtained by processing the speech-to-text result is added to the first speech recognition object.

[0312] If the flag is true, the note application can determine the type of the speech-to-text result. If the type of the speech-to-text result is part, the note application can put the speech-to-text result into the queue. If the type of the speech-to-text result is final, the note application can not only put the speech-to-text result into the queue, but also process the speech-to-text result to obtain the corresponding text data. Furthermore, the note application can determine whether the original text data in the first speech recognition object is the text data obtained based on the speech-to-text result of the corresponding type part, or the text data obtained based on the speech-to-text result of the corresponding type final. If the original text data in the first speech recognition object is the text data obtained based on the speech-to-text result of the corresponding type part, the note application can write the corresponding text data obtained by processing the speech-to-text result into the first speech recognition object and overwrite the original text data in the first speech recognition object. It can also be understood that the original text data in the first speech recognition object is replaced with the corresponding text data obtained by processing the speech-to-text result. If the original text data in the first speech recognition object is text data obtained by processing a speech-to-text result of the corresponding type "final", the note-taking application can append the corresponding text data obtained by processing the speech-to-text result to the end of the original text data in the first speech recognition object. In this case, the text data in the first speech recognition object = the original text data in the first speech recognition object + the corresponding text data obtained by processing the speech-to-text result.

[0313] For example, as shown in Figures 6E and 6F, the original text data in the first speech recognition object is "123456", which is the text data obtained based on the speech-to-text result of the corresponding type being final. The text data obtained by processing the newly received speech-to-text result may be "123456789". The note-taking application can write "123456789" into the first speech recognition object, overwriting the original text data in the first speech recognition object. The text data in the first speech recognition object then changes from "123456" to "123456789".

[0314] In some embodiments of the present application, if the flag bit is true and the type of the speech-to-text result is final, the note-taking application can process the speech-to-text result, obtain corresponding text data, and write the corresponding text data into the first speech recognition object, overwriting the original text data in the first speech recognition object.

[0315] In some embodiments of the present application, if the flag bit is true, the note application may execute step S204.

[0316] It should be noted that, after the note application exits the editing state (for example, the flag is reset from true to false), if the queue is not empty, the note application may execute step S213.

[0317] S216: The note application transmits the text data in the first voice recognition object to the display screen.

[0318] The note application can transmit the text data in the first voice recognition object to the display screen. Correspondingly, the display screen can receive the text data in the first voice recognition object transmitted by the note application, that is, the text data.

[0319] S217: The display screen displays the text data in the first voice recognition object.

[0320] After the display screen receives the text data (i.e., text data) in the first voice recognition object transmitted by the note application, the text data can be displayed in the display area corresponding to the first voice recognition object (for example, the recognition text display area 1033 shown in Figures 6F-6H).

[0321] For example, as shown in FIG6E , when the user has not triggered editing, after the note-taking application receives the speech-to-text result of the corresponding type of part, it can process the speech-to-text result as shown in step S204, and send the text data in the first speech recognition object obtained after processing to the display screen. Accordingly, the display screen can display the text data included in the speech-to-text in the display area corresponding to the first speech recognition object (for example, the recognition text display area 1033), that is, "123456". As shown in FIG6G , when the user has triggered editing, the note-taking application receives the speech-to-text result of the corresponding type of final, and after processing it, the corresponding text data, that is, "123456789", can be obtained. Since the original text data in the first voice recognition object is "123456", which is text data obtained based on the speech-to-text result of the corresponding type part, the note-taking application can directly write "123456789" into the first voice recognition object, overwrite the "123456" in the first voice recognition object, and send "123456789" to the display screen. Correspondingly, as shown in Figure 6G, the text data displayed on the display screen in the display area corresponding to the first voice recognition object can be changed to "123456789".

[0322] It should be noted that when the flag bit is true, the electronic device can also process and display the speech-to-text results in other ways. Similarly, when the flag bit is not true and the queue is not empty, the electronic device can also process the speech-to-text results in the queue in other ways. For the specific processing and display methods for the speech-to-text results when the flag bit is true and when the flag bit is not true and the queue is not empty, please refer to Table 1:

[0323] Table 1

[0324] As shown in Table 1, in some embodiments of the present application, when the flag bit is true (i.e., during the editing process), the note-taking application can place the received speech-to-text results of type part (or speech-to-text results of type part) and speech-to-text results of type final (or speech-to-text results of type final) into a queue. Furthermore, when the flag bit is true (i.e., during the editing process), the note-taking application can adopt any of the following three processing and display methods for the received speech-to-text results:

[0325] Processing and display method 1 during the editing process: When the flag bit is true (i.e., during the editing process), the note-taking application may not process the received final type speech-to-text data, but instead process the received part type speech-to-text results in the order in which they were received, write the processed text data into the first speech recognition object, and overwrite the original text data in the first speech recognition object, and then display the text data in the first speech recognition object on the display screen. It is understood that the specific implementation method of method 1 can be referred to the relevant description of steps S204-S206.

[0326] For example, as shown in FIG9A , during the editing process, the note application successively receives a speech-to-text result of the part type including the text “12”, a speech-to-text result of the part type including the text “1234”, a speech-to-text result of the part type including the text “123456”, a speech-to-text result of the final type including the text “1234567”, a speech-to-text result of the part type including the text “8”, a speech-to-text result of the part type including the text “89”, a speech-to-text result of the final type including the text “890”, and a speech-to-text result of the part type including the text “abc”, and successively puts the above-received speech-to-text results into a queue. It can be understood that during the editing process, the note application only processes the speech-to-text results of the part type and displays the processed text data on the display screen. Specifically, the note application first receives the speech-to-text result of the part type including the text “12”, puts it into the queue, and processes it to obtain the text “12”. The text “12” can be displayed in the text display area 1033 during recognition. The note-taking application then receives a speech-to-text result of the part type, including the text "1234," places it in a queue, and processes it to obtain the text "1234." The text "1234" can replace the text "12" and be displayed in the text display area 1033 during recognition. The note-taking application then receives a speech-to-text result of the part type, including the text "123456," places it in a queue, and processes it to obtain the text "123456." The text "123456" can replace the text "1234" and be displayed in the text display area 1033 during recognition. The note-taking application then receives a speech-to-text result of the final type, including the text "1234567," places it in a queue, but does not process it. The note-taking application then receives a speech-to-text result of the part type, including the text "8," places it in a queue, and processes it to obtain the text "8." The text "8" can replace the text "123456" and be displayed in the text display area 1033 during recognition. The note application then receives a speech-to-text result of the part type including the text "89", puts it into a queue, and processes it to obtain the text "89". The text "89" can replace the text "8" and be displayed in the text display area 1033 during recognition. After the note application receives a speech-to-text result of the final type including the text "890", it puts it into a queue without processing it. The note application then receives a speech-to-text result of the part type including the text "abc", puts it into a queue, and processes it to obtain the text "abc". The text "abc" can replace the text "89" and be displayed in the text display area 1033 during recognition.

[0327] Processing and display method 2 during the editing process: When the flag bit is true (i.e., during the editing process), the note application may not process the received part-type speech-to-text data, but may process the received final-type speech-to-text results in the order of reception to obtain text data. If the original text data in the first speech recognition object is obtained based on the final-type speech-to-text result, the note application may add the text data obtained by processing the received final-type speech-to-text result to the back of the original text data in the first speech recognition object. If the original text data in the first speech recognition object is obtained based on the part-type speech-to-text result, the note application may use the text data obtained by processing the received final-type speech-to-text result to overwrite the original text data in the first speech recognition object. After the note application completes the above processing, it may display the text data in the first speech recognition object on the display screen. It is understandable that the specific implementation method of method 2 can refer to the relevant description of steps S215-step S217.

[0328] For example, the order in which the note-taking application receives speech-to-text results is the same as that shown in FIG9B and FIG9A , and the order in which the speech-to-text results are queued is also the same as that shown in FIG9B and FIG9A . However, as shown in FIG9B , during the editing process, the note-taking application only processes received speech-to-text results of the final type, and does not process received speech-to-text results of the part type. Specifically, the note-taking application may receive a speech-to-text result of the part type including the text "12," a speech-to-text result of the part type including the text "1234," and a speech-to-text result of the part type including the text "123456," and then directly queue them without processing them. When the note-taking application subsequently receives a speech-to-text result of the final type including the text "1234567," it may queue it and process it to obtain the text "1234567." The text "1234567" may be displayed in the recognition text display area 1033. The note application then receives a speech-to-text result of the part type including the text "8" and a speech-to-text result of the part type including the text "89", puts them into a queue, and does not process them. After the note application subsequently receives a speech-to-text result of the final type including the text "890", it puts it into a queue, processes it, and obtains the text "890". The text "890" can be displayed in the recognition text display area 1033 immediately after the text "1234567", that is, "1234567890" is displayed in the recognition text display area 1033. The note application can then receive a speech-to-text result of the part type including the text "abc", put it into a queue, and do not process it.

[0329] Processing and display method three during the editing process: When the flag bit is true (i.e., during the editing process), the note-taking application can process the received part-type speech-to-text results and final-type speech-to-text results in the order in which they were received, write the processed text data into the first speech recognition object, and overwrite the original text data in the first speech recognition object, and then display the text data in the first speech recognition object on the display screen. It is understood that the specific implementation method of overwriting the original text data can be referred to the relevant description of steps S204-S206.

[0330] For example, the order in which the note application receives speech-to-text results is the same as that shown in Figures 9A-9B, and the order in which the speech-to-text results are queued is also the same as that shown in Figures 9A-9B. However, as shown in Figure 9C, during the editing process, the note application processes both the received final-type speech-to-text results and the received part-type speech-to-text results, and sends the processed text to the display screen for display. Specifically, as shown in Figure 9C, the note application can first receive a part-type speech-to-text result including the text "12" and process it to obtain the text "12". The text "12" can be displayed in the text display area 1033 during recognition. Similarly, the note application can then receive a part-type speech-to-text result including the text "1234" and process it to obtain the text "1234". The text "1234" can replace the text "12" and be displayed in the text display area 1033 during recognition. Similarly, the note-taking application can receive a part-type speech-to-text result including the text "123456" and process it to obtain the text "123456." The text "123456" can replace the text "1234" and be displayed in the recognition text display area 1033. Similarly, the note-taking application can receive a final-type speech-to-text result including the text "1234567" and process it to obtain the text "1234567." The text "1234567" can replace the text "123456" and be displayed in the recognition text display area 1033. Similarly, the note-taking application can receive a part-type speech-to-text result including the text "8" and process it to obtain the text "8." The text "8" can replace the text "1234567" and be displayed in the recognition text display area 1033. Similarly, the note-taking application can receive a part-type speech-to-text result including the text "89" and process it to obtain the text "89." The text "89" can replace the text "8" and be displayed in the text display area 1033 during recognition. Similarly, the note-taking application can receive a speech-to-text result of the final type including the text "890" and process it to obtain the text "890". The text "890" can replace the text "89" and be displayed in the text display area 1033 during recognition. Similarly, the note-taking application can receive a speech-to-text result of the part type including the text "abc" and process it to obtain the text "abc". The text "abc" can replace the text "890" and be displayed in the text display area 1033 during recognition.

[0331] Furthermore, when the note application adopts any of the above processing and display methods (processing and display method one during editing, processing and display method two during editing, or processing and display method three during editing) for the speech-to-text results received during the editing process, after exiting the editing process (specifically, the flag bit is not true and the queue is not empty), the note application can adopt any of the following two processing and display methods for the speech-to-text results in the queue:

[0332] Processing and display method 1 after exiting editing: When the mark bit is not true and the queue is not empty, the note-taking application can discard the part type speech-to-text results in the queue, and process the final type speech-to-text results in the queue in the order of entry. For details of this processing method and subsequent display method, please refer to the relevant descriptions of steps S207-S212.

[0333] Exemplarily, as shown in FIG9D , during the editing process, the note application processes both the received final type speech-to-text results and the part type speech-to-text results, and sends the processed text to the display screen for display. The specific processing and display process can refer to the relevant description of FIG9C , and this application will not repeat them here. As shown in FIG9D , after exiting the editing, the speech-to-text results in the queue are as follows: a speech-to-text result of the part type including the text "12", a speech-to-text result of the part type including the text "1234", a speech-to-text result of the part type including the text "123456", a speech-to-text result of the final type including the text "1234567", a speech-to-text result of the part type including the text "8", a speech-to-text result of the part type including the text "89", a speech-to-text result of the final type including the text "890", and a speech-to-text result of the part type including the text "abc". After exiting the editing, the note application can discard the speech-to-text results of the part type in the queue and process the speech-to-text results of the final type in the queue in the order in which they were queued. Specifically, the note-taking application can determine that the first enqueued speech-to-text result including the text "12" is of type part and discard the speech-to-text result. Similarly, the note-taking application can discard the speech-to-text result including the text "1234" of type part and the speech-to-text result including the text "123456" of type part. The note-taking application can determine that the next enqueued speech-to-text result including the text "1234567" is of type final and process it to obtain the text "1234567". In this case, the note-taking application can determine whether the speaker corresponding to the speech-to-text result is the speaker corresponding to the second speech recognition object. If the speaker corresponding to the speech-to-text result is the speaker corresponding to the second speech recognition object, the note-taking application can add the text "1234567" to the end of the original text in the second speech recognition object. If the speaker corresponding to the speech-to-text result is not the speaker corresponding to the second speech recognition object, the note-taking application can create another speech recognition object (such as a third speech recognition object) and add the text "1234567" to the created speech recognition object. Similarly, the note-taking application can then discard the speech-to-text results of the part type including the text "8" and the speech-to-text results of the part type including the text "89", process the speech-to-text results of the final type including the text "890", and discard the speech-to-text results of the part type including the text "abc".For the processing and display process of the final type speech-to-text result including the text "890", you can refer to the processing and display process of the final type speech-to-text result including the text "1234567", and this application will not repeat it here.

[0334] For example, as shown in FIG9E , during the editing process, the note application can process the received speech-to-text results of type part and send the processed text to the display screen for display, but does not process the speech-to-text results of type final. The specific processing and display process can refer to the relevant description of FIG9A , which is not repeated here in this application. The processing and display process of the speech-to-text results in the queue after exiting the editing process as shown in FIG9E can refer to the relevant description of FIG9D , which is not repeated here in this application.

[0335] Second processing and display method after exiting editing: When the mark bit is not true and the queue is not empty, the note application can process the part type speech-to-text results and final type speech-to-text results in the queue according to the order of entry. For details of this processing method and subsequent display method, please refer to the relevant description of steps S204-step S212.

[0336] Exemplarily, as shown in Figures 9A-9C, after exiting the editing, the note application can process the part-type speech-to-text results in the queue and the final-type speech-to-text results in the queue in the order of entry, that is, process the following speech-to-text results in sequence: a speech-to-text result of the part type including the text "12", a speech-to-text result of the part type including the text "1234", a speech-to-text result of the part type including the text "123456", a speech-to-text result of the final type including the text "1234567", a speech-to-text result of the part type including the text "8", a speech-to-text result of the part type including the text "89", a speech-to-text result of the final type including the text "890", and a speech-to-text result of the part type including the text "abc". After the note application obtains the text included in the corresponding speech-to-text each time, the text obtained by the processing can be sent to the display screen for display. The specific processing and display process can refer to steps S204-S212, and this application will not be repeated here.

[0337] In some other embodiments of the present application, when the mark bit is true (i.e., during the editing process), the note application can put the received speech-to-text result of the final type into the queue, but not put the speech-to-text result of the part type into the queue. Further, when the mark bit is true (i.e., during the editing process), the note application can adopt the processing and display method one in the above-mentioned editing process, or the processing and display method two in the above-mentioned editing process, or the processing and display method three in the above-mentioned editing process for the received speech-to-text result. It should be noted that after the note application uses any of the above-mentioned processing and display methods in the editing process to process the speech-to-text result of the part type, the speech-to-text result of the part type can be discarded and not put into the queue. Further, after exiting the editing (specifically, the mark bit is not true and the queue is not empty), the queue only includes the speech-to-text result of the final type, and does not include the speech-to-text result of the part type. In this case, the note application can process the speech-to-text results in the queue in the order of enqueuing, and send the processed text to the display screen for display. It can be understood that the processing and display process can be specifically referred to the relevant description of steps S207 to S212.

[0338] For example, when the mark bit is true, the note application can use the processing and display method 1 in the above-mentioned editing process to process the received part type speech-to-text result, and send the processed text to the display screen for display, but discard the part type speech-to-text result after the processing and do not put it into the queue. Specifically, when the mark bit is true, for the received part type speech-to-text result, the note application can process it and discard it, and write the processed text data into the first speech recognition object, and overwrite the original text data in the first speech recognition object, and then display the text data in the first speech recognition object on the display screen. For the received final type speech-to-text result, the note application does not process it, but puts it into the queue. Further, subsequently, when the mark bit is not true and the queue is not empty, the queue only includes final type speech-to-text results, and the note application can process the final type speech-to-text results in the queue in the order of entry.

[0339] For example, FIG9F illustrates the same order in which the note application receives speech-to-text results as shown in FIG9A-9E , and FIG9F illustrates the same process for processing and displaying the received speech-to-text results as shown in FIG9A . However, as shown in FIG9F , after processing a speech-to-text result of the part type, the note application can directly discard it without placing it in the queue. Specifically, after processing a speech-to-text result of the part type including the text "12," a speech-to-text result of the part type including the text "1234," and a speech-to-text result of the part type including the text "123456," the note application can discard them without placing them in the queue. The note application can then place a speech-to-text result of the final type including the text "1234567" in the queue without processing it. The note application can then process a speech-to-text result of the part type including the text "8" and a speech-to-text result of the part type including the text "89," and discard them after processing without placing them in the queue. The note-taking application can then queue the final-type speech-to-text result containing the text "890" without processing it. The note-taking application can then process the part-type speech-to-text result containing the text "abc" and discard it after processing without queueing it. It is understood that the processing and display of part-type speech-to-text results during the editing process can be seen in the description of FIG. 9A .

[0340] Similarly, when the flag bit is true, the note-taking application can use the second processing and display method in the above editing process to process the received final type speech-to-text results, and send the processed text to the display screen for display, and discard the part type speech-to-text results, without processing them or putting them into the queue.

[0341] For example, the order in which the note application receives speech-to-text results is the same as that shown in Figures 9A-9F, and the processing and display process of the received speech-to-text results in Figure 9G is the same as that shown in Figure 9B. The difference is that, as shown in Figure 9G, after processing the part-type speech-to-text results, the note application can directly discard them without placing them in the queue. Specifically, the note application directly discards the received part-type speech-to-text results including the text "12", the part-type speech-to-text results including the text "1234", and the part-type speech-to-text results including the text "123456", without processing them or placing them in the queue. The note application can then place the received final-type speech-to-text results including the text "1234567" in the queue and process them to obtain the text "1234567". The text "1234567" can be displayed in the recognition text display area 1033. The note application can then directly discard the received speech-to-text results of the part type including the text "8" and the speech-to-text results of the part type including the text "89", without processing them or putting them into the queue. The note application can then put the received speech-to-text results of the final type including the text "890" into the queue, process them, and obtain the text "890". The text "890" can be displayed in the recognition text display area 1033 immediately after the text "1234567", that is, the text "1234567890" is displayed in the recognition text display area 1033. The note application can then directly discard the received speech-to-text results of the part type including the text "abc", without processing them or putting them into the queue.

[0342] Similarly, when the flag bit is true, the note-taking application can use the third processing and display method in the above editing process to process the received final type speech-to-text results and part type speech-to-text results, and send the processed text to the display for display, but discard the part type speech-to-text results and do not put them into the queue.

[0343] For example, FIG9H illustrates the same order in which the note application receives speech-to-text results as shown in FIG9A-9F , and FIG9H illustrates the same process for processing and displaying the received speech-to-text results as shown in FIG9C . However, as shown in FIG9G , after processing a speech-to-text result of the part type, the note application may directly discard it without placing it in a queue. Specifically, the note application may process a speech-to-text result of the part type that includes the text "12," a speech-to-text result of the part type that includes the text "1234," and a speech-to-text result of the part type that includes the text "123456," send the processed text to the display screen for display, and then discard it after processing without placing it in a queue. The note application may then process a speech-to-text result of the final type that includes the text "1234567," place it in the queue, and send the processed text to the display screen for display. The note application may then process a speech-to-text result of the part type that includes the text "8," and a speech-to-text result of the part type that includes the text "89," and discard it after processing without placing it in a queue. The note-taking application can then process the final-type speech-to-text result containing the text "890," queue it, and send the processed text to the display screen for display. The note-taking application can then process the part-type speech-to-text result containing the text "abc," discard it after processing, and not queue it. It is understood that the processing and display process of the part-type speech-to-text result during the editing process can be referred to the relevant description of Figure 9C.

[0344] Exemplarily, as shown in Figures 9F to 9H, during the editing process, the note application only puts the speech-to-text result of the final type including the text "1234567" and the speech-to-text result of the final type including the text "890" into the queue. After exiting the editing, the queue only contains the speech-to-text result of the final type including the text "1234567" and the speech-to-text result of the final type including the text "890". The note application can first process the speech-to-text result of the final type including the text "1234567" in the order of enqueuing, and then process the speech-to-text result of the final type including the text "890". It can be understood that the processing and display process of the speech-to-text result of the final type can refer to the relevant description of Figure 9D and steps S207 to S212.

[0345] It should be noted that in some embodiments of the present application, if the display speaker function is enabled, the electronic device may perform the steps shown in Figure 7B. In some embodiments of the present application, if the display speaker function is not enabled, the electronic device may not perform steps S208-S209 and steps S211-S212.

[0346] It is worth noting that the first voice recognition object, the second voice recognition object, and the third voice recognition object involved in this application can be understood as different voice recognition objects. In some embodiments of the present application, the text stored in the first voice recognition object can be displayed in the text display area 1033 in recognition. It can also be understood that the display area corresponding to the first voice recognition object is the text display area 1033 in recognition. In some embodiments of the present application, the text stored in the second voice recognition object can be displayed in the second speaker text display area 1031. It can also be understood that the display area corresponding to the second voice recognition object is the second speaker text display area 1031. In some embodiments of the present application, the text stored in the third voice recognition object can be displayed in the third speaker text display area 1035. It can also be understood that the display area corresponding to the third voice recognition object is the third speaker text display area 1035.

[0347] It is understandable that the text included in the first speech recognition object, the second speech recognition object, and the third speech recognition object can be the same or different. In some embodiments of the present application, the first speech recognition object, the second speech recognition object, and the third speech recognition object can all be string objects. In the process of implementing speech-to-text conversion in the note application, the note application can also create more speech recognition objects, and this application does not limit this.

[0348] It is understandable that the note application can create an object for storing all recognized texts. For the sake of ease of description, this application will record the object for storing all recognized texts as a recognized text object. When the display speaker function is not turned on, if the type of the speech-to-text result is final, the note application can add the text data obtained by processing the speech-to-text result directly to the recognized text object. Specifically, it can be spliced ​​behind the original text data in the recognized text object, and then the text data in the recognized text object can be displayed on the display.

[0349] In some embodiments of the present application, when a user triggers the voice-to-text function or the electronic device automatically turns on the voice-to-text function, the note-taking application can create a recognized text object.

[0350] Optionally, after executing step S217, the electronic device may further execute at least one of steps S301 to S308.

[0351] 3. The user edits the recognized text and puts the newly received speech-to-text result into the queue (as shown in Figure 7C)

[0352] S301: In response to a user operation triggering editing of the recognized text, the note application sets a flag bit to true.

[0353] The user can trigger editing of the recognized text by clicking on the display area corresponding to the recognized text, thereby triggering the note-taking application to enter the editing state. For example, as shown in FIG6E , the user can trigger editing of the text in the second speaker text display area 1031 by clicking on the second speaker text display area 1031. Accordingly, the note-taking application can receive the user action of clicking on the second speaker text display area 1031 and enter the editing state.

[0354] Of course, users can also trigger editing of recognized text through gestures or voice control, and this application does not limit this.

[0355] After the note-taking application detects a user operation that triggers editing of the recognized text, the note-taking application may set the flag bit from false to true in response to the user operation.

[0356] For example, in response to the user operation of editing the recognized text, the note-taking application may set isEditByUser=true.

[0357] S302: The speech recognition server sends the speech-to-text result to the note-taking application via the modem.

[0358] According to the above, the speech recognition server can continuously send the speech-to-text results to the note-taking application via the modem. Correspondingly, the note-taking application can continuously receive the speech-to-text results sent by the speech recognition server via the modem.

[0359] S303: The note application determines that the flag is true and puts the speech-to-text result into a queue.

[0360] According to step S201, after receiving the speech-to-text result sent by the speech recognition server, the note-taking application can determine whether the flag bit is true. It is understandable that because the note-taking application sets the flag bit to true after the user triggers editing of the recognized text, the note-taking application can determine that the flag bit is true after receiving the speech-to-text result and place the speech-to-text result in a queue. The specific implementation method can be referred to in step S214, and this application will not be repeated here.

[0361] S304: When the type of the speech-to-text result is final, the note application processes the speech-to-text result to obtain text data. If the original text data in the first speech recognition object is obtained based on the speech-to-text result of type part, the text data obtained by processing the speech-to-text result is used to overwrite the original text data in the first speech recognition object. If the original text data in the first speech recognition object is obtained based on the speech-to-text result of type final, the text data obtained by processing the speech-to-text result is added to the first speech recognition object.

[0362] Since the note application sets the flag bit to true after the user triggers editing of the recognized text, the note application can determine that the flag bit is true after receiving the speech-to-text result. In this case, if the type of the speech-to-text result is part, the note application can put the speech-to-text result into the queue. If the type of the speech-to-text result is final, the note application can not only put the speech-to-text result into the queue, but also process the speech-to-text result to obtain the corresponding text data. If the original text data in the first speech recognition object is the text data obtained based on the speech-to-text result of the corresponding type of part, the note application can overwrite the original text data in the first speech recognition object with the corresponding text data. If the original text data in the first speech recognition object is the text data obtained based on the speech-to-text result of the corresponding type of final, the note application can add the corresponding text data to the end of the original text data in the first speech recognition object.

[0363] It is understandable that the specific implementation of step S304 can refer to the relevant description of step S215, and this application will not go into details here.

[0364] S305: The note application transmits the text data in the first voice recognition object to the display screen.

[0365] S306: The display screen displays the text data in the first voice recognition object.

[0366] It can be understood that the specific implementation of steps S305-S306 can refer to the relevant description of steps S216-S217, and this application will not repeat them here.

[0367] S307: The note application determines the target editing text, editing position, editing method and editing content, and edits the target editing text based on the editing position, editing method and editing content.

[0368] After the note-taking application detects a user operation that triggers editing of the recognized text, it can obtain the corresponding input event and input location in response to the user operation to edit the recognized text, and the display area where the input location is located can obtain focus. Furthermore, the note-taking application can determine the editing location based on the input location and determine that the text included in the display area where the input location is located is the target editing text. The editing location can also be understood as the cursor position.

[0369] It is understandable that when the note application detects a user operation that triggers the editing of the recognized text, the note application can enter the editing state. Later, in the process of the user editing the recognized text, the note application can continuously detect the corresponding user operation (for example, the user operation on the key in the keyboard). In response to the corresponding user operation, the note application can determine the editing method (for example, deleting text, adding text, etc.) and the editing content. It is understandable that the editing content refers to the text targeted by the corresponding editing method, such as the text that needs to be deleted or added. The editing content can be part or all of the text in the target editing text, or it can be other text to be added to the target editing text.

[0370] After the note application determines the target editing text, editing position, editing method and editing content according to the above method, it can edit the target editing text based on the editing position, editing method and editing content, obtain the edited text, and send the edited text to the display screen for display.

[0371] For example, in response to the user operation of clicking the second speaker text display area 1031 as shown in (1) in FIG6E , the note application can determine the cursor position and the target edit text. As shown in FIG6F , the cursor position is between “6” and “9” in the text included in the second speaker text display area 1031, and the target edit text is the text included in the second speaker text display area 1031, i.e., “abcde3336666999”. Further, the note application can receive the user operation of clicking the delete button 161 as shown in FIG6F . After receiving the user operation of clicking the delete button 161, the note application can determine that the editing mode is deletion and the editing content is the number “6” to the left of the cursor. Accordingly, the note application can delete the number “6” in the text included in the second speaker text display area 1031, obtain the edited text “abcde333666999”, and send the edited text to the display screen for display. As shown in FIG6G , the display screen can display the edited text, i.e., “abcde333666999”, in the second speaker text display area 1031.

[0372] In some embodiments of the present application, after determining the input location, the note-taking application may determine the object (e.g., the second speech recognition object) corresponding to the display area (e.g., the second speaker text display area 1031) where the input location is located, and determine that the text data in the object is the target edit text. In this case, the note-taking application edits the target edit text based on the edit location, edit method, and edit content, and after obtaining the edited text, the edited text may be written into the object and the text data in the object may be sent to the display screen for display.

[0373] It is understandable that this application does not limit the order in which step S302 and step S307 are executed.

[0374] S308: If the note application is in the editing state and no user operation is detected within T1, or a user operation acting on a non-editing area is detected, the note application sets the flag bit to false.

[0375] It is understood that the specific value of T1 can be set according to actual needs, and this application does not limit this. For example, T1 can be 10 seconds.

[0376] It is understood that the non-editing area may include areas other than the display area where the recognized text is located (e.g., the recognized text display area 130) and the display area where the keyboard is located (e.g., the display area where the keyboard 16 is located as shown in FIG6G ). For example, the non-editing area may include a blank area as shown in FIG6H .

[0377] It is understandable that the above-mentioned conditions for triggering the note application to set the mark bit to false are merely examples provided by this application and should not be regarded as limitations of this application.

[0378] Optionally, after executing step S308 , the electronic device may further execute at least one of steps S401 to S405 .

[0379] 4. After editing is completed, the speech-to-text results in the queue and the newly received speech-to-text results are processed (as shown in Figure 7D)

[0380] S401: In response to a user operation triggering the end of recording, the note-taking application requests the audio framework to stop collecting audio data.

[0381] The user can trigger the end of recording in the note application. Exemplarily, as shown in FIG6E , the user can trigger the end of recording by clicking the stop recording control 101 , and accordingly, the note application can receive the user operation of clicking the stop recording control 101 .

[0382] Of course, the user can also trigger the note-taking application to end recording through other methods (for example, gestures, voice control, etc.), and this application does not limit this.

[0383] After the note-taking application detects a user operation that triggers the end of recording, in response to the user operation that triggers the end of recording, the note-taking application may request the audio framework to stop collecting audio data.

[0384] In some embodiments of the present application, in response to the user operation that triggers the stop of recording, the note-taking application may use the stopRecording() method to stop recording.

[0385] S402: The audio framework notifies the microphone to stop collecting audio data through the audio HAL and the audio driver.

[0386] After the note-taking app requests the audio framework to stop collecting audio data, the audio framework can notify the microphone to stop collecting audio data through the audio HAL and audio driver.

[0387] Accordingly, the microphone can receive a notification from the audio framework through the audio HAL and the audio driver to stop collecting audio data.

[0388] S403: In response to the user operation that triggers the end of recording, the note-taking application determines whether the queue is empty.

[0389] After the note-taking application detects the user operation that triggers the end of recording, the note-taking application can determine whether the queue is empty in response to the user operation that triggers the end of recording. If the queue is not empty, the note-taking application can execute steps S404 and S405. If the queue is empty, the note-taking application can only execute step S405.

[0390] S404: The note application processes the speech-to-text results in the queue in the order in which they were queued.

[0391] If the queue is not empty after the note application stops recording, the note application can process the speech-to-text results in the queue in the order in which the speech-to-text results enter the queue. The specific implementation method can refer to the relevant description of step S213, and this application will not repeat it here.

[0392] S405: If there is an unprocessed speech-to-text result outside the queue, the note-taking application processes the speech-to-text result.

[0393] It is understandable that if the queue is empty and there are unprocessed speech-to-text results outside the queue, the note application can process the speech-to-text results outside the queue. The specific implementation method can refer to steps S204-S212, and this application will not repeat them here.

[0394] If the queue is not empty and there are unprocessed speech-to-text results outside the queue, the note-taking application can first process the speech-to-text results in the queue, and then process the speech-to-text results outside the queue after the processing is completed. The specific implementation method can refer to the relevant description of step S213, and this application will not repeat it here.

[0395] Based on the above display method provided in the embodiment of the present application, the electronic device can edit the recognized text during the ASR process, and neither the user editing content nor the recognized content will be lost during the editing process.

[0396] For example, as shown in Figure 10, the text originally stored in the speech recognition object (i.e., the recognized text) is 123456. The user can click the corresponding area on the display to edit the recognized text. In response to this click, at time t5, the electronic device can begin editing (or enter the editing state). After the electronic device begins editing, at time t6, the electronic device can read the recognized text, i.e., 123456, from the speech recognition object. The user can edit the recognized text, such as adding the number 0 to the front of the recognized text. Accordingly, the electronic device can obtain the edited text, i.e., 012345, and write the edited text to the speech recognition object at time t7. In other words, starting at time t7, the recognized text in the speech recognition object becomes 0123456. After completing the above editing, the user can exit the editing process at time t8. It should be noted that after the electronic device begins editing, at time t9, the electronic device can queue the acquired ASR data, such as 7 and 8. The electronic device does not process the data in the queue until it exits the editing process at time t8. After the electronic device exits editing, at time t10, the electronic device can read the recognized text (i.e., 0123456) from the voice recognition object, and obtain ASR data (including 7 and 8) from the queue, and add the obtained ASR data to the read recognized text to obtain the updated text, i.e., 012345678. Further, at time t11, the electronic device can write the updated text to the voice recognition object, and the text in the voice recognition object changes from 0123456 to 012345678. In this way, the electronic device can implement ASR and user editing independently of each other, and the ASR data obtained by the electronic device will not be lost due to user editing, nor will the user edited content be lost due to processing of the ASR data.

[0397] It should be noted that, according to the display method provided in the embodiment of the present application, the thread responsible for ASR in the electronic device will perform orderly and frequent operations on the object (the object storing the recognized text), while the thread responsible for user editing will only operate on the object (the object storing the recognized text) when the user triggers editing, that is, the object is occasionally operated and the operation is disordered. The electronic device can put the pending events (i.e., ASR data) of the thread responsible for ASR for the object storing the recognized text into a queue during the process of the thread responsible for user editing operating on the object storing the recognized text. After the thread responsible for user editing completes the operation on the object storing the recognized text, if there is ASR data in the queue, the electronic device first processes the ASR data in the queue through the thread responsible for ASR, and after the processing is completed, operates on the object storing the recognized text through the thread responsible for ASR. However, if there is no ASR data in the queue, the electronic device operates on the object storing the recognized text through the thread responsible for ASR, that is, writes the text included in the newly received ASR data into the object storing the recognized text.

[0398] It can be understood that according to the above method, as shown in Figure 11, when two threads of an electronic device jointly operate the same object (wherein thread A operates in an orderly and frequent manner, and thread B operates in an unordered and occasional manner), the electronic device can put the pending events of thread A for the object into a queue while thread B is operating on the object. After thread B completes the operation on the object, if there is a pending event for the object by thread A in the queue, the electronic device will first process the pending event in the queue through thread A, and after the processing is completed, the electronic device will operate on the object through thread A. However, if there is no pending event for the object by thread A in the queue, the electronic device will operate on the object through thread A.

[0399] It is understandable that thread A can also be a thread other than the thread responsible for ASR, and thread B can also be a thread other than the thread responsible for user editing. This application does not impose specific restrictions on thread A and thread B.

[0400] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.

Claims

1. A display method, characterized in that: Applied to a first device, the method includes: Displaying a first interface, wherein the first interface includes a first control; In response to a first operation on the first control, displaying a second interface, and starting recording after the first operation; During a first time period, first audio data obtained through recording is acquired, and first text content corresponding to the first audio data is displayed in a first display area of ​​the second interface. During the first time period, the first display area does not include a cursor. After the first time period, in response to a second operation on the first display area, a cursor is displayed in the first display area; After the second operation, an editing operation is received, and the first text content modified by the editing operation is displayed in the first display area; After the second operation, second audio data obtained through recording is acquired, and second text content corresponding to the second audio data is not displayed in the first display area.

2. The method according to claim 1, wherein The second interface includes a second display area; after the second operation, the second text content is displayed in the second display area.

3. The method according to claim 1 or 2, wherein: The method further comprises: In response to an event of exiting editing, the second text content and the modified first text content are displayed in the first display area. After the event of exiting editing, the cursor is not displayed in the first display area.

4. The method according to any one of claims 1 to 3, wherein The second interface includes a second control; and starting recording after the first operation includes: In response to the first operation, starting recording; The step of obtaining first audio data obtained through recording in the first time period includes: In response to a third operation on the second control, the first audio data and the first text content are acquired during the first time period.

5. The method according to claim 2, wherein After obtaining the first audio data obtained through recording, the method further includes: sending the first audio data to a second device; After sending the first audio data to the second device, receiving a plurality of audio recognition results sent by the second device; the text in the first type of audio recognition results included in the plurality of audio recognition results constitutes the first text content; the plurality of audio recognition results include the first recognition result; After receiving the first recognition result, the first recognition result is processed to obtain third text content; when the first recognition result is the first type of audio recognition result, the third text content is displayed in the first display area, and the first text content includes the third text content; when the first recognition result is the second In the case of a similar audio recognition result, the third text content replaces the content originally displayed in the second display area; In the process of processing the first recognition result, the first mark bit is the first content, and the first queue does not include the audio recognition result.

6. The method according to claim 5, wherein The method further comprises: In response to the first operation, the first queue is created, and the first mark bit is set to the first content.

7. The method according to claim 5 or 6, wherein: After obtaining the second audio data obtained through recording, the method further includes: sending the second audio data to the second device; After sending the second audio data to the second device, receiving multiple audio recognition results sent by the second device, and processing the multiple audio recognition results to obtain the second text content; In which, each of the multiple audio recognition results includes part or all of the second text content; in the process of receiving the multiple audio recognition results sent by the second device and processing the multiple audio recognition results, the first marker is the second content; after sending the second audio data to the second device and receiving the multiple audio recognition results sent by the second device, the first queue includes the multiple audio recognition results received after sending the second audio data to the second device, or the first queue includes the first category of audio recognition results among the multiple audio recognition results received after sending the second audio data to the second device.

8. The method according to claim 7, wherein The multiple audio recognition results received after sending the second audio data to the second device include a second recognition result and a third recognition result, the second recognition result and the third recognition result are first-category audio recognition results, the second recognition result is received earlier than the third recognition result; the second recognition result includes fourth text content, and the third recognition result includes fifth text content; The receiving the multiple audio recognition results sent by the second device and processing the multiple audio recognition results to obtain the second text content includes: receiving the second recognition result, and after receiving the second recognition result, processing the second recognition result to obtain the fourth text content; the fourth text content is displayed in the second display area; The third recognition result is received, and after receiving the third recognition result, the third recognition result is processed to obtain the fifth text content; the fifth text content and the fourth text content form the second text content and are displayed in the second display area.

9. The method according to claim 7, wherein The multiple audio recognition results received after sending the second audio data to the second device include a fourth recognition result and a fifth recognition result, the fourth recognition result is received earlier than the fifth recognition result; the fourth recognition result includes the sixth text content, and the fifth recognition result includes the second text content; The receiving the multiple audio recognition results sent by the second device and processing the multiple audio recognition results to obtain the second text content includes: receiving the fourth recognition result, and after receiving the fourth recognition result, processing the fourth recognition result to obtain sixth text content; the sixth text content is displayed in the second display area; The fifth recognition result is received. After receiving the fifth recognition result, the fifth recognition result is processed to obtain the second text content; the second text content replaces the sixth text content and is displayed in the second display area.

10. The method according to claim 9, wherein The fourth recognition result and the fifth recognition result are the first category audio recognition results; or, the fourth recognition result and the fifth recognition result are the second category audio recognition results; or, the fourth recognition result is the first category audio recognition result, and the fifth recognition result is the second category audio recognition result; or, the fourth recognition result is the second category audio recognition result, and the fifth recognition result is the first category audio recognition result.

11. The method according to any one of claims 7 to 10, wherein: After the first time period, the method further includes: In response to the second operation, setting the first flag bit to the second content; After obtaining the second text content, the method further includes: In response to an event of exiting editing, the first mark bit is set to the first content.

12. An electronic device, characterized in that: The electronic device includes one or more memories and one or more processors; the one or more memories are coupled to the one or more processors, the memories are used to store computer program code, the computer program code includes computer instructions, and the processor calls the computer instructions to execute the method described in any one of claims 1 to 11.

13. A computer-readable storage medium, characterized in that Used to store computer instructions, when the computer instructions are executed on an electronic device, the electronic device executes the method according to any one of claims 1 to 11.

Citation Information

Patent Citations

  • Audio editing method and device, electronic equipment, storage medium and program product

    CN116954442A

  • Voice information sending method and device and electronic equipment

    CN117041409A

  • Note generation method based on voice recognition, terminal equipment and storage medium

    CN117666917A

  • Content Editing Method and Terminal

    US20220005241A1