Voice text processing method and electronic equipment
By synchronously adding annotation information during the speech conversion process, the inefficiency problem caused by the separation of speech recording and annotation information is solved, multimodal synchronous recording is realized, and the efficiency and integrity of information recording are improved.
Patent Information
- Application Number
- CN202510464546.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-07-18
AI Technical Summary
In the prior art, voice recording and annotation information are each independent, and users need to manually organize and compare them through multiple applications, resulting in low information recording efficiency.
During the speech conversion process, multimodal synchronous recording of annotation and speech recording is realized by receiving interface input to the speech conversion text.
It improves the efficiency of information recording, ensures the integrity and accuracy of records, avoids information omissions, and simplifies user operation processes.
Smart Images

Figure CN120337883A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of electronic devices, and particularly relates to a method for processing speech texts and an electronic device. Background Art
[0002] With the continuous enrichment of the functions of electronic devices, note-taking and voice recording have become important tools for people's daily work, study, and communication.
[0003] For example, in scenarios such as meetings or lectures, users can first record the voice signals to be recorded through the recording application program in the electronic device, then convert the recorded audio file into the corresponding text content through the speech-to-text application program in the electronic device, and finally add annotations to the text content through the note application program in the electronic device, so that users can trace back the recorded information.
[0004] However, in the process of recording information as described above, the annotations and voice recordings are independent of each other and the information is scattered. Users need to manually organize and compare each piece of recorded information through multiple application programs to integrate the annotations and voice recordings, resulting in low efficiency of information recording. Summary of the Invention
[0005] The purpose of the embodiments of this application is to provide a method for processing speech texts and an electronic device, which can improve the efficiency of information recording.
[0006] In a first aspect, the embodiments of this application provide a method for processing speech texts, which includes: when performing speech conversion processing on a first speech, receiving a first input to a first interface; the first interface includes the speech conversion text of the first speech; in response to the first input, adding annotation information to the speech conversion text of the first speech.
[0007] In a second aspect, the embodiments of this application provide a device for processing speech texts, which includes: a receiving module and a processing module; the receiving module is used to receive a first input to a first interface when performing speech conversion processing on a first speech; the first interface includes the speech conversion text of the first speech; the processing module is used to add annotation information to the speech conversion text of the first speech in response to the first input received by the receiving module.
[0008] In a third aspect, the embodiments of this application provide an electronic device, which includes a processor and a memory, and the memory stores a program or instruction that can run on the processor. When the program or instruction is executed by the processor, the steps of the method described in the first aspect are implemented.
[0009] Fourthly, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the method described in the first aspect are implemented.
[0010] Fifthly, an embodiment of the present application provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor, and the processor is configured to run a program or instruction to implement the method described in the first aspect.
[0011] Sixthly, an embodiment of the present application provides a computer program / program product, which is stored in a storage medium and is executed by at least one processor to implement the method described in the first aspect.
[0012] In an embodiment of the present application, when performing speech conversion processing on a first speech, a first input to a first interface can be received; the first interface includes the speech conversion text of the first speech; and in response to the first input, annotation information is added to the speech conversion text of the first speech. Through this solution, since when performing speech conversion processing on the first speech, annotation information can be added to the speech conversion text of the first speech through the first input to the first interface including the speech conversion text of the first speech, annotation can be added synchronously during the process of performing speech conversion processing on the speech, and it is not necessary for the user to manually organize and compare each piece of information recorded through multiple application programs, so that the annotation and the speech record can be integrated, thereby improving the efficiency of information recording. Description of the Drawings
[0013] Figure 1 is a flowchart of a speech text processing method provided by some embodiments of the present application;
[0014] Figure 2 is a schematic diagram of a first interface provided by some embodiments of the present application;
[0015] Figure 3 is a schematic diagram of adding annotation information provided by some embodiments of the present application;
[0016] Figure 4 is a schematic diagram of adding annotation information provided by some embodiments of the present application;
[0017] Figure 5 is a schematic diagram of adding annotation information provided by some embodiments of the present application;
[0018] Figure 6 is a schematic diagram of adding annotation information provided by some embodiments of the present application;
[0019] Figure 7Schematic diagram of adding annotation information provided by some embodiments of the present application;
[0020] Figure 8 Schematic diagram of adding annotation information provided by some embodiments of the present application;
[0021] Figure 9 Schematic diagram of adding annotation information provided by some embodiments of the present application;
[0022] Figure 10 Schematic diagram of generating annotation information provided by some embodiments of the present application;
[0023] Figure 11 Schematic diagram of generating annotation information provided by some embodiments of the present application;
[0024] Figure 12 Schematic diagram of the voice text processing device provided by the embodiments of the present application;
[0025] Figure 13 Schematic diagram of the electronic device provided by the embodiments of the present application;
[0026] Figure 14 Schematic diagram of the hardware of the electronic device provided by the embodiments of the present application. Detailed implementation manners
[0027] Next, the technical solutions in the embodiments of the present application will be clearly described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, rather than all, of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art belong to the scope of protection of the present application.
[0028] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order different from those illustrated or described herein, and the objects distinguished by "first", "second", etc. generally belong to the same category, and the number of objects is not limited. For example, the first object can be one or more. In addition, "and / or" in the specification and claims means at least one of the connected objects, and the character " / " generally indicates an "or" relationship between the associated objects before and after.
[0029] The term "indication" in this application can be either a direct indication (or an explicit indication) or an indirect indication (or an implicit indication). Among them, a direct indication can be understood as that the sender clearly tells the receiver specific information, operations to be performed, request results, etc. in the sent indication; an indirect indication can be understood as that the receiver determines the corresponding information according to the indication sent by the sender, or makes a judgment and determines the operations to be performed or request results, etc. according to the judgment result.
[0030] The terms "at least one (item)", "at least one of", etc. in the specification and claims of this application refer to any one, any two or more combinations of the included objects. For example, at least one (item) of a, b, and c can represent: "a", "b", "c", "a and b", "a and c", "b and c", and "a, b, and c", where a, b, and c can be single or multiple. Similarly, "at least two (items)" means two or more, and its expressed meaning is similar to that of "at least one (item)".
[0031] Some concepts and terms involved in the speech text processing method, device, electronic device, and readable storage medium provided by the embodiments of this application are explained below.
[0032] Speech-to-text is a technology that converts spoken language into written language. It uses speech recognition technology to convert speech signals into readable text output. With the development of artificial intelligence and natural language processing, the application scope of this technology has been continuously expanded, becoming an important tool to improve work efficiency.
[0033] The basic principle of speech-to-text:
[0034] Speech-to-text technology uses speech recognition and natural language processing technologies to convert speech content into preliminary text, and corrects grammar and adds punctuation through post-text processing to make the text more professional and easy to read.
[0035] Application scenarios of speech-to-text:
[0036] 1. Meeting and interview records: Automatically generate meeting minutes, saving time and effort.
[0037] 2. Subtitle generation: Automatically generate subtitles for videos, providing a barrier-free viewing experience.
[0038] 3. Content creation: Transcribe podcast audio into blogs or reports to enrich the content.
[0039] 4. Professional records: Transcribe court trial records and doctor diagnoses into text files for convenient archiving.
[0040] 5. Education and Training: Convert classroom recordings into lecture notes for easy review by students.
[0041] The following will, in conjunction with the accompanying drawings, elaborate on the speech text processing method, apparatus, electronic device, and readable storage medium provided by the embodiments of the present application through specific embodiments and their application scenarios.
[0042] The speech text processing method provided by the embodiments of the present application can be applied to scenarios where notes are added during recording. Among them, one specific application scenario is adding notes to the recorded meeting content during a meeting; another specific application scenario is adding notes to the recorded course content during training.
[0043] It should be noted that for the speech text processing method provided by the embodiments of the present application, the execution subject can be a speech text processing device, an electronic device, or a functional module in the electronic device, etc. In some embodiments of the present application, the case where the electronic device executes the speech text processing method is taken as an example to illustrate the speech text processing method provided by the embodiments of the present application.
[0044] Figure 1 The flowchart of the speech text processing method provided by some embodiments of the present application is shown. As Figure 1 shown, the speech text processing method provided by some embodiments of the present application may include the following steps 101 and 102.
[0045] Step 101: When performing speech conversion processing on the first speech, the electronic device receives a first input to the first interface.
[0046] Among them, the first interface includes the speech conversion text of the first speech.
[0047] In some embodiments of the present application, the above-mentioned first speech may include a human voice speech signal. For example, the first speech may be the speech signal of the speaker in a meeting, or may be the speech signal of the lecturer in training, or may be the speech signal of the speaker in a speech, etc., which is not limited in the embodiments of the present application.
[0048] In some embodiments of the present application, the above-mentioned speech conversion processing may include any possible speech conversion processing such as speech recording processing or speech playback processing.
[0049] In some embodiments of the present application, the first interface may be an interface integrated with functions such as a recording function, a speech playback function, a speech-to-text function, and a note function.
[0050] Exemplarily, as Figure 2As shown in the figure, the first interface 10 includes a title bar 11, a speech-to-text area 12, an input area 13, and a speech progress control 14. Among them, the title bar 11 carries the functions of searching for notes, withdrawing the current operation, and exporting documents; the speech-to-text area 12 is used for real-time voice recording and conversion into editable text, while keeping synchronized with the handwritten content and can be expanded to full screen for browsing; the input area 13 supports real-time handwritten recording through a smart pen or a touch device, saved as an image or converted into text, and marked synchronously in the speech progress control; the speech progress control 14 is used to integrate and present multi-modal content in chronological order, and supports users to perform quick backtracking, marking, and retrieval through the speech progress control 14.
[0051] For example, when the user needs to record a voice signal, the recording control 15 in the above-mentioned speech progress control 14 can be clicked to make the electronic device start recording the audio signal, and the recorded voice signal is converted into text content in real time and the converted text content is displayed in the speech-to-text area 12.
[0052] In some embodiments of the present application, the speech conversion text of the first speech, that is, the text content obtained by converting the recorded voice signal through speech-to-text. For example, if the first speech is "Welcome everyone to today's meeting", the corresponding speech conversion text is the text "Welcome everyone to today's meeting".
[0053] Exemplarily, in the scenario of adding notes to the recorded meeting content during a meeting, assuming that attendee A is speaking during the meeting and the user needs to add notes to the speech content of attendee A, then the voice signal of attendee A can be recorded first through the electronic device, and during the process of recording the voice signal of attendee A, the recorded voice signal is converted into the corresponding speech conversion text in real time, and the speech conversion text is synchronously displayed on the first interface.
[0054] Exemplarily, in the scenario of adding notes to the recorded course content during a training, assuming that lecturer B is giving a lecture during the training and the user needs to add notes to the speech content of lecturer B, then the voice signal of lecturer B can be recorded first through the electronic device, and during the process of recording the voice signal of lecturer B, the recorded voice signal is converted into the corresponding speech conversion text in real time, and the speech conversion text is synchronously displayed on the first interface.
[0055] In some embodiments of the present application, the first input is used to perform operations related to adding annotation information to the text content.
[0056] In some embodiments of the present application, the first input can be any possible form of input such as touch input, voice input, or physical button input.
[0057] For example, taking the above-mentioned first input as a touch input, the first input may include, but is not limited to: single-click input, double-click input, triple-click input, long-press input, hard-press input, swipe input, special trajectory input, or other feasible inputs made by the user on the first interface using a finger or a stylus. Specifically, it can be determined according to actual usage requirements, and the embodiments of the present application do not make any limitations; among them, the special trajectory input may include, but is not limited to: any trajectory input such as circular trajectory input, elliptical trajectory input, heart-shaped trajectory input, wavy trajectory input, square trajectory input, rectangular trajectory input, diamond-shaped trajectory input, or trapezoidal trajectory input.
[0058] For example, taking the above-mentioned first input as a voice input, the first input may include, but is not limited to: any possible voice input such as the voice "Add a note to text content 1" or the voice "Add a comment to text content 2".
[0059] For example, taking the above-mentioned first input as a physical button input, the first input may include, but is not limited to, any one of the following: the input of pressing the lock screen button and the volume "+" button simultaneously, the input of pressing the lock screen button and the volume "-" button simultaneously, the input of pressing the volume "+" button and the volume "-" button simultaneously, and the input of pressing the lock screen button, the volume "+" button, and the volume "-" button simultaneously.
[0060] Step 102: In response to the first input, the electronic device adds annotation information to the speech conversion text of the first voice.
[0061] Among them, the above-mentioned annotation information includes at least one of text and image.
[0062] In some embodiments of the present application, the speech conversion text of the first voice may be at least part of the speech conversion text corresponding to the first voice.
[0063] In some embodiments of the present application, the above-mentioned annotation information is the annotation content of the user for the text content. For example, assuming that the speech conversion text of the first voice is the discussion content of the topic "Product Value" in a meeting, then the annotation information may be the central idea or key points summarized by the user for the discussion content of this topic.
[0064] In some embodiments of the present application, the first interface may further include an input area. Exemplarily, the above-mentioned step 101 may be specifically implemented by the following step 101a, and the above-mentioned step 102 may be specifically implemented by the following step 102a.
[0065] Step 101a: When performing speech conversion processing on the first voice, the electronic device receives the first input to the input area.
[0066] In some embodiments of the present application, the first input is specifically used to edit annotation information.
[0067] For example, as Figure 2 shown, after receiving the first input, the electronic device can, in response to the first input, edit annotation information within the input area 13 in the first interface 10.
[0068] In some embodiments of the present application, after the electronic device edits the annotation information, it can, based on the edited annotation information, determine first text content related to the annotation information from the text converted from the first voice, and then add the annotation information to the first text content.
[0069] In some embodiments of the present application, the first text content can be the text content in the text converted from the first voice that is most strongly associated with the above-mentioned annotation information.
[0070] In some embodiments of the present application, after the electronic device edits the above-mentioned annotation information, it can first perform semantic analysis on the text converted from the first voice, and then, according to the result of the semantic analysis, determine the first text content related to the annotation information from the text converted from the voice.
[0071] In some embodiments of the present application, after the electronic device adds the above-mentioned annotation information to the first text content, it can display an annotation mark at a certain position in the first text content and display the annotation information in the blank area near the first text content.
[0072] Exemplarily, assuming that in the discussion content of the topic of "product value" during a meeting recording, the user summarizes the central idea of this content after listening and wants to record it. Then, as Figure 3 shown, at 02:00 of the recording, the electronic device, through the user's first input, edits annotation information in the input area 13. During the editing process, an annotation mark 16 is defaultly displayed at the beginning of the text content where 02:00 is located in the speech-to-text area 12. At the same time, during the process of editing the annotation, it will be synchronously displayed in real time in the bubble control 17 of the speech-to-text area 12. The annotation edited at 02:08 is "The core is", which will be synchronously displayed in the bubble control 17. As Figure 4 shown, if the editing of the annotation information is completed at 03:32 of the recording, the complete paragraph content can be combined with the current mark, and through semantic analysis, the text content with the strongest association is determined as the first text content. An annotation mark 18 is displayed at the beginning of the first text content, and the annotation information 19 is displayed in the blank area of the speech-to-text area 12, thus completing the annotation of the first text content.
[0073] In some embodiments of the present application, since the electronic device can first input through the first input and edit annotation information on the first interface, and then the electronic device automatically determines the first text content related to the annotation information from the text converted from the first voice, and adds the annotation information to the first text content, there is no need for the user to manually select the text content to which the annotation needs to be added, thereby simplifying the process of adding annotations and improving the efficiency of adding annotations.
[0074] In some embodiments of the present application, the first input can specifically be used to select text content.
[0075] Optionally, in the embodiments of the present application, the electronic device can first respond to the first input and select the first text content from the text converted from the first voice. Then edit the annotation information on the first interface or generate annotation information based on the first text content. Finally, add the annotation information to the first text content.
[0076] For example, taking the first input as a touch input as an example, the user can first click on the start position of the text content to be selected, and then drag along with the hand to the end position of the text content to be selected, so that the electronic device can determine the text content selected by the user as the first text content, that is, select the first text content from the above text content.
[0077] In some embodiments of the present application, after the electronic device selects the first text content from the text converted from the first voice, it can first display an annotation addition control. The user can first click on the annotation addition control, and then write annotation information on the first interface.
[0078] Exemplarily, as Figure 5 shown, after long pressing and selecting the first text content 51 in the text converted from the first voice, click on the control 52 in the pop-up bubble menu, and an annotation style is displayed in the voice text. At this time, write annotation information in the input area 53, and then as Figure 6 shown, complete the operation of adding an annotation.
[0079] In some embodiments of the present application, the above input area can display the historically added annotation information. If the user wants to add a certain historical annotation information to the first text content, the user can first retrieve the historical annotation information to be added through input, and then directly drag the historical annotation information to the first text content through input, thereby completing the addition of the historical annotation information to the first text content.
[0080] Exemplarily, if the user needs to add the historical annotation information "The core is how to endow the server with capabilities" to the first text content, the user can first find the first text content by swiping up and down in the above voice-to-text area, and then as Figure 7As shown, long press the historical annotation information 54 in the above input area and drag it to the first text content, then release it, so that as Figure 8 shown, adding the historical annotation information to the first text content is completed. While adding is completed, the corresponding annotation information marker will also be updated by superposition according to the addition times of the annotation information. For example, as Figure 9 shown, the marker indicating the quantity 1 changes to the marker 55 indicating the quantity 2, and at the same time, a white dot marker 56 is added to the voice progress control to mark the time of adding the annotation information.
[0081] In some embodiments of the present application, since after the electronic device receives the first input, it can first select the first text content from the text converted from the first voice, then edit the annotation or generate an annotation based on the first text content on the first interface, and complete adding the annotation to the first text content, so any text content can be selected for adding an annotation according to the user's needs, thereby improving the flexibility of adding an annotation.
[0082] Step 102a: The electronic device responds to the first input and adds annotation information to the text converted from the first voice according to the input time of the first input.
[0083] In some embodiments of the present application, since in the case of performing voice conversion processing on the first voice, the electronic device can, through the first input to the input area, add annotation information to the text converted from the first voice according to the input time of the first input, so the accurate position for adding the annotation information can be automatically determined.
[0084] In some embodiments of the present application, the electronic device can respond to the first input, determine the position for adding annotation information according to the input event of the first input, and then add annotation information to the text converted from the first voice at this position.
[0085] In some embodiments of the present application, the above annotation information includes first annotation information and second annotation information. Exemplarily, the above step 102a can be specifically implemented through the following steps 102a1 to 102a4.
[0086] Step 102a1: The electronic device responds to the first input and determines the first annotation information.
[0087] In some embodiments of the present application, the first annotation information can be an image hand-drawn by the user or input text.
[0088] Step 102a2: The electronic device receives a fourth input to the text converted from the first voice.
[0089] In some embodiments of the present application, the fourth input is used to input text content.
[0090] For other descriptions of the fourth input, reference may be made to the above specific description of the first input. To avoid repetition, it will not be elaborated here.
[0091] Step 102a3: The electronic device responds to the fourth input and determines the second annotation information according to the text content corresponding to the fourth input.
[0092] In some embodiments of the present application, the second annotation information may be an image determined according to the text content corresponding to the fourth input.
[0093] Step 102a4: The electronic device adds the first annotation information and the second annotation information to the speech conversion text of the first speech according to the input time of the first input.
[0094] Exemplarily, as Figure 10 shown, the speech conversion text of the first speech is "According to the hierarchical structure in the pyramid principle, it can be divided into three hierarchical structures A, B, and C". Then, after the electronic device selects this text content, the user can drag the text content to the input area 13. Then, as Figure 11 shown, the electronic device can automatically generate an image annotation 20 according to this text content through an artificial intelligence model and display an image 21 in the speech-to-text area 12, thereby completing the annotation of the first text content.
[0095] In some embodiments of the present application, if the electronic device edits a text annotation on the first interface, it can also automatically generate a corresponding image annotation based on this text annotation, and then add this image annotation to the first text content.
[0096] In some embodiments of the present application, since the electronic device can determine the second annotation information through the fourth input after determining the first annotation information, and add the first annotation information and the second annotation information to the speech conversion text of the first speech according to the input time of the first input, the annotation information can be updated, thereby further improving the flexibility of adding annotation information.
[0097] In some embodiments of the present application, the second annotation information includes a third image. Exemplarily, the above step 102a3 can be specifically implemented through the following steps A to C.
[0098] Step A: The electronic device responds to the fourth input and determines a fourth image according to the text content corresponding to the fourth input.
[0099] In some embodiments of the present application, the fourth input is used to generate an image.
[0100] For other descriptions of the fourth input, reference may be made to the above relevant description of the first input. To avoid repetition, it will not be elaborated here.
[0101] In some embodiments of the present application, the fourth image is an image generated based on the text content corresponding to the fourth input.
[0102] Step B: The electronic device receives a fifth input.
[0103] In some embodiments of the present application, the fifth input is used to process an image.
[0104] In some embodiments of the present application, the fifth input may be any possible form of input such as a touch input, a voice input, or a physical button input.
[0105] For example, taking the above fifth input as a touch input, the fifth input may include, but is not limited to: a single - click input, a double - click input, a triple - click input, a long - press input, a hard - press input, a slide input, a special trajectory input, or other feasible inputs made by the user with a finger or a stylus on the fourth image. Specifically, it can be determined according to actual usage requirements, and the embodiments of the present application do not make limitations; among them, the special trajectory input may include, but is not limited to: any trajectory input such as a circular trajectory input, an elliptical trajectory input, a heart - shaped trajectory input, a wavy trajectory input, a square trajectory input, a rectangular trajectory input, a diamond - shaped trajectory input, or a trapezoidal trajectory input.
[0106] For example, taking the above fifth input as a voice input, the fifth input may include, but is not limited to: any possible voice input such as the voice "Supplement the fourth image" or the voice "Supplement and beautify the fourth image".
[0107] For example, taking the above fifth input as a physical button input, the fifth input may include, but is not limited to any one of the following: an input of pressing the lock screen button and the volume "+" button simultaneously, an input of pressing the lock screen button and the volume "-" button simultaneously, an input of pressing the volume "+" button and the volume "-" button simultaneously, an input of pressing the lock screen button, the volume "+" button, and the volume "-" button simultaneously.
[0108] Step C: In response to the fifth input, the electronic device determines a third image based on the fourth image.
[0109] In some embodiments of the present application, the third image may be an image obtained by processing the fourth image.
[0110] In some embodiments of the present application, the above - mentioned image processing may include, but is not limited to, at least one of the following: adding text processing, image beautification processing.
[0111] In some embodiments of the present application, the electronic device may use image processing parameters to perform image blackcurrant on the fourth image to obtain the third image.
[0112] In some embodiments of the present application, the above image processing parameters can be the system default, or can be arbitrarily set by the user according to actual usage requirements, and the embodiments of the present application do not make limitations.
[0113] In some embodiments of the present application, the above image processing parameters may include, but are not limited to, at least one of the following: image color, image size, image resolution, image transparency, and image shape.
[0114] In some embodiments of the present application, when the above image processing includes adding text processing, the electronic device may first determine the text related to the fourth image from the text converted from the first voice, and then perform the adding text processing on the fourth image to add the determined text to the fourth image; then, according to the image processing parameters, perform image beautification processing on the fourth image with added text to obtain the third image.
[0115] In some embodiments of the present application, after determining the third image, the electronic device may update the fourth image to the third image.
[0116] In some embodiments of the present application, when the electronic device updates the fourth image to the third image, it can be understood that: the electronic device cancels the display of the fourth image and displays the third image at the position where the fourth image is located.
[0117] In some embodiments of the present application, since the electronic device can obtain the third image based on the fourth image through the user's input, after adding the fourth image, the fourth image can be further supplemented and / or beautified, thereby improving the display effect and accuracy of the annotation.
[0118] In some embodiments of the present application, the first interface may further display an artificial intelligence note assistant control for real-time optimizing and marking content, and personalized generating summaries and recommendations.
[0119] In some embodiments of the present application, the first interface may further include a sharing control. Through this sharing control, multi-user collaborative recording and remote sharing of note content can be achieved, which is applicable to scenarios such as team meetings and online teaching.
[0120] In some embodiments of the present application, the first interface may further include a synchronization control. Through this synchronization control, integration with third-party applications, cloud synchronization, and seamless cross-device collaboration can be achieved.
[0121] In the speech text processing method provided by some embodiments of the present application, since when performing speech conversion processing on the first speech, annotation information can be added to the speech conversion text of the first speech through the first input of the first interface including the speech conversion text of the first speech, it is possible to synchronously add annotations during the process of performing speech conversion processing on the speech, achieving multimodal synchronous recording of annotations and speech records, enabling different recording methods to complement each other, and allowing the integration of annotations and speech records without the user manually sorting and comparing each piece of information recorded through multiple applications. Thus, the efficiency of information recording can be improved, and the integrity of the recorded information can be ensured, avoiding omission of information.
[0122] In some embodiments of the present application, the above annotation information may include annotation content and a first annotation identifier; the first interface may further include a speech progress control. Exemplarily, the speech text processing method provided by some embodiments of the present application may further include the following step 103.
[0123] Step 103: The electronic device adds a second annotation identifier to the speech progress control according to the input time of the first input.
[0124] In some embodiments of the present application, the electronic device determines the position for adding the second annotation identifier according to the input time of the first input, and then adds the second annotation identifier at this position on the speech progress control.
[0125] In some embodiments of the present application, the above speech progress control is used to indicate the progress of recording a speech signal.
[0126] In some embodiments of the present application, the electronic device may display the second annotation identifier at a first position on the speech progress control.
[0127] Wherein, the first position is the position corresponding to the time when the above annotation information starts to be edited or generated, and any annotation identifier on the above speech progress control is used to trigger the display of an annotation and the text content corresponding to the annotation.
[0128] In some embodiments of the present application, the display parameters of the speech progress control may be any display parameters.
[0129] For example, as Figure 2 shown, the speech progress control 14 may include a recording control 15. When the user clicks the recording control 15, recording of the speech signal starts, and the progress of recording the speech signal can be indicated through this speech progress control 14.
[0130] It can be understood that after each addition of annotation information, the electronic device can display the annotation identifier corresponding to the annotation information at the position corresponding to the time when the above annotation information starts to be edited or generated on the speech progress control, so as to facilitate the user to quickly locate or trace back the annotation information.
[0131] In some embodiments of the present application, different forms of annotation information may correspond to different annotation identifiers. For example, text annotations may correspond to white dot identifiers, and image annotations may correspond to green dot identifiers. In this way, through different annotation identifiers, the content of the corresponding form of annotation information can be quickly found.
[0132] Exemplarily, in a meeting, if you want to return to a previous paragraph, you can drag the slider of the bottom voice progress control, and browse back through the voice, text, and all the annotation information and markings made at that time point.
[0133] For example, by long-pressing and dragging the slider on the voice progress control, move forward and backward to complete the positioning and backtracking of the meeting record; during the process of long-pressing and dragging the slider, the content in the above-mentioned voice-to-text area 12 and the above-mentioned input area 13 are simultaneously associated with the time node and the content is updated and displayed in real time; first, the content in the voice-to-text area 12 will be the text content recorded at that time according to the positioning on the voice progress control, and the input area 13 will display the corresponding annotation information content according to the recorded white and green dots.
[0134] When the user wants to backtrack the previous text content and annotation information, they can long-press and drag from the position of 07:22 on the voice progress control to the position marked by the second green dot. The above-mentioned voice-to-text area 12 backtracks to the position where the voice record content is located at 03:25. The above-mentioned annotation information input area 13 is then positioned to the position where the picture is recorded and automatically positioned to the center of the panel for display, and other record elements are dimly displayed. Continuing to drag the slider forward, when positioned to the position marked by the white dot, that is, at 02:08, the input area 13 is positioned to the position where the annotation text is recorded, and other elements are dimly displayed, and they resume display after the drag is released.
[0135] In some embodiments of the present application, since the electronic device can add a second annotation identifier to the voice progress control according to the input time of the first input, it is convenient for the user to quickly locate and display the voice conversion text of the first voice and the annotation information through this annotation identifier during subsequent use, so that it is convenient for the user to backtrack the annotation information and improve the functionality of the annotation.
[0136] In some embodiments of the present application, the voice text processing method provided by some embodiments of the present application may further include the following step 104 and step 105.
[0137] Step 104: When performing voice conversion processing on the first voice, the electronic device receives a second input to the first text in the voice conversion text of the first voice.
[0138] In some embodiments of the present application, the second input is used to add marker information.
[0139] For other descriptions of the second input, reference may be made to the relevant descriptions of the first input above. To avoid repetition, they will not be elaborated here.
[0140] Step 105: The electronic device responds to the second input and adds marking information to the first text.
[0141] In some embodiments of the present application, the above marking information is the marking content of the user for the text content. For example, assuming the first text is "The core is how to endow the server with capabilities", then the marking information can be a bold marking, a highlighting marking, an underlining marking, etc. for the first text.
[0142] In some embodiments of the present application, after the electronic device selects the first text, it can display a marking control for marking the text content.
[0143] In some embodiments of the present application, the display parameters of the above marking control can be any possible display parameters.
[0144] In some embodiments of the present application, the above display parameters may include but are not limited to at least one of the following: display color, display size, display shape, display transparency.
[0145] In some embodiments of the present application, the electronic device can add marking information to the first text by inputting the above marking control and adopting preset marking parameters.
[0146] In some embodiments of the present application, the preset marking parameters can be the system default, or can be arbitrarily set by the user according to actual usage requirements, which are not limited in the embodiments of the present application.
[0147] In some embodiments of the present application, the preset marking parameters may include but are not limited to at least one of the following: marking color, marking size, marking shape, marking transparency.
[0148] In some embodiments of the present application, since in the case of performing speech conversion processing on the first speech, the electronic device can add marking information to the first text by the second input to the first text in the speech conversion text of the first speech, the key content in the text content can be displayed in a special display style, so that it can be convenient for the user to quickly locate and review the key content.
[0149] In some embodiments of the present application, the speech text processing method provided by some embodiments of the present application may further include the following steps 106 to 108.
[0150] Step 106: The electronic device responds to the first input and determines a second image.
[0151] In some embodiments of the present application, the first input is specifically used for drawing an image.
[0152] In some embodiments of the present application, the second image may be a sketch hand-drawn by a user.
[0153] Step 107: The electronic device receives a third input.
[0154] In some embodiments of the present application, the third input is used for processing an image.
[0155] Step 108: In response to the third input, the electronic device determines a first image according to the second image.
[0156] In some embodiments of the present application, the first image may be an image obtained by beautifying the second image. Exemplarily, the user can first directly draw a sketch (the second image) in the above input area 13 as a note to assist in comparing text with graphics. And based on the context of the first text content, the user can use artificial intelligence technology to beautify or redraw the sketch to obtain the first image. In the hand-drawing, setting items can be displayed at the bottom for functions and interactions.
[0157] For example, after discussing "the hierarchical structure in the pyramid principle can be divided into three hierarchical structures A, B, and C" in a meeting, a triangle is drawn in the above input area 13. At this time, through context semantic judgment, it is related to the content of the pyramid principle. After the sketch is completed, the sketch will be recorded at the beginning position marker of the speech content about "the hierarchical structure in the pyramid principle can be divided into three hierarchical structures A, B, and C", and a green dot will be marked on the speech progress control 14 at the time point when the sketch starts to be drawn; when the sketch is in the selected state and the AI function button in the bottom function area is clicked, the bottom function area supports copying, AI processing, and deletion; among them, AI processing includes generating a finished picture from the sketch and adding annotations by matching text content, and the two functions can be used superimposed. Among them, only checking "fill the sketch with content" will add ABC annotations in the sketch by associating the first text content "can be divided into three hierarchical structures A, B, and C" in the sketch; if both "AI generation" and "fill the sketch with content" are checked, a beautiful picture will be generated on the basis of the filled content.
[0158] For example, after it is mentioned in the training that "from the inner layer to the outer layer are E, F, and G in sequence", a circle is drawn in the above input area 13. At this time, through context semantic judgment, it is related to the layering from the inside out. After the sketch is completed, the sketch will be recorded at the starting position marker of the speech content that mentions "from the inner layer to the outer layer are E, F, and G in sequence", and a green dot will be marked on the speech progress control 14 at the time point when the sketch starts to be drawn; when the sketch is in the selected state and the AI function button in the bottom function area is clicked, the bottom function area supports copying, AI processing, and deletion; among them, AI processing includes generating a finished drawing from the sketch and adding annotations by matching text content, and the two functions can be used superimposed. Among them, only checking "fill the sketch with content", the first text content "in sequence are E, F, and G" is associated in the sketch, and then two circles (with different sizes and the same center) are nested and drawn in the drawn circle in sequence, and E, F, and G annotations are added respectively; if both "AI generation" and "fill the sketch with content" are checked, a beautiful picture will be generated on the basis of the filled content.
[0159] In some embodiments of the present application, since the electronic device can determine a third image based on the second image through another input after obtaining the second image through the user's input, after the user draws a sketch, the sketch can be beautified through the input, so that the effect and convenience of the annotation can be improved.
[0160] In some embodiments of the present application, the first interface further includes a first control. Exemplarily, the speech text processing method provided in some embodiments of the present application may further include the following step 109 and step 110.
[0161] Step 109: The electronic device receives a sixth input to the first control.
[0162] In some embodiments of the present application, the first control can be used to turn on the first display mode.
[0163] In some embodiments of the present application, the first display mode may include any possible display modes such as full-screen display mode, split-screen display mode, or small-window display mode.
[0164] In some embodiments of the present application, the display parameters of the first control can be any display parameters.
[0165] In some embodiments of the present application, the first control can be displayed at any position in the first interface.
[0166] For example, the first control can be displayed in the above-mentioned speech-to-text area 12 in the first interface. When the user clicks the first control, the electronic device can display text and annotation information according to the first display mode, so as to facilitate the user to view or search for notes.
[0167] In some embodiments of the present application, the sixth input is used to trigger the display of text and annotation information according to the first display mode.
[0168] For other descriptions of the sixth input, reference may be made to the relevant descriptions of the first input above. To avoid repetition, they will not be elaborated here.
[0169] Step 110: The electronic device responds to the sixth input and displays the speech conversion text and annotation information of the first speech according to the first display mode.
[0170] For example, assuming the first display mode is the full-screen display mode, after receiving the sixth input, the electronic device can respond to the sixth input and display the speech conversion text and annotation information of the first speech in full screen.
[0171] In some embodiments of the present application, since the electronic device can display the speech conversion text and annotation information of the first speech according to the first display mode by the sixth input to the first control, the note display effect can be improved, facilitating the user to view or search for notes.
[0172] In some embodiments of the present application, the above step 110 can be specifically implemented by the following step 110a, and the speech text processing method provided by some embodiments of the present application may further include the following steps 111 and 112.
[0173] Step 110a: The electronic device responds to the sixth input and displays the speech conversion text, annotation information, and speech progress control of the first speech according to the first display mode.
[0174] Step 111: The electronic device receives the seventh input to the speech progress control.
[0175] In some embodiments of the present application, the seventh input is used to select a position in the speech progress control.
[0176] In some embodiments of the present application, the seventh input may be an input to the annotation identifier on the speech progress control.
[0177] Step 112: The electronic device responds to the seventh input and displays the speech conversion text and annotation information corresponding to the speech position of the seventh input.
[0178] In some embodiments of the present application, the electronic device can display the speech conversion text and annotation information corresponding to the speech position of the seventh input at any position on the first interface. For example, the speech conversion text and annotation information corresponding to the speech position of the seventh input can be displayed around the speech progress control, or the speech conversion text and annotation information corresponding to the speech position of the seventh input can be superimposed on the speech progress control.
[0179] In some embodiments of the present application, since the electronic device can display the speech conversion text, annotation information, and speech progress control of the first speech according to the sixth input of the user, and display the speech conversion text and annotation information corresponding to the speech position of the seventh input through the seventh input of the user, it is convenient to quickly search for historical content and annotation information in different display modes.
[0180] In some embodiments of the present application, the above-mentioned speech text processing method can be applied to virtual reality (VR) devices, extended reality (XR) devices, augmented reality (AR) devices, etc. On these devices, multi-modal content can be displayed and operated in real time, providing a more intuitive annotation interaction experience.
[0181] In some embodiments of the present application, the above-mentioned speech text processing method can combine intelligent error correction and semantic analysis technologies to improve the accuracy and neatness of handwritten and speech content.
[0182] The speech text processing method provided by some embodiments of the present application integrates handwritten notes, speech-to-text conversion, and combines speech progress controls for unified display, and has the following beneficial effects:
[0183] 1. Improve the efficiency and integrity of speech text processing
[0184] Multi-modal synchronous recording of handwriting, speech, and automatically generated content is realized, enabling different recording methods to complement each other. Users do not need to switch between multiple tools, significantly improving the efficiency of speech text processing. At the same time, synchronous display through the timeline ensures the time consistency and integrity of the recorded content, avoiding omission of key information.
[0185] 2. Enhance the convenience of content retrieval and review
[0186] The system integrates all multi-modal content in chronological order through the timeline module. Users can quickly locate the note content at a specific time point through simple interaction operations such as sliding and zooming. This timeline display method makes information review more intuitive. Users can easily go back to any time point to view relevant content, improving the convenience of information retrieval.
[0187] 3. Intelligent association and marking to improve information acquisition efficiency
[0188] Using semantic analysis and timestamp technologies, handwriting and speech content are automatically associated, and key content is highlighted on the timeline. Users can quickly identify and obtain the required information. In addition, users are supported to manually add annotations and marks, further improving the efficiency of personalized note organization.
[0189] Each of the above method embodiments, or various possible implementation manners in each method embodiment, can be executed independently, or, on the premise of no contradiction, can also be executed in combination with each other. Specifically, it can be determined according to actual usage requirements, and the embodiments of the present application do not limit this.
[0190] In the voice text processing method provided by the embodiments of the present application, the execution subject can be a voice text processing device. In some embodiments of the present application, taking the voice text processing device executing the voice text processing method as an example, the voice text processing device provided by the embodiments of the present application is described.
[0191] As Figure 12 shown, some embodiments of the present application provide a voice text processing device 10, and the voice text processing device 10 may include: a receiving module 11 and a processing module 12.
[0192] Among them, the receiving module 11 can be used to receive a first input to a first interface when performing voice conversion processing on a first voice; the first interface includes a voice conversion text of the first voice. The processing module 12 can, in response to the first input received by the receiving module 11, add annotation information to the voice conversion text of the first voice.
[0193] In some embodiments of the present application, the first interface further includes an input area. Exemplarily, the receiving module 11 can specifically be used to receive a first input to the input area. The processing module 12 can specifically be used to, in response to the first input received by the receiving module 11, add annotation information to the voice conversion text of the first voice according to the input time of the first input.
[0194] In some embodiments of the present application, the above annotation information includes annotation content and a first annotation identifier; the first interface further includes a voice progress control. Exemplarily, the processing module 12 can further be used to add a second annotation identifier to the voice progress control according to the input time of the first input.
[0195] In some embodiments of the present application, the receiving module 11 can further be used to receive a second input to a first text in the voice conversion text of the first voice when performing voice conversion processing on the first voice. The processing module 12 can further be used to, in response to the second input received by the receiving module 11, add marking information to the first text.
[0196] In some embodiments of the present application, the above annotation information includes a first image; the processing module 12 can further be used to determine a second image in response to the first input received by the receiving module 11. The receiving module 11 can further be used to receive a third input. The processing module 12 can further be used to, in response to the third input received by the receiving module 11, determine the first image according to the second image.
[0197] In some embodiments of the present application, the above-mentioned annotation information includes first annotation information and second annotation information. The processing module 12 can specifically be used to determine the first annotation information in response to the first input received by the receiving module 11. The receiving module 11 can also be used to receive a fourth input for the speech conversion text of the first speech. The processing module 12 can also be used to determine the second annotation information according to the text content corresponding to the fourth input in response to the fourth input received by the receiving module 11; and add the first annotation information and the second annotation information to the speech conversion text of the first speech according to the input time of the first input.
[0198] In some embodiments of the present application, the second annotation information includes a third image. The processing module 12 can specifically be used to determine a fourth image according to the text content corresponding to the fourth input in response to the fourth input received by the receiving module 11. The receiving module 11 can also be used to receive a fifth input. The processing module 12 can also be used to determine the third image according to the fourth image in response to the fifth input received by the receiving module 11.
[0199] In some embodiments of the present application, the first interface further includes a first control. The receiving module 11 can also be used to receive a sixth input for the first control. The processing module 12 can also be used to display the speech conversion text and annotation information of the first speech according to the first display mode in response to the sixth input received by the receiving module 11.
[0200] In some embodiments of the present application, the processing module 12 can specifically be used to display the speech conversion text, annotation information, and speech progress control of the first speech according to the first display mode in response to the sixth input. The receiving module 11 can also be used to receive a seventh input for the speech progress control. The processing module 12 can also be used to display the speech conversion text and annotation information corresponding to the speech position of the seventh input in response to the seventh input received by the receiving module 11.
[0201] In the speech text processing device provided in some embodiments of the present application, since in the case of performing speech conversion processing on the first speech, annotation information can be added to the speech conversion text of the first speech through the first input of the first interface including the speech conversion text of the first speech, annotation can be added synchronously during the speech conversion processing of the speech, realizing multimodal synchronous recording of annotation and speech recording, enabling different recording methods to complement each other, and enabling integration of annotation and speech recording without the user manually organizing and comparing each piece of information recorded through multiple applications. Thus, the efficiency of information recording can be improved, and the integrity of the recorded information can be ensured, avoiding omission of information.
[0202] The voice text processing device in some embodiments of the present application may be an electronic device or a component in an electronic device, such as an integrated circuit or a chip. The electronic device may be a terminal or other devices other than terminals. Exemplarily, the electronic device may be a mobile phone, a tablet computer, a laptop computer, a handheld computer, an in-vehicle electronic device, a Mobile Internet Device (MID), an augmented reality (AR) / virtual reality (VR) device, a robot, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc. It may also be a server, a Network Attached Storage (NAS), a personal computer (PC), a television (TV), a teller machine, or a self-service machine, etc. The embodiments of the present application do not make specific limitations.
[0203] The voice text processing device in some embodiments of the present application may be a device with an operating system. The operating system may be the Android operating system, the IOS operating system, or other possible operating systems. The embodiments of the present application do not make specific limitations.
[0204] The voice text processing device provided in some embodiments of the present application can implement each process implemented in the above method embodiments and achieve the same technical effects. To avoid repetition, details are not described here again.
[0205] As Figure 13 shown, an electronic device 100 is further provided in an embodiment of the present application, including a processor 101 and a memory 102. A program or instruction that can run on the processor 101 is stored on the memory 102. When the program or instruction is executed by the processor 101, each step of the voice text processing method embodiment as described above is implemented, and the same technical effects can be achieved. To avoid repetition, details are not described here again.
[0206] It should be noted that the electronic devices in the embodiments of the present application include mobile electronic devices and non-mobile electronic devices.
[0207] Figure 14 It is a schematic diagram of the hardware structure of an electronic device for implementing some embodiments of the present application.
[0208] As Figure 14As shown, the electronic device 1000 includes, but is not limited to, components such as a radio frequency unit 1001, a network module 1002, an audio output unit 1003, an input unit 1004, a sensor 1005, a display unit 1006, a user input unit 1007, an interface unit 1008, a memory 1009, and a processor 1010.
[0209] Those skilled in the art can understand that the electronic device 1000 may further include a power source (such as a battery) for supplying power to each component. The power source can be logically connected to the processor 1010 through a power management system, so as to implement functions such as management of charging, discharging, and power consumption management through the power management system. Figure 14 The structure of the electronic device shown in the figure does not limit the electronic device. The electronic device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements, which will not be elaborated here.
[0210] Among them, the user input unit 1007 can be used to receive a first input to the first interface when performing voice conversion processing on the first voice; the first interface includes the voice conversion text of the first voice. The processor 1010 can, in response to the first input received by the user input unit 1007, add annotation information to the voice conversion text of the first voice.
[0211] In some embodiments of the present application, the first interface further includes an input area. Exemplarily, the user input unit 1007 can specifically be used to receive a first input to the input area. The processor 1010 can specifically be used to, in response to the first input received by the user input unit 1007, add annotation information to the voice conversion text of the first voice according to the input time of the first input.
[0212] In some embodiments of the present application, the above annotation information includes annotation content and a first annotation identifier; the first interface further includes a voice progress control. Exemplarily, the processor 1010 can further be used to add a second annotation identifier to the voice progress control according to the input time of the first input.
[0213] In some embodiments of the present application, the user input unit 1007 can also be used to receive a second input to the first text in the voice conversion text of the first voice when performing voice conversion processing on the first voice. The processor 1010 can also be used to, in response to the second input received by the user input unit 1007, add marker information to the first text.
[0214] In some embodiments of the present application, the above-mentioned annotation information includes a first image; the processor 1010 may further be configured to determine a second image in response to a first input received by the user input unit 1007. The user input unit 1007 may further be configured to receive a third input. The processor 1010 may further be configured to determine the first image according to the second image in response to the third input received by the user input unit 1007.
[0215] In some embodiments of the present application, the above-mentioned annotation information includes first annotation information and second annotation information. The processor 1010 may specifically be configured to determine the first annotation information in response to a first input received by the user input unit 1007. The user input unit 1007 may further be configured to receive a fourth input of the speech conversion text of the first speech. The processor 1010 may further be configured to determine the second annotation information according to the text content corresponding to the fourth input in response to the fourth input received by the user input unit 1007; and add the first annotation information and the second annotation information to the speech conversion text of the first speech according to the input time of the first input.
[0216] In some embodiments of the present application, the second annotation information includes a third image. The processor 1010 may specifically be configured to determine a fourth image according to the text content corresponding to the fourth input in response to the fourth input received by the user input unit 1007. The user input unit 1007 may further be configured to receive a fifth input. The processor 1010 may further be configured to determine the third image according to the fourth image in response to the fifth input received by the user input unit 1007.
[0217] In some embodiments of the present application, the first interface further includes a first control. The user input unit 1007 may further be configured to receive a sixth input to the first control. The processor 1010 may further be configured to display the speech conversion text of the first speech and the annotation information according to the first display mode in response to the sixth input received by the user input unit 1007.
[0218] In some embodiments of the present application, the processor 1010 may specifically be configured to display the speech conversion text of the first speech, the annotation information, and the speech progress control according to the first display mode in response to the sixth input. The user input unit 1007 may further be configured to receive a seventh input to the speech progress control. The processor 1010 may further be configured to display the speech conversion text and the annotation information corresponding to the speech position of the seventh input in response to the seventh input received by the user input unit 1007.
[0219] In some embodiments of the present application, in the electronic device provided, when performing speech conversion processing on the first speech, annotation information can be added to the speech conversion text of the first speech through the first input of the first interface including the first speech. Therefore, annotations can be added synchronously during the process of speech conversion processing on the speech, realizing multimodal synchronous recording of annotations and speech records, enabling different recording methods to complement each other, and allowing the integration of annotations and speech records without the user manually sorting and comparing each piece of information recorded through multiple applications. Thus, the efficiency of information recording can be improved, and the integrity of the recorded information can be ensured, avoiding information omission.
[0220] It should be understood that in the embodiments of the present application, the input unit 1004 may include a Graphics Processing Unit (GPU) 10041 and a microphone 10042. The graphics processor 10041 processes the image data of static pictures or videos obtained by an image capture device (such as a camera) in a video capture mode or an image capture mode. The display unit 1006 may include a display panel 10061, and the display panel 10061 can be configured in the form of a liquid crystal display, an organic light-emitting diode, etc. The user input unit 1007 includes at least one of a touch panel 10071 and other input devices 10072. The touch panel 10071 is also called a touch screen. The touch panel 10071 may include two parts: a touch detection device and a touch controller. The other input devices 10072 may include, but are not limited to, a physical keyboard, function keys (such as volume control keys, power-on keys, etc.), a trackball, a mouse, a joystick, which will not be elaborated here.
[0221] The memory 1009 can be used to store software programs and various data. The memory 1009 may mainly include a first storage area for storing programs or instructions and a second storage area for storing data. Among them, the first storage area can store an operating system, application programs or instructions required for at least one function (such as a sound playback function, an image playback function, etc.). In addition, the memory 1009 can include volatile memory or non-volatile memory, or the memory 1009 can include both volatile and non-volatile memory. Among them, the non-volatile memory can be a Read-Only Memory (ROM), a Programmable ROM (PROM), an Erasable PROM (EPROM), an Electrically Erasable PROM (EEPROM), or a flash memory. The volatile memory can be a Random Access Memory (RAM), a Static RAM (SRAM), a Dynamic RAM (DRAM), a Synchronous DRAM (SDRAM), a Double Data Rate SDRAM (DDR SDRAM), an Enhanced SDRAM (ESDRAM), a Synch link DRAM (SLDRAM), and a Direct Rambus RAM (DRRAM). The memory 1009 in the embodiments of the present application includes but is not limited to these and any other suitable types of memory.
[0222] The processor 1010 may include one or more processing units; optionally, the processor 1010 integrates an application processor and a modem processor. Among them, the application processor mainly processes operations related to the operating system, user interface, and application programs, etc., and the modem processor mainly processes wireless communication signals, such as a baseband processor. It can be understood that the above modem processor may not be integrated into the processor 1010 either.
[0223] The embodiments of the present application also provide a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, it implements each process of the above embodiment of the voice text processing method and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0224] Among them, the processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media such as computer read-only memory ROM, random access memory RAM, magnetic disks, or optical discs.
[0225] Another embodiment of the present application provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement each process of the above embodiment of the voice text processing method, and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.
[0226] It should be understood that the chip mentioned in the embodiments of the present application may also be referred to as a system-on-chip, system chip, chip system, or system-on-chip.
[0227] The embodiments of the present application provide a computer program / program product. The program / program product is stored in a storage medium and is executed by at least one processor to implement each process of the above embodiment of the voice text processing method, and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.
[0228] It should be noted that in this article, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in a reverse order according to the functions involved. For example, the described methods may be performed in an order different from that described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0229] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described example methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in various embodiments of the present application.
[0230] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific implementation manners. The above specific implementation manners are merely illustrative and not restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them belong to the protection scope of the present application.
Claims
1. A method for processing speech texts, characterized in that, The method includes: When performing speech conversion processing on a first speech, receiving a first input to a first interface; the first interface includes the speech conversion text of the first speech; In response to the first input, adding annotation information to the speech conversion text of the first speech.
2. The method according to claim 1, wherein The first interface further includes an input area; The receiving the first input to the first interface includes: Receiving a first input to the input area; The responding to the first input and adding annotation information to the speech conversion text of the first speech includes: In response to the first input, adding annotation information to the speech conversion text of the first speech according to the input time of the first input.
3. The method according to claim 1 or 2, characterized in that, The annotation information includes annotation content and a first annotation identifier; the first interface further includes a speech progress control; The method further includes: Adding a second annotation identifier to the speech progress control according to the input time of the first input.
4. The method according to claim 1, wherein The method further includes: When performing speech conversion processing on a first speech, receiving a second input to a first text in the speech conversion text of the first speech; In response to the second input, adding marking information to the first text.
5. The method according to claim 2, wherein The annotation information includes a first image; the method further includes: In response to the first input, determining a second image; Receiving a third input; In response to the third input, determining the first image according to the second image.
6. The method according to claim 2, wherein The annotation information includes first annotation information and second annotation information; The responding to the first input and adding annotation information to the speech conversion text of the first speech according to the input time of the first input includes: In response to the first input, determining the first annotation information; Receiving a fourth input to the speech conversion text of the first speech; In response to the fourth input, determining the second annotation information according to the text content corresponding to the fourth input; Adding the first annotation information and the second annotation information to the speech conversion text of the first speech according to the input time of the first input.
7. The method according to claim 6, wherein The second annotation information includes a third image; The responding to the fourth input and determining the second annotation information according to the text content corresponding to the fourth input includes: In response to the fourth input, determining a fourth image according to the text content corresponding to the fourth input; Receiving a fifth input; In response to the fifth input, determining the third image according to the fourth image.
8. The method according to claim 1, characterized in that The first interface further includes a first control; the method further includes: Receiving a sixth input to the first control; In response to the sixth input, displaying the speech conversion text of the first speech and the annotation information according to a first display mode.
9. The method according to claim 8, wherein The responding to the sixth input and displaying the speech conversion text of the first speech and the annotation information according to a first display mode includes: In response to the sixth input, displaying the speech conversion text of the first speech, the annotation information, and the speech progress control according to the first display mode; The method further includes: Receiving a seventh input to the speech progress control; In response to the seventh input, display the speech conversion text corresponding to the speech position of the seventh input and the annotation information.
10. An electronic device, characterized in that, It includes a processor and a memory, and the memory stores programs or instructions that can run on the processor. When the programs or instructions are executed by the processor, the steps of the speech text processing method according to any one of claims 1-9 are implemented.