Text display method and device
By converting voice into text on the server side and extracting tone characteristics, and determining the speaker's logo with the tone mark, the problem of hearing-impaired users being unable to distinguish the speaker in multiple communications is solved, and the accurate acquisition of information is achieved.
Patent Information
- Application Number
- CN202510395738.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-06-17
AI Technical Summary
In multi-person offline meetings or social gatherings, hearing-impaired users are unable to determine who said each sentence, resulting in the missed important information.
By collecting voice, sending voice to the server, and receiving the converted text and tone marks, the speaker logo is determined based on the tone mark, and the text and speaker logo are displayed.
It enables hearing-impaired users to quickly distinguish the speakers, avoid information omissions, and improve the accuracy of information acquisition.
Smart Images

Figure CN120164468A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of electronic devices, and specifically relates to a text display method and device. Background Art
[0002] As more and more hearing-impaired users participate in modern social work, how to improve the sense of participation of hearing-impaired users in multi-person offline meetings or social gatherings has become an urgent problem to be solved.
[0003] In the related art, the main method is to convert the voice in a multi-person offline meeting or social gathering into text in real time and display it for the hearing-impaired users to read by themselves. However, since the hearing-impaired users need to look at the text on the screen and there is no distinction of the speaker on the screen, the hearing-impaired users cannot determine who said each sentence and may need to look around to align who is speaking, resulting in missing some important information. Summary of the Invention
[0004] The purpose of the embodiments of this application is to provide a text display method and device that can avoid information omission.
[0005] In a first aspect, the embodiments of this application provide a text display method, which includes: collecting a first voice and sending the first voice to a server; receiving a first text and a first tone color identifier from the server, where the first text is the text obtained by converting the first voice, and the first tone color identifier corresponds to first tone color feature information, and the first tone color feature information is used to characterize the tone color of the first voice; determining a first speaker identifier based on the first tone color identifier, where the first speaker indicated by the first speaker identifier is the speaker who issued the first voice; displaying the first text and the first speaker identifier.
[0006] In a second aspect, the embodiments of this application provide a text display method, which includes: receiving a first voice from an electronic device; converting the first voice into a first text and extracting first tone color feature information of the first voice, where the first tone color feature information is used to characterize the tone color of the first voice; determining a first tone color identifier based on the first tone color feature information; sending the first text and the first tone color identifier to the electronic device.
[0007] In a third aspect, an embodiment of the present application provides a text display device, which includes an acquisition module, a sending module, a receiving module, a processing module, and a display module. The acquisition module is configured to acquire a first voice. The sending module is configured to send the first voice to a server. The receiving module is configured to receive a first text and a first voice color identifier from the server, where the first text is the text obtained by converting the first voice, and the first voice color identifier corresponds to first voice color feature information, and the first voice color feature information is used to characterize the voice color of the first voice. The processing module is configured to determine a first speaker identifier based on the first voice color identifier, and the first speaker indicated by the first speaker identifier is the speaker who emits the first voice. The display module is configured to display the first text and the first speaker identifier.
[0008] In a fourth aspect, an embodiment of the present application provides a text display device, which includes a receiving module, a processing module, and a sending module. The receiving module is configured to receive a first voice from an electronic device. The processing module is configured to convert the first voice into a first text, and extract first voice color feature information of the first voice, where the first voice color feature information is used to characterize the voice color of the first voice; and determine a first voice color identifier based on the first voice color feature information. The sending module is configured to send the first text and the first voice color identifier to the electronic device.
[0009] In a fifth aspect, an embodiment of the present application provides an electronic device, which includes a processor and a memory. The memory stores a program or instruction that can run on the processor, and when the program or instruction is executed by the processor, the steps of the text display method described in the first aspect are implemented.
[0010] In a sixth aspect, an embodiment of the present application provides a server, which includes a processor and a memory. The memory stores a program or instruction that can run on the processor, and when the program or instruction is executed by the processor, the steps of the text display method described in the second aspect are implemented.
[0011] In a seventh aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored, and when the program or instruction is executed by a processor, the steps of the text display method described in the first aspect or the second aspect are implemented.
[0012] In an eighth aspect, an embodiment of the present application provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor, and the processor is configured to run a program or instruction to implement the text display method described in the first aspect or the second aspect.
[0013] In a ninth aspect, an embodiment of the present application provides a computer program product, which is stored in a storage medium and executed by at least one processor to implement the text display method as described in the first aspect or the second aspect.
[0014] In an embodiment of the present application, a first voice is collected and sent to a server; a first text and a first voice color identifier are received from the server. The first text is the text obtained by converting the first voice, and the first voice color identifier corresponds to first voice color feature information, which is used to characterize the voice color of the first voice; based on the first voice color identifier, a first speaker identifier is determined, and the first speaker indicated by the first speaker identifier is the speaker who issued the first voice; then the first text and the first speaker identifier are displayed. In this way, by displaying the speaker identifier of the speaker who issued the voice while displaying the text after voice conversion, the hearing-impaired user can quickly distinguish who the speaker is when seeing the text, without the hearing-impaired user having to look up to judge who is speaking, so that important information in the displayed text will not be missed. Therefore, this solution can avoid information omission and improve the accuracy of information acquisition. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 is one of the flow diagrams of the text display method provided by an embodiment of the present application;
[0016] Figure 2 is a schematic diagram of text and speaker identifier provided by an embodiment of the present application;
[0017] Figure 3 is a schematic diagram of starting the speaker discrimination function in the speech recognition interface provided by an embodiment of the present application;
[0018] Figure 4 is one of the schematic diagrams of the speech recognition interface provided by an embodiment of the present application;
[0019] Figure 5 is another schematic diagram of the speech recognition interface provided by an embodiment of the present application;
[0020] Figure 6 is another flow diagram of the text display method provided by an embodiment of the present application;
[0021] Figure 7 is another schematic diagram of the speech recognition interface provided by an embodiment of the present application;
[0022] Figure 8 is one of the schematic diagrams of the speech recognition setting interface provided by an embodiment of the present application;
[0023] Figure 9 is a schematic diagram of the sound setting interface provided by an embodiment of the present application;
[0024] Figure 10 It is the second schematic diagram of the voice recognition setting interface provided by the embodiments of the present application;
[0025] Figure 11 It is the schematic diagram of the window in the voice recognition setting interface provided by the embodiments of the present application;
[0026] Figure 12 It is the third schematic diagram of the process of the text display method provided by the embodiments of the present application;
[0027] Figure 13 It is the fourth schematic diagram of the voice recognition interface provided by the embodiments of the present application;
[0028] Figure 14 It is the first schematic diagram of the window in the voice recognition interface provided by the embodiments of the present application;
[0029] Figure 15 It is the fifth schematic diagram of the voice recognition interface provided by the embodiments of the present application;
[0030] Figure 16 It is the second schematic diagram of the window in the voice recognition interface provided by the embodiments of the present application;
[0031] Figure 17 It is the sixth schematic diagram of the voice recognition interface provided by the embodiments of the present application;
[0032] Figure 18 It is the seventh schematic diagram of the voice recognition interface provided by the embodiments of the present application;
[0033] Figure 19 It is the eighth schematic diagram of the voice recognition interface provided by the embodiments of the present application;
[0034] Figure 20 It is the schematic diagram of the atomic note provided by the embodiments of the present application;
[0035] Figure 21 It is the fourth schematic diagram of the process of the text display method provided by the embodiments of the present application;
[0036] Figure 22 It is the fifth schematic diagram of the process of the text display method provided by the embodiments of the present application;
[0037] Figure 23 It is the sixth schematic diagram of the process of the text display method provided by the embodiments of the present application;
[0038] Figure 24 It is the seventh schematic diagram of the process of the text display method provided by the embodiments of the present application;
[0039] Figure 25 It is the eighth schematic diagram of the process of the text display method provided by the embodiments of the present application;
[0040] Figure 26 is one of the schematic diagrams of the text display device provided by the embodiments of the present application;
[0041] Figure 27 is the second of the schematic diagrams of the text display device provided by the embodiments of the present application;
[0042] Figure 28 is the third of the schematic diagrams of the text display device provided by the embodiments of the present application;
[0043] Figure 29 is the fourth of the schematic diagrams of the text display device provided by the embodiments of the present application;
[0044] Figure 30 is the structural schematic diagram of the electronic device provided by the embodiments of the present application;
[0045] Figure 31 is the hardware structural schematic diagram of the electronic device provided by the embodiments of the present application;
[0046] Figure 32 is the hardware structural schematic diagram of the server provided by the embodiments of the present application. Detailed implementation manners
[0047] Next, the technical solutions in the embodiments of the present application will be clearly described with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application belong to the scope of protection of the present application.
[0048] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such used data may be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order different from those illustrated or described herein, and the objects distinguished by "first", "second", etc. generally belong to the same category, and do not limit the number of objects. For example, the first object may be one or multiple. In addition, "and / or" in the specification and claims means at least one of the connected objects, and the character " / " generally means an "or" relationship between the associated objects before and after.
[0049] The terms "at least one (item)", "at least one of", etc. in the description and claims of this application refer to any one, any two or a combination of two or more of the objects it contains. For example, at least one (item) of a, b, and c can represent: "a", "b", "c", "a and b", "a and c", "b and c", and "a, b, and c", where a, b, and c can be single or multiple. Similarly, "at least two (items)" means two or more, and its meaning is similar to that of "at least one (item)".
[0050] The identifiers in this application are words, symbols, images, etc. used to indicate information, and can use controls or other containers as the carriers for displaying information, including but not limited to text identifiers, symbol identifiers, and image identifiers.
[0051] The following will combine the accompanying drawings and, through specific embodiments and their application scenarios, elaborate in detail on the text display method, device, equipment, storage medium, and program product provided by the embodiments of this application.
[0052] The text display method provided by the embodiments of this application can be applied to scenarios where hearing-impaired users communicate with people offline. Specifically, it can include scenarios such as hearing-impaired users participating in multi-person offline meetings, hearing-impaired users participating in social gatherings, and hearing-impaired users communicating with family members in daily life.
[0053] Taking the scenario where a hearing-impaired user participates in a multi-person offline meeting as an example, when a hearing-impaired user has an offline meeting with multiple people, the hearing-impaired user uses the voice recognition software on the mobile phone to record the speeches of multiple people in the meeting. The voice recognition software can convert the recorded speech content into text and display it on the mobile phone screen. Then, the hearing-impaired user participates in the meeting by looking at the text displayed on the mobile phone screen.
[0054] Taking the scenario where a hearing-impaired user communicates with multiple people in daily life as an example, when a hearing-impaired user has a meal with family members, the hearing-impaired user uses the voice recognition software on the mobile phone to record the voices of family members speaking. The voice recognition software can convert the recorded voices of family members speaking into text and display it on the mobile phone screen. Then, the hearing-impaired user chats with family members by looking at the text displayed on the mobile phone screen.
[0055] However, since the attention of hearing-impaired users is all on the mobile phone, they don't have time to look up to see who is speaking when multiple people are talking. Therefore, although hearing-impaired users can see the text content of others' speeches, they can't tell who is speaking, and can't distinguish who said each paragraph of text. Hearing-impaired users still need to look around to align with who is speaking, which makes it easy to miss some information displayed on the mobile phone screen.
[0056] To this end, the embodiments of the present application provide a text display method, apparatus, device, storage medium, and program product, which collect a first voice and send the first voice to a server; receive a first text and a first voice color identifier from the server, where the first text is the text obtained by converting the first voice, and the first voice color identifier corresponds to first voice color feature information, and the first voice color feature information is used to characterize the voice color of the first voice; determine a first speaker identifier based on the first voice color identifier, where the first speaker indicated by the first speaker identifier is the speaker who emits the first voice; and then display the first text and the first speaker identifier. In this way, by displaying the speaker identifier of the speaker who emits the voice while displaying the text after voice conversion, the hearing-impaired user can quickly distinguish who the speaker is when seeing the text, without the need for the hearing-impaired user to look up to judge who is speaking, so that important information in the displayed text will not be missed. Therefore, this solution can avoid information omission and improve the accuracy of information acquisition.
[0057] The execution subject of the text display method provided by the embodiments of the present application may be a text display apparatus. Exemplarily, the text display apparatus may be an electronic device, or a functional component or functional entity in the electronic device. Hereinafter, taking the execution subject as an electronic device as an example, the text display method provided by the embodiments of the present application will be described by way of example.
[0058] Figure 1 is a schematic flowchart of the text display method provided by the embodiments of the present application, as Figure 1 shown, the text display method provided by the embodiments of the present application may include the following steps 101 to 104.
[0059] Step 101: The electronic device collects a first voice and sends the first voice to the server.
[0060] In some embodiments of the present application, the above-mentioned first voice may be any voice collected by the electronic device during the process of collecting voices.
[0061] In some embodiments of the present application, a hearing assistance program may be installed in the electronic device, and the hearing assistance program may assist hearing-impaired users in communicating with others. Among them, the hearing assistance program may be a hearing assistance application program (APP), a hearing assistance component, or a hearing assistance applet.
[0062] In some embodiments of the present application, the above-mentioned server may be a server that establishes a communication connection with the electronic device. Exemplarily, the above-mentioned server may be a server that provides background services for the hearing assistance program.
[0063] Exemplarily, taking the electronic device as a mobile phone as an example, in a scenario where the user communicates with multiple people, the user can start the listening and speaking assistance program installed in the mobile phone, collect the voices of others through the listening and speaking assistance program, and send the collected voices to the server.
[0064] Step 102, the electronic device receives the first text and the first tone color identifier from the server.
[0065] In some embodiments of the present application, the above-mentioned first text may be the text obtained by converting the first voice.
[0066] In some embodiments of the present application, the above-mentioned first tone color identifier corresponds to the first tone color feature information, and the first tone color feature information is used to characterize the tone color of the first voice. In other words, the first tone color feature information can be used to characterize the tone color or voice characteristics of the speaker who emits the first voice.
[0067] Exemplarily, the above-mentioned first tone color identifier may be a text identifier, a digital identifier, etc. Among them, the text identifier may be tone color 1, tone color 2, etc. The digital identifier may be 0001, 0002, 0003, etc.
[0068] In some embodiments of the present application, the above-mentioned first tone color feature information includes but is not limited to Mel-frequency cepstral coefficients (MFCC), Linear prediction coefficients (LPCC), fundamental frequency, formant, spectral envelope, etc.
[0069] In some embodiments of the present application, after the server receives the first voice, it can convert the first voice into the first text, extract the first tone color feature information of the first voice, then determine the first tone color identifier based on the first tone color feature information, and then send the first text and the first tone color identifier to the electronic device.
[0070] It should be noted that for the specific implementation of the server converting the first voice into the first text, extracting the first tone color feature information, and determining the first tone color identifier, reference can be made to the relevant descriptions in the following embodiments.
[0071] In some embodiments of the present application, the first text and the first tone color identifier received by the electronic device may be represented in an associated form. Exemplarily, the first text and the first tone color identifier are represented in the form of "first text - first tone color identifier". For example, if a certain user says "What to eat at noon today", the first text and the first tone color identifier received by the electronic device may be represented in the form of "What to eat at noon today - 0001".
[0072] Step 103: The electronic device determines a first speaker identifier based on the first tone identifier.
[0073] In some embodiments of the present application, the first speaker indicated by the above first speaker identifier is the speaker who emits the first voice. Exemplarily, the above first speaker identifier may be a text identifier, an image identifier, a symbol identifier, etc. Among them, the text identifier may be the name, nickname, full name, etc. of the speaker. The image identifier may be the avatar, photo, etc. of the speaker.
[0074] In some embodiments of the present application, the above first speaker identifier corresponds to the first tone identifier.
[0075] In some embodiments of the present application, the electronic device may randomly generate a first speaker identifier, or may look up the first speaker identifier corresponding to the first tone identifier from the pre-stored correspondence between tone identifiers and speaker identifiers.
[0076] In some embodiments of the present application, the above step 103 may be specifically implemented by the following steps 1031 to 1033.
[0077] Step 1031: The electronic device obtains the correspondence between tone identifiers and speaker identifiers.
[0078] In some embodiments of the present application, the above correspondence between tone identifiers and speaker identifiers may be a set of correspondences or multiple sets of correspondences. Among them, each set of correspondences includes a tone identifier and a speaker identifier.
[0079] In some embodiments of the present application, the correspondence between tone identifiers and speaker identifiers may be pre-stored in the electronic device. Alternatively, the correspondence between tone identifiers and speaker identifiers may be pre-stored in the server, and the electronic device may obtain the correspondence between tone identifiers and speaker identifiers from the server.
[0080] It should be noted that the above correspondence between tone identifiers and speaker identifiers may be pre-established and stored in the electronic device, and the establishment process may refer to the relevant descriptions of the following embodiments.
[0081] Step 1032: When there is a first correspondence in the correspondence between tone identifiers and speaker identifiers, the electronic device determines the speaker identifier in the first correspondence as the first speaker identifier.
[0082] In some embodiments of the present application, the above first correspondence includes a first tone identifier and the speaker identifier corresponding to the first tone identifier.
[0083] In some embodiments of the present application, the electronic device may search in the correspondence relationship between the tone identifier and the speaker identifier to find whether there is a first correspondence relationship including the first tone identifier. If so, it indicates that the electronic device has previously stored the correspondence relationship between the first tone identifier and a certain speaker identifier. Then, the electronic device may determine the speaker identifier in the found first correspondence relationship, that is, the speaker identifier corresponding to the first tone identifier, as the first speaker identifier.
[0084] Exemplarily, assume that the first tone identifier is Tone 1, and the electronic device has previously stored the correspondence relationship between Tone 1 and "Teacher Zhang". Then, the electronic device may determine that the first speaker identifier is "Teacher Zhang".
[0085] Step 1033: In the case where there is no first correspondence relationship in the correspondence relationship between the tone identifier and the speaker identifier, the electronic device generates a first speaker identifier.
[0086] In some embodiments of the present application, the above first correspondence relationship includes the first tone identifier and the speaker identifier corresponding to the first tone identifier.
[0087] In some embodiments of the present application, the electronic device may search in the correspondence relationship between the tone identifier and the speaker identifier to find whether there is a first correspondence relationship including the first tone identifier. If not, it indicates that the electronic device has not previously stored the correspondence relationship between the first tone identifier and a certain speaker identifier. Then, the electronic device may generate a first speaker identifier.
[0088] In some embodiments of the present application, when generating the first speaker identifier, the electronic device also needs to ensure that the generated first speaker identifier does not repeat with other speaker identifiers currently displayed on the electronic device, and also needs to ensure that the generated first speaker identifier does not repeat with the speaker identifiers in the correspondence relationship between the tone identifier and the speaker identifier.
[0089] For example, assume that the correspondence relationship between the tone identifier and the speaker identifier includes Zhang and Li, and the speaker identifiers displayed on the electronic device are Wang. Then, the electronic device cannot use Zhang, Li, and Wang, and can use Speaker 1 or Speaker 2, etc. as the first speaker identifier.
[0090] It should be noted that the above Step 1032 and Step 1033 are executed alternatively.
[0091] In this way, regardless of whether the electronic device has previously stored the correspondence relationship between the tone identifier and the speaker representation, the electronic device can quickly and accurately determine the first speaker identifier of the first speaker who emits the first voice.
[0092] In some embodiments of the present application, after the above Step 1033, the text display method provided by the embodiments of the present application may further include the following Step 1034.
[0093] Step 1034: The electronic device establishes and stores the correspondence between the first tone identifier and the first speaker identifier.
[0094] Exemplarily, assume that the first tone identifier is Tone 1 and the first speaker identifier is "Teacher Zhang". The electronic device can establish and store the correspondence between "Tone 1" and "Teacher Zhang".
[0095] In this way, when the first speaker speaks next time, after receiving the tone identifier from the server, the electronic device can quickly identify that the speaker is the first speaker based on the established correspondence between the first tone identifier and the first speaker identifier, improving the efficiency of determining the speaker identifier.
[0096] Step 104: The electronic device displays the first text and the first speaker identifier.
[0097] In some embodiments of the present application, the electronic device can display the first speaker identifier before the first text, so that when the user sees the screen, they can quickly determine that the first text is spoken by the first speaker indicated by the first speaker identifier.
[0098] In some embodiments of the present application, the electronic device can display the first text and the first speaker identifier on the speech recognition interface of the hearing assistance program, and display the first speaker identifier before the first text. The speech recognition interface is an interactive interface in the hearing assistance program, which is used to collect speech and display the text obtained by converting the speech.
[0099] Exemplarily, Figure 2 is a schematic diagram of the text and the speaker identifier provided by the embodiments of the present application. As Figure 2 shown, in the scenario where the user communicates with a friend in daily life, the electronic device displays text 21 and speaker 22 on the speech recognition interface 2. The content of text 21 is "Let's go to the movies this afternoon", and speaker 22 is "Friend A". In this way, the user can quickly confirm that Friend A is inviting themselves to watch a movie.
[0100] The text display method provided by the embodiments of the present application collects the first voice and sends the first voice to the server; receives the first text and the first voice color identifier from the server, where the first text is the text obtained by converting the first voice, and the first voice color identifier corresponds to the first voice feature information, and the first voice feature information is used to characterize the voice color of the first voice; determines the first speaker identifier based on the first voice color identifier, and the first speaker indicated by the first speaker identifier is the speaker who emits the first voice; then displays the first text and the first speaker identifier. In this way, by displaying the speaker identifier of the speaker who emits the voice while displaying the text after voice conversion, the hearing-impaired user can quickly distinguish who the speaker is when seeing the text, without the hearing-impaired user having to look up to judge who is speaking, so that important information in the displayed text will not be missed. Therefore, this solution can avoid information omission and improve the accuracy of information acquisition.
[0101] In some embodiments of the present application, when the user uses the hearing assistance program to collect voice, if the user selects to enable the function of distinguishing speakers, after the electronic device sends the collected voice to the server, the server sends the text after voice conversion and the voice color identifier to the electronic device, and then the electronic device determines the speaker identifier based on the voice color identifier and displays the speaker identifier and the text on the voice recognition interface to distinguish which person each text is spoken by.
[0102] Exemplarily, the above voice recognition interface may include a multi-functional control. After the user triggers the multi-functional control, the electronic device can display a speaker distinguishing control on the voice recognition interface, and the user can turn on the speaker distinguishing function by triggering the speaker distinguishing control.
[0103] For example, as Figure 3 shown, the voice recognition interface 3 includes text 31, text 32, text 33 and a multi-functional control 34. After the user clicks the multi-functional control 34, the electronic device displays a speaker distinguishing control 35, a clear conversation record control 36 and a small window control 37 on the voice recognition interface 3. When the user clicks the speaker distinguishing control 35, the electronic device can correspondingly display each text and the speaker identifier of each text. As Figure 4 shown, text 31 corresponds to speaker 1, text 32 corresponds to speaker 2, and text 33 corresponds to speaker 3. And, as Figure 4 shown in the voice recognition interface 3, the time points when the speakers indicated by the speaker identifiers say each sentence can also be displayed, such as text 31 spoken by speaker 1 at 12:00:05, text 32 spoken by speaker 2 at 12:00:06, and text 33 spoken by speaker 3 at 12:00:06.
[0104] Further, after the user selects to enable the speaker separation function, the electronic device can add speaker identifiers to the previously displayed text, and then the electronic device continues to collect speech. At this time, the collected speech carries the speaker identifier when it is converted into text for display.
[0105] For example, in combination with Figure 4 , as Figure 5 shown, after the user enables the speaker separation function, the electronic device displays the speaker identifier and text in the speech recognition interface 3. When new text 38 of speaker 3 is collected, the electronic device can directly display speaker 3 and new text 38. For example, new text 38 is "By the way, Mr. Li, regarding the reception of the guests, do we need to prepare a special reception plan?".
[0106] In this way, it is determined whether to carry the speaker identifier when displaying the text according to the user's needs, improving the flexibility of text display.
[0107] In some embodiments of the present application, in combination with Figure 1 , as Figure 6 shown, before the above-mentioned step 101, the text display method provided by the embodiments of the present application may further include the following steps 10 to 12, and specifically, the above-mentioned step 104 may be implemented by the following step 1041.
[0108] Step 10: The electronic device displays a speech recognition interface, and the speech recognition interface includes a collection control.
[0109] In some embodiments of the present application, the above-mentioned speech recognition interface may be an interface for performing speech recognition and text display, and the speech recognition interface may be an interactive interface in the hearing assistance program installed in the electronic device. The above-mentioned collection control can be understood as a control for triggering the electronic device to collect speech.
[0110] Step 11: The electronic device receives a third input to the collection control.
[0111] In some embodiments of the present application, the above-mentioned third input may be any form of input by the user to the collection control through the touch area of the electronic device. For example, the above-mentioned touch area may be the screen of the electronic device.
[0112] In some embodiments of the present application, the above-mentioned third input includes but is not limited to: the user uses a touch device such as a finger or a stylus to perform a touch input on the collection control through the touch area of the electronic device, or is the user's voice input, or is a specific gesture input by the user, or is other feasible inputs. Specifically, it can be determined according to actual usage requirements, and the embodiments of the present application do not make limitations.
[0113] Exemplarily, the above touch input may be a single - click input, a double - click input, or a click input of any number of times, etc., or may also be a long - press input, a short - press input, a swipe input, a drag input, etc.
[0114] Exemplarily, the above - mentioned specific gesture may be any one of a single - click gesture, a swipe gesture, a drag gesture, a pressure - recognition gesture, a long - press gesture, an area - change gesture, a double - press gesture, and a double - click gesture.
[0115] Step 12: The electronic device starts to collect voice in response to the third input.
[0116] In some embodiments of the present application, when the user communicates with others offline, the listening and speaking assistance program installed in the electronic device can be started, and the voice recognition interface of the listening and speaking assistance program can be entered. By performing a third input on the collection control of the voice recognition interface, the electronic device is triggered to start collecting voice through the microphone or the voice collection module.
[0117] Exemplarily, taking the electronic device as a mobile phone and the listening and speaking assistance program as the listening and speaking assistance APP as an example, when a hearing - impaired user communicates with others offline, the listening and speaking assistance APP installed in the mobile phone can be opened, and the voice recognition interface of the listening and speaking assistance APP can be entered. The user clicks the collection control of the voice recognition interface, triggering the mobile phone to start collecting voice through the microphone.
[0118] Step 1041: The electronic device displays the first text and the first speaker identifier in the voice recognition interface.
[0119] In some embodiments of the present application, the electronic device starts to collect voice by triggering the collection control of the voice recognition interface, sends the collected voice to the server, then receives the text and the speaker identifier returned by the server, and then displays the received text and the speaker identifier in the voice recognition interface.
[0120] Exemplarily, Figure 7 is a schematic diagram of a voice recognition interface provided by an embodiment of the present application. As Figure 7 shown, in the scenario where the user participates in a multi - person offline meeting, the first text 41 and the first speaker identifier 42 are displayed in the voice recognition interface 4 of the electronic device. The first text 41 is "Hello everyone. Today we will have a short meeting to mainly discuss the preparations for the upcoming product launch. Manager Zhang, you can talk about the current preparations first.", and the first speaker identifier 42 is "Speaker 1".
[0121] In this way, through the user's input to the acquisition control of the speech recognition interface, the electronic device is triggered to start collecting speech, and the text obtained by converting the speech and the speaker identification of the speaker who uttered the speech are displayed on the speech recognition interface, enabling the user to quickly distinguish who said each sentence without having to look up, avoiding information omission, and improving the user experience.
[0122] In some embodiments of the present application, before step 101 above, the text display method provided by the embodiments of the present application may further include the following steps 13 to 18.
[0123] Step 13: When the speech recognition setting interface is displayed, the electronic device receives a first input to the speech recognition setting interface.
[0124] In some embodiments of the present application, the above speech recognition setting interface may be a setting interface in the hearing assistance program installed on the electronic device, and may be used to set the correspondence between the timbre feature information and the speaker identification.
[0125] In some embodiments of the present application, the above first input may be any form of input by the user to the speech recognition setting interface through the touch area of the electronic device. For example, the above touch area may be the screen of the electronic device.
[0126] In some embodiments of the present application, the above first input includes but is not limited to: the user performs a touch input on the speech recognition setting interface through the touch area of the electronic device by a touch device such as a finger or a stylus, or is the user's voice input, or is a specific gesture input by the user, or is other feasible input. Specifically, it can be determined according to the actual usage requirements, and the embodiments of the present application do not make limitations.
[0127] In some embodiments of the present application, the above speech recognition setting interface may include a speech acquisition control, and the above first input may be any form of input by the user to the speech acquisition control in the speech recognition setting interface.
[0128] Exemplarily, Figure 8 is a schematic diagram of a speech recognition setting interface provided by an embodiment of the present application. As Figure 8 shown, the speech recognition setting interface 5 includes a speech acquisition control 51. The information "long press to record" is displayed below the speech acquisition control 51. After the user long presses the speech acquisition control 51, the electronic device can start collecting speech. That is, the above first input may be the long press input by the user to the speech acquisition control 51.
[0129] In some embodiments of the present application, before displaying the voice recognition setting interface, the electronic device may display a sound setting interface, which includes a sound setting control. The user makes a fifth input to the sound setting control, and in response to the fifth input, the electronic device displays the above-mentioned voice recognition setting interface.
[0130] Exemplarily, Figure 9 is a schematic diagram of a sound setting interface provided by an embodiment of the present application. As Figure 9 shown, the sound setting interface 6 includes a sound setting control 61. The sound setting control 61 displays the words "Start Recording". The sound setting interface 6 also displays the precautions when setting the sound, such as "The recording process is about 15 seconds. Select a quiet recording environment: You can record in a room or in a car; Keep a distance of 20 cm: Avoid unclear recording due to being too far away from the mobile phone; Read in Mandarin: Relax as naturally as when talking to a friend." Then, the user clicks on the sound setting control 61, and the electronic device can display as Figure 8 shown in the voice recognition setting interface 5.
[0131] Step 14: The electronic device collects a second voice in response to the first input.
[0132] In some embodiments of the present application, after the user makes a first input to the voice recognition setting interface, the user can start speaking. The electronic device can start collecting the second voice of the user's speech and send the collected second voice to the server.
[0133] In some embodiments of the present application, the above-mentioned voice recognition setting interface may further include a recording text, and the user can directly record the voice according to the recording text.
[0134] Exemplarily, as Figure 8 shown, the voice recognition setting interface 5 also displays a recording text 52, such as "Put down the mobile phone and communicate face to face with family and friends. True happiness lies in these simple interactions." Further, the voice recognition setting interface 5 also displays a prompt message 53, such as "After long pressing the button below, you can read the above text aloud for recording." The prompt message 53 is used to prompt the user on how to record the voice.
[0135] In some embodiments of the present application, the above-mentioned voice recognition setting interface may further include a preview control, so that the user can hear the voice they recorded and determine whether to re-record. In this way, the quality of the recorded voice can be improved, the accuracy of the extracted timbre feature information can be improved, and further the accuracy of setting the speaker identification subsequently can be improved.
[0136] Exemplarily, as Figure 8As shown, a trial listening control 54 is also displayed in the speech recognition setting interface 5. The user can click on the trial listening control 54 to listen to the voice they recorded and determine whether to re-record.
[0137] In some embodiments of the present application, after the electronic device collects the second voice, it can analyze the second voice to determine the words with non-standard or unclear pronunciation in the second voice, mark the words with non-standard or unclear pronunciation in the speech recognition setting interface, and update the information in the collection control in the speech recognition setting interface to "long press to re-record" to prompt the user that they need to re-record the voice. Then, the user can perform a first input on the collection control again, and the electronic device responds to the first input and starts collecting the voice.
[0138] Exemplarily, Figure 10 is a schematic diagram of a speech recognition setting interface provided by an embodiment of the present application. As Figure 10 shown, a recorded text 52 is displayed in the speech recognition setting interface 5, and the words with non-standard or unclear pronunciation in the recorded text 52 are underlined. The information displayed in the collection control 51 in the speech recognition setting interface 5 is updated to "long press to re-record". After the user long presses the collection control 51, the electronic device can start collecting the voice.
[0139] Step 15: When the collection of the second voice ends, the electronic device displays a first window.
[0140] In some embodiments of the present application, the electronic device can determine that the collection of the second voice ends when no voice is collected for a long time or the first input of the user on the speech recognition setting interface ends, and then display the first window.
[0141] In some embodiments of the present application, the electronic device can superimpose and display the first window on the speech recognition setting interface, or the electronic device can display the first window in a new interface.
[0142] Exemplarily, the above-mentioned first window can be a floating window, a floating ball, a floating frame, etc., and the embodiments of the present application do not limit this.
[0143] Step 16: The electronic device receives the second speaker identifier input by the user in the first window.
[0144] In some embodiments of the present application, the second speaker indicated by the above-mentioned second speaker identifier is the speaker who emitted the second voice.
[0145] Exemplarily, the above-mentioned second speaker identifier can be a text identifier, an image identifier, a symbol identifier, etc. Among them, the text identifier can be the name, nickname, full name, etc. of the speaker. The image identifier can be the avatar, photo, etc. of the speaker.
[0146] In some embodiments of the present application, the above-mentioned first window may include an identification input area for a user to input a second speaker identification of a second speaker who emits a second voice.
[0147] Exemplarily, Figure 11 is a schematic diagram of a window in a voice recognition setting interface provided by an embodiment of the present application. As Figure 11 shown, a first window 55 is displayed in the voice recognition setting interface 5. The first window 55 includes an identification input area 551. The user inputs a second speaker identification, such as "Mom", in the identification input area 551. And the voice recognition interface 5 includes a voice collection control 51, and the information displayed below the voice collection control 51 is updated to "Recording completed".
[0148] Step 17: The electronic device sends the second voice to the server.
[0149] In some embodiments of the present application, the first window displayed on the electronic device further includes a determination control. The electronic device receives a trigger input from the user for the determination control. In response to the trigger input, the electronic device sends the second voice to the server, so that the server can extract second timbre feature information of the second voice, generate a second timbre identification corresponding to the second timbre feature information, and then establish and store a correspondence between the second timbre feature information and the second timbre identification.
[0150] Exemplarily, as Figure 11 shown, a determination control 552 and a cancellation control 553 are displayed in the first window 55. After the user inputs a second speaker identification in the identification input area 551, the user can click the determination control 552. Then, the electronic device sends the second voice to the server in response to the click input on the determination control 552.
[0151] It should be noted that the specific implementation of the above-mentioned server extracting the second timbre feature information of the second voice, generating a second timbre identification corresponding to the second timbre feature information, and establishing and storing a correspondence between the second timbre feature information and the second timbre identification can refer to the relevant descriptions in the following embodiments. In addition, the above-mentioned step 17 can be executed when the collection of the second voice ends, or can be executed after the above-mentioned step 15 or step 16. The embodiments of the present application do not make any limitations in this regard.
[0152] Step 18: The electronic device receives the second timbre identification from the server, and establishes and stores a correspondence between the second timbre identification and the second speaker identification.
[0153] In some embodiments of the present application, the above-mentioned second tone identifier corresponds to second tone feature information, and the second tone feature information is used to characterize the tone of the second speech. In other words, the second tone feature information can be used to characterize the tone or voice characteristics of the speaker who emits the second speech.
[0154] In some embodiments of the present application, the above-mentioned second speaker identifier is the speaker identifier of the speaker who emits the second speech, and the second tone identifier is the tone identifier corresponding to the second speech. Therefore, the electronic device can establish and store the correspondence between the second tone identifier and the second speaker identifier.
[0155] Exemplarily, assuming that the second tone identifier is Tone 2 and the second speaker identifier is "Mom", the electronic device can establish and store the correspondence between Tone 2 and "Mom".
[0156] In this way, by pre-recording the speaker's voice in the electronic device and recording the correspondence between the tone identifier of the speaker and the speaker identifier, when the voice is collected subsequently, the speaker identifier of the speaker who emits the voice can be quickly found from the stored correspondence, improving the efficiency of determining the speaker identifier.
[0157] It should be noted that the mobile phone can pre-collect multiple voices to establish multiple groups of correspondences between tone identifiers and speaker identifiers. The embodiments of the present application are only described by taking the establishment and storage of the correspondence between the second tone identifier and the second speaker identifier as an example, and do not constitute a limitation to the present application.
[0158] Next, taking the example of pre-recording Mom's voice in the hearing-impaired assistance APP installed on the mobile phone, the implementation process of the above steps 13 to 18 will be described in combination with Figures 8 to 11 the above.
[0159] Exemplarily, the user opens the hearing-impaired assistance APP on the mobile phone and clicks to add a voice in the settings interface of the hearing-impaired assistance APP. The mobile phone displays the voice settings interface 6 as shown in Figure 9 Then the user clicks the voice setting control 61 in the voice settings interface 6, and the mobile phone can display the speech recognition settings interface 5 as shown in Figure 8 The user long-presses the voice collection control 51 in the speech recognition settings interface 5. Mom speaks according to the recorded text 52 displayed in the speech recognition interface 5. The mobile phone collects the voice of Mom speaking. After the user cancels the long-press on the collection control 51, the voice collection ends, and the collected voice is analyzed. If the pronunciation of the collected voice is standard and clear, the mobile phone can extract the tone feature information of the collected voice, and the mobile phone can display as shown in Figure 11The voice recognition setting interface 5 as shown, and a first window 55 is displayed in the voice recognition setting interface 5, and an identification input area 551, an OK control 552 and a cancel control are displayed in the first window 55. The user inputs a speaker identification, such as "Mom", in the identification input area 551 and clicks the OK control 552. Then, the mobile phone sends the collected voice of "Mom" to the server, and the server extracts the timbre feature information of the voice of "Mom" to generate a timbre identification such as "Timbre 1", and establishes and stores the corresponding relationship between the timbre feature information and the timbre identification. Then the mobile phone receives the timbre identification from the server and establishes and stores the corresponding relationship between the timbre identification such as "Timbre 1" and the speaker identification "Mom". If the collected voice has the situation of unclear or inaccurate pronunciation, the mobile phone can display as Figure 10 The voice recognition setting interface 5 as shown, and the unclear or inaccurate words are underlined and marked in the recorded text 52 of the voice recognition setting interface 5, and the information displayed in the collection control 51 is updated to "Long press to re-record". The user long presses the collection control 51, and Mom speaks again according to the recorded text 52 displayed in the voice recognition interface 5. The mobile phone re-records Mom's voice, and then ends the voice collection after the user cancels the long press on the collection control 51 until a voice with standard and clear pronunciation is collected.
[0160] In some embodiments of the present application, in combination with Figure 1 , such as Figure 12 As shown, after the above step 104, the text display method provided by the embodiments of the present application may further include the following steps 105 to 108.
[0161] Step 105: The electronic device receives a second input for the first speaker identification.
[0162] In some embodiments of the present application, the above second input may be any form of input by the user on the first speaker identification through the touch area of the electronic device.
[0163] In some embodiments of the present application, the above second input includes but is not limited to: the user performs a touch input on the first speaker identification through the touch area of the electronic device by a touch device such as a finger or a stylus, or is the user's voice input, or is a specific gesture input by the user, or is other feasible input. Specifically, it can be determined according to the actual usage requirements, and the embodiments of the present application do not make a limitation.
[0164] In some embodiments of the present application, the user can modify the first speaker identification displayed in the voice recognition interface during the process of communicating with others, that is, the electronic device can receive a second input for the first speaker identification in the voice recognition interface during the process of collecting voice.
[0165] In some embodiments of the present application, after the user finishes communicating with others, the user can modify the first speaker identifier displayed on the speech recognition interface, that is, the electronic device can receive a second input for the first speaker identifier in the speech recognition interface after the speech collection ends.
[0166] Step 106: The electronic device displays a second window in response to the second input.
[0167] In some embodiments of the present application, in response to the second input, the electronic device can display a second window on the speech recognition interface, or, in response to the second input, the electronic device can display a second window in a new interface.
[0168] In some embodiments of the present application, after the user finishes communicating with others, the hearing aid assistance program records the speeches of each participant during the communication on the speech recognition interface. If the user wants to modify a certain speaker identifier, the user can perform a second input on the speaker identifier in the speech recognition interface. In response to the second input, the electronic device can display a second window on the speech recognition interface.
[0169] Exemplarily, Figure 13 is a schematic diagram of a speech recognition interface provided by an embodiment of the present application. Taking the example that the user modifies the first speaker identifier after attending an offline meeting, as Figure 13 shown, the speech recognition interface 7 includes the identifier 71 of Speaker 1 such as "Speaker 1", the text of Speaker 1, the identifier of Speaker 2 such as "Speaker 2", the text of Speaker 2, the identifier of Speaker 3 such as "Speaker 3", and the text of Speaker 3. Among them, the text of Speaker 1 is such as "Hello everyone, today we will have a short meeting to mainly discuss the preparations for the upcoming product launch. Manager Zhang, you can talk about the current preparations first.", the text of Speaker 2 is such as "Okay, Mr. Li. The venue has been reserved, and all the equipment has been confirmed, including the sound system and the projector. The promotional materials are being designed and are expected to be completed by next Monday.", and the text of Speaker 3 is such as "I am responsible for contacting the media and guests. Currently, five mainstream media have confirmed their participation. The invitations for the guests have also been sent out." and "By the way, Mr. Li, regarding the reception of the guests, do we need to prepare a special reception plan?". Assume that the user wants to modify the identifier of Speaker 1, then the identifier of Speaker 1 is the above-mentioned first speaker identifier. The user can click on the identifier 71 of Speaker 1, and the electronic device can Figure 14 as shown superimpose and display a second window 70 on the speech recognition interface 7.
[0170] In some embodiments of the present application, during the process of a user communicating with others, the hearing - impaired assistance program records the statements of each participant during the communication on the speech recognition interface. If the user wants to modify a speaker identifier, the user can perform a second input on the speaker identifier in the speech recognition interface, and in response to the second input, the electronic device can display a second window on the speech recognition interface.
[0171] Exemplarily, Figure 15 is a schematic diagram of a speech recognition interface provided by an embodiment of the present application. Taking the user modifying the first speaker identifier during a family gathering as an example, as Figure 15 shown, the speech recognition interface 8 includes the identifier of dad such as "dad", the text of dad, the identifier of grandma such as "grandma", the text of grandma, the identifier 81 of speaker 1 such as "speaker 1", and the text of speaker 1. Among them, the text of dad is such as "I'm going out to buy some groceries", the text of grandma is such as "Remember to buy something for sleeping, the child likes it" and "Oh, right, remember to bring some needles and thread back", and the text of speaker 1 is such as "I want to eat roast duck, dad". Assume that the user wants to modify the identifier 81 of speaker 1, then the identifier of speaker 1 is the above - mentioned first speaker identifier. The user can click on the identifier 81 of speaker 1, and the electronic device can, as Figure 16 shown, superimpose and display a second window 80 on the speech recognition interface 8.
[0172] Step 107: The electronic device receives the third speaker identifier input by the user in the second window.
[0173] In some embodiments of the present application, the above - mentioned second window may include an identifier input area, and the user can input the third speaker identifier in this identifier input area.
[0174] In some embodiments of the present application, the above - mentioned second window may include an identifier input area and a confirmation control. The user can input the third speaker identifier in the identifier input area and trigger the confirmation control. After the user triggers the confirmation control, it can be considered that the electronic device has received the third speaker identifier input by the user.
[0175] Exemplarily, taking the user modifying the first speaker identifier after an offline meeting as an example, as Figure 14 shown, a second window 70 is displayed on the speech recognition interface 7. The second window 70 includes an identifier input area 701, and the user can input the third speaker identifier, such as "General Manager Li", in the identifier input area 701. Further example, an input keyboard 72 is also displayed on the speech recognition interface 7, and the user can, through operations on the input keyboard 72, input the third speaker identifier, such as "General Manager Li", in the identifier input area 701.
[0176] Exemplarily, taking the user modifying the first speaker identifier during a family gathering as an example, asFigure 16 As shown, a second window 80 is displayed on the speech recognition interface 8. The second window 80 includes an identification input area 801 where the user can input a third speaker identification, such as "sister".
[0177] Step 108: The electronic device modifies the displayed first speaker identification to the third speaker identification.
[0178] In some embodiments of the present application, the electronic device can modify the first speaker identification displayed on the speech recognition interface to the third speaker identification input by the user.
[0179] Exemplarily, taking the example of the user modifying the first speaker identification after an offline meeting, as Figure 14 shown, a confirmation control 702 is also displayed on the speech recognition interface 7. After the user inputs the third speaker identification such as "General Manager Li" in the identification input area 701, the user clicks the confirmation control 702. Then, the electronic device can modify the identification 71 of speaker 1 displayed, such as "Speaker 1", to "General Manager Li". As Figure 17 shown, after the electronic device modifies the identification 71 of speaker 1 displayed, such as "Speaker 1", to "General Manager Li", the speech recognition interface 7 includes the identification 71 of General Manager Li, such as "General Manager Li", the text of General Manager Li, the identification of speaker 2, such as "Speaker 2", the text of speaker 2, the identification of speaker 3, such as "Speaker 3", and the text of speaker 3.
[0180] Exemplarily, taking the example of the user modifying the first speaker identification during a family gathering, as Figure 16 shown, a confirmation control 802 is also displayed on the speech recognition interface 8. After the user inputs the third speaker identification such as "sister" in the identification input area 801, the user clicks the confirmation control 802. Then, the electronic device can modify the identification 81 of speaker 1 displayed, such as "Speaker 1", to "sister". As Figure 18 shown, after the electronic device can modify the identification 81 of speaker 1 displayed, such as "Speaker 1", to "sister", the speech recognition interface 8 includes the identification of dad, such as "dad", the text of dad, the identification of grandma, such as "grandma", the text of grandma, the identification 81 of sister, such as "sister", and the text of sister.
[0181] In this way, during the process of collecting speech or after the speech collection is completed, the electronic device can modify the displayed speaker identification through the user's input. In this way, even if there is no pre-established correspondence between the speaker identification and the timbre feature information, each speaker's speech content can be quickly distinguished through the modification, which is convenient for the user to view and improves the user's operation efficiency.
[0182] In some embodiments of the present application, after the above step 108, the text display method provided by the embodiments of the present application may further include the following step 109 or step 110.
[0183] Step 109: When there is a second correspondence in the correspondence between the tone identifier and the speaker identifier, the electronic device modifies the first speaker identifier in the second correspondence to a third speaker identifier.
[0184] In some embodiments of the present application, the above second correspondence may be the correspondence between the first speaker identifier and the first tone identifier.
[0185] In some embodiments of the present application, when the second correspondence between the first tone identifier and the first speaker identifier is stored in the electronic device, since the user modifies the first speaker identifier displayed on the electronic device to a third speaker identifier, the electronic device may modify the first speaker identifier in the second correspondence to the third speaker identifier.
[0186] Exemplarily, assume that the first tone identifier is tone 3, the first speaker identifier is "Teacher Liu", and the user modifies the displayed "Teacher Liu" to "Director Liu". Then, the electronic device may modify "Speaker 3" in the second correspondence to "Director Liu". That is, the currently stored correspondence in the electronic device is the correspondence between "tone 3" and "Director Liu", and there is no longer the correspondence between "tone 3" and "Teacher Liu".
[0187] In some embodiments of the present application, taking the example that the user modifies the first speaker identifier after the offline meeting, the above second window may further include a memory option for whether to remember the tone. If this memory option is not selected, it means that the tone of the speaker indicated by the third speaker identifier is not to be remembered; if this memory option is selected, it means that the tone of the speaker indicated by the third speaker identifier can be remembered.
[0188] Exemplarily, after the user inputs the third speaker identifier in the second window, checks the memory option in the second window, and clicks the OK control in the second window. Then, in response to the click input on the OK control, the electronic device determines whether there is a second correspondence in the stored correspondence between the tone identifier and the speaker identifier. If so, the electronic device modifies the first speaker identifier in the stored second correspondence to the third speaker identifier, and the second correspondence is the correspondence between the first speaker identifier and the tone identifier.
[0189] For example, as Figure 14As shown, a second window 70 is displayed on the voice recognition interface 7. An identification input area 701, a confirmation control 702, and a memory option 703 are displayed in the second window 70. The user can input a third speaker identification, such as "General Manager Li", in the identification input area 701, check the memory option 703, and click the confirmation control 702. Then, the electronic device determines whether there is a second corresponding relationship in the stored corresponding relationship between the tone color identification and the speaker identification. If so, the electronic device modifies "Speaker 1" in the stored second corresponding relationship to "General Manager Li".
[0190] In some embodiments of the present application, taking the example that the user modifies the first speaker identification during a family gathering, after the user inputs the third speaker identification in the second window and clicks the confirmation control, it can be considered that it is necessary to memorize the tone color of the speaker indicated by the third speaker identification. Then, the electronic device can determine whether there is a second corresponding relationship in the stored corresponding relationship between the tone color identification and the speaker identification. If so, the electronic device modifies the first speaker identification in the second corresponding relationship to the third speaker identification.
[0191] Exemplarily, as Figure 16 shown, the second window 80 includes an identification input area 801 and a confirmation control 802, and also displays the words "Add to Voice Memory". The user can input a third speaker identification, such as "Sister", in the identification input area 801 and click the confirmation control 802. Then, the electronic device determines whether there is a second corresponding relationship in the stored corresponding relationship between the tone color identification and the speaker identification. If so, the electronic device modifies "Speaker 1" in the second corresponding relationship to "General Manager Li".
[0192] Step 110: When there is no second corresponding relationship in the corresponding relationship between the tone color identification and the speaker identification, the electronic device establishes and stores the corresponding relationship between the third speaker identification and the first tone color identification.
[0193] In some embodiments of the present application, when there is no stored corresponding relationship between the first tone color identification and any speaker identification in the electronic device, the electronic device can establish and store the corresponding relationship between the third speaker identification and the first tone color identification.
[0194] In some embodiments of the present application, the above first tone color identification may be the tone color identification corresponding to the first speaker identification stored in the electronic device. Although the electronic device does not store the corresponding relationship between the third speaker identification and the first tone color identification, the electronic device can store the first tone color identification.
[0195] In some embodiments of the present application, when there is no second corresponding relationship in the corresponding relationship between the timbre identifier and the speaker identifier, the electronic device may send the first voice to the server. The server extracts the first timbre feature information of the first voice, determines the first timbre identifier corresponding to the first timbre feature information, and then sends the first timbre identifier to the electronic device, and the electronic device receives the first timbre identifier from the server.
[0196] In some embodiments of the present application, taking the example that the user modifies the first speaker identifier after participating in an offline meeting, after the user inputs the third speaker identifier in the second window, checks the memory option in the second window, and clicks the OK control in the second window. Then, in response to the click input on the OK control, the electronic device determines whether there is a second corresponding relationship in the stored corresponding relationship between the timbre identifier and the speaker identifier. If not, the electronic device sends the first voice to the server, receives the timbre identifier from the server, and then establishes and stores the corresponding relationship between the timbre identifier and the first speaker identifier.
[0197] For example, as Figure 14 shown, a second window 70 is displayed on the voice recognition interface 7. An identifier input area 701, an OK control 702, and a memory option 703 are displayed in the second window 70. The user can input the third speaker identifier, such as "General Manager Li", in the identifier input area 701, check the memory option 703, and click the OK control 702. Then, the electronic device determines whether there is a second corresponding relationship in the stored corresponding relationship between the timbre identifier and the speaker identifier. If not, the electronic device sends the first voice to the server, receives the timbre identifier such as "Timbre 1" from the server, and then establishes and stores the corresponding relationship between "Timbre 1" and "General Manager Li".
[0198] In some embodiments of the present application, taking the example that the user modifies the first speaker identifier during a family gathering, after the user inputs the third speaker identifier in the second window and clicks the OK control, it can be considered that it is necessary to remember the timbre of the speaker indicated by the third speaker identifier. Then the electronic device determines whether there is a second corresponding relationship in the stored corresponding relationship between the timbre identifier and the speaker identifier. If not, the electronic device sends the first voice to the server, receives the timbre identifier from the server, and then establishes and stores the corresponding relationship between the timbre identifier and the first speaker identifier.
[0199] Exemplarily, as Figure 16As shown, the second window 80 includes an identification input area 801 and a determination control 802, and also displays the words "Add to Voice Memory". The user can input the third speaker identification, such as "sister", in the identification input area 801 and click the determination control 802. Then, the electronic device determines whether there is a second corresponding relationship in the corresponding relationship between the stored tone identifications and the speaker identifications. If not, the electronic device sends the first voice to the server and receives a tone identification from the server, such as "0001", and then establishes and stores the corresponding relationship between "0001" and "sister".
[0200] In this way, after the user modifies a certain speaker identification, the electronic device can establish and store the corresponding relationship between the tone identification and the modified speaker identification, which is convenient for the electronic device to quickly determine the speaker identification according to the pre-stored corresponding relationship between the tone identification and the speaker identification when collecting the voice of the same person later, and display the correct speaker identification for the text after voice conversion, improving the accuracy of determining the speaker identification.
[0201] In some embodiments of the present application, the text display method provided by the embodiments of the present application may further include the following step 111 and step 112.
[0202] Step 111: The electronic device receives a fourth input for the fourth speaker identification displayed on the voice recognition interface.
[0203] In some embodiments of the present application, the above-mentioned fourth speaker identification is any speaker identification displayed on the voice recognition interface.
[0204] In some embodiments of the present application, the above-mentioned fourth input may be any form of input by the user on the fourth speaker identification through the touch area of the electronic device.
[0205] In some embodiments of the present application, the above-mentioned fourth input includes but is not limited to: the user performs a touch input on the fourth speaker identification through the touch area of the electronic device by a touch device such as a finger or a stylus, or is the user's voice input, or is a specific gesture input by the user, or is other feasible input. Specifically, it can be determined according to the actual usage requirements, and the embodiments of the present application do not make limitations.
[0206] In some embodiments of the present application, after the voice collection is completed, the user can first perform a sixth input on any text on the voice recognition interface to make the voice recognition interface in an editable state, and then perform a fourth input on the fourth speaker identification on the voice recognition interface.
[0207] In some embodiments of the present application, the above-mentioned sixth input may be any form of input by the user on any text of the voice recognition interface through the touch area of the electronic device. Exemplarily, the above-mentioned sixth input may be a long-press input, a double-tap input, a specific gesture, etc., and the embodiments of the present application do not limit this.
[0208] In some embodiments of the present application, the above-mentioned fourth input may include a first sub-input and a second sub-input. When the voice recognition interface is in an editable state, the voice recognition interface may display a save control. The user makes a first input on the fourth speaker identifier of the voice recognition interface to select the fourth speaker identifier, and then the user makes a second sub-input on the save control to associate and store the fourth speaker identifier and the fourth text.
[0209] Step 112: The electronic device responds to the fourth input and associates and stores the fourth speaker identifier and the fourth text.
[0210] In some embodiments of the present application, the above-mentioned fourth text is the text obtained by converting the voice uttered by the fourth speaker indicated by the fourth speaker identifier.
[0211] Exemplarily, after the voice collection is completed, the user can long-press any text displayed on the voice recognition interface to make the voice recognition interface in an editable state and display the save control. Then, the user short-presses or clicks on the fourth speaker identifier of the voice recognition interface and short-presses or clicks on the save control. The electronic device associates and stores the fourth speaker identifier and the fourth text and displays them in the form of a note.
[0212] For example, Figure 19 is a schematic diagram of a voice recognition interface provided by an embodiment of the present application. As Figure 19 shown, the voice recognition interface 9 is in an editable state. Four texts and three speaker identifiers are displayed on the voice recognition interface 9, and a save control 91 is also displayed. Among them, the four texts are text 92, text 93, text 94, and text 95 respectively, and the three speaker identifiers are speaker identifier 96 such as "General Manager Li", speaker identifier 97 such as "Manager Zhang", and speaker identifier 98 such as "Supervisor Wang". If the user wants to save all the texts, the user can click on speaker identifier 96, speaker identifier 97, and speaker identifier 98, and then click on the save control 91. Then the electronic device can associate and save "General Manager Li" and text 92, associate and save "Manager Zhang" and text 93, associate and save "Supervisor Wang", text 94, and text 95, and display them in the form of an atomic note, as Figure 20 shown.
[0213] In this way, after the voice collection is completed, the user can associate and save the required text and speaker identifier in the atomic note for easy viewing later.
[0214] Further, the voice recognition interface may further include a copy control and a delete control. The user may select a speaker identifier, and then click the copy control to copy the text of the speaker indicated by the speaker identifier. Alternatively, the user may select a speaker identifier, and then click the delete control to delete the text of the speaker indicated by the speaker identifier.
[0215] Exemplarily, as Figure 19 shown, the voice recognition interface 9 further includes a copy control 90 and a delete control 99. The user may select the speaker identifier 96, and then click the copy control 90 to copy the text 92. Alternatively, the user may select the speaker identifier 98, and then click the delete control 99 to delete the text 94 and the text 95.
[0216] It should be noted that the above step 111 and step 112 may be executed after the above step 104, or may be executed after the above step 108. The embodiments of the present application do not limit this.
[0217] The embodiments of the present application further provide a text display method. The execution subject of the text display method may be a text display device. Exemplarily, the text display device may be a server, or a functional module or entity in the server. Hereinafter, taking the execution subject as the server as an example, the text display method provided by the embodiments of the present application will be described exemplarily.
[0218] Figure 21 is a schematic flowchart of the text display method provided by the embodiments of the present application. As Figure 21 shown, the text display method provided by the embodiments of the present application may include the following steps 201 to 204.
[0219] Step 201: The server receives a first voice from the electronic device.
[0220] In some embodiments of the present application, the above first voice may be any voice collected by the electronic device during the voice collection process.
[0221] In some embodiments of the present application, an auditory-visual assistance program may be installed in the electronic device, and the auditory-visual assistance program may assist hearing-impaired users in communicating with others. The above server may be a server that provides background services for the auditory-visual assistance program.
[0222] It should be noted that the specific implementation of step 201 may refer to the relevant description of step 101. To avoid repetition, it will not be elaborated here.
[0223] Step 202: The server converts the first voice into a first text and extracts the first timbre feature information of the first voice.
[0224] In some embodiments of the present application, the above-mentioned first timbre feature information can be used to characterize the timbre of the first speech. In other words, the first timbre feature information can be used to characterize the timbre or voice characteristics of the speaker who emits the first speech.
[0225] In some embodiments of the present application, the above-mentioned first timbre feature information includes, but is not limited to, MFCC, LPCC, fundamental frequency, formant, spectral envelope, and so on.
[0226] In some embodiments of the present application, the server can perform speaker recognition on the first speech through a speaker recognition model to extract the first timbre feature information of the first speech.
[0227] Exemplarily, the above-mentioned speaker recognition model can include a Gaussian Mixture Model (GMM), a Hidden Markov Model (HMM), a Vector Quantization (VQ) model, a Dynamic Time Warping (DTW) model, and a trained deep learning model capable of performing speaker recognition, and so on. Among them, the deep learning model can include a Convolutional Neural Network (CNN), a Recurrent Neural Network (RNN), a Long Short-Term Memory (LSTM) network, a Gated Recurrent Unit (GRU), and so on.
[0228] In some embodiments of the present application, the server can convert the collected first speech into a first text through a trained speech-to-text model.
[0229] Exemplarily, the server can use multiple speech-text sample pairs to train the speech-to-text model in advance, obtain the trained speech-to-text model and store it in the server. Then, when the server collects the first speech, the server can input the first speech into the trained speech-to-text model, and convert the first speech into a first text through the trained speech-to-text model to obtain the first text. Among them, the above-mentioned speech-to-text model can be a deep neural network, a convolutional neural network, a recurrent neural network, a Transformer model, and so on.
[0230] Exemplarily, taking an electronic device as a mobile phone as an example, in a scenario where a user communicates with multiple people, the user can start a hearing and speaking assistance program installed in the mobile phone, collect the voices of others through the hearing and speaking assistance program, and then the electronic device sends the collected voices to the server. After receiving the voices, the server uses a trained speech-to-text model to convert the voices into text and uses a speaker recognition model to extract the timbre feature information of the voices.
[0231] Step 203: The server determines a first timbre identifier based on the first timbre feature information.
[0232] In some embodiments of the present application, the above first timbre identifier corresponds to the first timbre feature information.
[0233] Exemplarily, the above first timbre identifier can be a text identifier, a digital identifier, etc. Among them, the text identifier can be timbre 1, timbre 2, etc. The digital identifier can be 0001, 0002, 0003, etc.
[0234] In some embodiments of the present application, the server can randomly generate the first timbre identifier, or can find the first timbre identifier corresponding to the first timbre feature information from the pre-stored correspondence between timbre feature information and timbre identifiers.
[0235] In some embodiments of the present application, in combination Figure 21 , such as Figure 22 shown, the above step 203 can be specifically implemented through the following steps 2031 to 2033.
[0236] Step 2031: The server obtains the correspondence between timbre feature information and timbre identifiers.
[0237] In some embodiments of the present application, the above correspondence between timbre feature information and timbre identifiers can be a set of correspondences or multiple sets of correspondences. Among them, each set of correspondences includes a timbre feature information and a timbre identifier.
[0238] In some embodiments of the present application, the correspondence between timbre feature information and timbre identifiers is pre-stored in the server.
[0239] It should be noted that the above correspondence between timbre feature information and timbre identifiers can be pre-established and stored in the server, and its establishment process can refer to the relevant descriptions of the following embodiments.
[0240] Step 2032: When there is a third correspondence in the correspondence between timbre feature information and timbre identifiers, the server determines the timbre identifier in the third correspondence as the first timbre identifier.
[0241] In some embodiments of the present application, the above third correspondence may include timbre feature information that matches the first timbre feature information, and a timbre identifier corresponding to the timbre feature information that matches the first timbre feature information.
[0242] In some embodiments of the present application, the server may match the first timbre feature information with each timbre feature information in the above correspondence between timbre feature information and timbre identifier. Then, the server determines the timbre identifier corresponding to the timbre feature information that matches the first timbre feature information as the first timbre identifier.
[0243] In some embodiments of the present application, by matching the first timbre feature information with each timbre feature information in the above correspondence between timbre feature information and timbre identifier, the server can obtain multiple matching degrees. When there is a matching degree greater than the matching degree threshold among the multiple matching degrees, the server determines that there is timbre feature information that matches the first timbre feature information in the above correspondence between timbre feature information and timbre identifier. Then, the server may determine the timbre identifier corresponding to the timbre feature information corresponding to the maximum matching degree as the first timbre identifier.
[0244] It should be noted that the above matching degree threshold can be set or adjusted according to actual needs, and the embodiments of the present application do not limit this. For example, the above matching degree threshold may be 95% or 98%.
[0245] Exemplarily, the server may use a similarity matching algorithm to calculate the matching degree of two timbre feature information. Or, the server may determine the matching degree of two timbre feature information through a pre-trained model. Or, the server may also use other algorithms to determine whether two timbre feature information match, and the embodiments of the present application do not limit this.
[0246] Step 2033: When there is no third correspondence in the correspondence between timbre feature information and timbre identifier, the server generates a first timbre identifier.
[0247] In some embodiments of the present application, the server may match the first timbre feature information with each timbre feature information in the above correspondence between timbre feature information and timbre identifier to obtain the matching degrees of the first timbre feature information and each timbre feature information, that is, obtain multiple matching degrees. When all the multiple matching degrees are less than the matching degree threshold, it means that there is no timbre feature information that matches the first timbre feature information in the above correspondence between timbre feature information and speaker identifier. Then, the server may generate a first timbre identifier.
[0248] It should be noted that the explanation of the matching degree threshold and the calculation of the matching degree of two timbre feature information can be referred to the relevant description in step 2032 above. To avoid repetition, it will not be elaborated here.
[0249] Exemplarily, the server can randomly generate a first timbre identifier and ensure that the first timbre identifier is different from the timbre identifiers in the corresponding relationship. For example, assuming that the timbre identifiers in the corresponding relationship include 0001, 0002, and 0003, the server cannot use 0001, 0002, 0003, and can use 0004 or 0005 as the first timbre identifier.
[0250] It should be noted that the above step 2032 and step 2033 are parallel solutions, and one of them is to be executed.
[0251] In the embodiment of the present application, when there is timbre feature information that matches the first timbre feature information in the corresponding relationship between the timbre feature information and the identifier, it indicates that the server has recorded the timbre of the first speaker who uttered the first voice, and the first timbre feature information can represent the timbre of the first speaker. Then, the timbre feature information that matches the first timbre feature information can also represent the timbre of the first speaker, and the server can determine the timbre identifier corresponding to the timbre feature information that matches the first timbre feature information as the first timbre identifier. When there is no timbre feature information that matches the first timbre feature information in the corresponding relationship between the timbre feature information and the timbre identifier, it indicates that the server has not recorded the timbre feature information of the first speaker who uttered the first voice, and the server can generate a first timbre identifier. In this way, the first timbre identifier can be determined quickly and accurately.
[0252] In some embodiments of the present application, in combination with Figure 22 , such as Figure 23 shown, after the above step 2033, the text display method provided by the embodiment of the present application may further include the following step 2034.
[0253] Step 2034: The server establishes and stores the corresponding relationship between the first timbre identifier and the first timbre feature information.
[0254] In some embodiments of the present application, when the server does not find timbre feature information that matches the first timbre feature information in the stored corresponding relationship between the timbre feature information and the timbre identifier, it indicates that the server has not stored the timbre feature information of the first speaker who uttered the first voice. Then, after generating the first timbre identifier, the server can establish and store the corresponding relationship between the first timbre identifier and the first timbre feature information.
[0255] In this way, when the first speaker speaks next time, the server can quickly determine the voiceprint identifier based on the correspondence between the voice of the speech and the established first identifier and the first voiceprint feature information, and send it to the electronic device, improving the efficiency of determining the voiceprint identifier.
[0256] It should be noted that the above step 2034 can be executed before step 204 or after step 204. Figure 23 Only for illustrative purposes when step 2034 is executed before step 204, it does not constitute a limitation to this application.
[0257] Step 204: The server sends the first text and the first voiceprint identifier to the electronic device.
[0258] In some embodiments of this application, the server can represent the first text and the first voiceprint identifier in an associated form. Exemplarily, the first text and the first voiceprint identifier are represented in the form of "first text - first voiceprint identifier".
[0259] In some embodiments of this application, after the server sends the first text and the first voiceprint identifier to the electronic device, the electronic device can determine the first speaker identifier based on the first voiceprint identifier, and then display the first text and the first speaker identifier, so that the user can quickly determine who said which text according to the displayed content.
[0260] For the text display method provided by the embodiments of this application, the server receives the first voice from the electronic device; converts the first voice into the first text, and extracts the first voiceprint feature information of the first voice, where the first voiceprint feature information is used to characterize the voiceprint of the first voice; determines the first voiceprint identifier based on the first voiceprint feature information; and sends the first text and the first voiceprint identifier to the electronic device. In this way, the server determines the voiceprint identifier by extracting the voiceprint feature information of the voice, and sends the voiceprint identifier to the electronic device when sending the text after voice conversion, so that the electronic device can determine the speaker identifier based on the voiceprint identifier, and while displaying the text after voice conversion, also display the speaker identifier of the speaker who uttered the voice, enabling the hearing-impaired user to quickly distinguish who the speaker is when seeing the text, without the hearing-impaired user having to look up to judge who is speaking, and thus not missing important information in the displayed text. Therefore, this solution can avoid information omission and improve the accuracy of information acquisition.
[0261] In some embodiments of this application, before the above step 203 or the above step 201, the text display method provided by the embodiments of this application may further include the following steps 205 to 209.
[0262] Step 205: The server receives the second voice from the electronic device.
[0263] In some embodiments of the present application, the second voice may be the voice collected by the user in the voice recognition setting interface. The second speaker indicated by the second speaker identifier is the speaker who emits the second voice.
[0264] In some embodiments of the present application, the electronic device may pre-collect voices in the voice recognition setting interface of the hearing assistance program and send the collected voices to the server, so that the server can pre-establish and store the corresponding relationship between the timbre feature information and the timbre identifier of the collected voices, facilitating the server to quickly determine the timbre identifier according to the corresponding relationship when collecting voices using the hearing assistance program subsequently.
[0265] It should be noted that the specific implementation of this step can refer to the relevant descriptions of steps 13 to 18 in the above embodiments. To avoid repetition, it will not be elaborated here.
[0266] Step 206: The server extracts the second timbre feature information of the second voice.
[0267] In some embodiments of the present application, the above second timbre feature information is used to characterize the timbre of the second voice. In other words, the above second timbre feature information can be used to characterize the timbre or voice characteristics of the speaker who emits the second voice.
[0268] In some embodiments of the present application, the above second timbre feature information includes but is not limited to MFCC, LPCC, fundamental frequency, formant, spectral envelope, and so on.
[0269] In some embodiments of the present application, the server can perform speaker recognition on the second voice through a speaker recognition model to extract the second timbre feature information of the second voice.
[0270] It should be noted that the specific implementation of this step can refer to the implementation process of the server extracting the first timbre feature information of the first voice in step 201 above. To avoid repetition, it will not be elaborated here.
[0271] Step 207: The server generates a second timbre identifier based on the second timbre feature information.
[0272] It should be noted that the specific implementation of this step can refer to the implementation process of the server determining the second timbre identifier based on the first timbre feature information in step 203 above. To avoid repetition, it will not be elaborated here.
[0273] Step 208: The server establishes and stores the corresponding relationship between the second timbre feature information and the second timbre identifier.
[0274] In some embodiments of the present application, the server can establish the corresponding relationship between the extracted second timbre feature information and the second timbre identifier and store the established corresponding relationship in the server.
[0275] It should be noted that the electronic device can pre-collect multiple voices and send the multiple voices to the server, so that the server can establish and store the corresponding relationship between multiple groups of timbre feature information and timbre identifiers. In the embodiments of the present application, only the establishment and storage of the corresponding relationship between the second timbre feature information and the second timbre identifier are taken as an example for illustration, which does not constitute a limitation to the present application.
[0276] Step 209: The server sends the second timbre identifier to the electronic device.
[0277] It should be noted that the above step 209 can be executed after the above step 208, or can be executed after the above step 207. The embodiments of the present application do not make a limitation on this.
[0278] In this way, by establishing and storing the corresponding relationship between the timbre identifier and the timbre feature information in the server in advance, when the voice is collected subsequently, the server can quickly find the timbre identifier of the speaker who uttered the voice from the stored corresponding relationship, improving the efficiency of determining the timbre identifier.
[0279] It should be noted that the above embodiments are all described by taking the storage of the corresponding relationship between the timbre feature information and the timbre identifier in the server and the storage of the corresponding relationship between the timbre identifier and the speaker identifier in the electronic device as an example for the text display method provided by the embodiments of the present application. In some other embodiments, the corresponding relationship between the timbre feature information and the timbre identifier and the corresponding relationship between the timbre identifier and the speaker identifier can both be stored in the server, or can both be stored in the electronic device. The embodiments of the present application do not make a limitation on this. And when the corresponding relationship between the timbre feature information and the timbre identifier and the corresponding relationship between the timbre identifier and the speaker identifier are both stored in the server, the server can determine the speaker identifier based on the voice from the electronic device and send it to the electronic device. When the corresponding relationship between the timbre feature information and the timbre identifier and the corresponding relationship between the timbre identifier and the speaker identifier are both stored in the electronic device, the electronic device can determine the speaker identifier based on the collected voice. The embodiments of the present application do not make a limitation on this.
[0280] Next, taking the electronic device as a terminal as an example, the interaction process between the electronic device and the server for the text display method provided by the embodiments of the present application will be described. Figure 24 It is a schematic flowchart of the text display method provided by the embodiments of the present application. As Figure 24 shown, the text display method may include the following steps 301 to step 307.
[0281] Step 301: The terminal collects the first voice.
[0282] Step 302: The terminal sends the first voice to the server.
[0283] Step 303: The server receives the first voice, converts the first voice into the first text, and extracts the first timbre feature information of the first voice.
[0284] In some embodiments of the present application, the above first timbre feature information is used to characterize the timbre of the first voice.
[0285] Step 304: The server determines the first timbre identifier based on the first timbre feature information.
[0286] Step 305: The server sends the first text and the first timbre identifier to the terminal.
[0287] Step 306: The terminal receives the first text and the first timbre identifier from the server, and determines the first speaker identifier based on the first timbre identifier.
[0288] In some embodiments of the present application, the above first text is the text obtained by converting the first voice, and the first speaker indicated by the first speaker identifier is the speaker who emits the first voice.
[0289] Step 307: The terminal displays the first text and the first speaker identifier.
[0290] In some embodiments of the present application, the above step 306 can be specifically implemented by the following steps 3061 to 3062.
[0291] Step 3061: The terminal obtains the correspondence between the timbre identifier and the speaker identifier.
[0292] Step 3062: When there is a first correspondence in the correspondence between the timbre identifier and the speaker identifier, the terminal determines the speaker identifier in the first correspondence as the first speaker identifier.
[0293] Wherein, the above first correspondence includes the first timbre identifier and the speaker identifier corresponding to the first timbre identifier.
[0294] Step 3063: When there is no first correspondence in the correspondence between the timbre identifier and the speaker identifier, the terminal generates the first speaker identifier.
[0295] In some embodiments of the present application, after the above step 3063, the text display method provided by the embodiments of the present application may further include the following step 3064.
[0296] Step 3064: The terminal establishes and stores the correspondence between the first timbre identifier and the first speaker identifier.
[0297] In some embodiments of the present application, before the above step 301, the text display method provided by the embodiments of the present application may further include the following steps 3a to 3f.
[0298] Step 3a: When the voice recognition setting interface is displayed, the terminal receives a first input to the voice recognition setting interface.
[0299] Step 3b: The terminal collects a second voice in response to the first input.
[0300] Step 3c: When the collection of the second voice ends, the terminal displays a first window.
[0301] Step 3d: The terminal receives a second speaker identifier input by the user in the first window, and the second speaker indicated by the second speaker identifier is the speaker who emits the second voice.
[0302] Step 3e: The terminal sends the second voice to the server.
[0303] Step 3f: The terminal receives a second tone identifier from the server, establishes and stores the correspondence between the second tone identifier and the second speaker identifier. The second tone identifier corresponds to the second tone feature information, and the second tone feature information is used to characterize the tone of the second voice.
[0304] In some embodiments of the present application, after the above step 307, the text display method provided by the embodiments of the present application may further include the following steps 308 to 311.
[0305] Step 308: The terminal receives a second input to the first speaker identifier.
[0306] Step 309: The terminal displays a second window in response to the second input.
[0307] Step 310: The terminal receives a third speaker identifier input by the user in the second window.
[0308] Step 311: The terminal modifies the displayed first speaker identifier to the third speaker identifier.
[0309] In some embodiments of the present application, the text display method provided by the embodiments of the present application may further include the following step 312 or step 313.
[0310] Step 312: When there is a second correspondence in the correspondence between the tone identifier and the speaker identifier, the terminal modifies the first speaker identifier in the second correspondence to the third speaker identifier.
[0311] Step 313: When there is no second correspondence in the correspondence between the tone identifier and the speaker identifier, the terminal establishes and stores the correspondence between the third speaker identifier and the first tone identifier.
[0312] Wherein, the second correspondence is the correspondence between the first speaker identifier and the first tone identifier.
[0313] In some embodiments of the present application, before the above step 301, the text display method provided by the embodiments of the present application may further include the following steps 401 to 403, and the above step 307 may be specifically implemented by the following step 3071.
[0314] Step 401: The terminal displays a voice recognition interface, and the voice recognition interface includes a collection control.
[0315] Step 402: The terminal receives a third input to the collection control.
[0316] Step 403: In response to the third input, the terminal starts to collect voice.
[0317] Step 3071: The terminal displays the first text and the first speaker identifier in the voice recognition interface.
[0318] In some embodiments of the present application, the above step 304 may be specifically implemented by the following steps 3041 to 3043.
[0319] Step 3041: The server obtains the correspondence between the tone feature information and the tone identifier.
[0320] Step 3042: When there is a third correspondence in the correspondence between the tone feature information and the tone identifier, the server determines the tone identifier in the third correspondence as the first tone identifier.
[0321] Wherein, the third correspondence includes the tone feature information matching the first tone feature information and the tone identifier corresponding to the tone feature information matching the first tone feature information.
[0322] Step 3043: When there is no third correspondence in the correspondence between the tone feature information and the tone identifier, the server generates the first tone identifier.
[0323] In some embodiments of the present application, after the above step 3043, the text display method provided by the embodiments of the present application may further include the following step 3044.
[0324] Step 3044: The server establishes and stores the correspondence between the first tone identifier and the first tone feature information.
[0325] In some embodiments of the present application, before step 304 above, the text display method provided by the embodiments of the present application may further include the following steps 404 to 408.
[0326] Step 404: The server receives a second voice from the terminal.
[0327] Step 405: The server extracts second voice color feature information of the second voice, and the second voice color feature information is used to characterize the voice color of the second voice.
[0328] Step 406: The server generates a second voice color identifier based on the second voice color feature information.
[0329] Step 407: The server establishes and stores a correspondence between the second voice color feature information and the second voice color identifier.
[0330] Step 408: The server sends the second voice color identifier to the terminal.
[0331] It should be noted that the specific implementation processes of the above steps 301 to 307 and other steps in this embodiment can refer to the relevant descriptions of the above embodiments. To avoid repetition, this embodiment will not be elaborated here.
[0332] In the text display method provided by the embodiments of the present application, when displaying the text converted from the voice, after distinguishing the speaker, the text and the speaker identifier are displayed to the user together, so that the hearing-impaired user can quickly distinguish who the speaker is when seeing the text, and the hearing-impaired user does not need to look up to judge who is speaking, and important information in the displayed text will not be missed. Therefore, this solution can avoid information omission and improve the accuracy of information acquisition. Moreover, the user can edit the displayed speaker identifier during the voice collection process, or record the speaker's voice in advance, establish the correspondence between the voice color identifier and the voice color feature information of the voice, and establish the correspondence between the voice color identifier and the speaker identifier, so that the speaker identifier can be directly determined in subsequent recognition. In addition, if the collected voice does not match the established voice color identifier, or the speaker identifier is not edited during a conversation, it can be edited after the conversation ends. The same speaker can be displayed as the same edited speaker identifier, which can improve the user experience.
[0333] Next, taking the electronic device as a mobile phone and the hearing assistance program as the hearing assistance APP as an example, the text display method provided by the embodiments of the present application will be described. As Figure 25 shown, the text display method may include the following steps 2301 to 2318.
[0334] Step 2301: The mobile phone is powered on and the hearing assistance APP is opened.
[0335] Step 2302: The mobile phone displays a voice recognition settings interface, receives a long - press input on the voice collection control in the voice recognition settings interface, and in response to the long - press input, the mobile phone collects a second voice.
[0336] Step 2303: When the collection of the second voice ends, the mobile phone sends the second voice to the server, displays a first window, and receives a second speaker identifier input by the user in the identifier input area of the first window.
[0337] In some embodiments of the present application, the second speaker indicated by the second speaker identifier is the speaker who emits the second voice.
[0338] Step 2304: The server receives the second voice, converts the second voice into a second text, extracts the second timbre feature information of the second voice, and determines a second timbre identifier based on the second timbre feature information.
[0339] In some embodiments of the present application, the server may establish a correspondence between the second timbre feature information and the second timbre identifier.
[0340] Step 2305: The server sends the second timbre identifier to the mobile phone.
[0341] Step 2306: The mobile phone receives the second timbre identifier, and establishes and stores a correspondence between the second timbre identifier and the second speaker identifier.
[0342] Step 2307: The mobile phone displays a voice recognition interface in the listening and speaking assistance APP, and the voice recognition interface includes a collection control and a multi - function control.
[0343] Step 2308: The mobile phone receives a trigger input on the multi - function control and displays a speaker differentiation control.
[0344] Step 2309: The mobile phone receives a trigger input on the speaker differentiation control and activates the speaker differentiation function.
[0345] Step 2310: The mobile phone receives a trigger input from the user on the collection control and starts collecting voice.
[0346] Step 2311: The mobile phone sends the collected first voice to the server.
[0347] Step 2312: The server receives the first voice, converts the first voice into a first text, extracts the first timbre feature information of the collected first voice, and determines a first timbre identifier based on the first timbre feature information.
[0348] Step 2313: The server sends the first timbre identifier and the first text to the mobile phone.
[0349] Step 2314: The mobile phone receives the first tone identifier and the first text, and determines the first speaker identifier based on the first tone identifier.
[0350] Step 2315: The mobile phone displays the first text and the first speaker identifier on the speech recognition interface.
[0351] Step 2316: The mobile phone receives a click input on the first speaker identifier displayed on the speech recognition interface, and in response to the click input, displays a second window.
[0352] Step 2317: The mobile phone receives the third speaker identifier input by the user in the second window.
[0353] Step 2318: The mobile phone modifies the displayed first speaker identifier to the third speaker identifier.
[0354] In some embodiments of the present application, when there is a second correspondence in the correspondence between the tone identifier and the speaker identifier, the mobile phone modifies the first speaker identifier in the second correspondence to the third speaker identifier. When there is no second correspondence in the correspondence between the tone identifier and the speaker identifier, the mobile phone establishes and stores the correspondence between the third speaker identifier and the first tone identifier. Wherein, the second correspondence is the correspondence between the first speaker identifier and the first tone identifier.
[0355] In some embodiments of the present application, after the voice collection is completed, the user can perform a selection input on the fourth speaker identifier among the multiple speaker identifiers displayed on the speech recognition interface, and the mobile phone, in response to the selection input, associates and stores the fourth speaker identifier and the fourth text as an atomic note.
[0356] The text display method provided by the embodiments of the present application, by distinguishing the speaker when displaying the text converted from the voice and then displaying the text and the speaker identifier to the user together, enables the hearing-impaired user to quickly distinguish who the speaker is when seeing the text, without the hearing-impaired user having to look up to judge who is speaking, so that important information in the displayed text will not be missed. Therefore, this solution can avoid information omission and improve the accuracy of information acquisition. Moreover, the user can edit the displayed speaker identifier during the voice collection process, or record the speaker's voice in advance, establish the correspondence between the tone identifier and the tone feature information of the voice, and establish the correspondence between the tone identifier and the speaker identifier, so that the speaker identifier can be directly determined in subsequent recognition. In addition, if the collected voice does not match the established tone identifier, or the speaker identifier is not edited during a conversation, it can be edited after the conversation ends, and the same speaker can be displayed as the same edited speaker identifier, which can improve the user experience.
[0357] It should be noted that the specific implementation processes of the above steps 2301 to 2318 can be referred to the relevant descriptions of the above embodiments. To avoid repetition, they will not be elaborated herein in this embodiment.
[0358] Next, for each scenario applicable to the embodiments of the present application, in combination with the above implementation solutions of the embodiments of the present application, specific examples are given below to illustrate the implementation processes in each scenario of the embodiments of the present application.
[0359] Scenario where a hearing-impaired user participates in a multi-person offline meeting: The hearing-impaired user opens the hearing assistance APP installed on the mobile phone, enters the speech recognition interface, and enables the function of differentiating speakers, and then starts recording the speech of the speakers at the meeting. The mobile phone sends the recorded speech to the server. The server converts the recorded speech into text, extracts the timbre feature information of the recorded speech, then determines the timbre identifier according to the timbre feature information, and sends the timbre identifier to the mobile phone. The mobile phone determines the speaker identifier according to the timbre identifier, and then displays the speaker identifier and the text on the speech recognition interface. During the recording process, the hearing-impaired user can modify the speaker identifier to "General Manager Li", "Manager Wang", etc. In this way, after the speech of "General Manager Li" is recorded subsequently, "General Manager Li" can be directly displayed in front of the text. After the recording is completed, the hearing-impaired user can also modify the speaker identifier and save the displayed text and the speaker identifier, so that the speech content of each person at the meeting can be quickly found subsequently.
[0360] Scene of a hearing - impaired user communicating with family members in daily life: The hearing - impaired user opens the hearing - aid APP installed on the mobile phone and enters the voice recognition setting interface. The mobile phone pre - records the voices of "mom" and "dad" in advance, and sends the voice of "mom" and the voice of "dad" to the server. The server extracts the timbre feature information 1 of the voice of "mom", generates the timbre identifier 1 corresponding to the timbre feature information 1, establishes and stores the corresponding relationship between the timbre feature information 1 and the timbre identifier 1, and extracts the timbre feature information 2 of the voice of "dad" by the server, generates the timbre identifier 2 corresponding to the timbre feature information 2, and establishes and stores the corresponding relationship between the timbre feature information 2 and the timbre identifier 2. Moreover, the server sends the timbre identifier 1 and the timbre identifier 2 to the mobile phone. The mobile phone locally stores the corresponding relationship between the timbre identifier 1 of "mom" and the speaker identifier "mom", and the corresponding relationship between the timbre identifier 2 of "dad" and the speaker identifier "dad". Then, when communicating with family members, enter the voice recognition interface of the hearing - aid APP and turn on the function of distinguishing speakers, and then start recording the voices of family members. The mobile phone sends the recorded voice to the server. The server converts the recorded voice into text, extracts the timbre feature information of the recorded voice, determines the timbre identifier based on the timbre feature information, and then sends the timbre identifier to the mobile phone. If the mobile phone determines that the speaker identifier is "mom" based on the timbre identifier, the mobile phone can display the text on the voice recognition interface and display "mom" in front of the text. If the mobile phone cannot determine the speaker identifier based on the timbre identifier, the speaker identifier can be recorded as speaker identifier 1. Then, during the recording process, if the user sees that speaker 1 is the younger sister, the speaker identifier 1 can be modified to "younger sister", and the corresponding relationship between "younger sister" and the timbre identifier is stored. After the recording ends, the hearing - impaired user can also save the displayed text and the speaker identifier, so that the content said by family members can be quickly found later.
[0361] It should be noted that each of the above - mentioned method embodiments, or various possible implementation manners in each method embodiment, can be executed independently, or any two or more of them can be combined with each other. It can be specifically determined according to actual usage requirements, and the embodiments of the present application do not limit this.
[0362] A text display method provided by an embodiment of the present application, and the execution subject can be a text display device. In the embodiments of the present application, taking the text display device executing the text display method as an example, the text display device provided by the embodiments of the present application is described.
[0363] Figure 26 It is a structural schematic diagram of the text display device provided by an embodiment of the present application. The text display device includes a collection module 2601, a sending module 2602, a receiving module 2603, a processing module 2604, and a display module 2605.
[0364] The acquisition module 2601 is used to acquire the first voice. The sending module 2602 is used to send the first voice to the server. The receiving module 2603 is used to receive the first text and the first tone color identifier from the server. The first text is the text obtained by converting the first voice. The first tone color identifier corresponds to the first tone color feature information, and the first tone color feature information is used to characterize the tone color of the first voice. The processing module 2604 is used to determine the first speaker identifier based on the first tone color identifier. The first speaker indicated by the first speaker identifier is the speaker who emits the first voice. The display module 2605 is used to display the first text and the first speaker identifier.
[0365] In some embodiments of the present application, the processing module 2604 is specifically configured to obtain the correspondence between the tone color identifier and the speaker identifier; and, in the case where there is a first correspondence in the correspondence between the tone color identifier and the speaker identifier, determine the speaker identifier in the first correspondence as the first speaker identifier; or, in the case where there is no first correspondence in the correspondence between the tone color identifier and the speaker identifier, generate the first speaker identifier; wherein, the first correspondence includes the first tone color identifier and the speaker identifier corresponding to the first tone color identifier.
[0366] In some embodiments of the present application, in combination Figure 26 , such as Figure 27 shown, the above text display device further includes a storage module 2606.
[0367] The processing module 2604 is further configured to establish a correspondence between the first tone color identifier and the first speaker identifier after generating the first speaker identifier. The storage module 2606 is used to store the correspondence between the first tone color identifier and the first speaker identifier.
[0368] In some embodiments of the present application, the receiving module 2603 is further configured to receive a first input to the voice recognition setting interface in the case of displaying the voice recognition setting interface before acquiring the first voice and sending the first voice to the server. The acquisition module 2601 is further configured to acquire the second voice in response to the first input. The display module 2605 is further configured to display the first window in the case where the acquisition of the second voice ends. The receiving module 2603 is further configured to receive the second speaker identifier input by the user in the first window. The second speaker indicated by the second speaker identifier is the speaker who emits the second voice. The sending module 2602 is used to send the second voice to the server. The receiving module 2603 is further configured to receive the second tone color identifier from the server. The processing module 2604 is further configured to establish a correspondence between the second tone color identifier and the second speaker identifier. The second tone color identifier corresponds to the second tone color feature information, and the second tone color feature information is used to characterize the tone color of the second voice. The storage module 2606 is used to store the correspondence between the second tone color identifier and the second speaker identifier.
[0369] In some embodiments of the present application, the receiving module 2603 is further configured to receive a second input to the first speaker identifier after the first text and the first speaker identifier are displayed. The display module 2605 is further configured to display a second window in response to the second input. The receiving module 2603 is further configured to receive a third speaker identifier input by the user in the second window. The display module 2605 is further configured to modify the displayed first speaker identifier to the third speaker identifier.
[0370] In some embodiments of the present application, the processing module 2604 is further configured to modify the first speaker identifier in the second correspondence to the third speaker identifier if there is a second correspondence in the correspondence between the voice color identifier and the speaker identifier; or, the processing module 2604 is further configured to establish a correspondence between the third speaker identifier and the first voice color identifier if there is no second correspondence in the correspondence between the voice color identifier and the speaker identifier. The storage module 2606 is used to store the correspondence between the third speaker identifier and the first voice color identifier. Wherein, the second correspondence is the correspondence between the first speaker identifier and the first voice color identifier.
[0371] In some embodiments of the present application, the display module 2605 is further configured to display a voice recognition interface before collecting the first voice and sending the first voice to the server, and the voice recognition interface includes a collection control. The receiving module 2603 is further configured to receive a third input to the collection control. The collection module 2601 is further configured to start collecting voice in response to the third input. The display module 2605 is specifically configured to display the first text and the first speaker identifier in the voice recognition interface.
[0372] The text display device provided by the embodiments of the present application can quickly distinguish who the speaker is when the hearing-impaired user sees the text by displaying the speaker identifier of the speaker who uttered the voice while displaying the text after voice conversion, without the hearing-impaired user having to look up to judge who is speaking, so that important information in the displayed text will not be missed. Therefore, this solution can avoid information omission and improve the accuracy of information acquisition.
[0373] The text display device in the embodiments of the present application may be an electronic device or a component in an electronic device, such as an integrated circuit or a chip. The electronic device may be a terminal or other devices other than terminals. Exemplarily, the electronic device may be a mobile phone, a tablet computer, a laptop computer, a handheld computer, a vehicle-mounted electronic device, a mobile Internet device, an augmented reality / virtual reality device, a robot, a wearable device, a super mobile personal computer, a netbook, or a personal digital assistant, etc. It may also be a Network Attached Storage (NAS), a personal computer (PC), a television (TV), a teller machine, or a self-service machine, etc. The embodiments of the present application do not make specific limitations.
[0374] The text display device in the embodiments of the present application may be a device with an operating system. The operating system may be an Android operating system, an iOS operating system, or other possible operating systems. The embodiments of the present application do not make specific limitations.
[0375] The text display device provided by the embodiments of the present application can implement each process implemented by each of the above-mentioned text display methods executed by the electronic device. To avoid repetition, it will not be elaborated here.
[0376] Figure 28 is a schematic structural diagram of the text display device provided by the embodiments of the present application. The text display device includes a receiving module 2801, a processing module 2802, and a sending module 2803.
[0377] The receiving module 2801 is configured to receive a first voice from an electronic device. The processing module 2802 is configured to convert the first voice into a first text, extract first tone color feature information of the first voice, and the first tone color feature information is used to characterize the tone color of the first voice; and, based on the first tone color feature information, determine a first tone color identifier. The sending module 2803 is configured to send the first text and the first tone color identifier to the electronic device.
[0378] In some embodiments of the present application, the processing module 2802 is specifically configured to obtain the correspondence between tone color feature information and tone color identifiers; and, in the case where there is a third correspondence in the correspondence between tone color feature information and tone color identifiers, determine the tone color identifier in the third correspondence as the first tone color identifier; or, in the case where there is no third correspondence in the correspondence between tone color feature information and tone color identifiers, generate a first tone color identifier; wherein, the third correspondence includes tone color feature information that matches the first tone color feature information and the tone color identifier corresponding to the tone color feature information that matches the first tone color feature information.
[0379] In some embodiments of the present application, in combination with Figure 28 , as Figure 29 shown, the text display device further includes a storage module 2804.
[0380] The processing module 2802 is further configured to establish a correspondence between the first tone identifier and the first tone feature information after generating the first tone identifier. The storage module 2804 is further configured to store the correspondence between the first tone identifier and the first tone feature information.
[0381] In some embodiments of the present application, the receiving module 2801 is further configured to receive a second voice from an electronic device before determining the first tone identifier based on the first tone feature information. The processing module 2802 is further configured to extract second tone feature information of the second voice, where the second tone feature information is used to characterize the tone of the second voice; and, based on the second tone feature information, generate a second tone identifier; and, establish a correspondence between the second tone feature information and the second tone identifier. The storage module 2804 is configured to store the correspondence between the second tone feature information and the second tone identifier. The sending module 2803 is further configured to send the second tone identifier to the electronic device.
[0382] The text display device provided by the embodiments of the present application determines a tone identifier by extracting the tone feature information of the voice, and also sends the tone identifier to the electronic device when sending the text after voice conversion, so that the electronic device determines the speaker identifier based on the tone identifier, and while displaying the text after voice conversion, also displays the speaker identifier of the speaker who uttered the voice, enabling the hearing-impaired user to quickly distinguish who the speaker is when seeing the text, without the hearing-impaired user having to look up to judge who is speaking, and thus not missing important information in the displayed text. Therefore, this solution can avoid information omission and improve the accuracy of information acquisition.
[0383] The text display device in the embodiments of the present application can be a server or a component in the server, such as an integrated circuit or a chip.
[0384] The text display device provided by the embodiments of the present application can implement each process implemented by each embodiment of the text display method executed by the above server. To avoid repetition, it will not be elaborated here.
[0385] Optionally, as Figure 30 shown, the embodiments of the present application further provide an electronic device 1200, including a processor 1201 and a memory 1202. A program or instruction that can run on the processor 1201 is stored on the memory 1202. When the program or instruction is executed by the processor 1201, it implements each step of the text display method embodiments executed by the above electronic device, and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0386] It should be noted that the electronic devices in the embodiments of the present application include mobile electronic devices and non-mobile electronic devices.
[0387] Figure 31 Schematic diagram of the hardware structure of an electronic device according to an embodiment of the present application.
[0388] The electronic device 1000 includes, but is not limited to: a radio frequency unit 1001, a network module 1002, an audio output unit 1003, an input unit 1004, a sensor 1005, a display unit 1006, a user input unit 1007, an interface unit 1008, a memory 1009, and a processor 1010 and other components.
[0389] Those skilled in the art can understand that the electronic device 1000 may further include a power source (such as a battery) for supplying power to each component. The power source can be logically connected to the processor 1010 through a power management system, so as to implement functions such as management of charging, discharging, and power consumption management through the power management system. Figure 29 The structure of the electronic device shown does not limit the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.
[0390] In some embodiments of the present application, the electronic device may be a terminal or other devices other than the terminal.
[0391] Among them, the input unit 1004 is used to collect the first voice. The above-mentioned radio frequency unit 1001 is used to send the first voice to the server; and receive the first text and the first tone color identifier from the server. The first text is the text obtained by converting the first voice, and the first tone color identifier corresponds to the first tone color feature information, and the first tone color feature information is used to characterize the tone color of the first voice. The processor 1010 is used to determine the first speaker identifier based on the first tone color identifier. The first speaker indicated by the first speaker identifier is the speaker who emits the first voice. The display unit 1006 is used to display the first text and the first speaker identifier.
[0392] In some embodiments of the present application, the processor 1010 is specifically used to obtain the correspondence between the tone color identifier and the speaker identifier; and, in the case where there is a first correspondence in the correspondence between the tone color identifier and the speaker identifier, determine the speaker identifier in the first correspondence as the first speaker identifier; or, in the case where there is no first correspondence in the correspondence between the tone color identifier and the speaker identifier, generate the first speaker identifier; wherein, the above-mentioned first correspondence includes the first tone color identifier and the speaker identifier corresponding to the first tone color identifier.
[0393] In some embodiments of the present application, the processor 1010 is further configured to establish a correspondence between the first timbre identifier and the first speaker identifier after generating the first speaker identifier. The memory 1009 is configured to store the correspondence between the first timbre identifier and the first speaker identifier.
[0394] In some embodiments of the present application, the user input unit 1007 is further configured to, before collecting the first voice and sending the first voice to the server, receive a first input to the voice recognition setting interface when the voice recognition setting interface is displayed. The input unit 1004 is further configured to collect a second voice in response to the first input. The display unit 1006 is further configured to display a first window when the collection of the second voice ends. The user input unit 1007 is further configured to receive a second speaker identifier input by the user in the first window, where the second speaker indicated by the second speaker identifier is the speaker who emits the second voice. The radio frequency unit 1001 is configured to send the second voice to the server. The radio frequency unit 1001 is further configured to receive a second timbre identifier from the server. The processor 1010 is further configured to establish a correspondence between the second timbre identifier and the second speaker identifier, where the second timbre identifier corresponds to second timbre feature information, and the second timbre feature information is used to characterize the timbre of the second voice. The storage module 2606 is configured to store the correspondence between the second timbre identifier and the second speaker identifier.
[0395] In some embodiments of the present application, the user input unit 1007 is further configured to receive a second input to the first speaker identifier after the first text and the first speaker identifier are displayed. The display unit 1006 is further configured to display a second window in response to the second input. The user input unit 1007 is further configured to receive a third speaker identifier input by the user in the second window. The display unit 1006 is further configured to modify the displayed first speaker identifier to the third speaker identifier.
[0396] In some embodiments of the present application, the processor 1010 is further configured to modify the first speaker identifier in the second correspondence to the third speaker identifier when there is a second correspondence in the correspondence between the timbre identifier and the speaker identifier; or, the processor 1010 is further configured to establish a correspondence between the third speaker identifier and the first timbre identifier when there is no second correspondence in the correspondence between the timbre identifier and the speaker identifier. The memory 1009 is configured to store the correspondence between the third speaker identifier and the first timbre identifier. Wherein, the second correspondence is the correspondence between the first speaker identifier and the first timbre identifier.
[0397] In some embodiments of the present application, the display unit 1006 is further configured to display a speech recognition interface before collecting the first speech and sending the first speech to the server. The speech recognition interface includes a collection control. The user input unit 1007 is further configured to receive a third input to the collection control. The input unit 1004 is further configured to start collecting speech in response to the third input. The display unit 1006 is specifically configured to display a first text and a first speaker identifier in the speech recognition interface.
[0398] The electronic device provided by the embodiments of the present application displays the speaker identifier of the speaker who issues the speech while displaying the text after speech conversion, so that the hearing-impaired user can quickly distinguish who the speaker is when seeing the text, without the need for the hearing-impaired user to look up to judge who is speaking, and important information in the displayed text will not be missed. Therefore, this solution can avoid information omission and improve the accuracy of information acquisition.
[0399] It should be understood that in the embodiments of the present application, the input unit 1004 may include a Graphics Processing Unit (GPU) 10041 and a microphone 10042. The graphics processor 10041 processes the image data of static pictures or videos obtained by an image capture device (such as a camera) in a video capture mode or an image capture mode. The display unit 1006 may include a display panel 10061, and the display panel 10061 may be configured in the form of a liquid crystal display, an organic light emitting diode, etc. The user input unit 1007 includes at least one of a touch panel 10071 and other input devices 10072. The touch panel 10071 is also referred to as a touch screen. The touch panel 10071 may include two parts: a touch detection device and a touch controller. The other input devices 10072 may include, but are not limited to, a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, and a joystick, which will not be elaborated here.
[0400] The memory 1009 can be used to store software programs and various data. The memory 1009 may mainly include a first storage area for storing programs or instructions and a second storage area for storing data. Among them, the first storage area may store an operating system, application programs or instructions required for at least one function (such as a sound playback function, an image playback function, etc.). In addition, the memory 1009 may include a volatile memory or a non-volatile memory, or the memory 1009 may include both a volatile memory and a non-volatile memory. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), a static random access memory (SRAM), a dynamic random access memory (DRAM), a synchronous dynamic random access memory (SDRAM), a double data rate synchronous dynamic random access memory (DDR SDRAM), an enhanced synchronous dynamic random access memory (ESDRAM), a synch link dynamic random access memory (SLDRAM), and a direct rambus random access memory (DRRAM). The memory 109 in the embodiments of the present application includes but is not limited to these and any other suitable types of memories.
[0401] The processor 1010 may include one or more processing units; optionally, the processor 1010 integrates an application processor and a modem processor. Among them, the application processor mainly processes operations related to the operating system, user interface, and application programs, etc., and the modem processor mainly processes wireless communication signals, such as a baseband processor. It can be understood that the above modem processor may not be integrated into the processor 1010 either.
[0402] The embodiments of the present application also provide an electronic device, which includes Figure 31 the processor 1010, the memory 1009 shown, and a computer program stored on the memory 1009 and executable on the processor 1010. When the computer program is executed by the processor 1010, it implements each process of the text display method embodiment executed by the above electronic device and can achieve the same technical effect, which will not be elaborated here.
[0403] Figure 32 A schematic diagram of the hardware structure of a server provided for some embodiments of the present application. The server may be the above media server. As Figure 32 shown, the server 1600 may include: one or more processors 1601, a memory 1602, a communication interface 1603, and a bus 1604.
[0404] Among them, the processor 1601 is configured to receive a first voice from an electronic device through the bus 1604 and the communication interface 1603. The processor 1601 is further configured to convert the first voice into a first text, and extract first timbre feature information of the first voice, where the first timbre feature information is used to characterize the timbre of the first voice; and, based on the first timbre feature information, determine a first timbre identifier. The processor 1601 is further configured to send the first text and the first timbre identifier to the electronic device through the bus 1604 and the communication interface 1603.
[0405] In some embodiments of the present application, the processor 1601 is specifically configured to obtain the correspondence between timbre feature information and timbre identifiers; and, in the case where there is a third correspondence in the correspondence between timbre feature information and timbre identifiers, determine the timbre identifier in the third correspondence as the first timbre identifier; or, in the case where there is no third correspondence in the correspondence between timbre feature information and timbre identifiers, generate a first timbre identifier; where the third correspondence includes timbre feature information matching the first timbre feature information, and the timbre identifier corresponding to the timbre feature information matching the first timbre feature information.
[0406] In some embodiments of the present application, the processor 1601 is further configured to establish a correspondence between the first timbre identifier and the first timbre feature information after generating the first timbre identifier. The memory 1602 is further configured to store the correspondence between the first timbre identifier and the first timbre feature information.
[0407] In some embodiments of the present application, the processor 1601 is further configured to receive a second voice from an electronic device through the bus 1604 and the communication interface 1603 before determining the first timbre identifier based on the first timbre feature information. The processor 1601 is further configured to extract second timbre feature information of the second voice, where the second timbre feature information is used to characterize the timbre of the second voice; and, based on the second timbre feature information, generate a second timbre identifier; and, establish a correspondence between the second timbre feature information and the second timbre identifier. The memory 1602 is configured to store the correspondence between the second timbre feature information and the second timbre identifier. The processor 1601 is further configured to send the second timbre identifier to the electronic device through the bus 1604 and the communication interface 1603.
[0408] Some embodiments of the present application provide a server that can implement each process implemented by the text display method embodiment executed by the above server and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.
[0409] For the beneficial effects of various implementation manners in this embodiment, reference may specifically be made to the beneficial effects of the corresponding implementation manners in the text display method embodiment executed by the above server. To avoid repetition, it will not be elaborated here.
[0410] In some embodiments of the present application, one or more processors 1601, a memory 1602, and a communication interface 1603 are interconnected through a bus 1604. Among them, the bus 1604 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The bus 1604 may be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, Figure 32 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus. In addition, the server 1600 may further include some functional modules not shown here, which will not be elaborated here.
[0411] The embodiment of the present application further provides a server, including Figure 32 the processor 1601, the memory 1602 shown in the figure, and a computer program stored on the memory 1602 and executable on the processor 1601. When the computer program is executed by the processor 1601, it implements each process of the text display method embodiment executed by the above server and can achieve the same technical effects, which will not be elaborated here.
[0412] The embodiment of the present application further provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, it implements each process of the text display method embodiment and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.
[0413] Among them, the processor is the processor in the electronic device or the processor in the server described in the above embodiment. The readable storage medium includes a computer-readable storage medium, such as a computer read-only memory ROM, a random access memory RAM, a magnetic disk, or an optical disc, etc.
[0414] Another embodiment of the present application provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is configured to run programs or instructions to implement each process of the above text display method embodiment, and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.
[0415] It should be understood that the chip mentioned in the embodiments of the present application may also be referred to as a system-on-chip, system chip, chip system, or system-on-chip, etc.
[0416] The embodiments of the present application provide a computer program product, which is stored in a storage medium and is executed by at least one processor to implement each process of the above text display method embodiment, and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.
[0417] It should be noted that in this article, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element. In addition, it should be pointed out that the methods and devices in the embodiments of the present application are not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in the reverse order according to the functions involved. For example, the described methods may be performed in an order different from that described, and various steps may be added, omitted, or combined. Additionally, the features described with reference to certain examples may be combined in other examples.
[0418] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases, the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc) and includes several instructions for causing a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present application.
[0419] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative rather than restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them fall within the protection scope of the present application.
Claims
1. A text display method, characterized in that: The method comprises: Collecting a first voice and sending the first voice to a server; Receiving a first text and a first timbre identifier from the server, wherein the first text is a text converted from the first voice, the first timbre identifier corresponds to first timbre feature information, and the first timbre feature information is used to characterize the timbre of the first voice; Determine a first speaker identifier based on the first timbre identifier, where the first speaker indicated by the first speaker identifier is the speaker who utters the first speech; The first text and the first speaker identifier are displayed.
2. The method according to claim 1, characterized in that The determining a first speaker identifier based on the first timbre identifier includes: Obtaining a correspondence between a timbre identifier and a speaker identifier; In the case where a first corresponding relationship exists in the corresponding relationship between the timbre identifier and the speaker identifier, determining the speaker identifier in the first corresponding relationship as the first speaker identifier; or, in the case where the first corresponding relationship does not exist in the corresponding relationship between the timbre identifier and the speaker identifier, generating the first speaker identifier; The first corresponding relationship includes the first timbre identifier and the speaker identifier corresponding to the first timbre identifier.
3. The method according to claim 1, characterized in that: Before collecting the first voice and sending the first voice to the server, the method further includes: When the voice recognition setting interface is displayed, receiving a first input to the voice recognition setting interface; In response to the first input, collecting a second voice; When the second voice collection is completed, displaying the first window; receiving a second speaker identifier input by a user in the first window, wherein the second speaker indicated by the second speaker identifier is a speaker who utters the second speech; Sending the second voice to the server; A second timbre identifier is received from the server, and a corresponding relationship between the second timbre identifier and the second speaker identifier is established and stored, wherein the second timbre identifier corresponds to second timbre feature information, and the second timbre feature information is used to characterize the timbre of the second voice.
4. The method according to claim 2, characterized in that: After displaying the first text and the first speaker identifier, the method further includes: receiving a second input of the first speaker identifier; In response to the second input, displaying a second window; receiving a third speaker identifier input by a user in the second window; The displayed first speaker identification is modified to the third speaker identification.
5. The method according to claim 4, characterized in that The method further comprises: In the case where a second corresponding relationship exists in the corresponding relationship between the timbre identifier and the speaker identifier, the first speaker identifier in the second corresponding relationship is modified to the third speaker identifier; or, When there is no second corresponding relationship in the corresponding relationship between the timbre identifier and the speaker identifier, establishing and storing the corresponding relationship between the third speaker identifier and the first timbre identifier; The second corresponding relationship is a corresponding relationship between the first speaker identifier and the first timbre identifier.
6. A text display method, characterized in that: The method comprises: receiving a first voice from an electronic device; Converting the first speech into a first text, and extracting first timbre feature information of the first speech, where the first timbre feature information is used to characterize the timbre of the first speech; Determining a first timbre identifier based on the first timbre characteristic information; The first text and the first timbre identifier are sent to the electronic device.
7. The method according to claim 6, characterized in that The determining of a first timbre identifier based on the first timbre feature information includes: Obtaining the corresponding relationship between the timbre characteristic information and the timbre identifier; In the case where a third corresponding relationship exists in the corresponding relationship between the timbre feature information and the timbre identifier, determining the timbre identifier in the third corresponding relationship as the first timbre identifier; or, in the case where the third corresponding relationship does not exist in the corresponding relationship between the timbre feature information and the timbre identifier, generating the first timbre identifier; The third corresponding relationship includes the timbre feature information matching the first timbre feature information and the timbre identifier corresponding to the timbre feature information matching the first timbre feature information.
8. The method according to claim 7, characterized in that Before determining the first timbre identifier based on the first timbre feature information, the method further includes: receiving a second voice from the electronic device; Extracting second timbre feature information of the second speech, where the second timbre feature information is used to characterize the timbre of the second speech; generating a second timbre identifier based on the second timbre characteristic information; Establishing and storing a correspondence between the second timbre characteristic information and the second timbre identifier; The second timbre identifier is sent to the electronic device.
9. A text display device, characterized in that: The device includes a collection module, a sending module, a receiving module, a processing module and a display module; The collection module is used to collect the first voice; The sending module is used to send the first voice to the server; The receiving module is used to receive a first text and a first timbre identifier from the server, wherein the first text is a text obtained by converting the first voice, the first timbre identifier corresponds to first timbre feature information, and the first timbre feature information is used to characterize the timbre of the first voice; The processing module is configured to determine a first speaker identifier based on the first timbre identifier, wherein the first speaker indicated by the first speaker identifier is a speaker who utters the first speech; The display module is used to display the first text and the first speaker identifier.
10. The device according to claim 9, characterized in that The processing module is specifically used for: Obtaining a correspondence between a timbre identifier and a speaker identifier; In the case where a first corresponding relationship exists in the corresponding relationship between the timbre identifier and the speaker identifier, determining the speaker identifier in the first corresponding relationship as the first speaker identifier; or, in the case where the first corresponding relationship does not exist in the corresponding relationship between the timbre identifier and the speaker identifier, generating the first speaker identifier; The first corresponding relationship includes the first timbre identifier and the speaker identifier corresponding to the first timbre identifier.
11. The device according to claim 9, characterized in that The device also includes a storage module; The receiving module is further configured to receive a first input to the voice recognition setting interface when the voice recognition setting interface is displayed before collecting the first voice and sending the first voice to the server; The acquisition module is further used to acquire a second voice in response to the first input; The display module is further configured to display the first window when the second voice collection is completed; The receiving module is further configured to receive a second speaker identifier input by a user in the first window, wherein the second speaker indicated by the second speaker identifier is a speaker who utters the second voice; The sending module is used to send the second voice to the server; The receiving module is further used to receive a second timbre identifier from the server; The processing module is further used to establish a corresponding relationship between the second timbre identifier and the second speaker identifier, the second timbre identifier corresponds to second timbre feature information, and the second timbre feature information is used to characterize the timbre of the second voice; The storage module is used to store the corresponding relationship between the second timbre identifier and the second speaker identifier.
12. The device according to claim 10, characterized in that The receiving module is further configured to receive a second input of the first speaker identifier after displaying the first text and the first speaker identifier; The display module is further configured to display a second window in response to the second input; The receiving module is further configured to receive a third speaker identifier input by a user in the second window; The display module is further configured to modify the displayed first speaker identifier to the third speaker identifier.
13. The device according to claim 12, characterized in that The device also includes a storage module; The processing module is further configured to modify the first speaker identifier in the corresponding relationship between the timbre identifier and the speaker identifier to the third speaker identifier when a second corresponding relationship exists in the corresponding relationship between the timbre identifier and the speaker identifier; or The processing module is further configured to establish a correspondence relationship between the third speaker identifier and the first timbre identifier when there is no second correspondence relationship in the correspondence relationship between the timbre identifier and the speaker identifier; The storage module is used to store the corresponding relationship between the third speaker identifier and the first timbre identifier; The second corresponding relationship is a corresponding relationship between the first speaker identifier and the first timbre identifier.
14. A text display device, characterized in that: The device comprises a receiving module, a processing module and a sending module; The receiving module is used to receive a first voice from an electronic device; The processing module is used to convert the first speech into a first text, and extract first timbre feature information of the first speech, where the first timbre feature information is used to characterize the timbre of the first speech; and determine a first timbre identifier based on the first timbre feature information; The sending module is used to send the first text and the first timbre identifier to the electronic device.
15. The device according to claim 14, characterized in that The processing module is specifically used for: Obtaining the corresponding relationship between the timbre characteristic information and the timbre identifier; In the case where a third corresponding relationship exists in the corresponding relationship between the timbre feature information and the timbre identifier, determining the timbre identifier in the third corresponding relationship as the first timbre identifier; or, in the case where the third corresponding relationship does not exist in the corresponding relationship between the timbre feature information and the timbre identifier, generating the first timbre identifier; The third corresponding relationship includes the timbre feature information matching the first timbre feature information and the timbre identifier corresponding to the timbre feature information matching the first timbre feature information.
16. The device according to claim 15, characterized in that The device also includes a storage module; The receiving module is further configured to receive a second voice from the electronic device before determining the first timbre identifier based on the first timbre feature information; The processing module is further used to extract second timbre feature information of the second voice, the second timbre feature information is used to characterize the timbre of the second voice; and, based on the second timbre feature information, generate a second timbre identifier; and, establish a corresponding relationship between the second timbre feature information and the second timbre identifier; The storage module is used to store the corresponding relationship between the second timbre characteristic information and the second timbre identifier; The sending module is further used to send the second timbre identifier to the electronic device.