Speech processing method, apparatus and system, and storage medium and program product
By acquiring speech features and repairing the speech, more intelligible speech is generated, solving the problem of communication difficulties for people with special speech impairments, realizing convenient and accurate voice communication, and improving the user experience.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2026-04-02
AI Technical Summary
People with speech impairments have defects in pronunciation, resulting in low intelligibility of their speech, which is difficult for listeners to understand. Current technologies that use text input for communication are not convenient enough.
By receiving settings input, the system obtains the user's voice characteristics and repairs the voice when a voice repair scenario or contact is enabled, generating a more intelligible repaired voice. It supports voice-to-text conversion and control, provides options to enable and disable the voice repair function, and optimizes voice recording and recognition.
It improves the convenience and intelligibility of communication for people with special speech impairments, enhances user experience, reduces call conflicts, and improves the accuracy and flexibility of communication.
Smart Images

Figure CN2025097957_02042026_PF_FP_ABST
Abstract
Description
Speech processing method, device, system, storage medium and program product
[0001] The present application claims priority to the Chinese patent application No. 202410808992.4, filed on June 20, 2024, entitled "A method and device for assisting communication", and the Chinese patent application No. 202411097752.4, filed on August 9, 2024, entitled "Speech processing method, device, system, storage medium and program product", the contents of which are incorporated herein by reference in their entirety. TECHNICAL FIELD
[0002] The present application relates to the technical field of terminals, and in particular to a speech processing method, device, system, storage medium and program product. BACKGROUND
[0003] With the development of the technical field of terminals, more and more functions are applied to terminals, such as user communication functions applied to terminals.
[0004] At present, there are some users who cannot communicate barrier-free with normal people due to congenital and acquired factors. For example, special speech barrier groups such as hearing impaired, frozen shoulder, and damaged vocal cords have defects in speaking pronunciation, which leads to low intelligibility of their speech and makes it difficult for listeners to understand, which greatly affects their daily social interaction. Therefore, for special speech barrier groups, they need to input text in an electronic device in the communication scene, and then the electronic device can convert the user input text into speech and output.
[0005] However, for special speech barrier groups, the way of inputting text in an electronic device for communication exists the situation that communication is not convenient enough. SUMMARY
[0006] The present application provides a speech processing method, device, system, storage medium and program product, which is beneficial to improve the convenience of communication between users.
[0007] In a first aspect, an embodiment of the present application provides a speech processing method, which can include:
[0008] A setting input is received, the setting input is used to indicate a scenario in which speech repair is started, a contact in which speech repair is started, or an application in which speech repair is started, and the scenario includes a face-to-face communication scenario or a remote communication scenario. Then, a first speech feature of a user is obtained through sound registration. The order of sound registration and setting input can be exchanged. Then, when the user speaks, a first speech input by the user can be received. Since the setting input is received in advance, the first speech can be repaired according to the setting input and the first speech feature.
[0009] In the embodiment of the present application, the first speech is repaired according to the setting input and the first speech feature, and since the setting input is used to indicate the scenario in which the speech repair is enabled, the contact person for which the speech repair is enabled, or the application for which the speech repair is enabled, the first speech is repaired according to the first speech feature when it is detected that the electronic device is in the scenario in which the speech repair is enabled, or when the contact person communicated by the electronic device is the contact person for which the speech repair is enabled, or when the application running on the electronic device is the application for which the speech repair is enabled, so that the repaired first speech can be played in the first speech feature, and the intelligibility of the repaired first speech is higher than that of the first speech before repair. The intelligibility can represent the accuracy of the expression of the user who wants to express, and can also be understood as the understanding degree of the listener to the speech signal delivered by the loudspeaker.
[0010] In the embodiment of the present application, by receiving the setting input, the setting input is used to indicate the scenario in which the speech repair is enabled, the contact person for which the speech repair is enabled, or the application for which the speech repair is enabled, and then the first speech feature of the user is obtained through voice registration, so that after receiving the first speech input by the user, the first speech can be repaired according to the setting input and the first speech feature, so that the speech-impaired user or the hearing-impaired user can communicate through the input speech, thereby improving the convenience of the speech-impaired user to communicate. In addition, the repaired first speech is generated by the first speech feature registered by the user, so that the repaired first speech can be played according to the first speech feature, thereby playing the repaired first speech according to the user's own speech feature, which is closer to the user's own tone, etc., thereby further improving the user experience.
[0011] In a possible implementation, the remote communication scenario includes a call scenario, and the method further includes:
[0012] sending first prompt information to the first contact person in the call, so as to prompt that the speech repair function has been enabled. And / or, sending second prompt information to the first contact person in the case of receiving the first speech, so as to prompt that the first speech is being repaired.
[0013] In the embodiment of the present application, the first prompt information is sent to the first contact person in the call to prompt the user that the speech repair function has been enabled, so that the first contact person can know that the user has enabled the speech repair function, thereby improving the experience of the call. In addition, the second prompt information is sent to the first contact person in the case of receiving the first speech to prompt that the first speech is being repaired, so that the first contact person can know that the delay is caused by the speech repair, and the speaking time conflict of the parties or multiple parties in the call is reduced, thereby improving the experience of the call.
[0014] In a possible implementation, the second prompt information comprises a prompt sound, and the sending of the second prompt information to the first contact in the case of receiving the first voice can comprise:
[0015] The sending of the prompt sound to the first contact is continued in the case of starting to receive the first voice. And the sending of the prompt sound to the first contact is stopped in the case of starting to send the repaired first voice to the first contact.
[0016] In the embodiments of the present application, the sending of the prompt sound to the first contact is continued in the case of starting to receive the first voice, so that the first contact can know that the user at the other end is speaking when hearing the prompt sound, and the sending of the prompt sound to the first contact is stopped in the case of starting to send the repaired first voice to the first contact, so that the first contact can subsequently hear the repaired first voice, thereby reducing the situation of conflict between the two or more parties speaking during the call, thereby improving the experience of the two parties in the call.
[0017] In a possible implementation, the remote communication scenario comprises a call scenario, and the method further comprises:
[0018] The repaired first voice is sent to the first contact in the call. Then, the first interface can be displayed, the first interface being an interface for calling with the first contact, and the first interface comprising a first control for controlling the closing of the voice repair function. Then, in response to a first operation on the first control, the voice repair function is closed to send a second voice to the first contact, the second voice comprising the voice of the user received after the closing of the voice repair function.
[0019] In the embodiments of the present application, the closing of the voice repair can also be controlled during the call, so that the user can control the closing of the voice repair when the voice repair is not needed, thereby improving the switching flexibility of the use or non-use of the voice repair. For example, the user can control the closing of the voice repair function through the first control when the electronic device has low power or the electronic device runs slowly, thereby improving the closing of the voice repair function during the call and improving the user experience.
[0020] In a possible implementation, in response to the first operation, the first control is further switched from the first state to the second state, the first state being used to indicate that the voice repair function has been started, and the second state being used to indicate that the voice repair function has been closed, and the method further comprises:
[0021] In response to a second operation on the first control, the first control is then switched from the second state to the first state, and the voice repair function is started, so that a third voice is sent to the first contact, the third voice comprising the voice of the user received after the starting of the voice repair function.
[0022] In the embodiment of the present application, the state of the first control can indicate whether the voice repair function is currently enabled, thereby improving the accuracy of the user's selection of using or not using the voice repair function. Moreover, the embodiment of the present application can not only disable the voice repair function during the call, but also enable the voice repair function again during the call, thereby improving the flexibility of enabling or disabling the voice repair. In addition, the same control is used to enable or disable the voice repair function, thereby improving the simplicity of the interface during the call.
[0023] It should be understood that the WeChat application and the face-to-face communication application can also support the disabling or enabling of the voice repair function in the interface of the application. The related description of how to disable or enable the voice repair function in the call interface can be referred to, and details are not described herein.
[0024] In a possible implementation, the method further includes:
[0025] The second interface is displayed, and the second interface includes a playback control of the repaired first voice and a first text corresponding to the repaired first voice. Then, if the user selects a first target character in the first text, a third operation is received, and the third operation is used to select the first target character in the first text. In response to the third operation, at least one candidate character related to the first target character is displayed. Then, if the user selects one of the candidate characters, a fourth operation can be received, and the fourth operation is used to select a second target character in the at least one candidate character. In response to the fourth operation, a second text and a playback control of a voice corresponding to the second text are displayed, and the second text is obtained by replacing the first target character in the first text with the second target character.
[0026] In the embodiment of the present application, the second interface is displayed, and the second interface includes a playback control of the repaired first voice and a first text corresponding to the repaired first voice. A third operation is received, and the third operation is used to select a first target character in the first text. In response to the third operation, at least one candidate character related to the first target character is displayed. A fourth operation is received, and the fourth operation is used to select a second target character in the at least one candidate character. In response to the fourth operation, a second text and a playback control of a voice corresponding to the second text are displayed, thereby enabling the user to manually and quickly correct when the result of the voice repair is inaccurate, thereby improving the accuracy of the communication through the voice.
[0027] In a possible implementation, the method further includes:
[0028] The display of the playback control of the repaired first voice is cancelled.
[0029] In the embodiment of the present application, by canceling the display of the playback control of the repaired first voice, only the playback control of the latest voice is displayed each time, so that the convenience of selecting a suitable voice for reporting can be improved.
[0030] In a possible implementation, in the remote communication scenario, the method further includes:
[0031] The third interface is displayed, the third interface being an interface for communicating with the second contact, the third interface including a second control and a third control, the second control being used to instruct sending the first voice to the second contact, and the third control being used to instruct sending the repaired first voice to the second contact. Then, the repaired first voice can be sent to the second contact in response to an operation on the third control. Alternatively, the first voice can be sent to the second contact in response to an operation on the second control.
[0032] In the embodiment of the present application, the user can selectively send the repaired voice or the unrepaired voice to the second contact through the second control and the third control, that is, even if the voice repair function is enabled, the user can still select to send the original voice to the second contact, thereby improving the flexibility of the user in remote communication.
[0033] In a possible implementation, in the face-to-face communication scenario, the method further includes:
[0034] The fourth interface is displayed, the fourth interface including a virtual keyboard. Then, the user can operate the virtual keyboard, and the third text can be displayed in response to the operation on the virtual keyboard. Then, the user can operate the fourth control, and a fifth operation on the fourth control in the fourth interface can be received, the fourth control being used to instruct generating a voice. In response to the fifth operation, the voice corresponding to the third text is generated according to the first voice feature.
[0035] In the embodiment of the present application, the fourth interface is displayed, the fourth interface including a virtual keyboard. In response to the operation on the virtual keyboard, the third text is displayed. A fifth operation on the fourth control in the fourth interface is received, the fourth control being used to instruct generating a voice. In response to the fifth operation, the voice corresponding to the third text is generated according to the first voice feature, that is, the user can also input text and then generate a voice according to the registered first voice feature, that is, the conversion from text to voice is realized, so that the user can select to input a voice or input text for communication, and the selectability and flexibility of the user in communication are improved.
[0036] In a possible implementation, the third text includes a punctuation mark and / or an emoticon, and when the voice corresponding to the third text is generated according to the first voice feature, the tone of the voice corresponding to the third text can also be controlled according to the punctuation mark and / or the emoticon.
[0037] In the embodiment of the present application, the tone of the voice corresponding to the third text can be controlled through the punctuation marks and / or emoticons in the third text input by the user, thereby adaptively adjusting the tone of the output voice according to the user input text, thereby improving the flexibility of voice broadcast and improving the user experience.
[0038] In a possible implementation, the first voice feature of the user is obtained through voice registration, including:
[0039] The fifth interface is displayed, the fifth interface being a voice registration interface, and the fifth interface including third prompt information and a fifth control, the third prompt information being used to prompt the recording content of the voice registration, and the fifth control being used to instruct recording of a voice. If the user operates the fifth control, the fourth voice can be recorded in response to the operation on the fifth control. The first voice feature is extracted from the fourth voice.
[0040] In the embodiment of the present application, the recording of the voice is triggered through the fifth control, thereby improving the accuracy and effectiveness of voice recording, and the subsequent comparison between the recording content and the voice recognition result is also more accurate, thereby improving the accuracy of the selection of the repair model.
[0041] In a possible implementation, the method further includes:
[0042] The fourth voice is subjected to voice recognition to obtain a fourth text. Then, the fourth text is compared with the recording content to obtain a similarity between the fourth text and the recording content. Then, the target repair model can be selected from the plurality of repair models based on the similarity, the target repair model being used to repair the first voice, the voice repair capabilities of the plurality of repair models being different from each other, and the voice repair capability of the target repair model being negatively correlated with the similarity.
[0043] In the embodiment of the present application, the text obtained by voice recognition on the voice registered by the user is compared with the recording content to obtain a similarity, and then the target repair model is selected from the plurality of repair models based on the similarity, so that a suitable repair model can be selected according to the degree of speech disorder of the user for repair, thereby balancing the accuracy of voice repair and the computing resources required for voice repair.
[0044] In a second aspect, the embodiment of the present application further provides another voice processing method. The method can include:
[0045] The setting input is received, and the setting input is used to indicate a scenario in which voice repair is enabled, a contact in which voice repair is enabled, or an application in which voice repair is enabled. The scenario includes a face-to-face communication scenario or a remote communication scenario. Then, a fifth voice from a target contact can be received. When voice repair is needed, a second voice feature can be acquired, and the second voice feature is a preset voice feature or a voice feature extracted from the fifth voice. Then, the fifth voice is repaired according to the setting input and the second voice feature.
[0046] In the embodiments of the present application, the fifth voice is repaired according to the setting input and the second voice feature, and because the setting input is used to indicate a scenario in which voice repair is enabled, a contact in which voice repair is enabled, or an application in which voice repair is enabled, when it is detected that the electronic device is in the scenario in which voice repair is enabled, the contact with which the electronic device communicates is the contact in which voice repair is enabled, or the application that is running on the electronic device is the application in which voice repair is enabled, the fifth voice is repaired according to the second voice feature, so that the intelligibility of the repaired fifth voice is higher than the intelligibility of the fifth voice before repair. The intelligibility can represent the degree of understanding when the voice is listened to, and can also be understood as the accuracy of the voice expression contact intended to express.
[0047] In a possible implementation, in the remote communication scenario, the method further includes:
[0048] The sixth interface is displayed, the sixth interface is an interface for communication with the target contact, and the sixth interface includes the fifth voice. Then, if the user operates the fifth voice, the sixth control and the seventh control are displayed in response to the operation on the fifth voice, the sixth control is used to indicate conversion of the fifth voice into text, and the seventh control is used to indicate conversion of the repaired fifth voice into text. Then, the user can operate the sixth control or the seventh control, and the text corresponding to the repaired fifth voice is displayed in response to the operation on the seventh control. Alternatively, the text corresponding to the fifth voice is displayed in response to the operation on the sixth control.
[0049] In the embodiments of the present application, after the voice from the target contact is received, the fifth voice can be selectively converted into text, which can also be understood as conversion of the original voice of the target contact into text. The repaired fifth voice can also be selectively converted into text, and the user can select the displayed text according to needs, thereby improving the flexibility and experience of user communication.
[0050] In a possible implementation, in the face-to-face communication scenario, the method further includes:
[0051] The seventh interface is displayed, the seventh interface is an interface for communication with the target contact, and the seventh interface includes an eighth control. Then, the user can operate the eighth control, and in response to a sixth operation on the eighth control, the fifth text can be displayed. Then, the user can operate the eighth control again, and in response to a seventh operation on the eighth control, the display of the fifth text can be stopped, the fifth text including text corresponding to the fifth voice or the repaired text corresponding to the fifth voice in a target time period, and the target time period including a time period from the response to the sixth operation to the response to the seventh operation.
[0052] In the embodiments of the present application, the voice of the target contact can also be converted into text, thereby improving the flexibility of the communication mode between the user and the contact.
[0053] In a third aspect, the embodiments of the present application also provide another voice processing method, which can include:
[0054] The target voice and the voice feature are obtained. Then, the target voice and the voice feature are input into a target repair model, the target repair model is used to extract the content of the target voice, the voice continuous feature is obtained according to the content of the target voice and the voice feature, and the repaired target voice is synthesized according to the voice continuous feature. Then, the repaired target voice output by the target repair model can be obtained.
[0055] In the embodiments of the present application, after the target voice is input, the electronic device can obtain the target voice and the voice feature, then input the target voice and the voice feature into the target repair model, the target repair model is used to extract the content of the target voice, the voice continuous feature is obtained according to the content of the target voice and the voice feature, and the repaired target voice is synthesized according to the voice continuous feature, so that the speech-impaired population can realize communication by inputting voice, thereby improving the convenience of communication of the speech-impaired population.
[0056] In a possible implementation, the target repair model includes a first module, a second module and a third module, the first module is used to extract the content of the target voice, the second module is used to obtain the voice continuous feature according to the content of the target voice and the voice feature, and the third module is used to synthesize the repaired target voice according to the voice continuous feature.
[0057] In a possible implementation, inputting the target voice and the voice feature into the target repair model includes:
[0058] The voice feature and a first part of the target voice are input into the target repairing model, the first module is configured to extract the content of the first part of the voice, the second module is configured to obtain target voice continuous features according to the voice feature and the content of the first part of the voice, and the third module is configured to synthesize the repaired first part of the voice according to the target voice continuous features. Then, the voice feature and a second part of the target voice are input into the target repairing model, the first module is configured to extract the content of the second part of the voice, the second module is configured to obtain second voice continuous features according to the voice feature and the content of the second part of the voice, and the third module is configured to synthesize the repaired second part of the voice according to the second voice continuous features.
[0059] In the embodiment of the present application, by repairing a part of the target voice first and then repairing another part of the target voice, the voice repairing can be started when a part of the voice is obtained, that is, the voice repairing can be started without obtaining the complete target voice, thereby improving the efficiency of voice repairing.
[0060] In a possible implementation, the target repairing model further includes a fourth module, the fourth module is configured to perform discretization processing on the first part of the voice to obtain target voice discrete features, and the second module is further configured to perform prediction according to the target voice discrete features to obtain repaired target voice discrete features, and obtain the second voice continuous features according to the repaired target voice discrete features, the voice feature and the content of the second part of the voice.
[0061] In the embodiment of the present application, the fourth module performs discretization processing on the first part of the voice to obtain target voice discrete features, and then the second module performs prediction according to the target voice discrete features to obtain repaired target voice discrete features, and then obtains the second voice continuous features according to the repaired target voice discrete features, the voice feature and the content of the second part of the voice, so that the second voice continuous features can be obtained in combination with the target voice discrete features, the voice feature and the content of the second part of the voice, thereby improving the accuracy of the obtained second voice continuous features and further improving the accuracy of voice repairing.
[0062] In a possible implementation, the fourth module is configured to perform discretization processing on the first part of the voice to obtain target voice discrete features, including:
[0063] The fourth module is configured to perform discretization processing on the first part of the voice according to the content of the first part of the voice to obtain target voice discrete features.
[0064] In the embodiment of the present application, the first part of speech is discretely processed according to the content of the first part of speech to obtain the target speech discrete feature, that is, the content of the first part of speech is used as a reference for the discretization processing, thereby improving the accuracy of the obtained target speech discrete feature, and further improving the accuracy of speech repair.
[0065] It should be noted that the scheme of the embodiment of the present application can not only be used in sound repair tasks, but also can be extended to tasks such as dialect to standard Chinese and cross-language translation.
[0066] In a fourth aspect, another speech processing apparatus is provided, including a processor coupled with a memory, and configured to execute instructions in the memory to implement the method in any possible implementation manner of the first aspect. Optionally, the apparatus further includes the memory. Optionally, the apparatus further includes a communication interface, and the processor is coupled with the communication interface.
[0067] In a fifth aspect, a processor is provided, including an input circuit, an output circuit and a processing circuit. The processing circuit is configured to receive a signal through the input circuit and transmit a signal through the output circuit, so that the processor executes the method in any possible implementation manner of the first aspect.
[0068] In the specific implementation process, the processor can be a chip, the input circuit can be an input pin, the output circuit can be an output pin, and the processing circuit can be a transistor, a gate circuit, a flip-flop and various logic circuits, etc. The input signal received by the input circuit can be received and input by, for example but not limited to, a receiver, the signal output by the output circuit can be output to and transmitted by, for example but not limited to, a transmitter, and the input circuit and the output circuit can be the same circuit, which is used as the input circuit and the output circuit at different times. The specific implementation of the processor and various circuits is not limited in the embodiment of the present application.
[0069] In a sixth aspect, a processing apparatus is provided, including a processor and a memory. The processor is configured to read instructions stored in the memory, and can receive a signal through a receiver and transmit a signal through a transmitter to execute the method in any possible implementation manner of the first aspect.
[0070] Optionally, the processor is one or more, and the memory is one or more.
[0071] Optionally, the memory can be integrated with the processor, or the memory and the processor are separately arranged.
[0072] In a specific implementation process, the memory can be a non-transitory memory, for example, a read only memory (ROM), which can be integrated on the same chip as the processor, or can be separately arranged on different chips. The embodiments of the present application do not limit the type of memory and the arrangement mode of the memory and the processor.
[0073] It should be understood that the related data interaction process, for example, the sending of the indication information, can be a process of outputting the indication information from the processor, and the receiving of the capability information can be a process of receiving the input capability information by the processor. Specifically, the processed output data can be output to a transmitter, and the input data received by the processor can come from a receiver. The transmitter and the receiver can be collectively referred to as a transceiver.
[0074] The processing device in the fourth aspect described above can be a chip, and the processor can be implemented by hardware or software. When implemented by hardware, the processor can be a logic circuit, an integrated circuit, etc. When implemented by software, the processor can be a general-purpose processor, which is implemented by reading software codes stored in a memory. The memory can be integrated in the processor or exist independently of the processor.
[0075] In a seventh aspect, a computer program product is provided, which includes a computer program (also referred to as code or instructions), which, when executed, causes a computer to perform the method in any possible implementation manner of the first aspect described above.
[0076] In an eighth aspect, a computer readable storage medium is provided, which stores a computer program (also referred to as code or instructions), which, when executed on a computer, causes the computer to perform the method in any possible implementation manner of the first aspect described above. BRIEF DESCRIPTION OF DRAWINGS
[0077] FIG. 1 is a schematic diagram of a system architecture provided by an embodiment of the present application;
[0078] FIG. 2 is a schematic diagram of an interface for sound registration provided by an embodiment of the present application;
[0079] FIG. 3 is a schematic diagram of an interface for configuring an application using a sound repair function provided by an embodiment of the present application;
[0080] FIG. 4 is a schematic diagram of an interface for configuring a white list of sound repair provided by an embodiment of the present application;
[0081] FIG. 5 is a schematic diagram of an interface for a remote real-time communication scenario provided by an embodiment of the present application;
[0082] FIG. 6 is a schematic diagram of a call scenario according to an embodiment of the present application;
[0083] FIG. 7 is a schematic diagram of an interface of a remote non-real-time communication scenario according to an embodiment of the present application;
[0084] FIG. 8 is a schematic diagram of an interface of a face-to-face communication scenario according to an embodiment of the present application;
[0085] FIG. 9 is a schematic diagram of another interface of a face-to-face communication scenario according to an embodiment of the present application;
[0086] FIG. 10 is a schematic diagram of a voice processing method according to an embodiment of the present application;
[0087] FIG. 11 is a schematic diagram of a sound registration method according to an embodiment of the present application;
[0088] FIG. 12 is a schematic diagram of an architecture of a repair model according to an embodiment of the present application;
[0089] FIG. 13 is a schematic diagram of another voice processing method according to an embodiment of the present application;
[0090] FIG. 14 is a schematic diagram of another voice processing method according to an embodiment of the present application;
[0091] FIG. 15 is a schematic diagram of another voice processing method according to an embodiment of the present application;
[0092] FIG. 16 is a schematic diagram of a structure of a voice processing apparatus according to an embodiment of the present application. DETAILED DESCRIPTION
[0093] To facilitate clear description of the technical solutions of the embodiments of the present application, the following briefly introduces some terms and technologies involved in the embodiments of the present application:
[0094] 1. Electronic device
[0095] The electronic device in the embodiments of the present application can also be referred to as a terminal device, a user equipment (UE), a mobile station (MS), a mobile terminal (MT), an access terminal, a subscriber unit, a subscriber station, a mobile station, a mobile terminal, a remote station, a remote terminal, a mobile device, a user terminal, a terminal, a wireless communication device, a user agent, or a user apparatus, etc.
[0096] The electronic device can be a device that provides voice / data connectivity to a user, for example, a handheld device having a wireless connection function, a vehicle-mounted device, etc. Currently, some terminals are exemplified by a mobile phone, a tablet computer, a notebook computer, a palm computer, a mobile internet device (MID), a wearable device, a virtual reality (VR) device, an augmented reality (AR) device, a wireless terminal in industrial control, a wireless terminal in self driving, a wireless terminal in remote medical surgery, a wireless terminal in a smart grid, a wireless terminal in transportation safety, a wireless terminal in a smart city, a wireless terminal in a smart home, a cellular phone, a cordless phone, a session initiation protocol (SIP) phone, a wireless local loop (WLL) station, a personal digital assistant (PDA), a handheld device having a wireless communication function, a computing device or other processing device connected to a wireless modem, a vehicle-mounted device, a wearable device, an electronic device in a 5G network, or an electronic device in a future evolved public land mobile network (PLMN), etc., and the present embodiments are not limited thereto.
[0097] By way of example and not limitation, in the present embodiments, the electronic device can also be a wearable device. The wearable device can also be referred to as a wearable smart device, which is a general term for devices that are designed and developed by applying wearable technology to daily wear, such as glasses, gloves, watches, clothing, and shoes. The wearable device is a portable device that is directly worn on the body or integrated into a user's clothes or accessories. The wearable device is not only a hardware device, but also achieves powerful functions through software support and data interaction and cloud interaction. The general wearable smart device includes a full function, a large size, and can realize complete or partial functions without relying on a smart phone, such as a smart watch or smart glasses, and focuses on a certain application function and needs to cooperate with other devices such as a smart phone, such as various smart wristbands and smart jewelry for monitoring vital signs.
[0098] In addition, in the embodiments of the present application, the electronic device can also be an electronic device in an internet of things (IoT) system. The IoT is an important component of future information technology development, and its main technical feature is to connect objects through communication technology and network, so as to realize the intelligent network of man-machine interconnection and object-object interconnection. The electronic device of the present application can also be a vehicle-mounted unit, a vehicle-mounted module, a vehicle-mounted component, a vehicle-mounted chip or a vehicle-mounted unit built in as one or more components or units. The vehicle can implement the method of the present application through the built-in vehicle-mounted unit, vehicle-mounted module, vehicle-mounted component, vehicle-mounted chip or vehicle-mounted unit. Therefore, the embodiments of the present application can be applied to the Internet of Vehicles, such as vehicle to everything (V2X), long term evolution-vehicle (LTE-V), vehicle-to-vehicle (V2V), etc. Optionally, the electronic device of the embodiments of the present application can also be referred to as a device.
[0099] 2. Access network device:
[0100] The access network device in the embodiments of the present application can also be referred to as a wireless access network device, which can be a transmission reception point (TRP), an evolved NodeB (eNB or eNodeB) in an LTE system, a home evolved NodeB or home NodeB (HNB), a baseband unit (BBU), a wireless controller in a cloud radio access network (CRAN) scenario, or a relay station, an access point, a vehicle-mounted device, a wearable device, an access network device in a 5G network, or an access network device in a future evolved PLMN network, etc. It can be an access point (AP) in a WLAN, a gNB in a new radio (NR) system, a satellite base station in a satellite communication system, etc., and the embodiments of the present application are not limited.
[0101] The access network device in the present embodiment can include a centralized unit (CU) node, or a distributed unit (DU) node, or an access network device including a CU node and a DU node, or an access network device including a control plane CU node (CU-CP node) and a user plane CU node (CU-UP node) and a DU node. The access network device including the CU node and the DU node can split the protocol layers of the access network device, and the functions of part of the protocol layers are controlled by the CU, and the functions of the remaining part or all of the protocol layers are distributed in the DU and controlled by the CU. As an implementation manner, the CU deploys the radio resource control (RRC) layer, the packet data convergence protocol (PDCP) layer, and the service data adaptation protocol (SDAP) layer in the protocol stack. The DU deploys the radio link control (RLC) layer, the media access control (MAC) layer, and the physical layer (PHY) layer in the protocol stack. Therefore, the CU has the processing capability of the RRC, the PDCP, and the SDAP. The DU has the processing capability of the RRC, the PDCP, and the SDAP. The above-mentioned splitting of functions is only an example and does not constitute a limitation on the CU and the DU. That is, there can be other ways of splitting functions between the CU and the DU, which are not described herein. The functions of the CU can be implemented by one entity or by different entities. For example, the functions of the CU can be further split, for example, the control plane (CP) and the user plane (UP) are separated, that is, the control plane of the CU (CU-CP) and the user plane of the CU (CU-UP). For example, the CU-CP and the CU-UP can be implemented by different functional entities, and the CU-CP and the CU-UP can be coupled with the DU to jointly complete the functions of the access network device. In a possible manner, the CU-CP is responsible for the control plane function and mainly includes the RRC and the PDCP-C, where the PDCP-C is mainly responsible for the encryption and decryption of the control plane data, the integrity protection, the data transmission, and the like. The CU-UP is responsible for the user plane function and mainly includes the SDAP and the PDCP-U, where the SDAP is mainly responsible for processing the data of the core network device and mapping the data flow to a bearer. The PDCP-U is mainly responsible for the encryption and decryption of the data plane, the integrity protection, the header compression, the sequence number maintenance, the data transmission, and the like. The CU-CP and the CU-UP are connected through an E1 interface. The CU-CP represents the access network device to connect with the core network device through the interface between the core network device and the access network device.Through F1-C (control plane) and DU connection. CU-UP is connected through F1-U (user plane) and DU. In addition, there is a possible implementation that PDCP-C is also in CU-UP, which is not limited in the present application. The access network device may, for example, be a base station.
[0102] 3. Artificial intelligence (AI):
[0103] AI is a new technical science that studies and develops theories, methods, technologies and application systems for simulating, extending and expanding human intelligence. AI involves fields including, but not limited to, at least one of automatic speech recognition (ASR), text to speech (TTS), image recognition or natural language processing. ASR is a technology that converts speech into text, which recognizes the language content in the speech signal by processing and analyzing the speech signal, and converts it into readable text form. TTS is a technology that converts text information into speech output, which converts text information into a speech feature vector, and then converts the speech feature into an audio signal, and can also provide personalized pronunciation services for enterprises and individuals through tone selection, custom volume, speech rate. Speech can also be referred to as sound, audio, etc.
[0104] 4. Other terms
[0105] In the embodiments of the present application, the same items or similar items with basically the same functions and effects are distinguished by using "first", "second", etc. For example, the first voice and the second voice are only used to distinguish different voices, and do not limit the order. Those skilled in the art can understand that "first", "second", etc. do not limit the quantity and execution order, and "first", "second", etc. also do not necessarily mean different.
[0106] It should be noted that in the embodiments of the present application, "exemplary" or "for example" is used to represent an example, illustration or description. Any embodiment or design scheme described as "exemplary" or "for example" in the present application should not be interpreted as more preferred or more advantageous than other embodiments or design schemes. Rather, the use of "exemplary" or "for example" is intended to present the relevant concept in a specific manner.
[0107] In the embodiments of the present application, "at least one" means one or more, and "multiple" means two or more. The association relationship of the associated objects is described by "and / or", which means that there can be three kinds of relationships, for example, A and / or B, which can represent the following three cases: A exists alone, A and B exist together, and B exists alone, where A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after it. "At least one of the following" or similar expressions means any combination of these items, including single or multiple combinations. For example, at least one of a, b, or c can represent a, b, c, a-b, a-c, b-c, or a-b-c, where a, b, and c can be single or multiple.
[0108] Currently, there are some users who cannot communicate with normal people without obstacles due to congenital and acquired factors. For example, people with special speech barriers such as hearing impairment, amyotrophic lateral sclerosis, and damaged vocal cords have defects in speaking pronunciation, which makes their speech less understandable and difficult for listeners to understand, which greatly affects their daily social interaction. For the above-mentioned special speech barrier groups, in order to solve the problem of daily communication, a system for assisting communication needs to be provided. The special speech barrier group can also be referred to as a speech barrier user.
[0109] For the above-mentioned special speech barrier groups, a system for assisting language understanding is needed to solve the problem of daily communication. In related technologies, for the above-mentioned special speech barrier groups, text can be input, and then the electronic device converts the input text into speech for output, so that the special speech barrier group can also communicate through sound.
[0110] However, special speech barrier groups need to input text to communicate through sound, and inputting text is cumbersome, which makes it inconvenient for special speech barrier groups to communicate through sound. For special speech barrier groups, their situation is that their speech is less understandable and difficult for listeners to understand, not that they cannot speak.
[0111] Therefore, the embodiments of the present application propose a speech processing method, device, system, storage medium and program product, which can repair the speech input by the speech barrier user, that is, dysarthric speech reconstruction (DSR), thereby assisting the communication of the speech barrier user. DSR can convert the speech with lower intelligibility of the special speech barrier group into normal speech with higher intelligibility. Speech can also be referred to as sound, audio or sound source, etc. The speech processing method of the embodiments of the present application can convert difficult-to-understand speech barrier speech into speech with higher intelligibility, and keep the timbre consistent with the person. In the near-range communication scene and the remote communication scene, the sound repair function has high use value.
[0112] In a possible implementation, the speech repairing manner can be, for example, that the electronic device performs speech repairing processing (referred to as speech repairing or repairing for short) on the speech input by the speech-impaired user, to obtain repaired speech, the intelligibility of the repaired speech is higher than that of the input speech, and then the electronic device can output the repaired speech, so that a party communicating with the speech-impaired user can better understand the expression of the speech-impaired user.
[0113] In another possible implementation, after the repaired speech is obtained, the repaired speech can be subjected to speech recognition processing to obtain recognized text, the recognized text converted from the repaired speech can be understood as repaired recognized text, and then the electronic device can output the repaired recognized text, so that a party communicating with the speech-impaired user can better understand the expression of the speech-impaired user.
[0114] In another possible implementation, the speech repairing manner can be, for example, that the electronic device performs speech recognition processing on the input speech to obtain recognized text, and then performs repairing processing on the recognized text to obtain repaired recognized text, the intelligibility of the repaired recognized text is higher than that of the recognized text, so that a party communicating with the speech-impaired user can better understand the expression of the speech-impaired user.
[0115] Therefore, by means of the speech repairing manner, the speech-impaired user can communicate through input speech, the convenience of communication of the speech-impaired user is improved, and the communication experience of the speech-impaired user is improved.
[0116] Please refer to FIG. 1, which is a schematic diagram of a system architecture provided by an embodiment of the present application.
[0117] As shown in (a) of FIG. 1, the system architecture of the embodiment of the present application can include a first device 101, a second device 102, an access network device 103, and a cloud 104. The first device 101 and the second device 102 can communicate through the access network device 103, and at least one of the first device 101 and the second device 102 is a device used by a speech-impaired user. The cloud 104 has a sound repairing function and can repair input speech to obtain output speech. The cloud 104 can be, for example, a server.
[0118] In a possible implementation, the first device 101 can be a device used by the speech-impaired user, and after receiving the speech input by the speech-impaired user, the first device 101 sends the input speech to the cloud 104, the cloud 104 can repair the input speech, and then the cloud 104 sends the repaired speech to the first device 101, and the first device 101 sends the repaired speech to the second device 102 through the access network device 103, and the second device 102 can play the repaired speech after receiving the repaired speech.
[0119] In another possible implementation, the second device 102 can be a device used by the speech-impaired user, and after receiving the speech input by the speech-impaired user, the second device 102 sends the input speech to the first device 101, and the first device 101 sends the input speech to the cloud 104 after receiving the input speech, the cloud 104 can repair the input speech, and then the cloud 104 sends the repaired speech to the first device 101, and the first device 101 can play the repaired speech after receiving the repaired speech.
[0120] It should be understood that the above speech repair processing can also be performed on at least one of the first device 101 or the second device 102, so that the cloud 104 can also not be needed.
[0121] It should be understood that the above speech repair processing can also be performed on at least one of the first device 101 or the second device 102, so that the cloud 104 can also not be needed.
[0122] As shown in (b) of FIG. 1, the system architecture of the embodiment of the application can include the first device 101 and the cloud 104.
[0123] In a possible implementation, the first device 101 can be a device used by the speech-impaired user, and after receiving the speech input by the speech-impaired user, the first device 101 sends the input speech to the cloud 104, the cloud 104 can repair the input speech, and then the cloud 104 sends the repaired speech to the first device 101, and the first device 101 can output the repaired speech, or convert the repaired speech into text and then output.
[0124] It should be understood that the above speech repair processing can also be performed on at least one of the first device 101 or the second device 102, so that the cloud 104 can also not be needed.
[0125] It should be understood that the above speech repair processing can also be performed on at least one of the first device 101 or the second device 102, so that the cloud 104 can also not be needed.
[0126] The usage scenario one is a remote communication scenario. The remote communication scenario can include, but is not limited to, a real-time or non-real-time remote communication scenario. The real-time remote communication scenario can include, but is not limited to, a real-time call scenario. The non-real-time remote communication scenario can include, but is not limited to, a non-real-time remote voice communication scenario.
[0127] The usage scenario two is a short-range communication scenario. The short-range communication scenario can include, but is not limited to, a face-to-face communication scenario.
[0128] That is, in the embodiment of the present application, the voice repair capability can be integrated into the interaction in the call scenario and the face-to-face communication scenario. In general, the embodiment of the present application can be applied to the special speech-impaired population whose speech pronunciation is unclear, to realize real-time repair of the input impaired speech and play the repaired audio to the listener or transmit the repaired audio to the opposite device in the call. The usage scenarios are the remote communication scenarios such as calling and video calling, and the short-range communication scenarios such as face-to-face communication.
[0129] The remote communication scenario and the face-to-face communication scenario will be described below.
[0130] First, the remote communication scenario will be described below. The embodiment of the present application takes the remote real-time communication scenario as an example.
[0131] It can be understood that the term "interface" and "user interface" of the present application is a medium interface for interaction and information exchange between an application or an operating system and a user, which realizes the conversion between the internal form of information and the form acceptable by the user. The commonly used form of the user interface is a graphic user interface (GUI), which refers to a user interface related to computer operation displayed in a graphic manner. It can be an icon, a window, a control element, and the like interface elements displayed in the display screen of an electronic device, wherein the control element can include an icon, a button, a menu, a tab, a text box, a dialog box, a status bar, a navigation bar, a widget, and the like visible interface elements.
[0132] Please refer to FIG. 2, which is a schematic diagram of a sound registration interface provided by the embodiment of the present application.
[0133] The interface shown in (a) of FIG. 2 displays a page on which application icons are placed, which can include a plurality of application icons (e.g., at least one of a settings application icon 201, a weather application icon, a calendar application icon, a photo album application icon, an email application icon, or an application store application icon, etc.). A page indicator can also be displayed below the plurality of application icons to indicate the positional relationship between the currently displayed page and other pages. Below the page indicator are a plurality of application icons (e.g., a camera application icon, a contacts application icon, an information application icon, a dialer application icon), which remain displayed when the page is switched.
[0134] It can be understood that the application icons in the embodiments of the present application are icons for starting an application program or icons for jumping to a certain interface. For example, the camera application icon is an icon of a camera application program (i.e., a camera application), that is, the camera application icon can be used to trigger the starting of the camera application program. For another example, the settings application icon 201 can be used to trigger the jumping to a settings interface.
[0135] In the embodiments of the present application, the electronic device can detect a user operation acting on the settings application icon 201, and in response to the user operation, the electronic device can display a user interface as shown in (b) of FIG. 2. The interface shown in (b) of FIG. 2 can be, for example, an interface of the settings application, on which the user can implement at least one function of querying information of the electronic device or configuring the user.
[0136] It can be understood that the user operation mentioned in the present application can include, but is not limited to, touch (e.g., click, etc.), voice control, gesture, etc., and the present application does not limit this.
[0137] The interface shown in (b) of FIG. 2 displays a page including settings options. The page can include one or more settings options. Among them, the settings options in the page can include at least one of an accessibility settings option 202, a battery and performance settings option, a security and privacy settings option, a display and brightness settings option, or a network connection management option. The accessibility settings option 202 can be designed for users with special needs, and help users use the electronic device more conveniently through a series of assistive functions. The functions that can be set by the accessibility settings option 202 in the embodiments of the present application can include, but are not limited to, at least one function of sound optimization, screen reading, or color correction. The sound optimization can also be referred to as sound repair.
[0138] It should be noted that in the interface shown in (b) of FIG. 2, for the setting option for configuration, such as setting option 2, an indication of whether the setting option is turned on or off can also be included. Optionally, when the indication corresponding to the setting option is "off", the configuration corresponding to the setting option is not turned on, and when the indication corresponding to the setting option is "on", the configuration corresponding to the setting option is turned on. It should be understood that the indication of whether the setting option is turned on or off can be different indications, and is not limited to the specific "off" or "on" text style indications. By displaying the indication of whether the setting option is turned on or off on the interface, the user can quickly know whether the setting option is turned on, and the user experience is improved. For the setting option for querying information, the setting option can not include the indication of whether the setting option is turned on or off. In addition, as shown in (b) of FIG. 2, the interface can also include an indication that the option has a next level interface, such as the ">" symbol. By displaying the indication that the interface has a next level interface on the interface, the user can know that there is a next level interface through the indication, and the user experience is improved.
[0139] In the embodiment of the present application, the electronic device can detect a user operation acting on the accessibility setting option 202, and in response to the operation, the electronic device can display a user interface as shown in (c) of FIG. 2. The interface as shown in (c) of FIG. 2 can be, for example, the interface of the accessibility setting option 202, and the user can implement the setting function of the accessibility setting option 202 on the interface.
[0140] As shown in (c) of FIG. 2, the interface displays a page including an accessibility function. The page can include one or more controls of the accessibility function, such as the sound repair function control 203, which can be used to turn on or off the sound repair function, that is, the sound repair function control 203 can be understood as a master switch of the sound repair function. Optionally, the interface can also include an indication of whether the sound repair function is turned on or off, and through the indication of whether the sound repair function is turned on or off, it can be known whether the sound repair function is turned on. As shown in (c) of FIG. 2, the sound repair function is in an off state, that is, the sound repair function is in an unturned-on state. Optionally, as shown in (c) of FIG. 2, the interface can also include an indication that the sound repair function has a next level interface. Optionally, as shown in (c) of FIG. 2, the interface can also include a control to return to the previous level interface, such as a control to return to the interface as shown in (b) of FIG. 2.
[0141] In the embodiments of the present application, the electronic device can detect a user operation of the control 203 acting on the sound repair function, and in response to the operation, the electronic device can display a user interface as shown in (d) of FIG. 2. The interface as shown in (d) of FIG. 2 may, for example, be an interface of the sound repair function, and can be understood as a detailed setting page of the sound repair. The user can implement relevant configurations of the sound repair function on the interface, for example, at least one of sound recording of the sound repair function or application or scene configuration supported by the sound repair function.
[0142] The interface as shown in (d) of FIG. 2 includes a sound recording control 204 and a sound repair application configuration control 205, and the sound repair application configuration control 205 can include, but is not limited to, a call application configuration control 2051, a face-to-face communication application configuration control 2052, and a WeChat application configuration control 2053.
[0143] The sound recording control 204 can be used for recording of the user's voice. Optionally, the interface as shown in (d) of FIG. 2 can further include a sound recording condition indication mark. The sound recording condition indication mark is used to indicate whether there is sound recorded. Optionally, if the sound recording condition indication mark is "to be recorded", it means that the user has not recorded the sound; if the sound recording condition indication mark is "recorded", it means that the user has recorded the sound. Optionally, in the case where the user has recorded the sound, the sound can be further added, so as to increase the number of recorded sounds. It should be understood that the indication mark indicating the recorded sound and the indication mark indicating the unrecorded sound can be different, and are not limited to the above examples. The sound recording condition indication mark as shown in (d) of FIG. 2 can be known that there is no sound recorded. The embodiments of the present application indicate whether the user has recorded the sound through the sound recording condition indication mark, so that the user can quickly know whether the sound is recorded through the sound recording condition indication mark, thereby improving the user experience. In another possible implementation, the sound recording condition indication mark can not be displayed, so as to reduce the content displayed on the interface and improve the efficiency of the interface display.
[0144] The sound repair application configuration control 205 is configured to open or close the sound repair function. The sound repair application configuration control 205 can also be understood as a switch of the sound repair function supported application. When the sound repair application configuration control 205 is in the open sound repair state, the sound repair function in the corresponding application or scene is activated. As shown in the interface (d) in FIG. 2, the application or scene that can be configured with the sound repair function can include, but is not limited to, at least one of the call application, the face-to-face communication application, or the WeChat application. The WeChat application can be understood as a non-real-time remote communication application. The (d) in FIG. 2 can include the sound repair application configuration control 205 of each application, such as the sound repair application configuration control 205 of the call application, the sound repair application configuration control 205 of the face-to-face communication application, and the sound repair application configuration control 205 of the WeChat application. Optionally, when the sound repair application configuration control 205 of an application or scene is in the closed sound repair state, the sound repair function of the application or scene cannot be used; when the sound repair application configuration control 205 of an application or scene is in the open sound repair state, the sound repair function of the application or scene can be used. The embodiments of the present application can configure each application or scene with the sound repair application configuration control 205, that is, the sound repair function of each application or scene can be individually opened or closed. In this way, the user can select to open or close the sound repair function of the application or scene according to the needs, thereby improving the user experience. In another possible implementation, the sound repair function of multiple applications or scenes can be opened or closed through one sound repair application configuration control 205. In this way, multiple applications or scenes can be opened or closed with one key, thereby improving the efficiency of opening or closing the sound repair function of multiple applications or scenes. In another possible implementation, the application or scene that supports the sound repair function configuration and the corresponding sound repair application configuration control 205 can not be displayed, that is, all applications or scenes can be opened by default. In this way, the content displayed on the interface can be reduced, and the efficiency of the interface display can be improved.
[0145] For example, as shown in the sound repair application configuration control 205 in (d) in FIG. 2, the sound repair function is in the closed state, that is, the sound repair function of each application cannot be used.
[0146] It should be noted that the sound recording and the application or scene configuration supported can be mutually constrained, and the constraint can be that one of the sound recording and the application or scene configuration supported can be completed before the other one of the sound recording and the application or scene configuration supported is operated. For example, the sound recording can be completed before the application or scene supported by the sound repair function is configured, that is, if the sound recording is not completed, any application or scene cannot start the sound repair function through the sound repair application configuration control 205. According to the embodiment of the present application, the sound recording and the application or scene configuration supported can be mutually constrained, and the user can be prompted to perform the sound recording and the configuration of the application or scene in which the sound repair function is started, thereby improving the user experience.
[0147] In the embodiment of the present application, the electronic device can detect a user operation acting on the sound recording control 204, and in response to the operation, the electronic device can display an interface as shown in (e) of FIG. 2. In the embodiment of the present application, since the sound optimization function needs the user to register the sound, after the sound repair is started, the interface for sound registration can be automatically jumped to, for example, the interface shown in (e) of FIG. 2 can be a kind of sound registration interface. The interface shown in (e) of FIG. 2 can include a new voiceprint control 206. The new voiceprint control 206 can be used to trigger the collection of voiceprint features (referred to as voiceprint for short). The voiceprint can be used to indicate the timbre of the registered user's voice. Optionally, the interface shown in (e) of FIG. 2 can also include a control for returning to the previous interface, for example, a control for returning to the interface shown in (d) of FIG. 2.
[0148] Then, the electronic device can detect a user operation acting on the new voiceprint control 206, and in response to the operation, the electronic device can display an interface as shown in (f) of FIG. 2. The interface shown in (f) of FIG. 2 can include a sound recording content 207, a recording control 208, and a save control. Among them, the sound recording content 207 can indicate the content that the user needs to record, for example, the sound recording content 207 can indicate the text content "Giant pandas do not have a fixed sleep time, they sleep wherever they go". The recording control 208 can be used to trigger sound collection. Optionally, it can be that the sound collection is triggered by long pressing the recording control 208, and the sound collected by the electronic device is the sound entered by the user during the long pressing of the recording control 208; or the user clicks the recording control 208 to trigger the sound collection, and then the electronic device starts to collect the sound, and the user clicks the recording control 208 again to stop the sound collection. The save control is used to include the collected sound. Optionally, the user can select to record the sound in segments, and when the electronic device detects a user operation acting on the save control, the sound recorded in segments is spliced in the order of the recording time to obtain a recording.
[0149] In the embodiment of the present application, the electronic device detects a user operation acting on the recording control 208, and in response to the operation, starts collecting the voice of the user.
[0150] Then, as shown in (g) of FIG. 2, the electronic device detects a user operation acting on the save control, and in response to the operation, saves the collected voice and extracts the voiceprint of the collected voice. It should be understood that the process of extracting the voiceprint can be performed at any time period after the voice is collected, and is not limited to being performed immediately after the voice is saved.
[0151] Optionally, the electronic device can compare the content of the recording with the voice recording content 207, and if the similarity between the content of the recording and the voice recording content 207 is higher than a similarity threshold, the voiceprint of the recording is taken as the voiceprint of the subsequent broadcast voice, and if the similarity between the content of the recording and the voice recording content 207 is not higher than the similarity threshold, the user is prompted to re-record the voice according to the guidance of the voice recording content 207. Optionally, the electronic device in the embodiment of the present application supports recording 1-n sentences of text, that is, after recording the text content as shown in (g) of FIG. 2, another segment of text content can be displayed for the user to record, and n is an integer not less than 2. After recording is completed, the user can save through the save control. It should be understood that the user can also choose to save after recording m sentences of text, 1≤m<n, and m is an integer.
[0152] Then, as shown in (h) of FIG. 2, the electronic device detects a user operation acting on the control for returning to the previous interface, and in response to the operation, the interface as shown in (i) of FIG. 2 can be displayed. In the interface as shown in (i) of FIG. 2, the voice recording status indication is "recorded". The similar parts of (i) of FIG. 2 and (d) of FIG. 2 can refer to the description of (d) of FIG. 2, which will not be repeated here.
[0153] In another possible implementation, the interface shown in (f) in FIG. 2 can also not include the sound recording content 207, so that the content displayed by the interface can be reduced, and the efficiency of the interface display can be improved. By displaying the sound recording content 207, the embodiments of the present application can guide the user to record the sound according to the sound recording content 207 when recording the sound, so that the effectiveness of the sound recording can be determined by the electronic device, and the accuracy of the voiceprint extraction can be improved. In addition, the content of the collected sound can be compared with the sound recording content 207 of the interface shown in (f) in FIG. 2, and then a suitable repair model can be selected according to the similarity of the comparison to repair the subsequent user input voice. The similarity of the comparison can also be referred to as intelligibility. The greater the similarity of the comparison, the more similar the content of the collected sound is to the sound recording content 207 of the interface shown in (f) in FIG. 2, and the less the impairment of the user's speaking ability. That is, the smaller the similarity of the comparison, the less similar the content of the collected sound is to the sound recording content 207 of the interface shown in (f) in FIG. 2, and the greater the impairment of the user's speaking ability. Generally speaking, the more complex the architecture of the repair model, the stronger the voice repair capability of the repair model, and the more accurate the voice repair, but the more computing resources required by the electronic device. Therefore, when selecting a suitable repair model according to the similarity of the comparison, the smaller the similarity of the comparison, the more complex the architecture of the selected repair model, and the greater the similarity of the comparison, the simpler the architecture of the selected repair model. For example, when the similarity of the comparison is a first similarity, a first repair model is selected; when the similarity of the comparison is a second similarity, a second repair model is selected; wherein the first similarity is higher than the second similarity, and the complexity of the model architecture of the first repair model is lower than the complexity of the second repair model, and accordingly the capability of the first repair model is lower than the capability of the second repair model.
[0154] In another possible implementation, the interface shown in (f) in FIG. 2 can also not include the sound recording content 207, so that the content displayed by the interface can be reduced, and the efficiency of the interface display can be improved. By displaying the sound recording content 207, the embodiments of the present application can guide the user to record the sound according to the sound recording content 207 when recording the sound, so that the effectiveness of the sound recording can be determined by the electronic device, and the accuracy of the voiceprint extraction can be improved. In addition, the content of the collected sound can be compared with the sound recording content 207 of the interface shown in (f) in FIG. 2, and then a suitable repair model can be selected according to the similarity of the comparison to repair the subsequent user input voice. The similarity of the comparison can also be referred to as intelligibility. The greater the similarity of the comparison, the more similar the content of the collected sound is to the sound recording content 207 of the interface shown in (f) in FIG. 2, and the less the impairment of the user's speaking ability. That is, the smaller the similarity of the comparison, the less similar the content of the collected sound is to the sound recording content 207 of the interface shown in (f) in FIG. 2, and the greater the impairment of the user's speaking ability. Generally speaking, the more complex the architecture of the repair model, the stronger the voice repair capability of the repair model, and the more accurate the voice repair, but the more computing resources required by the electronic device. Therefore, when selecting a suitable repair model according to the similarity of the comparison, the smaller the similarity of the comparison, the more complex the architecture of the selected repair model, and the greater the similarity of the comparison, the simpler the architecture of the selected repair model. For example, when the similarity of the comparison is a first similarity, a first repair model is selected; when the similarity of the comparison is a second similarity, a second repair model is selected; wherein the first similarity is higher than the second similarity, and the complexity of the model architecture of the first repair model is lower than the complexity of the second repair model, and accordingly the capability of the first repair model is lower than the capability of the second repair model.
[0155] In another possible implementation, the interface shown in (f) in FIG. 2 can also not include the sound recording content 207, so that the content displayed by the interface can be reduced, and the efficiency of the interface display can be improved. By displaying the sound recording content 207, the embodiments of the present application can guide the user to record the sound according to the sound recording content 207 when recording the sound, so that the effectiveness of the sound recording can be determined by the electronic device, and the accuracy of the voiceprint extraction can be improved. In addition, the content of the collected sound can be compared with the sound recording content 207 of the interface shown in (f) in FIG. 2, and then a suitable repair model can be selected according to the similarity of the comparison to repair the subsequent user input voice. The similarity of the comparison can also be referred to as intelligibility. The greater the similarity of the comparison, the more similar the content of the collected sound is to the sound recording content 207 of the interface shown in (f) in FIG. 2, and the less the impairment of the user's speaking ability. That is, the smaller the similarity of the comparison, the less similar the content of the collected sound is to the sound recording content 207 of the interface shown in (f) in FIG. 2, and the greater the impairment of the user's speaking ability. Generally speaking, the more complex the architecture of the repair model, the stronger the voice repair capability of the repair model, and the more accurate the voice repair, but the more computing resources required by the electronic device. Therefore, when selecting a suitable repair model according to the similarity of the comparison, the smaller the similarity of the comparison, the more complex the architecture of the selected repair model, and the greater the similarity of the comparison, the simpler the architecture of the selected repair model. For example, when the similarity of the comparison is a first similarity, a first repair model is selected; when the similarity of the comparison is a second similarity, a second repair model is selected; wherein the first similarity is higher than the second similarity, and the complexity of the model architecture of the first repair model is lower than the complexity of the second repair model, and accordingly the capability of the first repair model is lower than the capability of the second repair model.
[0156] It should be understood that the interface of the embodiments of the present application can be increased or reduced as needed, and is not limited to the number of interfaces and the style of the interface shown in FIG. 2, which is not limited herein.
[0157] In the embodiments of the present application, by allowing the user to register the voice, the voiceprint of the user can be extracted from the collected voice, so that the repaired voice can be played in the user's voiceprint, that is, the user can select a specific voiceprint to play the repaired voice as needed, thereby improving the user experience.
[0158] In another possible implementation, the voice recording process of FIG. 2 can also not be required, so that the electronic device can use the built-in voiceprint to play the repaired voice, and the ease of communication is also improved due to the reduction of the voice recording process.
[0159] Next, the configuration of the application using the voice repair function is exemplified.
[0160] Referring to FIG. 3, FIG. 3 is a schematic diagram of an interface for configuring an application using a voice repair function according to an embodiment of the present application.
[0161] As shown in the interface of (a) in FIG. 3, the interface includes a voice recording control 204 and a voice repair application configuration control 205. The interface shown in (a) in FIG. 3 can refer to the description of the interface shown in (i) in FIG. 2, which is not repeated here.
[0162] In the embodiments of the present application, the electronic device can detect a user operation acting on the voice repair application configuration control 205, for example, detect a user operation acting on the call application configuration control 2051, and in response to the operation, the interface shown in (b) in FIG. 3 can be displayed. In the interface shown in (b) in FIG. 3, the call application configuration control 2051 in the voice repair application configuration control 205 switches from the closed voice repair state to the open voice repair state. For example, in the interface shown in (b) in FIG. 3, that is, the call application configuration control 2051 is in the open voice repair state, the call application can use the voice repair function; and the face-to-face communication application configuration control 2052 and the WeChat application configuration control 2053 are both in the closed voice repair state, so the face-to-face communication application and the WeChat application cannot use the voice repair function, but can control the opening or closing of the voice repair function through a control in the interface of the face-to-face communication application or the WeChat application when using the face-to-face communication application or the WeChat application. In the embodiments of the present application, optionally, the face-to-face communication application, the WeChat application and the call application can control the opening or closing of the voice repair function through the control in the interface corresponding to each application, and the present embodiment does not describe how to control the opening or closing of the voice repair function through the control in the interface of the application, which will be further described in the following embodiments.
[0163] It should be noted that the electronic device can detect the user operation on the sound repair application configuration control 205, and the sound repair application configuration control 205 can also switch from the sound repair enabled state to the sound repair disabled state.
[0164] In the embodiment of the present application, the sound repair control is configured for each application, and the sound repair function of each application can be individually controlled, and the user can select the application whose sound repair function is enabled or disabled according to the need, thereby improving the user experience.
[0165] In the interface shown in FIG. 3, the sound repair function of the call application has been enabled. Therefore, the following embodiment is exemplified in the call scenario.
[0166] FIG. 4 is a schematic diagram of an interface for configuring a white list of sound repair according to an embodiment of the present application. In the embodiment of the present application, for the special speech-impaired population, there can be some friends who are used to understanding their pronunciation; for these friends, as shown in FIG. 4, the special speech-impaired user can set a white list in the phone address book to make the sound repair not enabled by default when talking to these friends.
[0167] As shown in (a) of FIG. 4, a page on which an application icon is placed is displayed, and the page can include a plurality of application icons (for example, a phonebook application 401). The interface shown in (a) of FIG. 4 can refer to the description of the interface shown in (a) of FIG. 2, and will not be repeated here.
[0168] In the embodiment of the present application, the electronic device can detect the user operation on the phonebook application 401, and in response to the operation, the electronic device can display the interface shown in (b) of FIG. 4. The interface shown in (b) of FIG. 4 can be, for example, the main interface of the phonebook. The interface shown in (b) of FIG. 4 can include at least one of a contact configuration control 403, a contact search bar, or a contact list. The contact configuration control 403 is used to configure the same operation for the selected contact, for example, adding the selected contact to the sound repair white list, or for example, enabling the call recording function by default when talking to the selected contact. The contact search bar can be used to quickly search for contacts. The contact list includes one or more contacts. Optionally, the contacts in the contact list can be classified according to certain rules, for example, classified according to the first letter of the contact name, and the interface shown in (b) of FIG. 4 further includes a control for quickly locating the classification. Optionally, the interface shown in (b) of FIG. 4 can further include a dialing control and a favorite control. Optionally, the interface shown in (b) of FIG. 4 can further include a contact adding control.
[0169] Then, as shown in (c) of FIG. 4, the electronic device can detect a user operation acting on the interface, and in response to the operation, the electronic device can display a contact selection control 402. The contact selection control 402 can be used to select a contact to add the selected contact to the whitelist. In the embodiment of the present application, each contact corresponds to a contact selection control 402. As shown in (c) of FIG. 4, the contact selection control 402 of the contact "B2" is in an unselected state.
[0170] In the embodiment of the present application, the electronic device can detect a user operation acting on the contact selection control 402, and in response to the operation, the electronic device can display an interface as shown in (d) of FIG. 4. In the interface shown in (d) of FIG. 4, the acted contact selection control 402 is in a selected state, for example, the contact selection control 402 of the contact "B2" is in a selected state.
[0171] Then, the electronic device can detect a user operation acting on the contact configuration control 403, and in response to the operation, the electronic device can display an interface as shown in (e) of FIG. 4. The contact configuration control 403 of the embodiment of the present application can be used to configure the selected contact. In the interface shown in (e) of FIG. 4, a sound repair whitelist option 404 and a call recording option can also be included. The call recording option is used to set the contact for which the contact selection control 402 is in a selected state as a contact for saving call recording. The sound repair whitelist option 404 is used to add the contact for which the contact selection control 402 is in a selected state to the sound repair whitelist, that is, to add the selected contact to the sound repair whitelist.
[0172] Then, as shown in (f) of FIG. 4, the electronic device can detect a user operation acting on the sound repair whitelist option 404, and then the electronic device can add the contact for which the contact selection control 402 is in a selected state to the sound repair whitelist. Optionally, the sound repair whitelist of the embodiment of the present application is a list of not starting sound repair, that is, the contact in the sound repair whitelist does not start the sound repair function when making a call; or the sound repair whitelist is a list of starting sound repair, that is, the contact in the sound repair whitelist starts the sound repair function when making a call.
[0173] The embodiment of the present application can improve user experience by configuring the sound repair whitelist. The user can select the contact to use the sound repair function or not to use the sound repair function in the call scenario according to the actual situation, so that the user can flexibly adjust the contact for which the sound repair is needed or not needed in the call scenario, thereby improving the user experience.
[0174] It should be noted that the sound repair whitelist can not be configured, and thus sound repair can be performed or not performed for all contacts.
[0175] After the sound repair whitelist is configured, a call can be initiated with one of the contacts. For ease of understanding, the following embodiments are described by taking a contact in the sound repair whitelist as an example, which is not subjected to sound repair.
[0176] The embodiments of the present application can use the sound repair function or not use the sound repair function for different contacts by setting the sound repair whitelist, and thus the use flexibility of the sound repair function is improved.
[0177] Referring to FIG. 5, FIG. 5 is an interface diagram of a remote real-time communication scenario provided by an embodiment of the present application. The embodiments of the present application take a contact in the sound repair whitelist as a contact by default which is not subjected to sound repair.
[0178] As shown in the interface of (a) in FIG. 5, a page on which an application icon is placed is displayed, and the page can include a plurality of application icons (for example, a telephone application 501). The interface shown in (a) in FIG. 5 can refer to the description of the interface shown in (a) in FIG. 2, and thus is not described herein.
[0179] In the embodiment of the present application, the electronic device can detect a user operation acting on the telephone application 501, and in response to the operation, the interface as shown in (b) of FIG. 5 can be displayed. The interface as shown in (b) of FIG. 5 displays the user's recent call records, which can include call contacts and corresponding call times. In the interface as shown in (b) of FIG. 5, a plurality of contacts can be included. Then, the electronic device can detect a user operation acting on one of the contacts, for example, a user operation acting on the contact "B2", and in response to the operation, a call with the contact "B2" can be initiated. As shown in (d) of FIG. 4, since the contact "B2" is added to the sound repair whitelist, that is, the contact "B2" is configured to not use the sound repair function, when the call with the contact "B2" is initiated, the call is made by default in the original sound, that is, the call is made with the repaired voice, as shown in (c) of FIG. 5. Optionally, in the interface as shown in (c) of FIG. 5, the sound switching control 502, the contact's avatar of the call, the sound shielding control, the keyboard calling control, the sound external playing control and the end call control are also included. The sound switching control 502 as shown in (c) of FIG. 5 is in the original sound playing state, that is, the original sound collected by the end is played at the other end of the call. The sound switching control 502 can be used to control the switching between the original sound and the repaired voice. As shown in (d) of FIG. 5, the electronic device can detect a user operation acting on the sound switching control 502, and in response to the operation, the sound switching control 502 of the electronic device is in the repaired voice playing state, that is, the repaired sound is played at the other end of the call, and the contact "B2" of the call hears the repaired sound. In the embodiment of the present application, the repaired sound can be obtained by repairing the original sound. The original sound can also be referred to as the original pronunciation, and the original pronunciation can be the collected sound uttered by the user.
[0180] Optionally, if the electronic device detects again a user operation acting on the sound switching control 502, the state of the sound switching control 502 switches to the original sound playing state, and the corresponding user makes the call by the original sound. In the embodiment of the present application, by setting the sound switching control 502 in the call interface, the user can select to use the original sound or the repaired sound through the sound switching control 502, which can improve the user experience.
[0181] It should be noted that if the electronic device initiates a call with a contact outside the sound repair whitelist, the call interface defaults to enable the sound repair function.
[0182] It should be noted that if the call is initiated by the electronic device at the other end, the electronic device at the current end can also display the interface shown in (c) of FIG. 5 or (d) of FIG. 5, so that the current end can control whether to play the original sound collected by the other end or the sound after repairing the original sound collected by the other end through the sound switching control 502.
[0183] In another possible implementation, if the contact of the call is a contact outside the sound repairing whitelist, that is, the contact of the call is configured to turn on the sound repairing function, the sound switching control 502 is in the repaired voice playing state when the call is connected.
[0184] In another possible implementation, the sound switching control 502 can also not be needed, that is, according to the configuration of the sound repairing whitelist, the call is made using the original sound or the repaired voice.
[0185] In some example cases, when the electronic device repairs the sound, a certain time is needed, so there is a certain time delay from when the user at the current end speaks to when the user at the other end hears the speaking content of the user at the current end. The time delay is related to the time needed by the electronic device to repair the sound, so that the user at the other end needs a long time to receive the repaired voice of the user at the current end, and the user at the other end may think that the user at the current end does not speak or that the signal is not good enough for the user at the other end to hear.
[0186] It should be noted that the sound transmitted by the contact in the sound repairing whitelist can also be repaired and played back when the call is made with the contact in the sound repairing whitelist, that is, the sound transmitted by the other end is repaired and played back by the current end device, for example, the sound of the contact "B2" is repaired and played back.
[0187] Please refer to FIG. 6, which is a call scenario provided by an embodiment of the present application. In the embodiment of the present application, there are a current end user and an opposite end user, the electronic device used by the current end user is referred to as a current end device, and the electronic device used by the opposite end user is referred to as an opposite end device. The current end device can also be referred to as a first electronic device, and the opposite end device can also be referred to as a second electronic device. The current end user can also be referred to as a first user, and the opposite end user can also be referred to as a second user. In the embodiment of the present application, the call audio after the sound repairing model will increase the call time delay, which can affect the call effect of multiple parties. Therefore, the embodiment of the present application is described in the call scenario to improve the call effect of multiple parties.
[0188] As shown in (a) of FIG. 6, if the voice input by the local user is "hello", the local device needs to perform the repairing, and then sends the repaired voice "hello" to the remote device. The remote device receives the repaired voice "hello" and plays it. Thus, the repaired voice can be heard by the remote user. However, it takes a long time from the voice input by the local user to the repaired voice heard by the remote user, because the local device needs a certain time to perform the voice repairing, i.e., processing delay. Then, the remote device can send the voice input by the remote user "hello, your takeout, I am downstairs" to the local device, and the local device plays the voice input by the remote user. At this time, if the local user wants to reply, the local user can input the voice "thank you, it can be placed downstairs". At this time, the local device needs to process the voice, but it takes a certain time in the processing, i.e., processing delay. At this time, the remote user finds that the local user does not reply for a long time, and thus can think that the local user does not hear the voice, and can repeat "hello, your takeout, I am downstairs". However, the local user has heard the voice, but it takes a certain time to repair, and thus the remote user can not hear the feedback of the local user for a long time.
[0189] Therefore, in a possible implementation, the local device can send a prompt to the remote device during the call, and the remote device plays the prompt. Thus, the remote user can know that the local device needs to process the voice input by the local user, so that the remote user can know that the local device needs to repair the voice.
[0190] As shown in (b) of FIG. 6, after the local device establishes the call channel with the remote device, the local device can send a prompt to the remote device. The remote device receives the prompt and plays it. Thus, the remote user can hear the prompt. The prompt can be used to prompt that the local device needs to repair the voice input by the local user. For example, the prompt can be "the opposite party has enabled the sound optimization function, the call may have delay, please wait for the prompt to end and then reply".
[0191] It should be understood that the local device can send the prompt to the remote device immediately after establishing the call channel with the remote device, or after a certain time, which is not limited herein.
[0192] In the embodiments of the present application, the local device can send the prompt to the remote device, so as to improve the experience of the two parties in the call.
[0193] As shown in (b) of FIG. 6, although the local device can send the prompt to the opposite device, in the process of communication, there is still a case that the opposite user thinks that the local user does not hear his voice.
[0194] In the embodiments of the present application, in a call scenario, such as a mobile phone call, a video call scenario, the voice changing function, the audio repairing function and the like are started through the switch of the call interface. After being started, the receiving party receives a corresponding prompt to prompt that the sound heard by the opposite party is optimized and a time delay is caused to improve the experience of the call between the two parties.
[0195] Therefore, in a possible implementation, the local device sends a prompt sound to the opposite device after receiving each piece of voice input by the local user, and then the opposite device can broadcast the prompt sound to prompt that the local user has input voice.
[0196] As shown in (c) of FIG. 6, for example, the local device can send a prompt to the opposite device after establishing a call channel with the opposite device. Then, after detecting that the local user inputs the voice of "Hello, you", the local device sends a prompt sound to the opposite device, and then the opposite device broadcasts the prompt sound. Then, after detecting that the local user inputs the voice of "Thank you, it can be placed on the ground floor", the local device sends a prompt sound to the opposite device, and then the opposite device broadcasts the prompt sound. Then, the opposite user can know from the prompt sound that the local user has input voice, and then waits until the voice of "Thank you, it can be placed on the ground floor" is received to reply the voice of "OK, goodbye". For example, the prompt sound can be a prompt sound such as "ding", "drop" or "tick". In the call scenario, the AI model converts the sound to cause a call delay. After the local user speaks, the opposite party is prompted by the "tick" prompt sound to show the waiting and guide to reduce the situation of overlapping speech of the two parties. Optionally, in the process of the call, when the sound repairing result is not given, the opposite party is guided to wait and reply patiently through the call prompt sound.
[0197] In the embodiments of the present application, the timing of sending the prompt sound by the local device can be sending the prompt sound after starting to receive the voice input by the user, or sending the prompt sound after receiving a piece of voice, which is not limited herein.
[0198] It should be noted that the prompt in the embodiments can also be used to prompt the role of the prompt sound in the process of the call. For example, the prompt can include "In the process of the call, the prompt sound will be heard, which indicates that the user is inputting voice at this time, please wait patiently" and the like.
[0199] In the embodiments of the present application, the local device sends a prompt tone to the remote device after receiving each piece of voice input by the local user, which can improve the experience of both parties when using the voice repair function and reduce the problem of inefficient communication caused by time delay.
[0200] In a possible implementation, only one of the prompt text and the prompt tone can be sent, which is not limited herein.
[0201] The above embodiments are exemplified in the remote real-time communication scenario. The following embodiments are exemplified in the remote non-real-time communication scenario.
[0202] Please refer to FIG. 7, which is an interface schematic diagram of a remote non-real-time communication scenario provided by the embodiments of the present application.
[0203] As shown in the interface of (a) in FIG. 7, a page with application icons is displayed, which can include multiple application icons (for example, a chat application 701). The chat application 701 in the embodiments of the present application can be a non-real-time remote chat application 701. The interface shown in (a) in FIG. 7 can refer to the description of the interface shown in (a) in FIG. 2, which is not repeated herein.
[0204] In the embodiments of the present application, the electronic device can detect a user operation acting on the chat application 701, and in response to the operation, the electronic device can display the interface shown in (b) in FIG. 7. Optionally, the interface shown in (b) in FIG. 7 can include a search bar, a chat function control, an address book function control, a sending function control and a personal center function control. The search bar can be used to quickly search information, for example, to quickly search contacts, etc. In the interface shown in (b) in FIG. 7, the chat function control is in a selected state, so the interface shown in (b) in FIG. 7 can be understood as the interface corresponding to the chat function, which includes one or more contacts and the contact information corresponding to each contact. Optionally, the contact information can include the contact name and the chat content, which can be, for example, the latest chat content. Exemplarily, the name of one of the contacts can be, for example, “Xiaoming”, and the latest chat record with him can be, for example, “What did you do today?”.
[0205] In the embodiment of the present application, the electronic device can detect a user operation acting on one of the contacts, and in response to the operation, the electronic device can display an interface as shown in (c) of FIG. 7, which can be a chat interface with the selected contact. The chat interface can include each piece of chat content and the user corresponding to the chat content. Optionally, the chat content can include, but is not limited to, at least one of text, voice, video or file, etc. The interface shown in (c) of FIG. 7 can include a chat content input control, which can include a voice input control 702. Optionally, the chat content input control can also include a text input control, a file / video transmission control, and a video shooting control, etc. The voice input control 702 is used to trigger the recording of voice. How the voice input control 702 triggers the recording of voice can refer to the description of how the sound recording control triggers the recording of sound in the embodiment of FIG. 2, which will not be repeated here.
[0206] In the embodiment of the present application, the electronic device can detect a user operation acting on the voice input control 702, and in response to the operation, the voice input by the user can be recorded. As shown in (d) of FIG. 7, the electronic device can also display a prompt information that the recording is in progress during the recording of the voice. Optionally, the prompt information that the recording is in progress can include text prompt information such as "recording in progress", and can also include the text content converted from the recorded voice. It should be understood that by displaying the text content converted from the recorded voice during the recording of the voice, the user can know whether the electronic device collects the voice of the user, and whether the text content converted by the electronic device matches the recorded voice, thereby improving the user experience.
[0207] Then, as shown in (e) of FIG. 7, the electronic device can display a prompt information that the recording is completed after the recording of the voice is completed. The prompt information that the recording is completed can include text prompt information such as "recording completed", and can also include the text content converted from the recorded voice. In the interface shown in (e) of FIG. 7, a direct sending option 703 and a repair and sending option 704 are included. The direct sending option 703 is used to trigger the sending of the original voice of the user collected, that is, if the electronic device detects a user operation acting on the direct sending option 703, the uncorrected voice is sent in the chat interface. The repair and sending option 704 is used to trigger the sending of the repaired voice 705, that is, if the electronic device detects a user operation acting on the repair and sending option 704, the repaired voice 705 is sent in the chat interface.
[0208] In the embodiments of the present application, the electronic device can detect a user operation acting on the repair and sending option 704, and in response to the operation, the repaired voice 705 can be sent in the chat interface. As shown in the interface of (f) in FIG. 7, the repaired voice 705 is sent. The electronic device can detect a user operation acting on the repaired voice 705, and play the repaired voice 705.
[0209] Optionally, if the user selects to send the original voice, the electronic device can also play the original voice after detecting a user operation acting on the original voice.
[0210] The embodiments of the present application can improve the experience of non-real-time remote communication between the user and the contact by allowing the user to send the repaired voice 705 to the contact.
[0211] Then, as shown in the interface of (g) in FIG. 7, the voice 706 from the contact is received.
[0212] In the embodiments of the present application, as shown in (h) in FIG. 7, the electronic device can detect a user operation acting on the voice 706 of the contact, and in response to the operation, the voice repair and conversion to text option 707 and the voice conversion to text option 708 can be displayed. The voice repair and conversion to text option 707 is used to trigger the display of the text converted from the selected voice after repair. The voice conversion to text option 708 is used to trigger the display of the text converted from the selected voice.
[0213] It should be noted that, optionally, the voice can be played by clicking the voice, and the voice repair and conversion to text option 707 and the voice conversion to text option 708 can be displayed by long-pressing the voice.
[0214] Then, as shown in (i) in FIG. 7, the electronic device can detect a user operation acting on the voice repair and conversion to text option 707, and in response to the operation, the text converted from the selected voice after repair can be sent in the chat interface.
[0215] It should be noted that, the electronic device can repair the voice and then convert the repaired voice into text for display after detecting a user operation acting on the voice repair and conversion to text option 707, or the electronic device can repair the voice after receiving the voice to obtain the repaired voice, and then display the text converted from the selected voice after repair by detecting a user operation acting on the voice repair and conversion to text option 707.
[0216] In the embodiment of the present application, if the received voice message of the opposite party is the audio of a special speech-impaired user, the user at the local end can operate the user operation of the voice 706 of the contact person, display the voice repair to text option 707 and the voice to text option 708, and then operate the user operation of the voice repair to text option 707, and obtain more accurate text.
[0217] In another possible implementation, after detecting the user operation of the voice 706 of the contact person, the electronic device can also display a voice repair option, which is used to trigger the playback of the repaired voice 705. After detecting the user operation of the voice repair option, the electronic device can play back the repaired voice 705, that is, the repaired voice 706 from the contact person.
[0218] The embodiment of the present application can repair the voice 706 of the contact person after receiving the voice 706 from the contact person, so that the speech-impaired user can also communicate normally, thereby improving the experience of non-real-time remote communication between the user and the contact person.
[0219] In general, in the remote communication scenario, the opposite party can hear the repaired voice with higher intelligibility and the same timbre in real time, so that the communication is smoother. In addition, in the remote communication scenario, the opposite party will receive the prompt broadcast of the voice repair function enabled at the local end, and the local end can also switch between the original audio source and the repaired audio source at any time, so that the user at the opposite end can obtain the original sound information, which can include at least one of the original timbre, the environmental sound or the original pronunciation.
[0220] In another possible implementation, in (h) of FIG. 7, a play voice option and a play repaired voice option can also be displayed. If the user operation of the play voice option is detected, the voice from the contact person is played back. If the user operation of the play repaired voice option is detected, the voice from the contact person is repaired and then played back. The embodiment of the present application can repair the voice from the contact person by displaying the play voice option and the play repaired voice option, thereby improving the experience of communication between the two parties.
[0221] The above describes the remote communication scenario. The following illustrates the near communication scenario. Unlike real-time communication, face-to-face repair is used by special speech-impaired people when communicating with others face-to-face offline.
[0222] Please refer to FIG. 8, which is an interface schematic diagram of a face-to-face communication scenario provided by an embodiment of the present application.
[0223] The interface shown in (a) of FIG. 8 displays a page on which an application icon is placed, and the page can include a plurality of application icons (for example, the face-to-face repair application 801). The face-to-face repair application 801 of the embodiment can be used for voice repair in a face-to-face communication scenario. The interface shown in (a) of FIG. 8 can refer to the description of the interface shown in (a) of FIG. 2, and details are not described herein.
[0224] In the embodiment of the present application, the electronic device can detect a user operation acting on the face-to-face repair application 801, and in response to the operation, the interface shown in (b) of FIG. 8 can be displayed. The interface shown in (b) of FIG. 8 can include a recording control 802 for triggering the collection of voice. For example, the recording control can be in the form of a recording ball. In the embodiment of the present application, the electronic device can detect a user operation acting on the recording control 802, and in response to the operation, the recording of voice is started. Optionally, the user operation of the embodiment can be, for example, a long press operation, and during the long press of the recording control 802, an upward swipe cancellation prompt and a release repair prompt can be displayed. The upward swipe cancellation prompt is used to prompt that the voice recording can be cancelled by upward swiping. The release repair prompt is used to prompt that the recording can be completed by releasing the recording control 802.
[0225] As shown in (c) of FIG. 8, after the electronic device detects the release operation acting on the recording control 802, a playback control 803 of a repaired voice is displayed, and the playback control 803 of the repaired voice is the voice obtained by repairing the recorded voice. In addition, the interface shown in (c) of FIG. 8 can further include a text display area 804 for displaying the text content corresponding to the playback control 803 of the repaired voice. For example, the text content displayed in the interface shown in (c) of FIG. 8 can be, for example, “today afternoon go or not go”. It should be noted that if the electronic device detects a user operation acting on the playback control 803 of the repaired voice in the interface shown in (c) of FIG. 8, the repaired voice can be played, for example, the voice corresponding to “today afternoon go or not go” is played. Optionally, after detecting the release operation acting on the recording control 802, the recording control 802 is displayed for a short period of time, and then the repaired text is displayed and the repaired audio is automatically played. If the user is not satisfied with the current recording during the recording process, the user can swipe upward in the long press state to cancel the current repair.
[0226] However, since voice restoration may not be entirely accurate, this application embodiment also supports rapid correction of the restoration results to improve the effectiveness of face-to-face communication. For example, if the restored text contains misidentified words, the user can click on the word they want to modify, and a candidate word list 806 will appear, from which the user can select the desired word for quick replacement. Users can also directly edit the text using an input method. After completing the correction, the user can choose to play the corrected audio.
[0227] As shown in Figure 8(d), the electronic device can detect a user operation on a portion of the text content, word 805, and in response to this operation, can display the interface shown in Figure 8(e). The interface shown in Figure 8(e) includes a candidate word list 806 corresponding to the user-selected portion of word 805, which includes one or more candidate words. For example, if the portion of word 805 is "afternoon," then the corresponding candidate words could include "morning," "dance well," and "foggy," among others.
[0228] Then, as shown in Figure 8(f), the electronic device can detect a user action applied to one of the candidate words, and in response to the action, can display the interface shown in Figure 8(g). In the interface shown in Figure 8(g), the word 805 in the displayed text content is replaced with the candidate word applied by the user action. For example, if the electronic device detects a user action applied to the candidate word "morning," then in the displayed text content "Will you go this afternoon?", "afternoon" is replaced with "morning," and the corresponding displayed text content is replaced with "Will you go this morning?".
[0229] Then, as shown in Figure 8(h), the electronic device can detect the user operation on the playback control 803 of the repaired voice in the interface shown in Figure 8(h), and in response to the operation, can play the voice corresponding to "Are you going this morning or not?". In this way, the other party communicating with the user can know the user's voice expression.
[0230] In other words, in this embodiment of the application, the speech can be broadcast according to the text content displayed in the text display area 804.
[0231] In this embodiment, by recording and repairing the user's voice during face-to-face communication, the playback control 803 displays the repaired voice, improving the user's communication experience. Furthermore, after recording the user's input voice, the corresponding text content can be displayed. This allows the user to determine the accuracy of the voice repair result through the text content. If the user deems the repair result inaccurate, they can adjust the text content, thereby adjusting the content of the played voice. In other words, face-to-face communication scenarios can quickly correct semantic errors, making user operation more convenient and faster, and improving the accuracy of voice communication. Moreover, in this embodiment, voice input during face-to-face communication allows for the playback of the repaired audio, eliminating the need for plain text input, making it faster. Furthermore, for words with errors after repair, the system supports quick replacement of the identified words in the text and replays the audio, improving the efficiency of face-to-face communication.
[0232] It should be noted that in the interfaces shown in Figures 8(c)-(e), one possible implementation is as follows: When performing ASR recognition processing on the speech, multiple ASR recognition results may be obtained, each corresponding to a probability value. The electronic device can then display the ASR recognition result with the highest probability value among the multiple results. For example, the ASR recognition results obtained from the ASR recognition processing may include: "Today, shall we go to Shanwu?", "Today, shall we go to Xiawu?", "Today, shall we go to Shanwu?", and "Today, shall we go to Shanwu?". Since the probability value of the recognition result "Today, shall we go to Shanwu?" is the highest, this ASR recognition result is displayed. The words in the other ASR recognition results besides the displayed one are also saved. Then, the displayed ASR recognition results can be segmented to obtain one or more words. When a user action is detected acting on one of the words (i.e., part of word 805), words different from that part of word 805 are searched from the remaining ASR recognition results and displayed as candidate words, such as "morning," "good at dancing," and "foggy." This embodiment improves the efficiency of obtaining candidate words by saving the remaining ASR recognition results and then searching for words different from the user-selected word (i.e., part of word 805) from the remaining ASR recognition results.
[0233] In another possible implementation, the remaining ASR recognition results may not be saved. In this way, candidate words can be predicted directly from the words selected by the user. Since the remaining ASR recognition results do not need to be stored, the storage resources required to store the ASR recognition results can be reduced.
[0234] It should be noted that in the interface shown in FIG. 8, there can be at most one playback control 803 of the repaired voice, for example, there is a playback control 803 of the latest repaired voice, so that the simplicity of interface display can be improved. In another possible implementation, a new voice can be displayed each time a voice is recorded, and a new voice can be displayed each time the text content is modified, so that each voice of the user and the modification process of the user can be recorded, thereby improving the user experience.
[0235] In another possible implementation, as shown in the interface (c) in FIG. 8, the voice or text display area 804 can also be displayed, so that the simplicity of interface display can be improved. In the embodiment of the present application, by displaying the voice and the text display area 804, the user can quickly know whether the voice repair result of the electronic device is accurate through the text content displayed in the text display area 804, thereby improving the user experience.
[0236] In another possible implementation, the user can also select a punctuation mark in the text content, so that the candidate punctuation mark can be displayed, the candidate punctuation mark selected by the user is replaced in the punctuation mark selected by the user in the text content, and then the intonation of the voice can be changed according to the updated punctuation mark. For example, the “?” is replaced by “!”, and the intonation of the voice is changed from the intonation of the question to the intonation of the exclamation.
[0237] It should be noted that the speech-impaired user is generally accompanied by a hearing ability impaired to some extent, that is, the speech-impaired user can also be a hearing-impaired user. Therefore, the following embodiments are exemplarily described by how to improve the communication experience of the hearing-impaired or speech-impaired user.
[0238] Please refer to FIG. 9, which is an interface schematic diagram of another face-to-face communication scenario provided by the embodiment of the present application.
[0239] The interface shown in (a) in FIG. 9 can be, for example, the interface of the face-to-face repair application. In the interface shown in (a) in FIG. 9, the recording control 802 is included, and the recording control 802 can be used to record the voice. Optionally, the recording control 802 can be used to record the voice of the hearing-impaired / speech-impaired user, and can also be used to record the voice of the normal user. The normal user can be a user with normal hearing ability and a user with normal speaking ability. It should be noted that although the recording controls 802 in FIG. 9 and FIG. 8 are different in style, their functions are the same, and they are both used to trigger the recording of the voice.
[0240] In the embodiment of the present application, the electronic device can detect a click operation acting on the recording control 802, and in response to the operation, enter the continuous sound recording mode, at which time the voice of the opposite party is converted into text in real time. As shown in the interface of (a) in FIG. 9, during the recording, a prompt information for prompting that the recording is being performed can also be displayed, such as a text prompt content of "listening, can help you convert into text". Then, as shown in the interface of (a) in FIG. 9, the electronic device can display the text converted from the voice of the normal user, such as the text of "I am a normal user, I am speaking". Optionally, the content of the conversation can be displayed in a conversation box (chat box). The electronic device can cancel the continuous sound recording mode by detecting a click operation acting on the recording control 802 again. Optionally, the electronic device can also enter the continuous sound recording mode when detecting a long press operation acting on the recording control 802, and then cancel the continuous sound recording mode after detecting an operation leaving the recording control 802, that is, during the long press operation on the recording control 802, the continuous sound recording mode is in the continuous sound recording mode. Optionally, in the embodiment of the present application, the voice of the normal user can be repaired before being converted into text for display, thereby improving the experience of face-to-face communication.
[0241] Then, as shown in (b) in FIG. 9, the electronic device can detect a user operation acting on the interface, such as detecting a user operation acting on the blank area of the interface, and in response to the operation, display the interface as shown in (c) in FIG. 9. The interface as shown in (a) in FIG. 9 can be, for example, a voice recording function interface, and the voice recording function interface can refer to the description of the embodiment of FIG. 8, which is not repeated here. The interface as shown in (c) in FIG. 9 can be, for example, a text input function interface, and switching from the interface as shown in (a) in FIG. 9 to the interface as shown in (c) in FIG. 9 can be understood as switching from the voice recording function to the text input function. Optionally, the user operation can be, for example, an up sliding operation. In the interface as shown in (c) in FIG. 9, the virtual keyboard 808 for inputting text, the voice generation control 809 for generating voice, and the text display box 810 for displaying text are included. Optionally, the recording control 802 can also be used as a control for generating voice, that is, after inputting text through the virtual keyboard 808, the voice is displayed in the conversation box through a user operation acting on the recording control 802.
[0242] In another possible implementation, the interface as shown in (b) in FIG. 9 can also include a keyboard triggering control for triggering the display of the virtual keyboard 808. The electronic device can display the interface as shown in (c) in FIG. 9 after detecting a user operation acting on the keyboard triggering control.
[0243] In the embodiments of the present application, as shown in (c) of FIG. 9, the electronic device can detect a user operation acting on the virtual keyboard 808, and in response to the operation, display the text input by the user in the text display box 810.
[0244] Then, as shown in (d) of FIG. 9, the electronic device can detect a user operation acting on the voice generation control 809, and in response to the operation, generate a voice according to the text displayed in the text display box 810. As shown in (d) of FIG. 9, the interface can include the voice of the text displayed in the text display box 810. For example, if the text displayed in the text display box 810 is "Hello, I am Zhang San", the electronic device generates a voice of "Hello, I am Zhang San" after detecting the user operation acting on the voice generation control 809. In addition, when generating the voice of "Hello, I am Zhang San", the tone of the voice can also be generated according to the punctuation marks or emoticons of the text.
[0245] For example, if the text includes ". ", the voice of "Hello, I am Zhang San" can be broadcasted in a flat tone; if the text includes "! ", the voice of "Hello, I am Zhang San" can be broadcasted in an exclamation tone; and if the text includes "? ", the voice of "Hello, I am Zhang San" can be broadcasted in a questioning tone. In addition, if the text includes a smiley face, the voice of "Hello, I am Zhang San" can be broadcasted in a happy tone, and if the text includes a crying face, the voice of "Hello, I am Zhang San" can be broadcasted in a sad tone.
[0246] Then, as shown in (e) of FIG. 9, the electronic device can detect a user operation acting on the voice, and in response to the operation, the voice of "Hello, I am Zhang San" can be broadcasted.
[0247] In the embodiments of the present application, the user can continue to input the text. Then, as shown in (f) of FIG. 9, the electronic device can detect a user operation acting on the virtual keyboard 808, and in response to the operation, display the text input by the user in the text display box 810.
[0248] Then, as shown in (g) of FIG. 9, the electronic device can detect a user operation acting on the voice generation control 809, and in response to the operation, generate a voice according to the text displayed in the text display box 810, for example, display the interface as shown in (h) of FIG. 9. As shown in (h) of FIG. 9, the interface can include the voice of the text displayed in the text display box 810. For example, if the text displayed in the text display box 810 is "I want to ask, what time does the store close today", the electronic device generates a voice of "I want to ask, what time does the store close today" after detecting the user operation acting on the voice generation control 809.
[0249] Then, as shown in interface (i) in FIG. 9, the electronic device can detect a user operation acting on the interface, for example, detect a user operation acting on a blank area of the interface, and in response to the operation, can switch to an interface of the voice recording function. Optionally, the user operation may, for example, be a downward swipe operation.
[0250] In the embodiments of the present application, for the voice input by a normal user, the voice can be converted into text, and the hearing / speech impaired user can input the text, and then the electronic device displays the voice corresponding to the text, so that the communication between the hearing / speech impaired user and the normal user is facilitated.
[0251] In general, in the face-to-face communication scenario, the functions of converting the voice of the other party into text, converting the text of the local user into voice, and repairing the sound of the local user are supported.
[0252] In another possible implementation, the face-to-face repair application can also only include an interface of the text input function, and does not have the function of voice input, so that the function of the face-to-face repair application can be simplified, and the program package size of the face-to-face repair application can be reduced, thereby reducing the storage resources required by the electronic device. By configuring the voice recording function and the text input function in the face-to-face repair application, the user can select the communication mode according to the needs, thereby improving the flexibility of communication and improving the user experience.
[0253] It should be noted that when the face-to-face repair application has the voice recording function and the text input function, the default interface of the face-to-face repair application when entering the face-to-face repair application can be the interface of the voice recording function, or can be the interface of the text input function, which is not limited herein. Optionally, the default interface of the face-to-face repair application can be the interface when exiting the face-to-face repair application, for example, if the interface of the voice recording function is exited when the face-to-face repair application is exited, then the default interface of the face-to-face repair application when entering the face-to-face repair application again is the interface of the voice recording function; if the interface of the text input function is exited when the face-to-face repair application is exited, then the default interface of the face-to-face repair application when entering the face-to-face repair application again is the interface of the text input function.
[0254] It should be noted that the sound repair capability can be integrated into the system of the electronic device, and the sound repair capability can be provided for part or all of the applications or scenarios involving the voice input function, that is, the voice can be repaired and output for part or all of the applications or scenarios involving the voice input function, thereby improving the user experience.
[0255] The above embodiments describe the interface of the communication scenario. The following embodiments describe the specific implementation of how to process the voice.
[0256] Please refer to FIG. 10, which is a flowchart of a voice processing method according to an embodiment of the present application. The method shown in FIG. 10 can be performed by at least one of an electronic device or a cloud. The method shown in FIG. 10 can include the following steps.
[0257] S1001, obtaining a sound input.
[0258] In the embodiments of the present application, the sound input received can be a user's sound input. For example, the sound input mode can refer to the modes shown in FIG. 5, FIG. 7, FIG. 8 or FIG. 9, which will not be repeated here.
[0259] S1002, inputting the obtained sound input to an ASR model.
[0260] The ASR model is used to convert sound into text. In the embodiments of the present application, the ASR model can be a trained model. In this embodiment, by inputting the obtained sound to the ASR model, the ASR model can convert the sound input to the ASR model into text output.
[0261] S1003, obtaining the text output by the ASR model.
[0262] S1004, correcting the text.
[0263] In the embodiments of the present application, if the obtained sound is a sound input by a speech-impaired user, the converted text can be biased, so the text needs to be corrected. Optionally, the text correction mode can be automatic correction or manual correction. Optionally, the automatic correction mode can include but is not limited to statistical-based text correction or context-aware text correction. Statistical-based text correction can be to learn the frequency and context information of words by using a large amount of text data, and to predict the most likely correct text by a statistical model. Context-aware text correction can be to correct the text in combination with context information, which can be, for example, the context before and after the sentence. The manual correction mode can be, for example, the mode shown in FIG. 8.
[0264] S1005, obtaining a result of sound registration.
[0265] In the embodiments of the present application, the result of sound registration can include registered sound or a voiceprint feature extracted from the registered sound. Optionally, when the user uses the sound repair function, if it is detected that the user has not registered the sound, the interface for registration can be jumped to for registration; or the user can have registered the sound before using the sound repair function. For example, the registration mode of the sound can refer to the related description of FIG. 2.
[0266] S1006, performing TTS broadcasting based on the corrected text and the result of sound registration.
[0267] In the embodiment of the present application, after the corrected text is converted into speech, the speech can be broadcast according to the voiceprint of the registered voice. In this way, not only is the intelligibility of the output speech higher than that of the obtained voice, but also the speech can be broadcast according to the voiceprint of the registered user, which can improve the user experience. Optionally, the TTS can convert the text into audio in the user's voice through a neural network model and play it out.
[0268] It should be understood that, in the embodiment of the present application, the voiceprint feature can be obtained in advance by registering the voiceprint. When the speech is broadcast, the pre-registered voiceprint can be obtained. In this way, the time for extracting the voiceprint after obtaining the voice can be reduced, and then the time between obtaining the voice and broadcasting the speech can be reduced, thereby improving the efficiency of speech processing.
[0269] In a possible implementation, S1005 can also not be required, and accordingly, S1006 can broadcast the speech based on the corrected text. At this time, the voiceprint of the broadcast speech can be the voiceprint built-in the electronic device or the voiceprint of the obtained voice.
[0270] It should be noted that, since the obtained voice is converted into text in the embodiment of the present application, the scheme of the embodiment of the present application can be applied to a scenario in which text needs to be displayed, such as the scenarios shown in FIG. 7, FIG. 8, or FIG. 9. In addition, since conversion into text is also required, the scheme of the embodiment of the present application can be applied to a scenario in which the real-time requirement is not particularly high.
[0271] In another possible implementation, the obtained voice can be corrected and then converted into text through an ASR model.
[0272] Next, the specific implementation of voice repair is illustrated by way of example.
[0273] Referring to FIG. 11, FIG. 11 is a flowchart of a voice registration method provided by an embodiment of the present application. The method shown in FIG. 11 can be performed by at least one of an electronic device or a cloud. The method shown in FIG. 11 can include the following steps.
[0274] S1101, voiceprint registration.
[0275] In the embodiment of the present application, the electronic device can trigger the registration process of the voiceprint in response to the operation of the user for voiceprint registration. For example, the operation of voiceprint registration can be the operation shown in (e) of FIG. 2, which is not limited herein.
[0276] S1102, voice input.
[0277] In the embodiments of the present application, the electronic device can collect the voice input of the user after responding to the operation of the user for voiceprint registration. For example, the manner of receiving the voice input of the user can be the manner shown in (e) of FIG. 2, which is not limited herein.
[0278] S1103, ASR recognition rate grading determination.
[0279] In the embodiments of the present application, after the voice input of the user is collected, the electronic device can perform ASR recognition on the voice input to obtain an ASR recognition result. It should be understood that the ASR recognition result can be understood as the recognized text corresponding to the voice input of the user. Then, the ASR recognition result is compared with the preconfigured voice recording content to obtain an ASR recognition rate. The ASR recognition rate can also be referred to as the similarity of the comparison, which is used to indicate the degree of similarity between the ASR recognition result and the preconfigured voice recording content. For example, the voice recording content can be the text content shown in (f) of FIG. 2, which is not limited herein. After obtaining the ASR recognition rate, it can be determined which recognition rate range the ASR recognition rate is in. Optionally, a plurality of recognition rate ranges can be pre-divided, and there is no common recognition rate between the plurality of recognition rate ranges, for example, a plurality of recognition rate ranges such as (0, 70%) and (70%, 100%). If the ASR recognition rate is 50%, it belongs to the range of (0, 70%); if the ASR recognition rate is 80%, it belongs to the range of (70%, 100%).
[0280] In the embodiments of the present application, after the ASR recognition rate grading determination is performed, the repair model is selected, which can improve the efficiency of selecting the repair model.
[0281] S1104, repair model selection.
[0282] The repair model is used to repair the input voice to obtain the repaired voice. In the embodiments of the present application, the repair model can be selected according to the grading determination result of the ASR recognition rate. Optionally, a mapping relationship between each recognition rate range and the repair model is preconfigured, and then the repair model corresponding to the grading determination result of the ASR recognition rate can be obtained from the mapping relationship after the grading determination result of the ASR recognition rate is obtained. Optionally, the voice repair capabilities of the repair models corresponding to different recognition rate ranges are different. Optionally, the higher the recognition rate in the recognition rate range, the weaker the voice repair capability of the corresponding repair model, that is, the lower the recognition rate in the recognition rate range, the stronger the voice repair capability of the corresponding repair model.
[0283] For example, a user with an ASR recognition rate lower than 70% uses a repair model with stronger repair, and a user with an ASR recognition rate higher than 70% uses a repair model with weaker repair and a function biased towards sound beautification and accent correction.
[0284] According to the application, the repair model is selected according to the ASR recognition rate, that is, the repair model is selected according to the degree of impairment of the user's speech ability. For a user with less impaired speech ability, a repair model with weaker repair is selected, and accordingly, the repair model requires less computing power. For a user with more impaired speech ability, a repair model with weaker repair is selected, and accordingly, the speech repair effect is better. Thus, the speech repair effect and the computing power required for speech repair can be considered.
[0285] S1105, an acoustic feature extraction module.
[0286] The acoustic feature extraction module is configured to extract acoustic features (acoustic features). In the application, the acoustic feature extraction module is configured to extract the acoustic features of the input sound. Optionally, the acoustic feature extraction module can convert an input audio of any length into a fixed-length acoustic feature vector. The acoustic feature extraction module can be an algorithm or a neural network model. Optionally, the acoustic feature extraction module can select an acoustic feature extraction model. The acoustic feature extraction model of the application can include, but is not limited to, an emphasized channel attention, propagation and aggregation in time delay neural network (ECAPA-TDNN) structure.
[0287] S1106, a fixed-length feature vector.
[0288] In the application, the fixed-length feature vector is a vector obtained by extracting the acoustic features of the user's recorded sound, which is used to indicate the acoustic features of the user. In the application, the acoustic features are converted into a fixed-length acoustic feature vector when extracting the acoustic features, which is beneficial to reduce the storage resources required for storing the feature vector.
[0289] S1107, save.
[0290] In the application, the feature vector can be saved, so that the sound can be broadcast according to the acoustic features in the subsequent broadcast.
[0291] In another possible implementation, the length of the feature vector saved in S1106 can be related to the length of the input sound, rather than being a fixed-length feature vector, so that it can be adaptively saved according to the length of the input sound, improving the accuracy of the saved voiceprint feature vector.
[0292] In another possible implementation, in S1103, after obtaining the ASR recognition rate, the gear judgment can also not be performed, and the selection of the repair model can be directly based on the ASR recognition rate.
[0293] In another possible implementation, S1103 and S1104 can also not be required, and a pre-configured repair model can be used to repair the sound, so that the time required for ASR recognition rate grading judgment and selection of the repair model can be reduced.
[0294] The architecture of one of the repair models will be described below.
[0295] Please refer to FIG. 12, which is a schematic diagram of the architecture of a repair model provided in an embodiment of the present application. The repair model of the present embodiment can be deployed in at least one of an electronic device or a cloud, and the repair model shown in FIG. 12 can include a prosody encoder 1201, a streaming ASR encoder 1202, an audio discretization unit 1203, a streaming audio repair large model 1204, and a vocoder 1205.
[0296] The prosody encoder 1201 is configured to extract prosodic features (referred to as prosody for short). For example, the prosody encoder 1201 can extract prosodic features from the input audio signal. The prosodic features can include, but are not limited to, at least one of pitch, rhythm, or intonation. Optionally, the audio signal of a normal user can be obtained, and then the prosodic features are extracted from the audio signal of the normal user as the prosodic features when the audio is broadcasted.
[0297] The streaming ASR encoder 1202 is an encoder that can convert an audio signal into a text sequence in real time or near real time. In the present embodiment, the streaming ASR encoder 1202 is configured to convert the audio input by the user into text in real time or near real time. The streaming ASR encoder 1202 can include a neural network model.
[0298] The audio discretization unit 1203 is configured to convert continuous audio signals into discrete symbol or unit sequences. For example, a continuous audio signal (such as a waveform signal) is decomposed into a series of discrete, identifiable units or symbols through a certain method. These units can be based on phonemes, features, or obtained through other means (such as cluster analysis).
[0299] The streaming audio repair large model 1204 is configured to repair or optimize the audio features. The audio features can include, but are not limited to, at least one of content features, prosody features, or voiceprint features. Optionally, the streaming audio repair large model 1204 is configured to repair or optimize in a streaming LM-based vc manner, which combines the advantages of streaming and language model (LM) to achieve real-time and efficient audio conversion. Optionally, the streaming audio repair large model 1204 can include, but is not limited to, a streaming prosody repair large model 1204.
[0300] The vocoder 1205 is a kind of synthesizer configured to convert the audio features into playable audio signals.
[0301] In the embodiments of the present application, the repair model can be a pre-trained model.
[0302] In a possible implementation, when performing audio repair, the repair model can perform the following processing:
[0303] The prosody encoder 1201 is configured to extract prosody features from the built-in normal user voice. The streaming ASR encoder 1202 is configured to perform ASR recognition processing on the input audio to be repaired to obtain an ASR recognition result (which can also be referred to as content features or content). Then, the ASR recognition result, the prosody features, and the voiceprint features registered by the user are input into the streaming audio repair large model 1204. In addition, the streaming audio repair large model 1204 fuses the voiceprint features, the prosody features, and the ASR recognition result in the spliced result with the initial audio discretization features (which can also be referred to as discretization features), and generates repaired discretization features and repaired continuous audio features through autoregressive inference. The discretization features generated by autoregressive inference will be used as the input of subsequent autoregressive inference, and will be fused and inferred by the streaming audio repair large model 1204, and the generated continuous audio features will be used as the input of the vocoder 1205. Then, the vocoder 1205 converts the audio continuous features into repaired audio, and the repaired audio can be played through an audio playing device (such as a loudspeaker), the voiceprint of the played audio can be the voiceprint registered by the user, and the prosody of the played audio can be the prosody extracted by the prosody encoder 1201.
[0304] The to-be-repaired audio can be audio input by a user and collected by a sound collection device (e.g., a microphone) of the electronic device. For example, the to-be-repaired audio of the embodiment of the present application can include, but is not limited to, voice input in the scenarios shown in FIG. 5, FIG. 7, FIG. 8, or FIG. 9. Optionally, the initial audio discretization feature can be obtained by the audio discretization unit 1203 based on at least one of the ASR recognition result or the to-be-repaired audio.
[0305] In the embodiment of the present application, since the ASR recognition result is obtained by performing ASR recognition on the to-be-repaired audio, and the audio discretization feature is also obtained by performing discretization on the to-be-repaired audio, the ASR recognition result and the audio discretization feature are fused, that is, the processing results of the to-be-repaired audio in different dimensions are fused, thereby improving the accuracy of the obtained audio continuous feature, and then improving the accuracy of the content of the obtained to-be-repaired audio.
[0306] Optionally, the input audio can be repaired when the user inputs the audio, which can improve the real-time performance of the audio repair processing. Optionally, the audio can be repaired after a certain amount of time of audio is obtained. For example, when the user inputs the audio, if 1 second of audio has been obtained, the 1 second of audio is repaired, and then the next 1 second of audio is repaired, until the complete audio input by the user is repaired.
[0307] In another possible implementation, the repaired discretization feature can also not be generated, which can simplify the architecture of the repair model and improve the efficiency of sound repair. By generating the repaired discretization feature, the repaired discretization feature can be used as the input of subsequent autoregressive reasoning, and the streaming audio repair large model 1204 can be used for fusion and reasoning, thereby improving the accuracy of sound repair.
[0308] Optionally, the length of each segment of to-be-repaired audio can be consistent, which can improve the accuracy of audio repair; or the length of different segments of to-be-repaired audio can be inconsistent, which can improve the flexibility of audio repair.
[0309] For example, assuming that the audio input by the user is "today afternoon, go or not", after obtaining "today afternoon", "today afternoon" is taken as the audio to be repaired. Then, the stream ASR encoder 1202 performs ASR recognition processing on the input "today afternoon" audio, obtains an ASR recognition result, and then inputs the spliced result of the ASR recognition result, prosody features, and the voiceprint features registered by the user into the stream audio repair large model 1204. In addition, the audio discretization unit 1203 performs discretization processing on the "today afternoon" audio based on the ASR recognition result to obtain audio discretization features, and then inputs the audio discretization features into the stream audio repair large model 1204. The stream audio repair large model 1204 can output the repaired audio discretization features and the repaired audio continuous features corresponding to the "today afternoon" audio, and the vocoder 1205 can output the repair result of the "today afternoon" audio based on the repaired audio continuous features. In addition, the stream audio repair large model 1204 can take the repaired audio discretization features corresponding to the "today afternoon" audio as a reference for repairing the "go or not" audio, for example, as a reference for discretization processing of the "go or not" audio.
[0310] Next, the training of the repair model is exemplarily described.
[0311] A training sample set is obtained, the training sample set including one or more training samples, each training sample including an audio sample, a text sample corresponding to the audio sample, and a repaired audio sample corresponding to the audio sample.
[0312] During training, the audio sample is taken as the input of the stream ASR encoder 1202, the stream ASR encoder 1202 outputs a predicted text, the predicted text is compared with the text sample corresponding to the audio sample to calculate a first training loss, if the first training loss meets a first training end condition, the stream ASR encoder 1202 ends the training; if the first training loss does not meet the first training end condition, the parameters of the stream ASR encoder 1202 are updated and the training is continued until the first training loss meets the first training end condition. Optionally, the first training end condition can include that the first training loss is less than a first threshold.
[0313] Then, continue to train the repair model, take the audio sample as the input of the streaming ASR encoder 1202, or take the text sample corresponding to the audio sample as the input of the streaming audio repair large model 1204, obtain the repaired audio output by the vocoder 1205, and then compare the repaired audio output by the vocoder 1205 with the repaired audio sample corresponding to the audio sample to calculate a second training loss. If the second training loss satisfies a second training end condition, the repair model ends the training; if the second training loss does not satisfy the second training end condition, the parameters of the repair model are updated and the training is continued until the second training loss satisfies the second training end condition. Optionally, the second training end condition can include that the second training loss is less than a second threshold. Optionally, at least one of the vocoder 1205 or the prosody encoder 1201 in the embodiment of the present application can be pre-trained, and updating the parameters of the repair model can be updating at least one of the streaming audio repair large model 1204 or the audio discretization unit 1203.
[0314] Optionally, the audio sample can also include a first segment of audio sample and a second segment of audio sample. For example, the audio sample can include "today the weather is very sunny", and the first segment of audio sample can include "today the weather", and the second segment of audio sample can include "very sunny". Correspondingly, the text sample corresponding to the audio sample includes the text of the first segment of audio sample and the text of the second segment of audio sample, and the repaired audio sample corresponding to the audio sample includes the repaired audio sample corresponding to the first segment of audio sample and the repaired audio sample corresponding to the second segment of audio sample.
[0315] Then, in the training, the repaired audio corresponding to the first segment of audio sample output by the repair model is compared with the repaired audio sample to calculate a second training loss, and the repaired audio discretization feature corresponding to the first segment of audio sample output by the repair model is compared with the audio discretization feature of the second segment of audio sample to calculate a third training loss. If the third training loss satisfies a third training end condition, the repair model ends the training; if the third training loss does not satisfy the third training end condition, the parameters of the repair model are updated and the training is continued until the third training loss satisfies the third training end condition. Optionally, the third training end condition can include that the third training loss is less than a third threshold.
[0316] It should be understood that the processing flow of the repair model in the embodiment of the present application in the training process can refer to the description of the processing flow of the sound repair in the above embodiment, which is not repeated here.
[0317] It should be noted that the user can enter the personalized portal to record a large amount of corpus as a training sample, thereby training a personalized repair model, which can improve the user's own experience.
[0318] In another possible implementation, the audio discretization unit 1203 can also not require the ASR recognition result when discretizing the audio to be repaired, which can improve the efficiency of discretization processing and in turn improve the efficiency of audio repair. By using the ASR recognition result to discretize the audio to be repaired by the audio discretization unit 1203, the accuracy of audio repair can be improved.
[0319] In another possible implementation, the repair model can include one of the audio discretization unit 1203 or the streaming ASR encoder 1202, which can improve the efficiency of discretization processing. By using the audio discretization unit 1203 and the streaming ASR encoder 1202 for audio repair, the accuracy of audio repair can be improved. It should be noted that if the repair model includes the audio discretization unit 1203 but does not include the streaming ASR encoder 1202, the result of splicing the audio discretization feature, the prosodic feature, and the user-registered voiceprint feature can be input to the streaming audio repair large model 1204.
[0320] In another possible implementation, the prosodic encoder 1201 can also not require the pre-stored prosodic feature when performing audio repair. By using the pre-stored prosodic feature to perform audio repair, the efficiency of audio repair can be improved.
[0321] In another possible implementation, the streaming audio repair large model 1204 can also not require the prediction of the audio discretization feature, which can improve the efficiency of audio repair. By predicting the audio discretization feature, the accuracy of audio repair can be improved.
[0322] It should be noted that the prosody encoder 1201 and the streaming prosody repair large model 1204 can both adopt a generative pre-trained transformer (GPT) structure, the speech discretization unit can use a vector-quantized variational autoencoder (VQVAE) model, the streaming ASR encoder 1202 can adopt a recurrent neural network transducer (RNNT), and the speech decoder 1205 can include an efficient and high-fidelity generative adversarial network (HiFi GAN). Among them, the VQVAE can convert audio into frame-level discretization units (discretization features). The HiFi GAN can convert audio features into audio sampling points (audio signals).
[0323] The following embodiments are based on the above embodiments and describe the flow of the embodiments of the present application.
[0324] Please refer to FIG. 13, which is a flowchart of another voice processing method provided by an embodiment of the present application. The method of the embodiment of the present application can be executed by an electronic device, a server, or a system including an electronic device and a server. The embodiment of the present application can repair the voice input by a user. The method shown in FIG. 13 can include:
[0325] S1301, receiving a setting input, the setting input being used to indicate a scenario in which voice repair is started, a contact in which voice repair is started, or an application in which voice repair is started, the scenario including a face-to-face communication scenario or a remote communication scenario.
[0326] The face-to-face communication scenario can be a scenario of face-to-face communication, for example, a scenario of communication between at least two users face to face. The remote communication scenario can be a scenario of communication by means of communication. In the embodiments of the present application, at least one of the following can be selected: a scenario in which voice repair is selectively enabled, a contact in which voice repair is enabled, or an application in which voice repair is enabled. Optionally, if the input is set for enabling the scenario of voice repair, if it is detected that the electronic device is in the scenario of voice repair enabled, for example, the face-to-face communication scenario or the remote communication scenario, the voice input by the user is repaired. For example, the face-to-face communication scenario can be the scenario described in the embodiments of FIG. 8 or FIG. 9, which is not repeated here. The remote communication scenario can be the scenario described in the embodiments of FIG. 5 or FIG. 7, which is not limited here. If the input is set for enabling the contact in which voice repair is enabled, when the contact communicated by the electronic device is the contact in which voice repair is enabled, the electronic device repairs the voice input by the user. For example, the manner of enabling the contact in which voice repair is enabled can be described in the embodiment of FIG. 4, which is not repeated here. If the input is set for enabling the application in which voice repair is enabled, when the application running on the electronic device is the application in which voice repair is enabled, the voice input by the user is repaired. For example, the manner of enabling the application in which voice repair is enabled can be the manner described in the embodiment of FIG. 3, which is not repeated here.
[0327] S1302, obtaining a first voice feature of the user through sound registration.
[0328] The first voice feature can be a voice feature of the user. In the embodiments of the present application, the first voice feature is used to represent the characteristics of the user's pronunciation. Optionally, the first voice feature can include at least one of a voiceprint feature or a prosody feature. The voiceprint feature can be a sound wave spectrum carrying speech information displayed by an electroacoustic instrument. The prosody feature can refer to those features in the voice that are not directly manifested as voice quality changes, but are manifested through changes in factors such as pitch, duration, and intensity. These features are manifested as suprasegmental components in the voice signal, and together with the voice quality components (such as vowels and consonants) form a complete voice signal. The first voice feature can be obtained by feature extraction on the voice registered by the user. For example, the manner of registering voice by the user can refer to the description of the embodiment of FIG. 2, which is not repeated here. It should be noted that the user in the embodiments of the present application can include a speech-impaired user or a hearing-impaired user. The speech-impaired user can be, for example, a user who has certain difficulties or defects in pronunciation. The hearing-impaired user can be, for example, a user who has certain difficulties or defects in hearing.
[0329] S1303, receiving a first voice input by the user.
[0330] S1304, repairing the first voice according to the setting input and the first voice feature.
[0331] In the embodiments of the present application, the first speech is repaired according to the setting input and the first speech feature, and since the setting input is used to indicate the scenario in which the speech repair is enabled, the contact person for which the speech repair is enabled, or the application for which the speech repair is enabled, the first speech is repaired according to the first speech feature when it is detected that the electronic device is in the scenario in which the speech repair is enabled, or when the contact person communicated by the electronic device is the contact person for which the speech repair is enabled, or when the application running on the electronic device is the application for which the speech repair is enabled, so that the repaired first speech can be played in the first speech feature, and the intelligibility of the repaired first speech is higher than that of the first speech before repair. The intelligibility can represent the accuracy of the expression of the user, and can also be understood as the understanding degree of the listener to the speech signal delivered by the loudspeaker.
[0332] In the embodiments of the present application, by receiving the setting input, the setting input is used to indicate the scenario in which the speech repair is enabled, the contact person for which the speech repair is enabled, or the application for which the speech repair is enabled, and then the first speech feature of the user is obtained through voice registration, so that after receiving the first speech input by the user, the first speech can be repaired according to the setting input and the first speech feature. In this way, the speech-impaired user or the hearing-impaired user can communicate through the input speech, thereby improving the convenience of the speech-impaired user in communication. In addition, the repaired first speech is generated by the first speech feature registered by the user, so that the repaired first speech can be played according to the first speech feature, thereby playing the repaired first speech according to the user's own speech feature, which is closer to the user's own tone, etc., thereby further improving the user experience.
[0333] In another possible implementation, the speech repair can also be performed by using the preset speech feature. For example, before the user registers the speech, the speech repair is performed according to the preset speech feature, so that the repaired speech is played according to the preset speech feature; and after the user registers the speech, the repair is performed according to the first speech feature of the user. That is, even if the user does not register the speech, the speech repair function can still be used, thereby improving the applicable scenario of the speech repair function. In addition, the processing of extracting the speech feature from the registered speech is also reduced, thereby improving the efficiency of the speech repair.
[0334] In a possible implementation, the remote communication scenario includes a call scenario, and in the call scenario, the method further includes:
[0335] sending first prompt information to the first contact person in the call, the first prompt information being used to prompt that the speech repair function has been enabled.
[0336] and / or,
[0337] The second prompt information is sent to the first contact in a case where the first voice is received, and the second prompt information is used to prompt that the first voice is being repaired.
[0338] The first prompt information can include prompt text or prompt sound. Optionally, the prompt text in the first prompt information can be displayed on a display screen of the terminal corresponding to the first contact, and the prompt sound in the first prompt information can be played through a loudspeaker of the terminal corresponding to the first contact. For example, the first prompt information can refer to the related description of the prompt text in the embodiment of FIG. 6, and details are not described herein again. The second prompt information can include prompt text or prompt sound. Optionally, the prompt text in the second prompt information can be displayed on a display screen of the terminal corresponding to the first contact, and the prompt sound in the first prompt information can be played through a loudspeaker of the terminal corresponding to the first contact. For example, the prompt sound in the second prompt information can be, for example, the prompt sound of "ding" in the embodiment of FIG. 6, and details are not described herein again.
[0339] In the embodiment of the present application, the first prompt information is sent to the first contact in the call to prompt that the voice repair function is started, so that the first contact can know that the voice repair function is started, thereby improving the experience of the call. In addition, the second prompt information is sent to the first contact in a case where the first voice is received to prompt that the first voice is being repaired, so that the first contact can know that the delay is caused by the voice repair, and the situation of the speaking time conflict of the two or more parties in the call is reduced, thereby improving the experience of the call.
[0340] In another possible implementation, at least one of the first prompt information or the second prompt information can not be sent, so that the occupation of the transmission resource in the call can be reduced.
[0341] In a possible implementation, the second prompt information includes prompt sound, and the second prompt information is sent to the first contact in a case where the first voice is received, including:
[0342] The prompt sound is continuously sent to the first contact in a case where the first voice is started to be received.
[0343] The method further includes:
[0344] The prompt sound is stopped to be sent to the first contact in a case where the repaired first voice is started to be sent to the first contact.
[0345] For example, the embodiment of the present application can refer to the description of (c) in FIG. 6, and details are not described herein again.
[0346] In the embodiment of the present application, the first contact can know that the user at the other end is speaking when hearing the prompt tone, and the first contact can hear the repaired first voice subsequently, so that the situation of conflict between the two or more parties during the call can be reduced, and the experience of the two parties during the call can be improved.
[0347] In another possible implementation, the prompt tone can be sent to the first contact when the first voice is received or after a certain time interval, so that the transmission resources required for transmitting the prompt tone can be reduced.
[0348] In a possible implementation, the remote communication scenario includes a call scenario, and the method further includes:
[0349] The repaired first voice is sent to the first contact in the call. Then, the first interface can be displayed, the first interface being an interface for the call with the first contact, and the first interface including a first control for controlling the closing of the voice repair function. Then, in response to a first operation on the first control, the voice repair function is closed, and a second voice is sent to the first contact, the second voice including the voice of the user received after the voice repair function is closed.
[0350] For example, the first interface can be the interface shown in FIG. 5. The first control can be the sound switching control 502. The first operation can be an operation of the user. In the embodiment of the present application, the repaired voice is sent to the first contact before the voice repair function is closed, that is, when the voice repair function is in an open state, and the voice before repair is sent to the first contact after the voice repair function is closed.
[0351] In the embodiment of the present application, the closing of the voice repair can be controlled during the call, so that the user can control the closing of the voice repair when the voice repair is not needed, and the switching flexibility of the use or non-use of the voice repair is improved. For example, the user can control the closing of the voice repair function through the first control when the electronic device has low power or the electronic device runs slowly, so that the closing of the voice repair function during the call can be improved, and the user experience can be improved.
[0352] In another possible implementation, the closing of the voice repair function can not be displayed on the first interface, that is, the opening of the voice repair function is not supported during the call, so that the display simplicity of the call interface can be improved.
[0353] In another possible implementation, the voice repair function can also be controlled by a voice instruction to be turned off.
[0354] In a possible implementation, in response to the first operation, the first control is also switched from the first state to a second state, the first state is used to indicate that the voice repair function is turned on, and the second state is used to indicate that the voice repair function is turned off. The method further includes:
[0355] In response to a second operation on the first control, the first control is switched from the second state to the first state, and the voice repair function is turned on to send the repaired third voice to the first contact, the third voice including the voice of the user received after the voice repair function is turned on.
[0356] The second operation can be an operation of the user. The first state can be, for example, a state of the sound switching control 502 shown in (d) of FIG. 5, and the second state can be, for example, a state of the sound switching control 502 shown in (c) of FIG. 5.
[0357] In the embodiments of the present application, whether the voice repair function is turned on or not can be known through the state of the first control, thereby improving the accuracy of the user's selection of using or not using the voice repair function. Moreover, the voice repair function can be turned on again during the call in addition to being turned off during the call, thereby improving the flexibility of turning on or off the voice repair function. In addition, the turning on or off of the voice repair function is realized by the same control, thereby improving the simplicity of the interface during the call.
[0358] It should be understood that the WeChat application and the face-to-face communication application can also support the turning off or turning on of the voice repair function in the interface of the application. The related description of how to turn off or turn on the voice repair function in the call interface can be referred to, and details are not described herein.
[0359] In another possible implementation, after the voice repair function is turned off during the call, the voice repair function can no longer be turned on again.
[0360] In another possible implementation, the turning on or off of the voice repair function can also be controlled by different controls, for example, one control is used to control the turning on of the voice repair function during the call, and another control is used to control the turning off of the voice repair function, thereby improving the control independence of the turning on or off of the voice repair function.
[0361] In a possible implementation, the method further includes:
[0362] The second interface is displayed, and the second interface includes a playing control of the repaired first voice and a first text corresponding to the repaired first voice. A third operation is received, and the third operation is used for selecting a first target character in the first text. In response to the third operation, at least one candidate character related to the first target character is displayed. A fourth operation is received, and the fourth operation is used for selecting a second target character in the at least one candidate character. In response to the fourth operation, a second text and a playing control of a voice corresponding to the second text are displayed, and the second text is obtained by replacing the first target character in the first text with the second target character.
[0363] For example, the second interface can be the interface shown in FIG. 8, the playing control of the repaired first voice can be the playing control 803, the first text corresponding to the first voice can be, for example, "today go or not go in the afternoon". The third operation can be an operation of the user. The first target character can be a character selected by the user. Optionally, the first target character can include, but is not limited to, a character, a punctuation mark, an emoticon, or the like. For example, the first target character can be, for example, "afternoon". The at least one candidate character can be displayed in a list. For example, the at least one candidate character can be, for example, the characters displayed in the candidate word list 806, such as candidate characters "morning", "good dance", and "fog". The fourth operation can be an operation of the user. The second target character can be a character selected by the user from the at least one candidate character. For example, the second target character can be, for example, "morning". For example, the second text can be, for example, "today go or not go in the morning". The playing control of the voice corresponding to the second text can also be, for example, the playing control 803.
[0364] In the embodiments of the present application, by displaying the second interface, the second interface includes a playing control of the repaired first voice and a first text corresponding to the repaired first voice. A third operation is received, and the third operation is used for selecting a first target character in the first text. In response to the third operation, at least one candidate character related to the first target character is displayed. A fourth operation is received, and the fourth operation is used for selecting a second target character in the at least one candidate character. In response to the fourth operation, a second text and a playing control of a voice corresponding to the second text are displayed, so that when the result of voice repair is inaccurate, the user can manually and quickly correct it, thereby improving the accuracy of communication through voice.
[0365] In another possible implementation, the user can also re-input the voice when finding that the result of voice repair is inaccurate.
[0366] In another possible implementation, the user can also generate the second text and the playing control corresponding to the second text by one key, and at this time, at least one character or word in the first text can be replaced to generate the second text and the playing control corresponding to the second text.
[0367] In a possible implementation, the method further includes:
[0368] displaying the playing control of the first voice after the repair is cancelled.
[0369] In the embodiments of the present application, by displaying the playing control of the first voice after the repair is cancelled, only the playing control of the latest voice is displayed each time, which can improve the convenience of selecting a suitable voice for reporting.
[0370] In another possible implementation, the displaying of the playing control of the first voice after the repair is not cancelled, that is, not only the playing control of the latest voice is reserved, but also the playing control of the historical voice is reserved.
[0371] In a possible implementation, in the remote communication scenario, the method further includes:
[0372] displaying a third interface, the third interface being an interface for communicating with the second contact, the third interface including a second control and a third control, the second control being used to instruct to send the first voice to the second contact, and the third control being used to instruct to send the first voice after the repair to the second contact; and then, in response to an operation on the third control, the first voice after the repair is sent to the second contact. Alternatively, in response to an operation on the second control, the first voice is sent to the second contact.
[0373] For example, the third interface may, for example, be the interfaces shown in (a) to (f) in FIG. 7. The second control may, for example, be the direct sending option 703, and the third control may, for example, be the repair and send option 704.
[0374] In the embodiments of the present application, the user can selectively send the voice after the repair or the voice before the repair to the second contact through the second control and the third control, that is, even if the voice repair function is enabled, the user can still select to send the original voice to the second contact, thereby improving the flexibility of the user in remote communication.
[0375] In another possible implementation, there may also be only one voice sending control, and the electronic device sends the voice after the repair after detecting an operation on the voice sending control, which can improve the efficiency of sending the voice after the repair in the remote communication scenario.
[0376] In a possible implementation, in the face-to-face communication scenario, the method further includes:
[0377] The fourth interface is displayed, and the fourth interface includes a virtual keyboard. In response to an operation on the virtual keyboard, the third text is displayed. A fifth operation on a fourth control in the fourth interface is received, and the fourth control is used to indicate generation of speech. In response to the fifth operation, speech corresponding to the third text is generated according to the first speech feature.
[0378] For example, the fourth interface may, for example, be the interfaces shown in (c)-(i) in FIG. 9. The third text may be text input by the user through the virtual keyboard, for example, the third text may be "Hello, I am Zhang San" or "I want to ask, what time is the closing today", etc. The fourth control may, for example, be the control 809. The fifth operation may be an operation of the user. In an embodiment of the present application, the speech corresponding to the third text is generated according to the first speech feature, and the speech corresponding to the third text can be broadcast according to the first speech feature.
[0379] In an embodiment of the present application, by displaying the fourth interface, the fourth interface includes a virtual keyboard. In response to an operation on the virtual keyboard, the third text is displayed. A fifth operation on a fourth control in the fourth interface is received, and the fourth control is used to indicate generation of speech. In response to the fifth operation, speech corresponding to the third text is generated according to the first speech feature, that is, the user can also input text and then generate speech according to the registered first speech feature, that is, the conversion from text to speech is realized, and the user can communicate according to the selection of input speech or input text, improving the selectivity and flexibility of the user to communicate.
[0380] In another possible implementation, the user can also automatically generate the speech corresponding to the input third text after the input of the third text through the virtual keyboard is completed, which can reduce the operation of the user. Optionally, it can be considered that the input of the third text is completed when no operation on the virtual keyboard is detected within a certain time.
[0381] In a possible implementation, the third text includes punctuation marks and / or emoticons, and when the speech corresponding to the third text is generated according to the first speech feature, the speech corresponding to the third text is also generated according to the punctuation marks and / or emoticons to control the tone of the speech.
[0382] In the embodiments of the present application, the tone of the voice corresponding to the third text is controlled according to the punctuation marks and / or emoticons in the third text, so that the tone of the voice corresponding to the third text matches the punctuation marks and / or emoticons. For example, if the punctuation mark of the last character of the third text is "!", the tone of the voice corresponding to the third text can be an exclamation tone; if the punctuation mark of the last character of the third text is "?", the tone of the voice corresponding to the third text can be a question tone. For example, if the emoticon in the third text includes a smiling face, the tone of the voice corresponding to the third text is a happy tone; if the emoticon in the third text includes a crying face, the tone of the voice corresponding to the third text is a sad tone.
[0383] In the embodiments of the present application, the tone of the voice corresponding to the third text can be controlled by the punctuation marks and / or emoticons in the third text input by the user, so that the tone of the output voice can be adaptively adjusted according to the user input text, thereby improving the flexibility of voice broadcasting and improving the user experience.
[0384] In another possible implementation, the punctuation marks or emoticons in the third text can not be considered, which can reduce the computing resources required to determine the tone matching the punctuation marks or the tone matching the emoticons. Alternatively, the mapping relationship between the punctuation marks and the tone, and the mapping relationship between the emoticons and the tone can be pre-set to determine the tone matching the punctuation marks or the tone matching the emoticons.
[0385] In a possible implementation, the first voice feature of the user is obtained by voice registration, including:
[0386] The fifth interface is displayed, the fifth interface being a voice registration interface, the fifth interface including third prompt information and a fifth control, the third prompt information being used to prompt the recording content of the voice registration, and the fifth control being used to instruct recording of a voice. In response to an operation on the fifth control, a fourth voice is recorded. The first voice feature is extracted from the fourth voice.
[0387] For example, the fifth interface can be, for example, the interface shown in FIG. 2. The third prompt information can be, for example, textual prompt information such as "where the giant panda goes, where it sleeps", or voice prompt information, which is not limited herein. The fifth control can be, for example, the recording control 208. The fourth voice can be a voice input by the user, which is used to register the voice of the user, so as to extract the first voice feature of the user, thereby enabling the voice to be broadcast according to the first voice feature of the user when the repaired voice is broadcast.
[0388] In a possible implementation, the method further includes:
[0389] The fourth speech is subjected to speech recognition to obtain a fourth text. Then, the fourth text is compared with the recording content to obtain a similarity between the fourth text and the recording content. Then, a target repair model can be selected from the plurality of repair models based on the similarity, the target repair model being used to repair the first speech, the speech repair capabilities of the plurality of repair models being different from each other, and the speech repair capability of the target repair model being negatively correlated with the similarity.
[0390] In the embodiments of the present application, the text obtained by subjecting the speech registered by the user to speech recognition is compared with the recording content, and then the similarity is obtained, and then the target repair model is selected from the plurality of repair models based on the similarity. In this way, the appropriate repair model can be selected according to the degree of speech barrier of the user to perform repair, so as to balance the accuracy of speech repair and the computing resource required for speech repair.
[0391] It should be understood that the plurality of repair models can be deployed in one server, and then the corresponding repair model can be obtained from the server through the model identifier to perform speech repair. In another possible implementation, the plurality of repair models can also be deployed in a plurality of servers respectively, and then the server corresponding to the target repair model can be called to perform speech repair according to the correspondence between the server and the repair model. In addition, the plurality of repair models can also be deployed in the electronic device, which is not limited herein.
[0392] In another possible implementation, only one repair model can also be configured, so that the similarity matching with the recording content can also not be performed, thereby improving the efficiency of speech repair.
[0393] The above embodiments are described with respect to the embodiments of repairing the speech of the user, and the following embodiments are exemplarily described with respect to the embodiments of repairing the speech of the contact person in contact with the user.
[0394] Please refer to FIG. 14, which is a flow diagram of another speech processing method provided by the embodiments of the present application. The method of the embodiments of the present application can be executed by an electronic device, a server or a system including the electronic device and the server. The embodiments can also be directed to repairing the speech of the contact person in communication with the user. The method shown in FIG. 14 can include:
[0395] 1401, receiving a setting input, the setting input being used to indicate a scenario in which speech repair is started, a contact person in which speech repair is started or an application in which speech repair is started, the scenario including a face-to-face communication scenario or a remote communication scenario.
[0396] Wherein, S1401 can refer to the description of S1301, which is not described herein.
[0397] 1402, receiving a fifth speech from a target contact person.
[0398] The target contact can be a contact with which the user communicates, for example, the target contact can be the first contact or the second contact.
[0399] 1403、obtaining a second voice feature, the second voice feature being a preset voice feature or a voice feature extracted from a fifth voice.
[0400] The second voice feature can include at least one of a voiceprint feature or a prosody feature. In another possible implementation, the second voice feature can also be a voice feature extracted from a voice registered by a user. For example, the user using the electronic device is a first user, and the first user wants to hear the voice of a second user during voice repair, the second user can also be registered for voice, and then a voice feature is extracted from the voice registered by the second user.
[0401] 1404、repairing the fifth voice according to the setting input and the second voice feature.
[0402] In the embodiments of the present application, the fifth voice is repaired according to the setting input and the second voice feature, and since the setting input is used to indicate a scenario in which voice repair is enabled, a contact in which voice repair is enabled, or an application in which voice repair is enabled, when it is detected that the electronic device is in the scenario in which voice repair is enabled, the contact with which the electronic device communicates is the contact in which voice repair is enabled, or the application running on the electronic device is the application in which voice repair is enabled, the fifth voice is repaired according to the second voice feature, so that the intelligibility of the repaired fifth voice is higher than that of the fifth voice before repair. The intelligibility can represent the degree of understanding when the voice is listened to, and can also be understood as the accuracy of the voice expression contact intended to express.
[0403] In a possible implementation, in the remote communication scenario, the method further includes:
[0404] displaying a sixth interface, the sixth interface being an interface for communicating with the target contact, the sixth interface including the fifth voice. Then, in response to an operation on the fifth voice, a sixth control and a seventh control are displayed, the sixth control being used to indicate conversion of the fifth voice into text, and the seventh control being used to indicate conversion of the repaired fifth voice into text. Then, in response to an operation on the seventh control, text corresponding to the repaired fifth voice is displayed. Alternatively, in response to an operation on the sixth control, text corresponding to the fifth voice is displayed.
[0405] For example, the sixth interface can be the interface shown in (g)-(i) in FIG. 7, the sixth control can be the voice-to-text option 708, and the seventh control can be the voice-repaired-to-text option 707.
[0406] In the embodiment of the present application, after receiving the voice from the target contact, the fifth voice can be selectively converted into text, which can also be understood as converting the original voice of the target contact into text; or the repaired fifth voice can be converted into text, and the user can select the displayed text according to the needs, thereby improving the flexibility and experience of user communication.
[0407] In another possible implementation, the sixth control can be used to indicate playing the fifth voice, and the seventh control can be used to indicate playing the repaired fifth voice. In the embodiment of the present application, the user can select to play the original voice or the repaired voice of the target contact, thereby improving the flexibility and experience of user communication.
[0408] In a possible implementation, in the face-to-face communication scenario, the method further includes:
[0409] The seventh interface is an interface for communication with the target contact, and the seventh interface includes an eighth control. Then, in response to a sixth operation on the eighth control, the fifth text is displayed. Then, in response to a seventh operation on the eighth control, the display of the fifth text is stopped, and the fifth text includes text corresponding to the fifth voice or text corresponding to the repaired fifth voice in a target time period, and the target time period includes a time period from the response to the sixth operation to the response to the seventh operation.
[0410] For example, the seventh interface can be, for example, the interface shown in (a)-(b) of FIG. 9, and the eighth control can be, for example, the recording control 802. Wherein, the sixth operation can be an operation input by the user. The fifth text can be, for example, “I am a hard-of-hearing person, and I am speaking”.
[0411] In the embodiment of the present application, the voice of the target contact can also be converted into text, thereby improving the flexibility of the communication mode between the user and the contact.
[0412] In another possible implementation, the eighth control can also not be set to control the conversion of the voice of the contact into text, for example, all voices of the contact are converted into text, which can reduce the operation of converting the voice of the contact into text.
[0413] The following is an exemplary description of how to implement voice repair based on the above embodiments.
[0414] Please refer to FIG. 15, which is a flowchart of another voice processing method provided by the embodiment of the present application. The method of the embodiment of the present application can be executed by an electronic device, a server, or a system including an electronic device and a server. The method shown in FIG. 15 can include:
[0415] S1501, obtaining a target voice and a voice feature.
[0416] The target voice can be a voice that has not been repaired and can be understood as an original voice or an audio to be repaired. For example, the target voice can include, but is not limited to, the first voice or the fifth voice. The voice feature can include, but is not limited to, the first voice feature or the second voice feature.
[0417] S1502, input the target voice and the voice feature into a target repair model, the target repair model is used to extract the content of the target voice, obtain the voice continuity feature according to the content of the target voice and the voice feature, and synthesize the repaired target voice according to the voice continuity feature.
[0418] Optionally, the content of the target voice can be obtained through voice recognition. In the embodiment of the application, the obtained target voice can be a broadcast of the extracted content of the target voice according to the voice feature.
[0419] S1503, obtain the repaired target voice output by the target repair model.
[0420] For example, the embodiment can refer to the related description of FIG. 12, which will not be repeated here.
[0421] In the embodiment of the application, after inputting the target voice, the electronic device can obtain the target voice and the voice feature, and then input the target voice and the voice feature into the target repair model. The target repair model is used to extract the content of the target voice, obtain the voice continuity feature according to the content of the target voice and the voice feature, and synthesize the repaired target voice according to the voice continuity feature. In this way, the speech-impaired population can achieve communication through input voice, thereby improving the convenience of communication of the speech-impaired population.
[0422] In a possible implementation, the target repair model includes a first module, a second module and a third module. The first module is used to extract the content of the target voice. The second module is used to obtain the voice continuity feature according to the content of the target voice and the voice feature. The third module is used to synthesize the repaired target voice according to the voice continuity feature.
[0423] For example, the first module can include a streaming ASR encoder. The second module can include a streaming audio repair large model. The third module can include a vocoder.
[0424] In a possible implementation, inputting the target voice and the voice feature into the target repair model includes:
[0425] The voice feature and the first part of the target voice are input into the target repairing model, the first module is configured to extract the content of the first part of the voice, the second module is configured to obtain target voice continuous features according to the voice feature and the content of the first part of the voice, and the third module is configured to synthesize the repaired first part of the voice according to the target voice continuous features. Then, the voice feature and the second part of the target voice are input into the target repairing model, the first module is configured to extract the content of the second part of the voice, the second module is configured to obtain second voice continuous features according to the voice feature and the content of the second part of the voice, and the third module is configured to synthesize the repaired second part of the voice according to the second voice continuous features.
[0426] For example, the first part of the voice and the second part of the voice can refer to the description of (c) in FIG. 6. For example, the first part of the voice includes "thank you, put it downstairs", and the first part of the voice can include "it is OK".
[0427] The repaired target voice includes the repaired first part of the voice and the repaired second part of the voice.
[0428] In the embodiment of the present application, by repairing a part of the target voice first and then repairing another part of the target voice, the voice repairing can be started when a part of the voice is obtained, that is, the voice repairing can be started without obtaining the complete target voice, thereby improving the efficiency of voice repairing.
[0429] In another possible implementation, the voice repairing can be started after obtaining the complete target voice, thereby reducing the situation that the target voice is obtained and the target voice is repaired at the same time, thereby reducing the resources required for voice repairing.
[0430] In a possible implementation, the target repairing model further includes a fourth module, the fourth module is configured to perform a discretization process on the first part of the voice to obtain target voice discrete features, the second module is further configured to perform prediction according to the target voice discrete features to obtain repaired target voice discrete features, and is configured to obtain the second voice continuous features according to the repaired target voice discrete features, the voice feature and the content of the second part of the voice.
[0431] For example, the fourth module can include an audio discretization unit.
[0432] In the embodiment of the present application, the fourth module is used to perform discretization processing on the first part of speech to obtain target speech discretization features, and then the second module is used to perform prediction according to the target speech discretization features to obtain repaired target speech discretization features, and then the second speech continuous features are obtained according to the repaired target speech discretization features, speech features and the content of the second part of speech. In this way, the second speech continuous features can be obtained by combining the target speech discretization features, speech features and the content of the second part of speech, thereby improving the accuracy of the obtained second speech continuous features and further improving the accuracy of speech repair.
[0433] In another possible implementation, the fourth module can also not be needed, which can improve the efficiency of speech repair.
[0434] In a possible implementation, the fourth module is configured to perform discretization processing on the first part of speech to obtain target speech discretization features, and the fourth module comprises:
[0435] The fourth module is configured to perform discretization processing on the first part of speech according to the content of the first part of speech to obtain target speech discretization features.
[0436] In the embodiment of the present application, the target speech discretization features are obtained by performing discretization processing on the first part of speech according to the content of the first part of speech, that is, the content of the first part of speech is used as a reference for discretization processing, thereby improving the accuracy of the obtained target speech discretization features and further improving the accuracy of speech repair.
[0437] In another possible implementation, the fourth module can also be used to perform discretization processing on the first part of speech according to the content of the first part of speech to obtain target speech discretization features, which can improve the efficiency of obtaining target speech discretization features and thereby improve the efficiency of speech repair.
[0438] It should be noted that the scheme of the embodiment of the present application can not only be used in sound repair tasks, but also can be extended to tasks such as dialect to standard Chinese conversion and cross-language translation.
[0439] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or authorized by all parties, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation portal for user to choose authorization or refusal.
[0440] It should also be understood that various steps of the above-mentioned various embodiments can also be coupled with each other, and the application does not limit this. The sequence of the above-mentioned processes does not mean the order of execution, and the execution order of the processes should be determined according to their functions and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the application.
[0441] The embodiments of the application also propose a speech processing apparatus, which can execute the steps of the above-mentioned method embodiments. For example, the speech processing apparatus includes a receiving module and an output module, wherein the receiving module is configured to receive a setting input, acquire a first speech feature of a user through voice registration, receive a first speech input by the user, etc.; and the output module is configured to repair the first speech according to the setting input and the first speech feature, etc.
[0442] In another possible implementation, the receiving module is configured to receive a setting input, receive a fifth speech from a target contact, acquire a second speech feature, etc., and the output module is configured to repair the fifth speech according to the setting input and the second speech feature, etc.
[0443] In another possible implementation, the receiving module is configured to acquire a target speech and a speech feature, and the output module is configured to input the target speech and the speech feature into a target repair model, the target repair model is configured to extract a content of the target speech, obtain a speech continuity feature according to the content of the target speech and the speech feature, and synthesize a repaired target speech according to the speech continuity feature; and acquire the repaired target speech output by the target repair model.
[0444] The steps performed by the speech processing apparatus can refer to the description of the method embodiments, which will not be repeated here.
[0445] It should be understood that the steps performed by the apparatus of the embodiments of the application can refer to the description of the above method embodiments, which will not be repeated here.
[0446] It should be understood that the apparatus herein is embodied in the form of functional modules. The term "module" herein can refer to an application specific integrated circuit (ASIC), an electronic circuit, a processor (for example, a shared processor, a dedicated processor, or a group processor, etc.) and a memory for executing one or more software or firmware programs, a combination logic circuit, and / or other suitable components supporting the described functions. In an optional example, those skilled in the art can understand that the apparatus can be specifically the first electronic device, the second electronic device, or the fourth electronic device in the above-mentioned embodiments, or the functions described in the above-mentioned embodiments can be integrated in the apparatus, and the apparatus can be used to execute the various processes and / or steps corresponding to the electronic device or the server in the above-mentioned method embodiments. To avoid repetition, they will not be repeated here.
[0447] The apparatus has functions of implementing corresponding steps of the electronic device or the server in the method; the functions can be implemented by hardware, or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the functions.
[0448] In embodiments of the present application, the apparatus can also be a chip or a chip system, for example, a system on chip (SoC).
[0449] Embodiments of the present application also provide another schematic block diagram of a voice processing apparatus. The apparatus includes a processor 1601, a transceiver 1602 and a memory 1603. The processor 1601, the transceiver 1602 and the memory 1603 communicate with each other through internal connection paths. The memory 1603 is used to store instructions, and the processor 1601 is used to execute the instructions stored in the memory 1603 to control the transceiver 1602 to send signals and / or receive signals.
[0450] It should be understood that the apparatus can be specifically the electronic device or the server in the above embodiments, and can be used to execute the respective steps and / or processes corresponding to the electronic device or the server in the above method embodiments. Optionally, the memory 1603 can include a read-only memory 1603 and a random access memory 1603, and provide the processor 1601 with instructions and data. Part of the memory 1603 can also include a non-volatile random access memory 1603. For example, the memory 1603 can also store device type information. The processor 1601 can be used to execute the instructions stored in the memory 1603, and when the processor 1601 executes the instructions stored in the memory 1603, the processor 1601 is used to execute the respective steps and / or processes of the above method embodiments. The transceiver 1602 can include a transmitter and a receiver. The transmitter can be used to implement the respective steps and / or processes corresponding to the transmitter of the transceiver 1602 for executing the sending action, and the receiver can be used to implement the respective steps and / or processes corresponding to the receiver of the transceiver 1602 for executing the receiving action.
[0451] It should be understood that in embodiments of the present application, the processor 1601 can be a central processing unit (CPU), and the processor 1601 can also be other general-purpose processors 1601, digital signal processors 1601 (DSP), application-specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The general-purpose processor 1601 can be a microprocessor or the processor 1601 can also be any conventional processor 1601, etc.
[0452] In the implementation process, each step of the above method can be completed by the integrated logic circuit of hardware in the processor 1601 or the instruction in the form of software. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as the execution completed by the hardware processor 1601, or executed by the combination of hardware and software modules in the processor 1601. The software module can be located in the storage medium in the field, such as random access memory 1603, flash memory, read-only memory 1603, programmable read-only memory 1603, or electrically erasable programmable memory 1603, register, etc. The storage medium is located in the memory 1603, and the processor 1601 executes the instruction in the memory 1603, and combines the hardware to complete the steps of the above method. To avoid repetition, it will not be described in detail here.
[0453] The embodiments of the present application also provide a voice processing system, comprising a terminal device and a server. The steps performed by the terminal device and the server can refer to the description of the above method embodiments, and will not be described here.
[0454] The embodiments of the present application also provide a computer readable storage medium for storing a computer program for implementing the method shown in the above method embodiments.
[0455] The embodiments of the present application also provide a computer program product, which includes a computer program (also can be called code or instruction), when the computer program runs on the computer, the computer can execute the method shown in the above method embodiments.
[0456] Those skilled in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be realized by electronic hardware or combination of computer software and electronic hardware. Whether the functions are executed in hardware or software mode depends on the specific application and design constraints of the technical solution. The skilled person can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0457] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the above described system, device and unit can refer to the corresponding process in the foregoing method embodiments, which will not be described here.
[0458] In several embodiments provided in the present application, it should be understood that the disclosed system, device and method can be implemented by other other ways. For example, the above-described device embodiments are merely illustrative, for example, the division of units is merely a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the displayed or discussed units can be indirect coupling or communication connection through some interfaces, devices or units, and can be electrical, mechanical or other other forms.
[0459] The units described as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, that is, can be located in one place, or can be distributed to multiple network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiments of the present application.
[0460] In addition, the functional units in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit.
[0461] If the functions are realized in the form of software function units and sold or used as independent products, they can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the embodiments of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various program code storage media.
Claims
1. A voice processing method, characterized by, The method comprises: receiving a setting input, the setting input being used to indicate a scenario in which voice repair is enabled, a contact in which voice repair is enabled, or an application in which voice repair is enabled, the scenario comprising a face-to-face communication scenario or a remote communication scenario; obtaining a first voice feature of a user through voice registration; receiving a first voice input by the user; repairing the first voice according to the setting input and the first voice feature.
2. The method of claim 1, wherein, The remote communication scenario comprises a call scenario, and the method further comprises: sending first prompt information to a first contact in a call, the first prompt information being used to prompt that a voice repair function has been enabled; and / or, sending second prompt information to the first contact in a case where the first voice is received, the second prompt information being used to prompt that the first voice is being repaired.
3. The method of claim 2, wherein, The second prompt information comprises a prompt tone, and the sending of the second prompt information to the first contact in a case where the first voice is received comprises: continuously sending the prompt tone to the first contact in a case where the first voice is started to be received; The method further comprises: stopping the sending of the prompt tone to the first contact in a case where the repaired first voice is started to be sent to the first contact.
4. The method according to any one of claims 1 to 3, characterized in that, The remote communication scenario comprises a call scenario, and the method further comprises: sending the repaired first voice to a first contact in a call; displaying a first interface, the first interface being an interface for a call with the first contact, the first interface comprising a first control, the first control being used to control the closing of a voice repair function; in response to a first operation on the first control, closing the voice repair function to send a second voice to the first contact, the second voice comprising a voice of the user received after the voice repair function is closed.
5. The method of claim 4, wherein, In response to the first operation, the first control is further switched from a first state to a second state, the first state being used to indicate that the voice repair function has been enabled, and the second state being used to indicate that the voice repair function has been closed, and the method further comprises: in response to a second operation on the first control, switching the first control from the second state to the first state, and enabling the voice repair function to send a repaired third voice to the first contact, the third voice comprising a voice of the user received after the voice repair function is enabled.
6. The method according to any one of claims 1-5, characterized in that, The method further comprises: displaying a second interface, the second interface comprising a play control of the repaired first voice and a first text corresponding to the repaired first voice; receiving a third operation, the third operation being used to select a first target character in the first text; in response to the third operation, displaying at least one candidate character related to the first target character; receiving a fourth operation, the fourth operation being used to select a second target character in the at least one candidate character; In response to the fourth operation, a second text and a playing control of a voice corresponding to the second text are displayed, the second text being obtained by replacing the first target character in the first text with the second target character.
7. The method of claim 6, wherein, The method further includes: displaying the playing control of the repaired first voice is cancelled.
8. The method according to any one of claims 1-7, characterized in that, In the remote communication scenario, the method further includes: displaying a third interface, the third interface being an interface for communicating with a second contact, the third interface including a second control and a third control, the second control being used to instruct sending the first voice to the second contact, and the third control being used to instruct sending the repaired first voice to the second contact; in response to an operation on the third control, sending the repaired first voice to the second contact; or in response to an operation on the second control, sending the first voice to the second contact.
9. The method according to any one of claims 1-8, characterized in that, In the face-to-face communication scenario, the method further includes: displaying a fourth interface, the fourth interface including a virtual keyboard; in response to an operation on the virtual keyboard, displaying a third text; receiving a fifth operation on a fourth control in the fourth interface, the fourth control being used to instruct generating a voice; in response to the fifth operation, generating a voice corresponding to the third text according to the first voice feature.
10. The method of claim 9, wherein, The third text includes punctuation marks and / or emoticons, and when the voice corresponding to the third text is generated according to the first voice feature, the tone of the voice corresponding to the third text is controlled according to the punctuation marks and / or emoticons.
11. The method according to any one of claims 1-10, characterized in that, The first voice feature of the user is obtained through voice registration, including: displaying a fifth interface, the fifth interface being an interface for voice registration, the fifth interface including third prompt information and a fifth control, the third prompt information being used to prompt recording content of voice registration, and the fifth control being used to instruct recording a voice; in response to an operation on the fifth control, recording a fourth voice; extracting the first voice feature from the fourth voice.
12. A voice processing method, characterized by, including: receiving a setting input, the setting input being used to instruct a scenario in which voice repair is enabled, a contact with which voice repair is enabled, or an application in which voice repair is enabled, the scenario including a face-to-face communication scenario or a remote communication scenario; receiving a fifth voice from a target contact; obtaining a second voice feature, the second voice feature being a preset voice feature or a voice feature extracted from the fifth voice; repairing the fifth voice according to the setting input and the second voice feature.
13. The method of claim 12, wherein, In the remote communication scenario, the method further includes: displaying a sixth interface, the sixth interface being an interface for communicating with the target contact, the sixth interface including the fifth voice; in response to an operation on the fifth voice, displaying a sixth control and a seventh control, the sixth control being used to instruct converting the fifth voice into a text, and the seventh control being used to instruct converting a repaired fifth voice into a text; in response to an operation on the seventh control, displaying a text corresponding to the repaired fifth voice; or In response to an operation on the sixth control, text corresponding to the fifth voice is displayed.
14. The method according to claim 12 or 13, characterized in that, In the face-to-face communication scenario, the method further includes: displaying a seventh interface, the seventh interface being an interface for communication with the target contact, the seventh interface including an eighth control; in response to a sixth operation on the eighth control, displaying a fifth text; in response to a seventh operation on the eighth control, stopping display of the fifth text, the fifth text including text corresponding to the fifth voice or the repaired text corresponding to the fifth voice in a target time period, the target time period including a time period from a response to the sixth operation to a response to the seventh operation.
15. A voice processing method, characterized by, comprising: obtaining a target voice and a voice feature; inputting the target voice and the voice feature into a target repair model, the target repair model being configured to extract content of the target voice, obtain a voice continuity feature according to the content of the target voice and the voice feature, and synthesize a repaired target voice according to the voice continuity feature; obtaining the repaired target voice output by the target repair model.
16. A speech processing device, characterized by comprising: a processor coupled to a memory, the memory storing computer-executable instructions, and the processor executing the computer-executable instructions stored in the memory, so that the processor executes the method of any one of claims 1-11, or executes the method of any one of claims 12-14, or executes the method of claim 15.
17. A speech processing system characterized by comprising an electronic device and a server, the electronic device executing the method of any one of claims 1-11, or executing the method of any one of claims 12-14, and the server executing the method of claim 15.
18. A computer-readable storage medium, characterized in that, a computer program for storing instructions for implementing the method of any one of claims 1-11, or for implementing the method of any one of claims 12-14, or for implementing the method of claim 15.
19. A computer program product comprising computer program code in said computer program product, characterised in that, when the computer program code is run on a computer, the computer executes the method of any one of claims 1-11, or executes the method of any one of claims 12-14, or executes the method of claim 15.