Speech processing method, apparatus and system, and storage medium and program product
By acquiring speech features to repair the speech of people with special speech impairments, more intelligible repaired speech is generated, solving the problem of inconvenient communication for people with special speech impairments and achieving a more efficient voice communication experience and flexible control.
Patent Information
- Application Number
- PCT/CN2025/097957
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-09
- Filing Date
- 2025-05-29
- Publication Date
- 2025-12-26
AI Technical Summary
People with speech impairments often have low speech intelligibility due to pronunciation defects, making it difficult for listeners to understand them. Existing technologies that convert text into speech are inconvenient.
By receiving setting input and acquiring user voice characteristics, the system repairs the user's voice based on the setting input and voice characteristics, generating a more intelligible repaired voice, and broadcasting it in communication scenarios. It supports the on/off control of the voice repair function and provides multiple communication interfaces and repair modes.
It improves the convenience and intelligibility of communication for people with special speech impairments, enhances the communication experience, reduces call conflicts, improves the flexibility and accuracy of voice restoration, and supports multiple communication methods.
Smart Images

Figure CN2025097957_26122025_PF_FP_ABST
Abstract
Description
Speech processing methods, devices, systems, storage media and software products
[0001] This application claims priority to Chinese Patent Application No. 202410808992.4, filed on June 20, 2024, entitled "A Method and Apparatus for Assisting Communication," and to Chinese Patent Application No. 202410808992.4, filed on August 9, 2024, entitled "A Voice Processing Method, Apparatus, System, Storage Medium, and Program Product," the entire contents of which are incorporated herein by reference. Technical Field
[0002] This application relates to the field of terminal technology, and in particular to a voice processing method, apparatus, system, storage medium, and program product. Background Technology
[0003] With the development of terminal technology, more and more functions are being applied to terminals, such as user communication functions.
[0004] Currently, some users, due to congenital or acquired factors, cannot communicate as freely as others. For example, people with hearing impairments, ALS, vocal cord damage, or other speech impairments have pronunciation defects, resulting in low intelligibility of their speech and difficulty in being understood by listeners, which significantly impacts their daily social interactions. Therefore, for people with speech impairments, in communication scenarios, they need to input text into electronic devices, which can then convert the user's input text into speech and output it.
[0005] However, for people with speech impairments, communicating by inputting text on electronic devices can be inconvenient. Summary of the Invention
[0006] This application provides a voice processing method, apparatus, system, storage medium, and program product, which helps to improve the convenience of communication between users.
[0007] In a first aspect, embodiments of this application provide a voice processing method, which may include:
[0008] The system receives settings input, which indicates the scenario, contact, or application for enabling voice repair. Scenarios include face-to-face or remote communication. Then, it obtains the user's initial voice characteristics through voice registration. The order of voice registration and settings input can be reversed. When the user speaks, the system receives their initial voice input. Since the settings input was received previously, the initial voice can be repaired based on the settings input and the initial voice characteristics.
[0009] In this embodiment, the first speech is repaired based on the setting input and the first speech feature. Since the setting input is used to indicate the scenario where speech repair is enabled, the contact person for enabling speech repair, or the application for enabling speech repair, when the electronic device is detected to be in a scenario where speech repair is enabled, or when the contact being communicated by the electronic device is the contact person for enabling speech repair, or when the application being run by the electronic device is the application for enabling speech repair, the first speech is repaired based on the first speech feature. This allows the repaired first speech to be played using the first speech feature, and the intelligibility of the repaired first speech is higher than that of the unrepaired first speech. Intelligibility can represent the accuracy of expressing what the user wants to say, or it can be understood as the listener's understanding of the speech signal transmitted by the speaker.
[0010] In this embodiment, by receiving setting input—which indicates the scenario, contact, or application for enabling voice restoration—and then obtaining the user's first voice feature through voice registration, the system can restore the first voice based on the setting input and the first voice feature after receiving the user's input. This allows users with speech or hearing impairments to communicate via voice input, improving their communication convenience. Furthermore, by generating the restored first voice using the user's registered first voice feature, the restored first voice can be played according to that feature, making it more closely resemble the user's own voice and further enhancing the user experience.
[0011] In one possible implementation, the remote communication scenario includes a call scenario, and in the call scenario, the method further includes:
[0012] Send a first notification message to the first contact in the call, indicating that voice repair has been enabled. And / or, upon receiving the first voice message, send a second notification message to the first contact, indicating that the first voice message is being repaired.
[0013] In this embodiment, by sending a first notification message to the first contact in the call to inform the user that the voice repair function has been enabled, the first contact can be informed that the user has enabled the voice repair function, thereby improving the call experience. Furthermore, by sending a second notification message to the first contact upon receiving the first voice message to indicate that the first voice message is being repaired, the first contact can be informed that the delay is due to voice repair, reducing the possibility of conflicting speaking times in two-party or multi-party calls, thereby improving the call experience.
[0014] In one possible implementation, if the second notification message includes a notification tone, then sending the second notification message to the first contact upon receiving the first voice message may include:
[0015] Upon receiving the first voice message, continue sending a notification tone to the first contact. Then, upon starting to send the repaired first voice message to the first contact, stop sending notification to the first contact.
[0016] In this embodiment of the application, by continuously sending a prompt tone to the first contact when the first voice is received, the first contact can know that the user on the other end is speaking when they hear the prompt tone. And when the prompt tone is stopped after the repaired first voice is sent to the first contact, the first contact can then hear the repaired first voice. This can reduce the situation of conflict between two or more parties during the call, thereby improving the experience of both parties in the call.
[0017] In one possible implementation, the remote communication scenario includes a call scenario, and in the call scenario, the method further includes:
[0018] A repaired first voice message is sent to the first contact in the call. Then, a first interface can be displayed, which is the interface for communicating with the first contact. The first interface includes a first control for disabling the voice repair function. Then, in response to a first operation on the first control, the voice repair function is disabled, and a second voice message is sent to the first contact. The second voice message includes the user's voice received after the voice repair function was disabled.
[0019] In this embodiment, during a call, the user can also control the voice repair function to be turned off. This allows the user to disable voice repair when it is not needed, improving the flexibility of switching between using and not using voice repair. For example, the user can control the voice repair function to be turned off via the first control when the electronic device's battery is low or the electronic device is running slowly, thereby improving the ability to disable voice repair during calls and enhancing the user experience.
[0020] In one possible implementation, in response to the first operation, the first control is also switched from a first state to a second state, where the first state indicates that the voice repair function is enabled and the second state indicates that the voice repair function is disabled. The method further includes:
[0021] In response to a second operation on the first control, the first control is switched from a second state to a first state, and the voice repair function is enabled, thereby sending a repaired third voice message to the first contact, the third voice message including the voice message received from the user after the voice repair function is enabled.
[0022] In this embodiment, the state of the first control indicates whether the voice repair function is currently enabled, thereby improving the accuracy of the user's choice to use or disable the voice repair function. Furthermore, this embodiment not only allows the voice repair function to be disabled during a call but also to be re-enabled, thus increasing the flexibility of enabling or disabling voice repair. In addition, enabling or disabling the voice repair function through the same control improves the simplicity of the interface during calls.
[0023] It should be understood that WeChat and face-to-face communication applications can also enable or disable the voice repair function within the application interface. Please refer to the relevant descriptions on how to enable or disable the voice repair function in the call interface, which will not be repeated here.
[0024] In one possible implementation, the method further includes:
[0025] A second interface is displayed, including playback controls for the repaired first audio and corresponding first text. Then, if the user selects a first target character from the first text, a third operation is received, selecting that first target character from the first text. In response to the third operation, at least one candidate character associated with the first target character is displayed. Then, if the user selects one of the candidate characters, a fourth operation is received, selecting a second target character from the at least one candidate character. In response to the fourth operation, second text and playback controls for the corresponding audio are displayed; the second text is obtained by replacing the first target character in the first text with the second target character.
[0026] In this embodiment, a second interface is displayed, including playback controls for the repaired first speech and first text corresponding to the repaired first speech. A third operation is received, whereby a first target character in the first text is selected. In response to the third operation, at least one candidate character related to the first target character is displayed. A fourth operation is received, whereby a second target character is selected from the at least one candidate character. In response to the fourth operation, second text and playback controls for the speech corresponding to the second text are displayed. This allows the user to quickly and manually correct inaccurate speech repair results, thereby improving the accuracy of speech communication.
[0027] In one possible implementation, the method further includes:
[0028] Cancel the display of the playback controls for the first audio clip after the repair.
[0029] In this embodiment of the application, by canceling the display of the playback control for the repaired first voice, only the latest voice can be played at any given time, which improves the convenience of selecting the appropriate voice for playback.
[0030] In one possible implementation, for remote communication scenarios, the method also includes:
[0031] A third interface is displayed, which is the interface for communicating with the second contact. The third interface includes a second control and a third control. The second control is used to instruct the sending of a first voice message to the second contact, and the third control is used to instruct the sending of a repaired first voice message to the second contact. Then, in response to an operation on the third control, the repaired first voice message can be sent to the second contact. Alternatively, in response to an operation on the second control, the first voice message can be sent to the second contact.
[0032] In this embodiment, the user can selectively send the repaired voice or the original voice to the second contact through the second and third controls. That is, even if the voice repair function is enabled, the user can still choose to send the original voice to the second contact, thereby improving the user's flexibility in remote communication.
[0033] In one possible implementation, for face-to-face communication scenarios, the method also includes:
[0034] A fourth interface is displayed, including a virtual keyboard. The user can then interact with the virtual keyboard, and in response to these interactions, third text is displayed. The user can then interact with a fourth control on the fourth interface, which instructs the generation of speech. In response to this fifth interaction, speech corresponding to the third text is generated based on first speech characteristics.
[0035] In this embodiment, a fourth interface is displayed, including a virtual keyboard. In response to an operation on the virtual keyboard, third text is displayed. A fifth operation is received on a fourth control in the fourth interface, the fourth control indicating the generation of speech. In response to the fifth operation, speech corresponding to the third text is generated based on a first speech feature. That is, the user can also input text and then generate speech based on the registered first speech feature, thus achieving text-to-speech conversion. Users can then communicate by choosing to input either speech or text, improving the selectivity and flexibility of user communication.
[0036] In one possible implementation, the third text includes punctuation marks and / or emojis. In this case, when generating the speech corresponding to the third text based on the first speech features, the tone of the speech corresponding to the third text can also be controlled based on the punctuation marks and / or emojis.
[0037] In this embodiment, the tone of the speech corresponding to the third text can be controlled by punctuation marks and / or emoticons in the third text input by the user. This allows the tone of the output speech to be adjusted adaptively according to the user input text, thereby improving the flexibility of voice broadcasting and enhancing the user experience.
[0038] In one possible implementation, the user's initial voice characteristics are obtained through voice registration, including:
[0039] The fifth interface is displayed, which is the sound registration interface. This fifth interface includes a third prompt message and a fifth control. The third prompt message indicates the recording content for sound registration, and the fifth control indicates the recording of audio. If the user interacts with the fifth control, a fourth audio recording can be initiated in response to the interaction with the fifth control. The first speech feature is then extracted from the fourth audio recording.
[0040] In this embodiment, the recording of voice is triggered by a fifth control, which can improve the accuracy and effectiveness of voice recording and make the subsequent comparison between the recorded content and the voice recognition result more accurate, thereby improving the accuracy of the repair model selection.
[0041] In one possible implementation, the method also includes:
[0042] Speech recognition is performed on the fourth speech to obtain the fourth text. Then, the fourth text is compared with the recorded content to obtain the similarity between the fourth text and the recorded content. Then, a target restoration model can be selected from multiple restoration models based on the similarity. The target restoration model is used to restore the first speech. The speech restoration capabilities of the multiple restoration models are different, and the speech restoration capability of the target restoration model is negatively correlated with the similarity.
[0043] In this embodiment of the application, the text obtained by speech recognition of the user's registered speech is compared with the recorded content to obtain a similarity score. Then, a target repair model is selected from multiple repair models based on the similarity score. This allows for the selection of a suitable repair model based on the user's speech impairment level, thereby balancing the accuracy of speech repair with the computing resources required for speech repair.
[0044] Secondly, embodiments of this application also provide another speech processing method. This method may include:
[0045] The system receives settings input, which indicates the scenario for enabling voice repair, the contact to enable voice repair, or the application to enable voice repair. Scenarios include face-to-face communication and remote communication. Then, it can receive a fifth voice message from the target contact. When voice repair is needed, a second voice feature can be obtained, which can be a preset voice feature or a voice feature extracted from the fifth voice message. Then, the fifth voice message is repaired based on the settings input and the second voice feature.
[0046] In this embodiment, the fifth voice is repaired based on the setting input and the second voice feature. Since the setting input is used to indicate the scenario where voice repair is enabled, the contact person who enables voice repair, or the application that enables voice repair, when the electronic device is detected to be in a scenario where voice repair is enabled, when the contact person being communicated with by the electronic device is the contact person who enables voice repair, or when the application being run by the electronic device is the application that enables voice repair, the fifth voice is repaired based on the second voice feature, so that the intelligibility of the repaired fifth voice is higher than that of the unrepaired fifth voice. Intelligibility can represent the degree of understanding when the voice is heard, or it can be understood as the accuracy with which the voice expresses what the contact person intends to convey.
[0047] In one possible implementation, for remote communication scenarios, the method also includes:
[0048] The sixth interface is displayed, which is the interface for communicating with the target contact. The sixth interface includes the fifth voice message. Then, if the user interacts with the fifth voice message, a sixth and seventh control can be displayed in response. The sixth control instructs the user to convert the fifth voice message to text, and the seventh control instructs the user to convert the repaired fifth voice message to text. The user can then interact with either the sixth or seventh control. Alternatively, an interaction with the seventh control can display the text corresponding to the repaired fifth voice message.
[0049] In this embodiment of the application, after receiving the voice from the target contact, the fifth voice can be selectively converted into text, which can also be understood as converting the original voice of the target contact into text; or the repaired fifth voice can be converted into text. Users can choose to display the text as needed, thereby improving the flexibility and experience of user communication.
[0050] In one possible implementation, for face-to-face communication scenarios, the method also includes:
[0051] The seventh interface is displayed, which is the interface for communicating with the target contact. The seventh interface includes an eighth control. The user can then interact with the eighth control, which, in response to a sixth action on the eighth control, displays the fifth text. The user can then interact with the eighth control again, which, in response to a seventh action on the eighth control, stops the display of the fifth text. The fifth text includes either the text corresponding to the fifth voice message within a target time period or the text corresponding to the repaired fifth voice message. The target time period includes the period between the responses to the sixth and seventh actions.
[0052] In this embodiment of the application, the voice of the target contact can also be converted into text, thereby improving the flexibility of communication between the user and the contact.
[0053] Thirdly, embodiments of this application also provide another speech processing method, which may include:
[0054] The target speech and its features are obtained. Then, the target speech and features are input into a target restoration model. This model extracts the content of the target speech, obtains continuous speech features based on the content and features, and synthesizes the restored target speech based on these features. Finally, the restored target speech output by the target restoration model can be obtained.
[0055] In this embodiment, after the target speech is input, the electronic device can acquire the target speech and speech features, and then input the target speech and speech features into the target restoration model. The target restoration model is used to extract the content of the target speech, obtain speech continuity features based on the content and speech features of the target speech, and synthesize the restored target speech based on the speech continuity features. In this way, people with speech impairments can communicate by inputting speech, thereby improving the convenience of communication for people with speech impairments.
[0056] In one possible implementation, the target restoration model includes a first module, a second module, and a third module. The first module is used to extract the content of the target speech, the second module is used to obtain continuous speech features based on the content and speech features of the target speech, and the third module is used to synthesize the restored target speech based on the continuous speech features.
[0057] In one possible implementation, the target speech and speech features are input into the target inpainting model, including:
[0058] The speech features and a first portion of the target speech are input into the target restoration model. The first module extracts the content of the first portion of speech, the second module obtains continuous features of the target speech based on the speech features and the content of the first portion, and the third module synthesizes the restored first portion of speech based on the continuous features of the target speech. Then, the speech features and a second portion of the target speech are input into the target restoration model. The first module extracts the content of the second portion of speech, the second module obtains continuous features of the second portion of speech based on the speech features and the content of the second portion, and the third module synthesizes the restored second portion of speech based on the continuous features of the second portion.
[0059] In this embodiment, by first repairing a portion of the target speech and then repairing another portion, speech repair can begin when only a portion of the speech is acquired. In other words, speech repair can begin without acquiring the complete target speech, thereby improving the efficiency of speech repair.
[0060] In one possible implementation, the target restoration model further includes a fourth module, which is used to discretize the first part of the speech to obtain discrete features of the target speech. The second module is also used to make predictions based on the discrete features of the target speech to obtain the restored discrete features of the target speech, and to obtain the second continuous features of the speech based on the restored discrete features of the target speech, the speech features, and the content of the second part of the speech.
[0061] In this embodiment, the first part of the speech is discretized by the fourth module to obtain the discrete features of the target speech. Then, the second module predicts based on the discrete features of the target speech to obtain the repaired discrete features of the target speech. Then, the second continuous features of the speech are obtained based on the repaired discrete features of the target speech, the speech features, and the content of the second part of the speech. In this way, the second continuous features of the speech can be obtained by combining the discrete features of the target speech, the speech features, and the content of the second part of the speech, thereby improving the accuracy of the obtained second continuous features of the speech and thus improving the accuracy of speech repair.
[0062] In one possible implementation, the fourth module is used to discretize the first part of the speech to obtain discrete features of the target speech, including:
[0063] The fourth module is used to discretize the first part of the speech based on its content to obtain the discrete features of the target speech.
[0064] In this embodiment of the application, the first part of the speech is discretized according to the content of the first part of the speech to obtain the discrete features of the target speech. That is, the content of the first part of the speech is used as a reference for the discretization process, thereby improving the accuracy of the obtained discrete features of the target speech and thus improving the accuracy of speech restoration.
[0065] It should be noted that the solution in this application embodiment can be used not only for sound restoration tasks, but also extended to tasks such as dialect to Mandarin conversion and cross-language translation.
[0066] Fourthly, another voice processing apparatus is provided, including a processor coupled to a memory for executing instructions in the memory to implement the methods in any of the possible implementations of the first aspect described above. Optionally, the apparatus further includes a memory. Optionally, the apparatus also includes a communication interface to which the processor is coupled.
[0067] Fifthly, a processor is provided, comprising: an input circuit, an output circuit, and a processing circuit. The processing circuit is used to receive signals through the input circuit and transmit signals through the output circuit, causing the processor to execute the method in any possible implementation of the first aspect described above.
[0068] In specific implementation, the processor can be a chip, the input circuit can be an input pin, the output circuit can be an output pin, and the processing circuit can be a transistor, gate circuit, flip-flop, and various logic circuits. The input signal received by the input circuit can be received and input by, for example, but not limited to, a receiver, and the signal output by the output circuit can be output to, for example, but not limited to, a transmitter and transmitted by the transmitter. Furthermore, the input circuit and the output circuit can be the same circuit, which is used as the input circuit and the output circuit at different times. This application does not limit the specific implementation of the processor and various circuits.
[0069] In a sixth aspect, a processing apparatus is provided, including a processor and a memory. The processor is configured to read instructions stored in the memory and to receive signals via a receiver and transmit signals via a transmitter to execute the method in any of the possible implementations of the first aspect described above.
[0070] Optionally, there may be one or more processors and one or more memories.
[0071] Alternatively, the memory can be integrated with the processor, or the memory can be set up separately from the processor.
[0072] In specific implementation, the memory can be a non-transitory memory, such as read-only memory (ROM), which can be integrated with the processor on the same chip or set on different chips. The embodiments of this application do not limit the type of memory or the way the memory and processor are set.
[0073] It should be understood that the relevant data interaction process, such as sending indication information, can be the process of outputting indication information from the processor, and receiving capability information can be the process of the processor receiving input capability information. Specifically, the processed output data can be output to the transmitter, and the input data received by the processor can come from the receiver. Here, the transmitter and receiver can be collectively referred to as a transceiver.
[0074] The processing device in the fourth aspect above can be a chip. The processor can be implemented in hardware or software. When implemented in hardware, the processor can be a logic circuit, integrated circuit, etc. When implemented in software, the processor can be a general-purpose processor that reads software code stored in memory. The memory can be integrated into the processor or located outside the processor and exist independently.
[0075] In a seventh aspect, a computer program product is provided, comprising: a computer program (also referred to as code or instructions) that, when executed, causes a computer to perform the method in any of the possible implementations of the first aspect described above.
[0076] Eighthly, a computer-readable storage medium is provided that stores a computer program (also referred to as code or instructions) that, when executed on a computer, causes the computer to perform the methods in any of the possible implementations of the first aspect described above. Attached Figure Description
[0077] Figure 1 is a schematic diagram of a system architecture provided in an embodiment of this application;
[0078] Figure 2 is a schematic diagram of a sound registration interface provided in an embodiment of this application;
[0079] Figure 3 is a schematic diagram of the interface for configuring an application using the sound restoration function according to an embodiment of this application;
[0080] Figure 4 is a schematic diagram of an interface for configuring a whitelist for sound restoration provided in an embodiment of this application;
[0081] Figure 5 is a schematic diagram of an interface for a remote real-time communication scenario provided in an embodiment of this application;
[0082] Figure 6 is a schematic diagram of a call scenario provided in an embodiment of this application;
[0083] Figure 7 is a schematic diagram of an interface for a remote non-real-time communication scenario provided in an embodiment of this application;
[0084] Figure 8 is a schematic diagram of an interface for a face-to-face communication scenario provided in an embodiment of this application;
[0085] Figure 9 is a schematic diagram of another face-to-face communication scenario provided in an embodiment of this application;
[0086] Figure 10 is a flowchart illustrating a speech processing method provided in an embodiment of this application;
[0087] Figure 11 is a flowchart illustrating a sound registration method provided in an embodiment of this application;
[0088] Figure 12 is a schematic diagram of the architecture of a repair model provided in an embodiment of this application;
[0089] Figure 13 is a flowchart illustrating another speech processing method provided in an embodiment of this application;
[0090] Figure 14 is a flowchart illustrating another speech processing method provided in an embodiment of this application;
[0091] Figure 15 is a flowchart illustrating another speech processing method provided in an embodiment of this application;
[0092] Figure 16 is a schematic diagram of the structure of a voice processing device provided in an embodiment of this application. Detailed Implementation
[0093] To facilitate a clear description of the technical solutions in the embodiments of this application, some terms and technologies involved in the embodiments of this application will be briefly introduced below:
[0094] 1. Electronic equipment:
[0095] The electronic devices in the embodiments of this application may also be referred to as: terminal equipment, user equipment (UE), mobile station (MS), mobile terminal (MT), access terminal, user unit, user station, mobile station, mobile station, remote station, remote terminal, mobile device, user terminal, terminal, wireless communication equipment, user agent, or user device, etc.
[0096] Electronic devices can be devices that provide users with voice / data connectivity, such as handheld devices with wireless connectivity, in-vehicle devices, etc. Currently, examples of terminals include: mobile phones, tablets, laptops, PDAs, mobile internet devices (MIDs), wearable devices, virtual reality (VR) devices, augmented reality (AR) devices, wireless terminals in industrial control, wireless terminals in self-driving vehicles, wireless terminals in remote medical surgery, wireless terminals in smart grids, wireless terminals in transportation safety, wireless terminals in smart cities, wireless terminals in smart homes, cellular phones, cordless phones, session initiation protocol (SIP) phones, wireless local loop (WLL) stations, personal digital assistants (PDAs), handheld devices with wireless communication capabilities, computing devices or other processing devices connected to wireless modems, in-vehicle devices, wearable devices, electronic devices in 5G networks, or future evolution of public land mobile communication networks. The embodiments of this application do not limit the scope of electronic devices in a network (PLMN).
[0097] By way of example and not limitation, in this embodiment, the electronic device can also be a wearable device. Wearable devices, also known as wearable smart devices, are a general term for devices that utilize wearable technology to intelligently design and develop everyday wearables, such as glasses, gloves, watches, clothing, and shoes. Wearable devices are portable devices that are worn directly on the body or integrated into the user's clothing or accessories. Wearable devices are not merely hardware devices, but also achieve powerful functions through software support, data interaction, and cloud interaction. Broadly speaking, wearable smart devices include those that are feature-rich, large in size, and can achieve complete or partial functions without relying on a smartphone, such as smartwatches or smart glasses, as well as those that focus on a specific type of application function and require the use of other devices such as smartphones, such as various smart bracelets and smart jewelry for vital sign monitoring.
[0098] Furthermore, in this application embodiment, the electronic device can also be an electronic device in an Internet of Things (IoT) system. IoT is an important component of future information technology development, and its main technical feature is connecting objects to networks through communication technology, thereby realizing an intelligent network of human-machine interconnection and object-to-object interconnection. The electronic device in this application can also be an on-board unit, on-board module, on-board component, on-board chip, or on-board unit built into a vehicle as one or more components or units. The vehicle can implement the methods of this application through the built-in on-board unit, on-board module, on-board component, on-board chip, or on-board unit. Therefore, the embodiments of this application can be applied to vehicle networking, such as vehicle-to-everything (V2X), long-term evolution-vehicle (LTE-V) communication, and vehicle-to-vehicle (V2V). Optionally, the electronic device in the embodiments of this application can also be simply referred to as a device.
[0099] 2. Access network equipment:
[0100] The access network device in this application embodiment can also be called a wireless access network device. It can be a transmission reception point (TRP), an evolved NodeB (eNB or eNodeB) in an LTE system, a home base station (e.g., home evolved NodeB or home Node B, HNB), a base band unit (BBU), a wireless controller in a cloud radio access network (CRAN) scenario, or the access network device can be a relay station, access point, vehicle-mounted equipment, wearable device, or access network device in a 5G network or an access network device in a future evolved PLMN network, etc. It can be an access point (AP) in a WLAN, a gNB in a new radio (NR) system, a satellite base station in a satellite communication system, etc. The embodiments of this application are not limited.
[0101] The access network equipment in this embodiment may include centralized unit (CU) nodes, distributed unit (DU) nodes, or access network equipment including CU nodes and DU nodes, or access network equipment including control plane CU nodes (CU-CP nodes), user plane CU nodes (CU-UP nodes), and DU nodes. Access network equipment including CU nodes and DU nodes can separate the protocol layers of the access network equipment, with some protocol layer functions centrally controlled by the CU, and the remaining part or all protocol layer functions distributed in the DU, which is centrally controlled by the CU. As one implementation, the CU deployment protocol stack includes a radio resource control (RRC) layer, a packet data convergence protocol (PDCP) layer, and a service data adaptation protocol (SDAP) layer. The DU deployment protocol stack includes a radio link control (RLC) layer, a media access control (MAC) layer, and a physical layer (PHY). Thus, the CU has RRC, PDCP, and SDAP processing capabilities. The DU has RLC, MAC, and PHY processing capabilities. The above functional division is merely an example and does not constitute a limitation on the CU and DU. That is, there can be other ways to divide functions between the CU and DU, which will not be elaborated upon in this embodiment. The functions of the CU can be implemented by one entity or by different entities. For example, the functions of the CU can be further divided, such as separating the control plane (CP) and user plane (UP), i.e., the CU control plane (CU-CP) and the CU user plane (CU-UP). For example, CU-CP and CU-UP can be implemented by different functional entities, and CU-CP and CU-UP can be coupled with the DU to jointly complete the functions of the access network device. In one possible approach, CU-CP is responsible for control plane functions, mainly including RRC and PDCP-C, where PDCP-C is mainly responsible for encryption / decryption, integrity protection, and data transmission of control plane data. CU-UP is responsible for user plane functions, mainly including SDAP and PDCP-U, where SDAP is mainly responsible for processing data from the core network device and mapping data flows to bearers. PDCP-U is primarily responsible for data plane encryption / decryption, integrity protection, header compression, sequence number maintenance, and data transmission. CU-CP and CU-UP are connected via an E1 interface. CU-CP represents the access network device connecting to the core network device through the interface between the core network device and the access network device.The CU-UP is connected to the DU via F1-C (control plane). The CU-UP is connected to the DU via F1-U (user plane). Alternatively, the PDCP-C may also be located in the CU-UP, but this application does not limit this implementation. The access network equipment can be, for example, a base station.
[0102] 3. Artificial Intelligence (AI):
[0103] AI is a new technical science that studies and develops theories, methods, technologies, and application systems for simulating, extending, and expanding human intelligence. AI encompasses fields including, but not limited to, at least one of automatic speech recognition (ASR), text-to-speech (TTS), image recognition, or natural language processing. ASR is a technology that converts speech into text, processing and analyzing speech signals to identify the linguistic content and converting it into readable text. TTS is a technology that converts text information into speech output, transforming text information into speech feature vectors, then converting the speech features into audio signals. It can also provide personalized pronunciation services for businesses and individuals through timbre selection, customizable volume, and speech rate. Speech can also be referred to as sound, audio, etc.
[0104] 4. Other terms
[0105] In the embodiments of this application, terms such as "first" and "second" are used to distinguish identical or similar items with essentially the same function and effect. For example, "first speech" and "second speech" are used only to distinguish different speech and do not limit their order. Those skilled in the art will understand that terms such as "first" and "second" do not limit the quantity or execution order, and that terms such as "first" and "second" do not necessarily imply that they are different.
[0106] It should be noted that, in the embodiments of this application, the terms "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design scheme described as "exemplary" or "for example" in this application should not be construed as being more preferred or advantageous than other embodiments or design schemes. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.
[0107] In this application embodiment, "at least one" refers to one or more, and "more than one" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can represent: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple.
[0108] Currently, some users, due to congenital or acquired factors, cannot communicate as freely as others. For example, people with hearing impairments, ALS, or vocal cord damage have speech defects, resulting in low intelligibility and difficulty in being understood, which significantly impacts their daily social interactions. To address these communication challenges, systems that assist with daily communication are needed for these individuals. These people with specific speech impairments can also be referred to as speech-impaired users.
[0109] For individuals with specific speech impairments, a language comprehension system is needed to address daily communication challenges. In related technologies, text can be input, and an electronic device can convert the text into speech for output, enabling these individuals to communicate through sound.
[0110] However, people with speech impairments need to input text to communicate via sound, which is cumbersome, resulting in insufficient convenience for them to communicate via sound. For people with speech impairments, their situation is that their speech has low intelligibility and is difficult for listeners to understand, not that they cannot produce sound.
[0111] In view of this, embodiments of this application propose a speech processing method, apparatus, system, storage medium, and program product that can assist speech-impaired users in communication by repairing the speech input by them, namely, dysarthric speech reconstruction (DSR). DSR can convert the low-intelligibility speech of a special speech-impaired group into higher-intelligibility normal speech. Speech can also be referred to as sound, audio, or sound source, etc. The speech processing method of this application embodiment can convert difficult-to-understand speech into more intelligible speech while maintaining the timbre consistent with the original speaker. This voice restoration function has high practical value in both short-range and long-range communication scenarios.
[0112] In one possible implementation, voice restoration can be achieved by an electronic device performing voice restoration processing (referred to as voice restoration or repair) on the voice input by a speech-impaired user to obtain restored voice. The restored voice is more intelligible than the input voice, so the electronic device can output the restored voice, which enables the party communicating with the speech-impaired user to better understand the speech-impaired user's expression.
[0113] In another possible implementation, after obtaining the repaired speech, the repaired speech can be processed by speech recognition to obtain the recognized text. The recognized text converted from the repaired speech can be understood as the repaired recognized text. Then, the electronic device can output the repaired recognized text, which also enables the party communicating with the speech-impaired user to better understand the speech-impaired user's expression.
[0114] In another possible implementation, speech restoration can be achieved by having an electronic device perform speech recognition processing on the input speech to obtain recognized text, and then perform restoration processing on the recognized text to obtain restored recognized text. The restored recognized text is more intelligible than the original recognized text, which also enables the party communicating with the speech-impaired user to better understand the expression of the speech-impaired user.
[0115] Therefore, it can be seen that voice restoration can enable users with speech impairments to communicate by inputting voice, thereby improving the convenience and experience of communication for them.
[0116] Please refer to Figure 1, which is a schematic diagram of a system architecture provided in an embodiment of this application.
[0117] As shown in Figure 1(a), the system architecture of this embodiment may include a first device 101, a second device 102, an access network device 103, and a cloud 104. The first device 101 and the second device 102 can communicate via the access network device 103. At least one of the first device 101 and the second device 102 is a device used by a user with a speech impairment. The cloud 104 has a voice restoration function, capable of restoring the input speech to obtain the output speech. The cloud 104 may be, for example, a server.
[0118] In one possible implementation, the first device 101 can be a device used by a user with a speech impairment. After receiving the voice input from the user with a speech impairment, the first device 101 sends the input voice to the cloud 104. The cloud 104 can then repair the input voice and send the repaired voice back to the first device 101. The first device 101 then sends the repaired voice to the second device 102 through the access network device 103. After receiving the repaired voice, the second device 102 can then play the repaired voice.
[0119] In another possible implementation, the second device 102 can be a device used by a user with a speech impairment. After receiving the voice input from the user with a speech impairment, the second device 102 sends the input voice to the first device 101. After receiving the input voice, the first device 101 sends the input voice to the cloud 104. The cloud 104 can then repair the input voice and send the repaired voice to the first device 101. After receiving the repaired voice, the first device 101 can play the repaired voice.
[0120] It should be understood that the above-mentioned voice repair processing can also be performed on at least one of the first device 101 or the second device 102, so the cloud 104 may not be necessary.
[0121] It should be noted that the cloud-based 104 can be configured with multiple speech restoration models. These models, also known as restoration models, are used to process the input speech to obtain restored speech. Optionally, the speech restoration capabilities of the multiple restoration models may differ.
[0122] As shown in Figure 1(b), the system architecture of this application embodiment may include a first device 101 and a cloud 104.
[0123] In one possible implementation, after receiving the voice input from a user with a speech impairment, the first device 101 sends the input voice to the cloud 104. The cloud 104 can then repair the input voice and send the repaired voice back to the first device 101. The first device 101 can then output the repaired voice or convert the repaired voice into text and output it.
[0124] It should be understood that the above-mentioned voice repair processing can also be performed on the first device 101, so the cloud 104 is not needed.
[0125] It should be noted that the solutions in this application embodiment can be used in scenarios including but not limited to the following:
[0126] Use Case 1: Remote Communication Scenarios. Remote communication scenarios can include, but are not limited to, real-time or non-real-time remote communication scenarios. Real-time remote communication scenarios can include, but are not limited to, real-time call scenarios. Non-real-time remote communication scenarios can include, but are not limited to, non-real-time remote voice communication scenarios.
[0127] Use Case 2: Close-range communication scenarios. Close-range communication scenarios can include, but are not limited to, face-to-face communication scenarios.
[0128] In other words, in this embodiment, sound restoration capabilities can be integrated into the interaction in both call and face-to-face communication scenarios. In general, this embodiment can be applied to users with speech impairments, enabling real-time restoration of input speech and playback of the restored audio to the listener or transmission to the other end of the call. Application scenarios include remote communication scenarios such as phone calls and video calls, as well as short-range communication scenarios such as face-to-face interactions.
[0129] The following examples illustrate remote communication scenarios and face-to-face communication scenarios respectively.
[0130] First, an example illustration will be provided for a remote communication scenario. This embodiment uses a remote real-time communication scenario as an example.
[0131] As understood, the terms "interface" and "user interface" used in this application refer to the medium through which an application or operating system interacts and exchanges information with the user. It facilitates the conversion between the internal form of information and a form acceptable to the user. A common form of user interface is the graphical user interface (GUI), which refers to a user interface related to computer operation displayed graphically. It can be an icon, window, control, or other interface element displayed on the screen of an electronic device. Controls can include visual interface elements such as icons, buttons, menus, tabs, text boxes, dialog boxes, status bars, navigation bars, and widgets.
[0132] Please refer to Figure 2, which is a schematic diagram of a sound registration interface provided in an embodiment of this application.
[0133] As shown in Figure 2(a), the interface displays a page with application icons. This page may include at least one of the following application icons: Settings app icon 201, Weather app icon, Calendar app icon, Photo app icon, Email app icon, or App Store app icon. Below these application icons, a page indicator may also be displayed to show the positional relationship between the currently displayed page and other pages. Below the page indicator are several application icons (e.g., Camera app icon, Contacts app icon, Messages app icon, Dialer app icon), which remain displayed when switching pages.
[0134] It is understood that the application icons in this application embodiment are icons for launching an application or icons for navigating to a certain interface. For example, the camera application icon is the icon of the camera application (i.e., the camera app), meaning that the camera application icon can be used to trigger the launch of the camera application. As another example, the settings application icon 201 can be used to trigger navigation to the settings interface.
[0135] In this embodiment, the electronic device can detect a user operation applied to the settings application icon 201. In response to the user operation, the electronic device can display the user interface shown in FIG2(b). The interface shown in FIG2(b) can be, for example, the settings application interface, on which the user can perform at least one of the following functions: querying information about the electronic device or configuring the user.
[0136] It is understood that the user operations mentioned in this application may include, but are not limited to, touch (e.g., click), voice control, gestures, etc., and this application does not limit them.
[0137] The interface shown in Figure 2(b) displays a page including settings options. This page may include one or more settings options. These settings options may include at least one of the following: accessibility settings option 202, battery and performance settings, security and privacy settings, display and brightness settings, or network connectivity management. Accessibility settings option 202 may be designed for users with special needs, helping them use electronic devices more conveniently through a series of assistive functions. The functions that accessibility settings option 202 in this embodiment can set may include, but are not limited to, at least one of the following: sound optimization, screen reading, or color correction. Sound optimization can also be referred to as sound restoration.
[0138] It should be noted that in the interface shown in Figure 2(b), for configuration options, such as setting option 2, an indication of whether the option is on or off may also be included. Optionally, when the indication for a setting option is "off," the configuration corresponding to that setting option is not enabled, and when the indication for a setting option is "on," the configuration corresponding to that setting option is enabled. It should be understood that the indications for enabling or disabling a setting option can be different, and are not limited to specific text styles such as "off" or "on." By displaying the indication of whether the option is on or off on the interface, users can quickly know whether the setting option is enabled or disabled, improving the user experience. For setting options used to query information, the setting option may not include an indication of whether the option is on or off. In addition, the interface shown in Figure 2(b) may also include an indication that the option has a next-level interface, such as the ">" symbol. By displaying the indication of having a next-level interface on the interface, users can know that there is a next-level interface through this indication, improving the user experience.
[0139] In this embodiment of the application, the electronic device can detect a user operation applied to the accessibility setting option 202, and in response to the operation, the electronic device can display the user interface shown in FIG2(c). The interface shown in FIG2(c) can be, for example, the interface of the accessibility setting option 202, on which the user can realize the setting function of the accessibility setting option 202.
[0140] The interface shown in Figure 2(c) displays a page that includes accessibility features. This page may include one or more accessibility controls, such as a sound restoration control 203. The sound restoration control 203 can be used to turn the sound restoration function on or off; in other words, the sound restoration control 203 can be understood as the master switch for the sound restoration function. Optionally, the interface may also include an indicator to show whether the sound restoration function is on or off. As shown in Figure 2(c), the sound restoration function is off, meaning it is not enabled. Optionally, the interface shown in Figure 2(c) may also include an indicator that the sound restoration function has a next-level interface. Optionally, the interface shown in Figure 2(c) may also include a control to return to the previous level interface, such as a control to return to the interface shown in Figure 2(b).
[0141] In this embodiment, the electronic device can detect user operations on the control 203 acting on the sound repair function. In response to the operation, the electronic device can display the user interface shown in Figure 2(d). The interface shown in Figure 2(d) can be, for example, the interface of the sound repair function, which can be understood as a detailed settings page for sound repair. The user can configure the sound repair function on this interface, such as at least one of the following: sound recording for the sound repair function or configuration of applications or scenarios supported by the sound repair function.
[0142] The interface shown in Figure 2(d) includes a sound recording control 204 and a sound repair application configuration control 205. The sound repair application configuration control 205 may include, but is not limited to, a call application configuration control 2051, a face-to-face communication application configuration control 2052, and a WeChat application configuration control 2053.
[0143] The sound recording control 204 can be used to record the user's voice. Optionally, the interface shown in Figure 2(d) may also include a sound recording status indicator. This indicator indicates whether sound has been recorded. Optionally, if the sound recording status indicator is "to be recorded," it means the user has not yet recorded sound; if it is "recorded," it means the user has already recorded sound. Optionally, even if the user has already recorded sound, additional sounds can be added, increasing the number of recorded sounds. It should be understood that the indicator indicating recorded sound can be different from the indicator indicating no recorded sound, and is not limited to the above example. As shown in Figure 2(d), the sound recording status indicator indicates that no sound has been recorded. This embodiment of the application uses a sound recording status indicator to indicate whether the user has recorded sound, allowing the user to quickly determine whether sound has been recorded, thereby improving the user experience. In another possible implementation, the sound recording status indicator may not need to be displayed, reducing the content displayed on the interface and improving display efficiency.
[0144] The sound repair application configuration control 205 is used to configure whether the sound repair function is enabled or disabled. The sound repair application configuration control 205 can also be understood as a switch for applications that support the sound repair function. When the sound repair application configuration control 205 is in the enabled state, the sound repair function in the corresponding application or scenario is activated. As shown in Figure 2(d), the applications or scenarios that can be configured with the sound repair function can include, but are not limited to, at least one of the following: call applications, face-to-face communication applications, or WeChat applications. The WeChat application can be understood as a non-real-time remote communication application. Figure 2(d) can include the sound repair application configuration controls 205 for various applications, such as the sound repair application configuration control 205 for call applications, the sound repair application configuration control 205 for face-to-face communication applications, and the sound repair application configuration control 205 for WeChat applications. Optionally, when the sound repair application configuration control 205 of an application or scenario is in the disabled state, the sound repair function cannot be used in that application or scenario; when the sound repair application configuration control 205 of an application or scenario is in the enabled state, the sound repair function can be used in that application or scenario. This embodiment of the application configures a sound repair application configuration control 205 for each application or scenario, meaning that the sound repair function of each application or scenario can be individually enabled or disabled. This allows users to select which application or scenario to enable or disable the sound repair function as needed, thereby improving the user experience. In another possible implementation, a single sound repair application configuration control 205 can be used to enable or disable the sound repair function for multiple applications or scenarios, allowing for one-click enabling or disabling of sound repair functions for multiple applications or scenarios, thus improving the efficiency of enabling or disabling sound repair functions for multiple applications or scenarios. In yet another possible implementation, the application or scenario supporting sound repair configuration and its corresponding sound repair application configuration control 205 may not be displayed; that is, the sound repair function may be enabled by default for all applications or scenarios. This reduces the content displayed on the interface and improves the efficiency of the interface display.
[0145] For example, as shown in Figure 2(d), the sound repair application configuration control 205 is in a disabled state, meaning that the sound repair function cannot be used by any application.
[0146] It should be noted that sound recording and supported application or scene configuration can be mutually constrained. This constraint can be that one of the sound recording and supported application or scene configurations can only be performed after the other operation is completed. For example, sound recording can be completed before configuring the application or scene supported by the sound repair function. In other words, if sound recording is not completed, no application or scene can enable the sound repair function through the sound repair application configuration control 205. This embodiment of the application allows sound recording and supported application or scene configuration to be mutually constrained, prompting the user to perform sound recording and configure the application or scene for enabling the sound repair function, thereby improving the user experience.
[0147] In this embodiment, the electronic device can detect user operations applied to the sound recording control 204. In response to these operations, the electronic device can display the interface shown in Figure 2(e). In this embodiment, since the sound optimization function requires user voice registration, after enabling sound repair, the user can automatically jump to the sound registration interface, such as the interface shown in Figure 2(e). The interface shown in Figure 2(e) may include a new voiceprint control 206. The new voiceprint control 206 can be used to trigger the acquisition of voiceprint features (referred to as voiceprint). The voiceprint can be used to indicate the timbre of the registered user's voice. Optionally, the interface shown in Figure 2(e) may also include a control to return to the previous interface, such as a control to return to the interface shown in Figure 2(d).
[0148] Then, the electronic device can detect the user operation applied to the newly created voiceprint control 206. In response to this operation, the electronic device can display the interface shown in Figure 2(f). The interface shown in Figure 2(f) can include sound recording content 207, recording control 208, and save control. The sound recording content 207 can guide the user to record the content, such as text like "Giant pandas don't have fixed sleeping times; they sleep wherever they go." The recording control 208 can be used to trigger sound acquisition. Optionally, sound acquisition can be triggered by long-pressing the recording control 208, and the sound acquired by the electronic device is the sound recorded by the user during the long-press process; or the user can click the recording control 208 to trigger sound acquisition, and the electronic device will start acquiring sound. If the user clicks the recording control 208 again, sound acquisition will stop. The save control is used to include the acquired sound. Optionally, the user can choose to record sound in segments. When the electronic device detects the user operation applied to the save control, it splices the segmented sound recordings according to the recording time sequence to obtain a single recording.
[0149] In this embodiment, the electronic device detects a user operation on the recording control 208 and, in response to the operation, begins to collect the user's voice.
[0150] Then, as shown in Figure 2(g), the electronic device detects a user operation on the save control, and in response to the operation, saves the collected sound and extracts the voiceprint of the collected sound. It should be understood that the voiceprint extraction process can be performed at any time after the sound is collected, and is not limited to extraction immediately after the sound is saved.
[0151] Optionally, the electronic device can compare the recorded content with the audio recording content 207. If the similarity between the recorded content and the audio recording content 207 is higher than a similarity threshold, the voiceprint of the recorded content will be used as the voiceprint for subsequent voice broadcast. If the similarity between the recorded content and the audio recording content 207 is not higher than the similarity threshold, the user will be prompted to re-record the audio according to the guidance of the audio recording content 207. Optionally, the electronic device in this embodiment supports recording 1 to n sentences of text. That is, after recording the text content shown in Figure 2(g), another piece of text content can be displayed for the user to record, where n is an integer not less than 2. After recording, the user can save the recording using the save control. It should be understood that the user can also choose to save the recording after recording m sentences of text, where 1 ≤ m < n, and m is an integer.
[0152] Then, as shown in Figure 2(h), the electronic device detects a user action performed on the control to return to the previous screen. In response to this action, the interface shown in Figure 2(i) can be displayed. In the interface shown in Figure 2(i), the sound recording status indicator is marked as "Recorded". The similar parts of Figure 2(i) and Figure 2(d) can be referred to the description in Figure 2(d), and will not be repeated here.
[0153] In another possible implementation, the interface shown in Figure 2(f) may also exclude the audio recording content 207. This reduces the content displayed on the interface and improves display efficiency. However, by displaying the audio recording content 207, this embodiment guides the user to record according to the audio recording content 207, which helps the electronic device determine the validity of the audio recording and improves the accuracy of voiceprint extraction. Furthermore, the content of the collected audio can be compared with the audio recording content 207 of the interface shown in Figure 2(f), and a suitable repair model can be selected based on the similarity to repair subsequent user input. The similarity can also be called intelligibility. A higher similarity indicates a greater similarity between the collected audio content and the audio recording content 207 of the interface shown in Figure 2(f), suggesting less impairment of the user's speaking ability. Conversely, a lower similarity indicates a less similarity between the collected audio content and the audio recording content 207 of the interface shown in Figure 2(f), suggesting greater impairment of the user's speaking ability. Generally speaking... The more complex the architecture of a restoration model, the stronger its speech restoration capability and the more accurate the speech restoration. However, this also increases the computational resources required by the electronic device. Therefore, when selecting a suitable restoration model based on the similarity of the comparison, a lower similarity allows for a more complex restoration model architecture, while a higher similarity allows for a simpler architecture. For example, if the similarity is at the first level, the first restoration model is selected; if the similarity is at the second level, the second restoration model is selected. Since the first similarity is higher than the second, the architecture complexity of the first restoration model is lower than that of the second restoration model, and consequently, the capability of the first restoration model is lower than that of the second restoration model.
[0154] In another possible implementation, the interface shown in Figure 2(f) may not include the recording control 208. In this case, sound collection begins when the interface shown in Figure 2(f) is entered, which reduces the content displayed on the interface and improves the efficiency of the interface display. However, in this embodiment, by displaying the recording control 208 and only starting sound collection after detecting that the recording control 208 is being used, the resource utilization rate of the electronic device can be reduced, and the service life of the electronic device can be improved.
[0155] In another possible implementation, the interface shown in Figure 2(f) may not include a save control. In this case, the recorded sound is saved directly after the recording is completed. For example, if no recording is detected within a certain period of time, the recording is considered complete. Or, if the length of the recorded content is the same as the length of the content shown in Figure 2(f), the recording is considered complete. However, by setting a save control, the embodiments of this application can enable users to record sound flexibly.
[0156] It should be understood that the interface of the embodiments of this application can be increased or decreased as needed, and is not limited to the number of interfaces and the style of the interfaces shown in FIG2, and is not limited herein.
[0157] In this embodiment of the application, by having users register their voices, the user's voiceprint can be extracted from the collected voice. In this way, during the broadcast, the repaired voice can be broadcast using the user's voiceprint. That is, the user can select a specific voiceprint to broadcast the repaired voice as needed, thereby improving the user experience.
[0158] In another possible implementation, the audio recording process in Figure 2 may not be necessary. In this way, the electronic device can use the built-in voiceprint to broadcast the restored speech, which also improves the ease of communication due to the reduction of the audio recording process.
[0159] Below, we will provide an example of how to configure an application that uses the sound restoration feature.
[0160] Please refer to Figure 3, which is a schematic diagram of the interface for configuring an application using the sound restoration function according to an embodiment of this application.
[0161] The interface shown in Figure 3(a) includes a sound recording control 204 and a sound repair application configuration control 205. The interface shown in Figure 3(a) can be described with reference to the interface shown in Figure 2(i), and will not be repeated here.
[0162] In this embodiment, the electronic device can detect user operations applied to the voice repair application configuration control 205, such as detecting user operations applied to the call application configuration control 2051. In response to this operation, the interface shown in Figure 3(b) can be displayed. In the interface shown in Figure 3(b), the call application configuration control 2051 in the voice repair application configuration control 205 switches from a voice repair off state to a voice repair on state. For example, in the interface shown in Figure 3(b), i.e., the call application configuration control 2051 is in the voice repair on state, the call application can use the voice repair function; while the face-to-face communication application configuration control 2052 and the WeChat application configuration control 2053 are both in the voice repair off state, the face-to-face communication application and the WeChat application cannot use the voice repair function. However, when using the face-to-face communication application or the WeChat application, the voice repair function can be turned on or off through controls within the face-to-face communication application or the WeChat application interface. In this embodiment, optionally, face-to-face communication applications, WeChat applications, and call applications can control the opening or closing of the sound repair function through controls in their respective application interfaces. This embodiment will not explain how to control the opening or closing of the sound repair function through controls in the application interface; this will be explained further in later embodiments.
[0163] It should be noted that if the electronic device can detect user operations on the sound repair application configuration control 205, the sound repair application configuration control 205 can also switch from the sound repair enabled state to the sound repair disabled state.
[0164] In this embodiment of the application, by configuring a sound repair control for each application, the enabling or disabling of the sound repair function of each application can be controlled independently. Users can choose to enable or disable the sound repair function of the application as needed, thereby improving the user experience.
[0165] The sound restoration function for the call application has been enabled in the interface shown in Figure 3. Therefore, the following example illustrates the call scenario.
[0166] Figure 4 is a schematic diagram of an interface for configuring a whitelist for voice restoration according to an embodiment of this application. In this embodiment, for people with special speech impairments, there may be some friends who are used to understanding their pronunciation; for these friends, when the voice restoration switch is turned on, as shown in Figure 4, the user with special speech impairments can set a whitelist in the phone's address book so that voice restoration is not turned on by default when talking to these friends.
[0167] The interface shown in Figure 4(a) displays a page with application icons, which may include multiple application icons (e.g., the Contacts application 401). The interface shown in Figure 4(a) can be described with reference to the interface shown in Figure 2(a), and will not be repeated here.
[0168] In this embodiment, the electronic device can detect user operations applied to the address book application 401. In response to the operation, the electronic device can display the interface shown in Figure 4(b). The interface shown in Figure 4(b) can be, for example, the main interface of the address book. The interface shown in Figure 4(b) can include at least one of the following: a contact configuration control 403, a contact search bar, or a contact list. The contact configuration control 403 is used to configure the same operation for selected contacts, such as adding the selected contact to the voice repair whitelist, or enabling the call log function by default for the selected contact during a call. The contact search bar can be used to quickly search for contacts. The contact list includes one or more contacts. Optionally, the contacts in the contact list can be categorized according to certain rules, such as categorizing them according to the first letter of their names, and the interface shown in Figure 4(b) also includes controls for quickly locating categories. Optionally, the interface shown in Figure 4(b) can also include dial controls and favorite controls. Optionally, the interface shown in Figure 4(b) can also include a contact addition control.
[0169] Then, as shown in Figure 4(c), the electronic device can detect the user operation applied to the interface, and in response to the operation, the electronic device can display the contact selection control 402. The contact selection control 402 can be used to select a contact to add the selected contact to a whitelist. In this embodiment, each contact has a corresponding contact selection control 402. As shown in Figure 4(c), the contact selection control 402 is in an unselected state.
[0170] In this embodiment of the application, the electronic device can detect a user operation applied to the contact selection control 402. In response to the operation, the electronic device can display the interface shown in Figure 4(d). In the interface shown in Figure 4(d), the applied contact selection control 402 is in a selected state, for example, the contact selection control 402 for the contact "B2" is in a selected state.
[0171] Then, the electronic device can detect the user operation applied to the contact configuration control 403, and in response to the operation, the electronic device can display the interface shown in Figure 4(e). The contact configuration control 403 of this embodiment can be used to configure selected contacts. The interface shown in Figure 4(e) may also include a voice repair whitelist option 404 and a call recording option. The call recording option is used to set the contact selected in the contact selection control 402 as the contact for saving call recordings. The voice repair whitelist option 404 is used to add the contact selected in the contact selection control 402 to the voice repair whitelist, that is, to add the selected contact to the voice repair whitelist.
[0172] Then, as shown in Figure 4(f), the electronic device can detect the user operation applied to the voice repair whitelist option 404, and the electronic device can add the contacts whose contact selection control 402 is in the selected state to the voice repair whitelist. Optionally, in this embodiment, the voice repair whitelist is a list where voice repair is not enabled, that is, the contacts in the voice repair whitelist do not enable the voice repair function when making a call; or, the voice repair whitelist is a list where voice repair is enabled, that is, the contacts in the voice repair whitelist enable the voice repair function when making a call.
[0173] This application embodiment allows users to configure a whitelist for voice repair, enabling them to select contacts for use or disuse of the voice repair function during calls, based on actual circumstances. This flexibility allows users to adjust which contacts require or do not require voice repair during calls, thereby improving the user experience.
[0174] It should be noted that you can also choose not to configure the voice repair whitelist. In this case, you can perform voice repair on all contacts or not.
[0175] After configuring the voice repair whitelist, you can initiate a call with one of the contacts. For ease of understanding, the following example illustrates how contacts on the voice repair whitelist are not subject to voice repair.
[0176] This application embodiment improves the flexibility of using the voice repair function by setting a voice repair whitelist, allowing different contacts to use or not use the voice repair function.
[0177] Please refer to Figure 5, which is a schematic diagram of an interface for a remote real-time communication scenario provided in an embodiment of this application. In this embodiment, contacts in the voice repair whitelist are those for whom voice repair is disabled by default.
[0178] The interface shown in Figure 5(a) displays a page with application icons, which may include multiple application icons (e.g., phone application 501). The interface shown in Figure 5(a) can be described with reference to the interface shown in Figure 2(a), and will not be repeated here.
[0179] In this embodiment, the electronic device can detect user operations applied to the telephone application 501, and in response to the operation, can display the interface shown in Figure 5(b). The interface shown in Figure 5(b) displays the user's recent call records, which may include call contacts and corresponding call times. The interface shown in Figure 5(b) may include multiple contacts. Then, if the electronic device can detect user operations applied to one of the contacts, for example, user operations applied to contact "B2", in response to the operation, a call with contact "B2" can be initiated. As shown in Figure 4(d), since contact "B2" is added to the voice repair whitelist, that is, contact "B2" is configured not to use the voice repair function, as shown in Figure 5(c), when initiating a call with contact "B2", the call is conducted in its original voice by default, that is, the call is conducted with contact "B2" using the voice before repair. Optionally, the interface shown in Figure 5(c) also includes a voice switching control 502, the avatar of the call contact, a voice mute control, a keyboard access control, a voice speaker control, and a call end control. As shown in Figure 5(c), the sound switching control 502 is in the original sound playback state, meaning the other end of the call plays the original sound captured by this device. The sound switching control 502 can be used to control the switching between the original sound and the restored audio. As shown in Figure 5(d), the electronic device can detect user operation on the sound switching control 502. In response to this operation, the electronic device's sound switching control 502 is in the restored audio playback state, meaning the other end of the call plays the restored sound. Therefore, the contact "B2" in the call hears the restored sound. In this embodiment, the restored sound can be obtained by restoring the original sound. The original sound can also be called the original pronunciation, which can be the sound captured from the user's mouth.
[0180] Optionally, if the electronic device detects another user operation on the sound switching control 502, the state of the sound switching control 502 switches to the original sound playback state, and the corresponding user conducts the call using the original sound. In this embodiment, by setting the sound switching control 502 in the call interface, the user can choose to use the original sound or the restored sound through the sound switching control 502, which can improve the user experience.
[0181] It should be noted that if an electronic device initiates a call with a contact not on the voice repair whitelist, the voice repair function will be enabled by default in the call interface.
[0182] It should be noted that if the call is initiated by the electronic device on the other end, the electronic device on the local end can also display the interface shown in Figure 5(c) or Figure 5(d). In this way, the local end can use the sound switching control 502 to control whether the local end plays the original sound collected by the other end or the sound after the original sound collected by the other end has been repaired.
[0183] In another possible implementation, if the caller is a contact outside the voice repair whitelist, that is, the caller is configured to enable the voice repair function, then the voice switching control 502 is in the repaired voice playback state when the call is connected.
[0184] In another possible implementation, the voice switching control 502 may not be necessary, meaning that the original voice or the repaired voice can be used for the call depending on the configuration of the voice repair whitelist.
[0185] In some cases, when electronic devices repair the sound, it takes a certain amount of time. Therefore, there will be a delay from when the local user speaks to when the remote user hears the local user's speech. This delay is related to the time required for the electronic device to repair the sound. As a result, it will take a long time for the remote user to receive the repaired voice from the local user. The remote user may then think that the local user did not speak or that the signal is poor, which is why the remote user cannot hear it.
[0186] It should be noted that it can also be used when making a call with a contact in the voice repair whitelist, repairing the voice transmitted by the contact in the voice repair whitelist and then playing it. In other words, the local device repairs the voice transmitted from the other end and then plays it, for example, repairing the voice of the contact "B2" and then playing it.
[0187] Please refer to Figure 6, which is a schematic diagram of a call scenario provided in an embodiment of this application. In this embodiment, the local user and the remote user are referred to as "local device" and "remote device," respectively. The electronic device used by the local user is called the local device, and the electronic device used by the remote user is called the remote device. The local device can also be called the first electronic device, and the remote device can also be called the second electronic device. The local user can also be called the first user, and the remote user can also be called the second user. In this embodiment, the call audio will increase the call latency after being processed by the sound restoration model, which may affect the call quality of multiple parties. Therefore, this embodiment describes improving the call quality of multiple parties in a call scenario.
[0188] As shown in Figure 6(a), if the user inputs "Hello," the local device needs to perform repair processing before sending the repaired "Hello" to the remote device. The remote device then receives and plays the repaired "Hello," allowing the user to hear it. However, there is a significant time delay between the user's input and the remote device hearing the repaired message. This is because the local device needs time to perform the repair, which is a processing latency. Alternatively, the remote device can send the user's input, "Hello, your takeout, I'm downstairs," to the local device, which will then play the message. When the user on this end wants to reply, they can type "Thank you, just leave it downstairs." The local device needs to process this voice message, but this process takes time, which is called processing delay. If the user on the other end doesn't hear the response from the local user for a long time, they might assume that the local user didn't hear their voice message and repeat "Hello, your takeout, I'm downstairs." However, the local user has already heard it, but because the processing takes time, the user on the other end may not hear the local user's feedback for an extended period.
[0189] Therefore, in one possible implementation, during a call, the local device can send a prompt to the remote device, which then plays the prompt, thereby informing the remote user that the local device needs to process the voice input from the user. This allows the remote user to know that the local device needs to repair the voice input.
[0190] As shown in Figure 6(b), after establishing a communication channel with the peer device, the local device can send a prompt to the peer device. Upon receiving the prompt, the peer device can broadcast it, allowing the peer user to hear it. This prompt can indicate that the local device needs to repair the voice input by the local user. For example, the prompt might include phrases such as "The other party has enabled voice optimization; there is a delay in the call. Please wait patiently for the prompt tone to finish before replying," to inform the other party that a call delay may occur.
[0191] It should be understood that the timing of the local device sending a prompt to the remote device can be either immediately after establishing a communication channel with the remote device or after a certain interval; there is no restriction on this.
[0192] In this embodiment of the application, sending a prompt message from the local device to the remote device can improve the experience for both parties in the call.
[0193] As shown in Figure 6(b), although the local device can send prompts to the remote device, there are still situations during the communication process where the remote user thinks that the local user did not hear their voice.
[0194] In this application embodiment, in call scenarios, such as mobile phone calls and video calls, the voice changing and audio repair functions are enabled by switching on the call interface. After being enabled, the receiving party will receive a corresponding prompt message to inform the other party that the sound they hear has been optimized and that there will be a delay, so as to improve the call experience between the two parties.
[0195] Therefore, in one possible implementation, after receiving each segment of voice input from the local user, the local device sends a prompt tone to the remote device, which can then play the prompt tone to indicate that the local user has input voice.
[0196] As shown in Figure 6(c), for example, after establishing a communication channel with the peer device, the local device can send a prompt to the peer device. Then, after detecting the user's voice input "Hello," the local device sends a prompt tone to the peer device, which is then played back by the peer device. Then, after detecting the user's voice input "Thank you, just leave it downstairs," the local device sends a prompt tone to the peer device, which is then played back by the peer device. The peer user can then know that the local user has input a voice message through the prompt tone, and after receiving the "Thank you, just leave it downstairs" message, can reply with "Okay, goodbye." For example, the prompt tone can be, for instance, a "ding," a "beep," or a "tick-tock." In a call scenario, the AI model's voice conversion can cause call delays. By first playing a "beep" prompt tone to the peer device after the local device speaks, it indicates a wait and reduces overlap in the conversation. Optionally, during the call, before the sound repair result is given, the other party can be guided to wait patiently and reply through a call prompt tone.
[0197] In this embodiment of the application, the timing of the prompt tone sent by the local terminal may be as follows: after receiving the user's voice input, or after receiving a segment of voice input. No limitation is imposed here.
[0198] It should be noted that the prompts in this embodiment can also be used to indicate the function of notification sounds during a call. For example, the prompts may include phrases such as "You will hear a notification sound during the call, indicating that the user is currently inputting voice; please wait patiently."
[0199] In this embodiment of the application, by sending a prompt tone to the peer device after receiving each segment of voice input by the local device, the experience of both parties in the call when using the voice repair function can be improved, and the problem of communication inefficiency caused by latency can be reduced.
[0200] In one possible implementation, it is also possible to send only one of the prompts or sounds; this is not a limitation.
[0201] The above examples illustrate remote real-time communication scenarios. The following examples illustrate remote non-real-time communication scenarios.
[0202] Please refer to Figure 7, which is a schematic diagram of an interface for a remote non-real-time communication scenario provided in an embodiment of this application.
[0203] The interface shown in Figure 7(a) displays a page with application icons, which may include multiple application icons (e.g., chat application 701). In this embodiment, chat application 701 may be a non-real-time remote chat application 701. The interface shown in Figure 7(a) can be described with reference to the interface shown in Figure 2(a), and will not be repeated here.
[0204] In this embodiment, the electronic device can detect user operations applied to the chat application 701. In response to the operation, the electronic device can display the interface shown in Figure 7(b). Optionally, the interface shown in Figure 7(b) may include a search bar, chat function controls, address book function controls, send function controls, and personal center function controls. The search bar can be used to quickly search for information, such as quickly searching for contacts. In the interface shown in Figure 7(b), the chat function controls are selected. Therefore, the interface shown in Figure 7(b) can be understood as the interface corresponding to the chat function, which includes one or more contacts and their corresponding contact information. Optionally, the contact information may include the contact name and chat content, which may be, for example, the latest chat message. For example, the name of one of the contacts may be "Xiaoming," and the latest chat message with him may be, for example, "What did you do today?"
[0205] In this embodiment, the electronic device can detect a user operation applied to one of the contacts. In response to the operation, the electronic device can display the interface shown in Figure 7(c), which can be a chat interface with the selected contact. The chat interface can include each chat message and the user corresponding to the chat message. Optionally, the chat message can include, but is not limited to, at least one of text, voice, video, or file. The interface shown in Figure 7(c) can include a chat message input control, which can include a voice input control 702. Optionally, the chat message input control can also include a text input control, a file / video transfer control, and a video recording control. The voice input control 702 is used to trigger voice recording. How the voice input control 702 triggers voice recording can be referred to the description of the recording control triggering sound recording in the embodiment of Figure 2, which will not be repeated here.
[0206] In this embodiment, the electronic device can detect user operations applied to the voice input control 702, and in response to the operation, can record the user's input voice. As shown in Figure 7(d), the electronic device can also display a recording in progress message during voice recording. Optionally, the recording in progress message may include text prompts such as "Recording in progress," and may also include text content converted from the recorded voice. It should be understood that by displaying the text content converted from the recorded voice during the recording process, the user can know whether the electronic device has captured their voice and whether the text content converted by the electronic device matches the recorded voice, thereby improving the user experience.
[0207] Then, as shown in Figure 7(e), the electronic device can display a recording completion message after the recording of the voice is finished. This recording completion message may include text prompts such as "Recording Complete," and may also include the text content converted from the recorded voice. The interface shown in Figure 7(e) includes a direct send option 703 and a repair and send option 704. The direct send option 703 is used to trigger the transmission of the captured user's original voice; that is, if the electronic device detects a user operation acting on the direct send option 703, it will send the unrepaired voice in the chat interface. The repair and send option 704 is used to trigger the transmission of the repaired voice 705; that is, if the electronic device detects a user operation acting on the repair and send option 704, it will send the repaired voice 705 in the chat interface.
[0208] In this embodiment, the electronic device can detect the user operation applied to the repair and send option 704, and in response to the operation, can send the repaired voice 705 in the chat interface. As shown in the interface in Figure 7(f), the repaired voice 705 is sent. The electronic device can detect the user operation applied to the repaired voice 705 and play the repaired voice 705.
[0209] Optionally, if the user chooses to send the original audio, the electronic device can also play the original audio after detecting a user action applied to the original audio.
[0210] This application embodiment allows users to send the repaired voice message 705 to their contacts, thereby improving the experience of non-real-time remote communication between users and their contacts.
[0211] Then, in the interface shown in Figure 7(g), a voice message 706 is received from the contact.
[0212] In this embodiment of the application, as shown in FIG7(h), the electronic device can detect user operation 706 applied to a contact's voice, and in response to the operation, can display a voice-to-text option 707 and a voice-to-text option 708. The voice-to-text option 707 is triggered to display the text converted from the selected voice after repair. The voice-to-text option 708 is triggered to display the text converted from the selected voice.
[0213] It should be noted that, optionally, clicking on the voice message will play the voice message, while long-pressing the voice message will display options 707 (voice-to-text after repair) and 708 (voice-to-text).
[0214] Then, as shown in Figure 7(i), the electronic device can detect the user operation on the voice-to-text option 707, and in response to the operation, can send the selected voice to be converted into text after repair in the chat interface.
[0215] It should be noted that the electronic device may detect a user operation that applies the voice repair to the text conversion option 707, repair the voice, and then convert it into text for display; or, the electronic device may receive the voice, repair the voice to obtain the repaired voice, and then, upon detecting a user operation that applies the voice repair to the text conversion option 707, display the selected voice converted into text.
[0216] In this embodiment of the application, if the received voice message is the audio of a user with a special speech impairment, the user on this end can perform the user operation 706 on the contact's voice to display the voice-to-text option 707 and the voice-to-text option 708. Then, by performing the user operation on the voice-to-text option 707, more accurate text can be obtained.
[0217] In another possible implementation, after the electronic device detects the user operation 706 applied to the contact's voice, it can also display a voice repair option. The voice repair option is used to trigger the playback of the repaired voice 705. Then, after the electronic device detects the user operation applied to the voice repair option, it can play the repaired voice 705, which is to say, play the repaired voice 706 from the contact.
[0218] This application embodiment can repair the voice 706 of a contact after receiving the voice 706 from the contact, so that even if the other party is a speech impaired user, normal communication can be achieved, thereby improving the experience of non-real-time remote communication between the user and the contact.
[0219] In summary, in remote communication scenarios, the other party can hear the restored, more intelligible, and identical voice in real time, making the call smoother. Furthermore, in remote communication scenarios, the other party will receive a notification that "the local device has enabled the voice restoration function." Simultaneously, the local device can switch between the original and restored audio sources at any time via buttons, allowing the other user to access the original audio information, which may include at least one of the original timbre, ambient sound, or original pronunciation.
[0220] In another possible implementation, in Figure 7(h), a play voice option and a play repaired voice option can also be displayed. If a user action is detected on the play voice option, the voice from the contact is played; if an action is detected on the play repaired voice option, the voice from the contact is repaired before playback. This embodiment of the application, by displaying the play voice option and the play repaired voice option, can repair the voice from the contact, thereby improving the communication experience between the two parties.
[0221] The above describes remote communication scenarios. Below, we will provide examples of near-field communication scenarios. Unlike real-time calls, face-to-face rehabilitation is used by individuals with speech impairments when communicating with others offline.
[0222] Please refer to Figure 8, which is a schematic diagram of an interface for a face-to-face communication scenario provided in an embodiment of this application.
[0223] The interface shown in Figure 8(a) displays a page with application icons, which may include multiple application icons (e.g., Face-to-Face Repair Application 801). The Face-to-Face Repair Application 801 in this embodiment can be used for voice repair in face-to-face communication scenarios. The interface shown in Figure 8(a) can be described with reference to the interface shown in Figure 2(a), and will not be repeated here.
[0224] In this embodiment, the electronic device can detect user actions applied to the face-to-face repair application 801, and in response to the action, can display the interface shown in Figure 8(b). The interface shown in Figure 8(b) may include a recording control 802, which is used to trigger voice recording. For example, the recording control may be in the form of a recording ball. In this embodiment, the electronic device can detect user actions applied to the recording control 802, and in response to the action, begin recording voice. Optionally, the user action in this embodiment may be, for example, a long press operation. During the long press of the recording control 802, a swipe-up cancel prompt and a release repair prompt may be displayed. The swipe-up cancel prompt indicates that voice recording can be canceled by swiping up. The release repair prompt indicates that recording can be completed by releasing the recording control 802.
[0225] As shown in Figure 8(c), after the electronic device detects a release operation on the recording control 802, it displays a playback control 803 for the repaired audio. The playback control 803 displays the repaired audio obtained by repairing the recorded audio. Furthermore, the interface shown in Figure 8(c) may also include a text display area 804, which displays the text content corresponding to the playback control 803 for the repaired audio. For example, the text content displayed on the interface shown in Figure 8(c) could be, for instance, "Are you going this afternoon?". It should be noted that if the electronic device detects a user operation on the playback control 803 for the repaired audio in the interface shown in Figure 8(c), it can play the repaired audio, such as the audio corresponding to "Are you going this afternoon?". Optionally, after detecting a release operation on the recording control 802, the animation of the recording control 802 scrolls for a short period before displaying the repaired text and automatically playing the repaired audio. If the user is not satisfied with the recording during the recording process, they can swipe up while holding down the recording to cancel the repair.
[0226] However, since voice restoration may not be entirely accurate, this application embodiment also supports rapid correction of the restoration results to improve the effectiveness of face-to-face communication. For example, if the restored text contains misidentified words, the user can click on the word they want to modify, and a candidate word list 806 will appear, from which the user can select the desired word for quick replacement. Users can also directly edit the text using an input method. After completing the correction, the user can choose to play the corrected audio.
[0227] As shown in Figure 8(d), the electronic device can detect a user operation on a portion of the text content, word 805, and in response to this operation, can display the interface shown in Figure 8(e). The interface shown in Figure 8(e) includes a candidate word list 806 corresponding to the user-selected portion of word 805, which includes one or more candidate words. For example, if the portion of word 805 is "afternoon," then the corresponding candidate words could include "morning," "dance well," and "foggy," among others.
[0228] Then, as shown in Figure 8(f), the electronic device can detect a user action applied to one of the candidate words, and in response to the action, can display the interface shown in Figure 8(g). In the interface shown in Figure 8(g), the word 805 in the displayed text content is replaced with the candidate word applied by the user action. For example, if the electronic device detects a user action applied to the candidate word "morning," then in the displayed text content "Will you go this afternoon?", "afternoon" is replaced with "morning," and the corresponding displayed text content is replaced with "Will you go this morning?".
[0229] Then, as shown in Figure 8(h), the electronic device can detect the user operation on the playback control 803 of the repaired voice in the interface shown in Figure 8(h), and in response to the operation, can play the voice corresponding to "Are you going this morning or not?". In this way, the other party communicating with the user can know the user's voice expression.
[0230] In other words, in this embodiment of the application, the speech can be broadcast according to the text content displayed in the text display area 804.
[0231] In this embodiment, by recording and repairing the user's voice during face-to-face communication, the playback control 803 displays the repaired voice, improving the user's communication experience. Furthermore, after recording the user's input voice, the corresponding text content can be displayed. This allows the user to determine the accuracy of the voice repair result through the text content. If the user deems the repair result inaccurate, they can adjust the text content, thereby adjusting the content of the played voice. In other words, face-to-face communication scenarios can quickly correct semantic errors, making user operation more convenient and faster, and improving the accuracy of voice communication. Moreover, in this embodiment, voice input during face-to-face communication allows for the playback of the repaired audio, eliminating the need for plain text input, making it faster. Furthermore, for words with errors after repair, the system supports quick replacement of the identified words in the text and replays the audio, improving the efficiency of face-to-face communication.
[0232] It should be noted that in the interfaces shown in Figures 8(c)-(e), one possible implementation is as follows: When performing ASR recognition processing on the speech, multiple ASR recognition results may be obtained, each corresponding to a probability value. The electronic device can then display the ASR recognition result with the highest probability value among the multiple results. For example, the ASR recognition results obtained from the ASR recognition processing may include: "Today, shall we go to Shanwu?", "Today, shall we go to Xiawu?", "Today, shall we go to Shanwu?", and "Today, shall we go to Shanwu?". Since the probability value of the recognition result "Today, shall we go to Shanwu?" is the highest, this ASR recognition result is displayed. The words in the other ASR recognition results besides the displayed one are also saved. Then, the displayed ASR recognition results can be segmented to obtain one or more words. When a user action is detected acting on one of the words (i.e., part of word 805), words different from that part of word 805 are searched from the remaining ASR recognition results and displayed as candidate words, such as "morning," "good at dancing," and "foggy." This embodiment improves the efficiency of obtaining candidate words by saving the remaining ASR recognition results and then searching for words different from the user-selected word (i.e., part of word 805) from the remaining ASR recognition results.
[0233] In another possible implementation, the remaining ASR recognition results may not be saved. In this way, candidate words can be predicted directly from the words selected by the user. Since the remaining ASR recognition results do not need to be stored, the storage resources required to store the ASR recognition results can be reduced.
[0234] It should be noted that, in this embodiment of the application, the interface shown in Figure 8 may contain at most one playback control 803 for a repaired voice message, such as the playback control 803 for the latest repaired voice message. This improves the simplicity of the interface display. In another possible implementation, a new voice message may be displayed after each recorded voice message, and a new voice message may also be displayed after each change to the text content. This allows for the recording of every voice message from the user and the user's modification process, thereby improving the user experience.
[0235] In another possible implementation, as shown in Figure 8(c), the interface can also display either voice or text in the display area 804, which improves the simplicity of the interface. In this embodiment, by displaying both voice and text in the display area 804, users can quickly determine the accuracy of the electronic device's voice repair results through the text content displayed in the text display area 804, thereby improving the user experience.
[0236] In another possible implementation, the user can select punctuation marks from the text content. In this case, candidate punctuation marks can be displayed, and the user-selected punctuation marks can be used to replace the selected punctuation marks in the text content. The tone of the speech can then be changed based on the updated punctuation. For example, replacing “?” with “!” changes the tone of the speech from a question to an exclamation.
[0237] It should be noted that users with speech impairments often also have some degree of hearing impairment; that is, users with speech impairments may also be hearing impaired. Therefore, the following examples illustrate how to improve the communication experience of users with hearing or speech impairments.
[0238] Please refer to Figure 9, which is a schematic diagram of another face-to-face communication scenario provided in an embodiment of this application.
[0239] The interface shown in Figure 9(a) can be, for example, the interface of a face-to-face repair application. The interface shown in Figure 9(a) includes a recording control 802, which can be used to record speech. Optionally, the recording control 802 can be used to record the speech of a hearing-impaired / speech-impaired user, or it can be used to record the speech of a normal user. Normal users can be users with normal hearing or normal speaking abilities. It should be noted that although the style of the recording control 802 in Figures 9 and 8 is different, its function is the same; both are used to trigger speech recording.
[0240] In this embodiment, the electronic device can detect a click operation on the recording control 802 and, in response to the operation, enter a continuous audio recording mode, during which the other party's voice is converted into text in real time. As shown in the interface of Figure 9(a), during the recording process, a prompt message indicating that recording is in progress can be displayed, such as "Listening, we can help you convert it to text." Then, in the interface of Figure 9(a), the electronic device can display the text converted from the user's speech, such as "I am a hearing person, I am speaking." Optionally, the content of the conversation can be displayed in a dialog box (chat box). The electronic device can cancel the continuous audio recording mode by detecting another click operation on the recording control 802. Optionally, the electronic device can also enter the continuous audio recording mode when it detects a long press operation on the recording control 802, and then cancel the continuous audio recording mode after detecting the operation of leaving the recording control 802; that is, it is in continuous audio recording mode while the recording control 802 is long-pressed. Optionally, in this embodiment, the user's speech can be repaired before being converted into text for display, thereby improving the face-to-face communication experience.
[0241] Then, as shown in Figure 9(b), the electronic device can detect user operations on the interface, such as user operations on a blank area of the interface, and in response to the operation, display the interface shown in Figure 9(c). The interface shown in Figure 9(a) can be, for example, the interface for a voice recording function, which can be described with reference to the embodiment in Figure 8, and will not be repeated here. The interface shown in Figure 9(c) can be, for example, the interface for a text input function, and switching from the interface shown in Figure 9(a) to the interface shown in Figure 9(c) can be understood as switching from the voice recording function to the text input function. Optionally, the user operation can be, for example, a swipe-up operation. The interface shown in Figure 9(c) includes a virtual keyboard 808 for inputting text, a voice generation control 809 for generating speech, and a text display box 810 for displaying text. Optionally, the recording control 802 can also be used as the speech generation control, that is, after inputting text through the virtual keyboard 808, the speech is displayed in the dialog box through user operations on the recording control 802.
[0242] In another possible implementation, the interface shown in Figure 9(b) may also include a keyboard trigger control, which is used to trigger the display of the virtual keyboard 808. After the electronic device detects a user operation on the keyboard trigger control, it can display the interface shown in Figure 9(c).
[0243] In this embodiment of the application, as shown in FIG9(c), the electronic device can detect user operations on the virtual keyboard 808 and, in response to the operation, display the text entered by the user in the text display box 810.
[0244] Then, as shown in Figure 9(d), the electronic device can detect user actions performed on the speech generation control 809 and, in response to these actions, generate speech based on the text displayed in the text display box 810. The interface shown in Figure 9(d) can include the speech of the text displayed in the text display box 810. For example, if the text displayed in the text display box 810 is "Hello, I am Zhang San," then the electronic device will generate the speech "Hello, I am Zhang San" after detecting a user action performed on the speech generation control 809. Furthermore, when generating the speech "Hello, I am Zhang San," the tone of the speech can also be generated based on punctuation marks or emoticons in the text.
[0245] For example, if the text includes a period (.), the message "Hello, I am Zhang San" can be announced in a neutral tone; if the text includes an exclamation mark (!), the message can be announced in an exclamatory tone; and if the text includes a question mark (?), the message can be announced in a questioning tone. Furthermore, if the text includes a smiley face, the message can be announced in a cheerful tone; and if the text includes a crying face, the message can be announced in a sad tone.
[0246] Then, as shown in the interface in Figure 9(e), the electronic device can detect the user's operation on the voice and, in response to the operation, can play the voice message "Hello, I am Zhang San".
[0247] In this embodiment, the user can continue to input text. Then, as shown in FIG9(f), the electronic device can detect the user operation on the virtual keyboard 808 and, in response to the operation, display the text entered by the user in the text display box 810.
[0248] Then, in the interface shown in Figure 9(g), the electronic device can detect the user operation applied to the voice generation control 809, and in response to the operation, generate voice according to the text displayed in the text display box 810, for example, displaying the interface shown in Figure 9(h). The interface shown in Figure 9(h) can include the voice of the text displayed in the text display box 810. For example, if the text displayed in the text display box 810 is "I want to ask what time it closes today," then after detecting the user operation applied to the voice generation control 809, the electronic device generates the voice of "I want to ask what time it closes today."
[0249] Then, as shown in Figure 9(i), the electronic device can detect user actions on the interface, such as user actions on a blank area of the interface, and in response to such actions, can switch to the voice recording function interface. Optionally, the user action can be, for example, a swipe down operation.
[0250] In this embodiment of the application, the voice input by a normal user can be converted into text, and the hearing / speech impaired user can input text, and then the electronic device will display the corresponding voice, which is beneficial to the communication between the hearing / speech impaired user and the normal user.
[0251] In general, in face-to-face communication scenarios, it supports functions such as speech-to-text conversion for the other party, text-to-speech conversion for the local user, and voice restoration for the local user.
[0252] In another possible implementation, the face-to-face repair application may only include a text input interface, omitting voice input functionality. This simplifies the application's functionality, reduces its package size, and consequently decreases the storage resources required by the electronic device. Furthermore, the embodiments of this application, by configuring both voice recording and text input functions within the face-to-face repair application, allow users to choose their preferred communication method, thereby enhancing communication flexibility and improving the user experience.
[0253] It should be noted that when the Face-to-Face Repair app has both voice recording and text input functions, the default interface upon entering the app can be either the voice recording interface or the text input interface; there is no restriction on this. Optionally, the default interface of the Face-to-Face Repair app can be the interface displayed when exiting the app. For example, if the exit interface was the voice recording interface, the default interface upon re-entry will be the voice recording interface; similarly, if the exit interface was the text input interface, the default interface upon re-entry will be the text input interface.
[0254] It should be noted that a system that can integrate voice restoration capabilities into electronic devices can provide voice restoration capabilities for some or all applications or scenarios involving voice input functions. In other words, it can restore the voice before it is emitted for some or all applications or scenarios involving voice input functions, thus comprehensively improving the user experience.
[0255] The above examples illustrate the interface for communication scenarios. The following examples illustrate the specific implementation of voice processing.
[0256] Please refer to Figure 10, which is a flowchart illustrating a voice processing method provided in an embodiment of this application. The method shown in Figure 10 can be executed by at least one of an electronic device or a cloud. The method shown in Figure 10 may include:
[0257] S1001, Obtain sound input.
[0258] In this embodiment of the application, the received input can be the user's voice. For example, the voice input method can be as shown in Figures 5, 7, 8, or 9, and will not be described in detail here.
[0259] S1002. Input the acquired sound into the ASR model.
[0260] The ASR model is used to convert sound into text. In this embodiment, the ASR model can be a pre-trained model. This embodiment inputs the acquired sound into the ASR model, which then converts the input sound into text output.
[0261] S1003. Obtain the text output by the ASR model.
[0262] S1004. Correct the text.
[0263] In this embodiment, if the acquired sound is input by a user with a speech impairment, the converted text may contain errors, thus requiring text correction. Optionally, text correction can be automatic or manual. Automatic text correction may include, but is not limited to, statistical text correction or context-aware text correction. Statistical text correction may utilize large amounts of text data to learn word frequencies and contextual information, using a statistical model to predict the most likely correct text. Context-aware text correction may combine contextual information, such as the preceding and following sentences, for example, for text correction. Manual correction may be, for example, as shown in Figure 8.
[0264] S1005. Obtain the result of the voice registration.
[0265] In this embodiment, the result of voice registration may include the registered voice or voiceprint features extracted from the registered voice. Optionally, when a user uses the voice repair function, if it is detected that the user has not registered a voice, the user is redirected to the voice registration interface to register; alternatively, the user may have already registered a voice before using the voice repair function. For example, the method of registering a voice can be referred to the relevant description in Figure 2.
[0266] S1006. Perform TTS broadcasting based on the corrected text and voice registration results.
[0267] In this embodiment, the corrected text can be converted into speech and then played back according to the registered voiceprint. This not only results in higher intelligibility of the output speech compared to the acquired speech, but also improves the user experience by playing back the voiceprint of the registered user. Optionally, TTS can use a neural network model to convert the text into audio with the user's voice and play it back.
[0268] It should be understood that, through the registration of voiceprints in this application embodiment, voiceprint features can be obtained in advance. When performing voice broadcasting, the pre-registered voiceprints can be obtained, which can reduce the time for extracting voiceprints after obtaining the sound, and thus reduce the time between obtaining the sound and broadcasting the speech, thereby improving the efficiency of speech processing.
[0269] In one possible implementation, S1005 may not be necessary. In this case, S1006 can be based on the corrected text to broadcast the speech. In this case, the voiceprint of the broadcast speech can be the voiceprint built into the electronic device or the voiceprint of the acquired sound.
[0270] It should be noted that, since the acquired sound is converted into text in this embodiment, the solution of this embodiment can be applied to scenarios that require text display, such as the scenarios shown in Figures 7, 8, or 9. Furthermore, because it also needs to be converted into text, the solution of this embodiment can be applied to scenarios where real-time requirements are not particularly high.
[0271] Another possible implementation is to correct the acquired sound and then convert it into text using an ASR model.
[0272] The following is an example illustrating the specific implementation of sound restoration.
[0273] Please refer to Figure 11, which is a flowchart illustrating a sound registration method according to an embodiment of this application. The method shown in Figure 11 can be executed by at least one of an electronic device or a cloud. The method shown in Figure 11 may include:
[0274] S1101, Voiceprint Registration.
[0275] In this embodiment of the application, the electronic device can trigger voiceprint registration processing in response to a user's voiceprint registration operation. For example, the voiceprint registration operation can be as shown in Figure 2(e), and is not limited thereto.
[0276] S1102, Sound Input.
[0277] In this embodiment of the application, after the electronic device responds to the user's voiceprint registration operation, it can collect the user's voice input. For example, the way to receive the user's voice input can be as shown in Figure 2(e), and is not limited here.
[0278] S1103, ASR recognition rate grading judgment.
[0279] In this embodiment, after collecting the user's voice input, the electronic device can perform ASR recognition on the voice input to obtain the ASR recognition result. It should be understood that the ASR recognition result can be interpreted as the recognized text corresponding to the user's voice input. Then, the ASR recognition result is compared with pre-recorded audio content to obtain the ASR recognition rate. The ASR recognition rate can also be called the similarity of the comparison, used to indicate the degree of similarity between the ASR recognition result and the pre-recorded audio content. For example, the audio content can be, for instance, the text content shown in Figure 2(f), and is not limited here. After obtaining the ASR recognition rate, it can be determined which recognition rate level it belongs to. Optionally, multiple recognition rate levels can be pre-defined, with no common recognition rate between the multiple recognition rate levels, such as (0, 70%) and (70%, 100%). If the ASR recognition rate is 50%, it belongs to the (0, 70%) level; if the ASR recognition rate is 80%, it belongs to the (70%, 100%) level.
[0280] This application embodiment improves the efficiency of repair model selection by classifying ASR recognition rates before selecting a repair model.
[0281] S1104, Repair Model Selection.
[0282] The repair model is used to repair the input speech to obtain the repaired speech. In this embodiment, the repair model can be selected based on the ASR recognition rate classification result. Optionally, a mapping relationship between each recognition rate level and the repair model is pre-configured. After obtaining the ASR recognition rate classification result, the repair model corresponding to the ASR recognition rate classification result can be obtained from the mapping relationship. Optionally, the speech repair capabilities of the repair models corresponding to different recognition rate levels are different. Optionally, the higher the recognition rate in a recognition rate level, the weaker the speech repair capability of the corresponding repair model; that is, the lower the recognition rate in a recognition rate level, the stronger the speech repair capability of the corresponding repair model.
[0283] For example, by calculating the recognition rate of user-recorded audio, the intelligibility of the user's speech can be indirectly determined, allowing for the selection of at least one suitable restoration model from a pool of alternative restoration models. For instance, users with an ASR recognition rate below 70% use a more robust restoration model, while users with an ASR recognition rate above 70% use a less robust restoration model that focuses on voice enhancement and accent correction.
[0284] This application embodiment selects a repair model based on the ASR recognition rate classification results. In other words, it can select a repair model based on the degree of impairment of the user's speaking ability. This allows for the selection of a weaker repair model for users with less impairment of speaking ability, resulting in less computational power required for speech repair. Conversely, it allows for the selection of a weaker repair model for users with more impairment of speaking ability, resulting in better speech repair performance. This approach balances the effectiveness of speech repair with the computational resources required for speech repair.
[0285] S1105, Voiceprint Feature Extraction Module.
[0286] The voiceprint feature extraction module is used to extract voiceprint features (referred to as voiceprint). In this embodiment, the voiceprint feature extraction module is used to extract the voiceprint of the user's input voice. Optionally, the voiceprint feature extraction module can convert audio of arbitrary length into a fixed-length voiceprint feature vector. This voiceprint feature extraction module can be an algorithm or a neural network model. Optionally, the voiceprint extraction module can select a voiceprint extraction model. In this embodiment, the voiceprint extraction model can include, but is not limited to, an emphasized channel attention, propagation and aggregation in time delay neural network (ECAPA-TDNN) structure.
[0287] S1106, a fixed-length eigenvector.
[0288] In this embodiment, the fixed-length feature vector is a vector obtained by extracting voiceprint features from the user's recorded voice, used to indicate the user's voiceprint characteristics. By converting the extracted voiceprint features into a fixed-length voiceprint feature vector, this embodiment helps reduce the storage resources required to store the feature vector.
[0289] S1107, Save.
[0290] In this embodiment of the application, the feature vector can be saved so that the voiceprint can be used for subsequent speech playback.
[0291] In another possible implementation, the length of the feature vector saved by S1106 can be related to the length of the input sound, rather than being a fixed-length feature vector. This allows for adaptive saving based on the length of the input sound, improving the accuracy of the saved voiceprint feature vector.
[0292] In another possible implementation, in S1103, after obtaining the ASR recognition rate, gear determination can be omitted, and the repair model can be selected directly based on the ASR recognition rate.
[0293] In another possible implementation, S1103 and S1104 may not be necessary. Instead, a pre-configured repair model can be used to repair the sound, which can reduce the time required for ASR recognition rate classification and repair model selection.
[0294] The architecture of one of the repair models is described below.
[0295] Please refer to Figure 12, which is a schematic diagram of the architecture of a repair model provided in an embodiment of this application. The repair model of this embodiment can be deployed on at least one of electronic devices or in the cloud. As shown in Figure 12, the repair model may include a prosodic encoder 1201, a streaming ASR encoder 1202, an audio discretization unit 1203, a streaming audio repair large model 1204, and a vocoder 1205.
[0296] The prosody encoder 1201 is used to extract prosodic features (hereinafter referred to as prosody). For example, the prosody encoder 1201 can extract prosodic features from an input audio signal. Prosodic features may include, but are not limited to, at least one of pitch, rhythm, or intonation. Optionally, an audio signal from a normal user can be acquired, and then prosodic features can be extracted from the normal user's audio signal as the prosodic features for playing audio.
[0297] The streaming ASR encoder 1202 is an encoder capable of converting audio signals into text sequences in real-time or near real-time. In this embodiment, the streaming ASR encoder 1202 is used to convert user-input audio into text in real-time or near real-time. The streaming ASR encoder 1202 may include a neural network model.
[0298] The audio discretization unit 1203 is used to convert a continuous audio signal into a discrete sequence of symbols or units. For example, a continuous audio signal (such as a waveform signal) is decomposed into a series of discrete, identifiable units or symbols by some method. These units can be phoneme-based, feature-based, or obtained through other means (such as cluster analysis).
[0299] The streaming audio restoration model 1204 is used to restore or optimize audio features. Audio features may include, but are not limited to, at least one of content features, prosodic features, or voiceprint features. Optionally, the streaming audio restoration model 1204 utilizes streaming language model-based audio conversion (streaming LM-based VC) for restoration or optimization, combining the advantages of streaming processing and language models (LM) to achieve real-time, efficient audio conversion. Optionally, the streaming audio restoration model 1204 may include, but is not limited to, the streaming prosodic restoration model 1204.
[0300] Among them, the vocoder 1205 is a synthesizer that converts audio features into playable audio signals.
[0301] In the embodiments of this application, the repair model can be a pre-trained model.
[0302] In one possible implementation, the restoration model can perform the following processing when performing audio restoration:
[0303] The prosodic encoder 1201 can extract prosodic features from the built-in normal user audio. The streaming ASR encoder 1202 performs ASR recognition processing on the input audio to be repaired, obtaining the ASR recognition result (also known as content features or content). Then, the result of concatenating the ASR recognition result, prosodic features, and user-registered voiceprint features is input into the streaming audio restoration model 1204. In addition, the streaming audio restoration model 1204 fuses the voiceprint features, prosodic features, and ASR recognition result in the concatenated result with the initial audio discretization features (also known as discretization features), and generates the repaired discretization features and repaired continuous audio features through autoregressive inference. The discretization features generated by autoregressive inference will be used as input for subsequent autoregressive inference, which will be fused and inferred by the streaming audio restoration model 1204, while the generated continuous audio features will be used as input for the vocoder 1205. Then, by converting the continuous features of the audio into the repaired audio through the vocoder 1205, the repaired audio can be broadcast through an audio broadcasting device (e.g., a loudspeaker). The voiceprint of the broadcast audio can be the voiceprint registered by the user, and the prosody of the broadcast audio can be the prosody extracted by the prosody encoder 1201.
[0304] The audio to be repaired can be user-input audio, acquired through a sound acquisition device (e.g., a microphone) of an electronic device. For example, the audio to be repaired in this embodiment may include, but is not limited to, the input speech in the scenarios shown in Figures 5, 7, 8, or 9. Optionally, the initial audio discretization features may be obtained by the audio discretization unit 1203 based on at least one of the ASR recognition results or the audio to be repaired.
[0305] In this embodiment of the application, since the ASR recognition result is obtained by performing ASR recognition processing on the audio to be repaired, and the audio discretization feature is also obtained by performing discretization processing on the audio to be repaired, the ASR recognition result and the audio discretization feature are fused together. That is, the processing results of the audio to be repaired from different dimensions can be fused together, thereby improving the accuracy of the obtained audio continuous features, and then improving the accuracy of the content of the obtained audio to be repaired.
[0306] Optionally, audio repair can be performed as soon as the user inputs audio, thus improving the real-time performance of audio repair. Alternatively, repair can begin after a certain amount of audio has been acquired. For example, if one second of audio has been acquired while the user is inputting audio, repair processing can begin on that one second, and then on the next second, until a complete segment of audio input by the user has been repaired.
[0307] In another possible implementation, generating the restored discretized features may not be necessary, which simplifies the architecture of the restoration model and improves the efficiency of sound restoration. However, the embodiments of this application generate restored discretized features, which can be used as input for subsequent autoregressive inference and fused and inferred by the streaming audio restoration large model 1204, thereby improving the accuracy of sound restoration.
[0308] Alternatively, each segment of the audio to be repaired can be of the same length, which can improve the accuracy of audio repair; or the segments of the audio to be repaired can be of different lengths, which can improve the flexibility of audio repair.
[0309] For example, assuming the user inputs the audio "Going or not this afternoon", after obtaining "this afternoon", it is used as the audio to be repaired. Then, the streaming ASR encoder 1202 performs ASR recognition processing on the input "this afternoon" audio segment to obtain the ASR recognition result. Then, the result of concatenating the ASR recognition result, prosodic features, and user-registered voiceprint features is input into the streaming audio repair model 1204. Furthermore, the audio discretization unit 1203 discretizes the "this afternoon" audio segment based on the ASR recognition result to obtain audio discretization features. Then, the audio discretization features are input into the streaming audio repair model 1204. The streaming audio repair model 1204 can output the repaired audio discretization features and repaired audio continuous features corresponding to the "this afternoon" audio segment. The vocoder 1205 can output the repair result of the "this afternoon" audio segment based on the repaired audio continuous features. Furthermore, the streaming audio restoration model 1204 can use the discretized features of the restored audio corresponding to the "this afternoon" audio segment as a reference for whether or not to remove this audio segment during restoration processing, for example, as a reference for whether or not to remove this audio segment during discretization processing.
[0310] The training of the repair model will be illustrated below.
[0311] Obtain a training sample set, which includes one or more training samples. Each training sample includes an audio sample, a text sample corresponding to the audio sample, and a repaired audio sample corresponding to the audio sample.
[0312] During training, audio samples are used as input to the streaming ASR encoder 1202, which outputs predicted text. The predicted text is compared with the text samples corresponding to the audio samples to calculate a first training loss. If the first training loss meets a first training termination condition, the streaming ASR encoder 1202 terminates training. If the first training loss does not meet the first training termination condition, the parameters of the streaming ASR encoder 1202 are updated, and training continues until the first training loss meets the first training termination condition. Optionally, the first training termination condition may include the first training loss being less than a first threshold.
[0313] Then, the restoration model continues to be trained. Audio samples are used as input to the streaming ASR encoder 1202, or the corresponding text samples are used as input to the streaming audio restoration model 1204. The restored audio output by the vocoder 1205 is obtained, and then the restored audio output by the vocoder 1205 is compared with the restored audio samples corresponding to the audio samples to calculate the second training loss. If the second training loss meets the second training termination condition, the restoration model ends training; if the second training loss does not meet the second training termination condition, the parameters of the restoration model are updated and training continues until the second training loss meets the second training termination condition. Optionally, the second training termination condition may include the second training loss being less than a second threshold. Optionally, at least one of the vocoder 1205 or the prosodic encoder 1201 in this embodiment may be pre-trained. In this case, updating the parameters of the restoration model may involve updating at least one of the streaming audio restoration model 1204 or the audio discretization unit 1203.
[0314] Optionally, the audio samples may include a first audio sample and a second audio sample. For example, if an audio sample could include "Today's weather is sunny," then the first audio sample could include "Today's weather," and the second audio sample could include "sunny." Correspondingly, the text samples corresponding to the audio samples include the text of the first audio sample and the text of the second audio sample, and the repaired audio samples corresponding to the audio samples include the repaired audio samples corresponding to the first and second audio samples.
[0315] Then, during training, the repaired audio corresponding to the first audio sample output by the repair model is compared with the repaired audio sample to calculate the second training loss. Furthermore, the discretized features of the repaired audio corresponding to the first audio sample output by the repair model are compared with the discretized features of the second audio sample to calculate the third training loss. If the third training loss satisfies the third training termination condition, the repair model training ends; if the third training loss does not satisfy the third training termination condition, the parameters of the repair model are updated and training continues until the third training loss satisfies the third training termination condition. Optionally, the third training termination condition may include the third training loss being less than a third threshold.
[0316] It should be understood that the processing flow of the repair model in the embodiment of this application during the training process can be referred to the description of the sound repair processing flow in the above embodiment, and will not be repeated here.
[0317] It should be noted that users can access the personalized portal and record a large amount of text as training samples to train a personalized repair model, which can improve the user experience.
[0318] In another possible implementation, the audio discretization unit 1203 may not require ASR recognition results when discretizing the audio to be repaired. This can improve the efficiency of discretization processing and thus improve the efficiency of audio repair. In this embodiment, the audio discretization unit 1203 uses ASR recognition results to discretize the audio to be repaired, which can improve the accuracy of audio repair.
[0319] In another possible implementation, the restoration model may include either the audio discretization unit 1203 or the streaming ASR encoder 1202. This improves the efficiency of discretization processing, while audio restoration using both the audio discretization unit 1203 and the streaming ASR encoder 1202 enhances the accuracy of audio restoration. It should be noted that if the restoration model includes the audio discretization unit 1203 but not the streaming ASR encoder 1202, the result of concatenating the audio discretization features, prosodic features, and user-registered voiceprint features can be input into the large streaming audio restoration model 1204.
[0320] In another possible implementation, the prosody encoder 1201 may not be necessary during audio restoration. Instead, by pre-storing prosodic features, audio restoration can be performed using these pre-stored prosodic features, thus improving the efficiency of audio restoration.
[0321] In another possible implementation, the streaming audio inpainting large model 1204 may not require the prediction of audio discretization features, which can improve the efficiency of audio inpainting, while the prediction of audio discretization features can improve the accuracy of audio inpainting.
[0322] It should be noted that both the prosodic encoder 1201 and the streaming prosodic restoration model 1204 can adopt a generative pre-trained transformer (GPT) structure. The speech discretization unit can use a vector-quantized variational autoencoder (VQVAE) model, and the streaming ASR encoder 1202 can adopt a recurrent neural network transducer (RNNT). The vocoder 1205 can include an efficient and high-fidelity generative adversarial network (HiFi GAN). Specifically, VQVAE can convert audio into frame-level discretized units (discrete features). HiFi GAN can convert audio features into audio sample points (audio signals).
[0323] The following examples illustrate the process of the embodiments of this application based on the above examples.
[0324] Please refer to Figure 13, which is a flowchart illustrating another voice processing method provided in an embodiment of this application. The method of this application embodiment can be executed by an electronic device, a server, or a system including both an electronic device and a server. This application embodiment can repair user-inputted voice. The method shown in Figure 13 may include:
[0325] S1301, Receive setting input, the setting input is used to indicate the scenario for enabling voice repair, the contact for enabling voice repair, or the application for enabling voice repair, the scenario includes face-to-face communication scenario or remote communication scenario.
[0326] The face-to-face communication scenario can be a scenario where at least two users are communicating face-to-face. The remote communication scenario can be a scenario where communication is conducted using communication means. In this embodiment, at least one of the following can be selectively enabled: a scenario for enabling voice repair, a contact for enabling voice repair, or an application for enabling voice repair. Optionally, if an input is set to enable a scenario for enabling voice repair, then if the electronic device is detected to be in a scenario where voice repair is enabled, such as a face-to-face communication scenario or a remote communication scenario, the user's input voice is repaired. For example, a face-to-face communication scenario can be, for example, the scenario described in the embodiments of FIG8 or FIG9, and will not be described in detail here. A remote communication scenario can be, for example, the scenario described in the embodiments of FIG5 or FIG7, and will not be limited here. If an input is set to enable a contact for enabling voice repair, then when the contact being communicated by the electronic device is the contact for which voice repair is enabled, the electronic device repairs the user's input voice. For example, the method of enabling a contact for enabling voice repair can be, for example, the description in the embodiment of FIG4, and will not be described in detail here. If an input is set to enable an application for enabling voice repair, then when the application being run by the electronic device is an application for which voice repair is enabled, the user's input voice is repaired. For example, the way to enable the voice repair application can be as shown in the embodiment of Figure 3, which will not be described in detail here.
[0327] S1302. Obtain the user's first voice characteristics through voice registration.
[0328] The first speech feature can be a user's speech characteristic. In this embodiment, the first speech feature is used to represent the characteristics of the user's pronunciation. Optionally, the first speech feature can include at least one of voiceprint features or prosodic features. Voiceprint features can be a sound wave spectrum carrying speech information displayed by an electroacoustic instrument. Prosodic features can refer to those features in speech that are not directly manifested as changes in timbre, but are reflected through changes in factors such as pitch, duration, and intensity. These features are manifested as suprasegmental components in the speech signal, which together with timbre components (such as vowels and consonants) constitute a complete speech signal. The first speech feature can be obtained by feature extraction from the user's registered speech. For example, the method of user registration of speech can be described with reference to the embodiment in Figure 2, and will not be repeated here. It should be noted that the users in this embodiment can include users with speech impairments or hearing impairments. Users with speech impairments can be, for example, users who have certain difficulties or defects in pronunciation. Users with hearing impairments can be, for example, users who have certain difficulties or defects in hearing.
[0329] S1303, Receive the first voice input from the user.
[0330] S1304. Repair the first speech based on the input settings and the first speech characteristics.
[0331] In this embodiment, the first speech is repaired based on the setting input and the first speech feature. Since the setting input is used to indicate the scenario where speech repair is enabled, the contact person for enabling speech repair, or the application for enabling speech repair, when the electronic device is detected to be in a scenario where speech repair is enabled, or when the contact being communicated by the electronic device is the contact person for enabling speech repair, or when the application being run by the electronic device is the application for enabling speech repair, the first speech is repaired based on the first speech feature. This allows the repaired first speech to be played using the first speech feature, and the intelligibility of the repaired first speech is higher than that of the unrepaired first speech. Intelligibility can represent the accuracy of expressing what the user wants to say, or it can be understood as the listener's understanding of the speech signal transmitted by the speaker.
[0332] In this embodiment, by receiving setting input—which indicates the scenario, contact, or application for enabling voice restoration—and then obtaining the user's first voice feature through voice registration, the system can restore the first voice based on the setting input and the first voice feature after receiving the user's input. This allows users with speech or hearing impairments to communicate via voice input, improving their communication convenience. Furthermore, by generating the restored first voice using the user's registered first voice feature, the restored first voice can be played according to that feature, making it more closely resemble the user's own voice and further enhancing the user experience.
[0333] In another possible implementation, voice repair can be performed using pre-defined voice features. For example, before a user registers a voice feature, the voice is repaired according to pre-defined features, allowing the repaired voice to be played according to those features. After a user registers a voice feature, the repair is performed based on the user's initial voice features. This means that the voice repair function can be used even if the user has not registered a voice feature, thus expanding the applicable scenarios for the voice repair function. Furthermore, it reduces the processing of extracting voice features from registered voice features, thereby improving the efficiency of voice repair.
[0334] In one possible implementation, the remote communication scenario includes a call scenario, and in the call scenario, the method further includes:
[0335] Send a first notification message to the first contact in the call. The first notification message is used to indicate that the voice repair function has been enabled.
[0336] And / or,
[0337] Upon receiving the first voice message, a second notification message is sent to the first contact, indicating that the first voice message is being repaired.
[0338] The first notification information may include notification text or a notification sound. Optionally, the notification text in the first notification information may be displayed on the screen of the terminal corresponding to the first contact person, and the notification sound in the first notification information may be played through the speaker of the terminal corresponding to the first contact person. For example, the first notification information can refer to the relevant description of the notification language in the embodiment of Figure 6, which will not be repeated here. The second notification information may include notification text or a notification sound. Optionally, the notification text in the second notification information may be displayed on the screen of the terminal corresponding to the first contact person, and the notification sound in the first notification information may be played through the speaker of the terminal corresponding to the first contact person. For example, the notification sound in the second notification information may be, for example, a "ding" sound as exemplified in the embodiment of Figure 6, which will not be repeated here.
[0339] In this embodiment, by sending a first notification message to the first contact in the call to inform the user that the voice repair function has been enabled, the first contact can be informed that the user has enabled the voice repair function, thereby improving the call experience. Furthermore, by sending a second notification message to the first contact upon receiving the first voice message to indicate that the first voice message is being repaired, the first contact can be informed that the delay is due to voice repair, reducing the possibility of conflicting speaking times in two-party or multi-party calls, thereby improving the call experience.
[0340] In another possible implementation, at least one of the first or second prompt messages may not be sent, which can reduce the consumption of transmission resources during the call.
[0341] In one possible implementation, the second notification message includes a notification tone, which is sent to the first contact upon receiving the first voice message, including:
[0342] Once the first voice message is received, continuously send notification to the first contact.
[0343] The method also includes:
[0344] Once the first restored voice message is sent to the first contact, stop sending notification sounds to the first contact.
[0345] For example, embodiments of this application can be described with reference to (c) in Figure 6, which will not be repeated here.
[0346] In this embodiment of the application, by continuously sending a prompt tone to the first contact when the first voice is received, the first contact can know that the user on the other end is speaking when they hear the prompt tone. And when the prompt tone is stopped after the repaired first voice is sent to the first contact, the first contact can then hear the repaired first voice. This can reduce the situation of conflict between two or more parties during the call, thereby improving the experience of both parties in the call.
[0347] In another possible implementation, a brief notification tone could be sent to the first contact when the first voice message is received, or after a certain interval following the receipt of the first voice message. This would reduce the transmission resources required to transmit the notification tone.
[0348] In one possible implementation, the remote communication scenario includes a call scenario, and in the call scenario, the method further includes:
[0349] A repaired first voice message is sent to the first contact in the call. Then, a first interface can be displayed, which is the interface for communicating with the first contact. The first interface includes a first control for disabling the voice repair function. Then, in response to a first operation on the first control, the voice repair function is disabled, and a second voice message is sent to the first contact. The second voice message includes the user's voice received after the voice repair function was disabled.
[0350] For example, the first interface can be, for instance, the interface shown in Figure 5. The first control can be, for example, the sound switching control 502. The first operation can be a user operation. In this embodiment, before the voice repair function is turned off, that is, when the voice repair function is on, the repaired voice is sent to the first contact, and after the voice repair function is turned off, the unrepaired voice is sent to the first contact.
[0351] In this embodiment, during a call, the user can also control the voice repair function to be turned off. This allows the user to disable voice repair when it is not needed, improving the flexibility of switching between using and not using voice repair. For example, the user can control the voice repair function to be turned off via the first control when the electronic device's battery is low or the electronic device is running slowly, thereby improving the ability to disable voice repair during calls and enhancing the user experience.
[0352] In another possible implementation, the voice repair function can be disabled on the first interface, meaning that the voice repair function is not supported during a call, thereby improving the simplicity of the call interface.
[0353] In another possible implementation, the voice repair function can be turned off via voice commands.
[0354] In one possible implementation, in response to the first operation, the first control is also switched from a first state to a second state, where the first state indicates that the voice repair function is enabled and the second state indicates that the voice repair function is disabled. The method further includes:
[0355] In response to a second operation on the first control, the first control is switched from a second state to a first state, and the voice repair function is enabled to send a repaired third voice message to the first contact, the third voice message including the voice message received from the user after the voice repair function is enabled.
[0356] The second operation can be a user operation. The first state can be, for example, the state of the sound switching control 502 shown in Figure 5(d), and the second state can be, for example, the state of the sound switching control 502 shown in Figure 5(c).
[0357] In this embodiment, the state of the first control indicates whether the voice repair function is currently enabled, thereby improving the accuracy of the user's choice to use or disable the voice repair function. Furthermore, this embodiment not only allows the voice repair function to be disabled during a call but also to be re-enabled, thus increasing the flexibility of enabling or disabling voice repair. In addition, enabling or disabling the voice repair function through the same control improves the simplicity of the interface during calls.
[0358] It should be understood that WeChat and face-to-face communication applications can also enable or disable the voice repair function within the application interface. Please refer to the relevant descriptions on how to enable or disable the voice repair function in the call interface, which will not be repeated here.
[0359] In another possible implementation, once the voice repair function is turned off during a call, it can no longer be re-enabled.
[0360] In another possible implementation, the voice repair function can be turned on or off through different controls. For example, one control can be used to turn on the voice repair function during a call, and another control can be used to turn it off. This can improve the independence of the control over turning the voice repair function on or off.
[0361] In one possible implementation, the method further includes:
[0362] A second interface is displayed, including playback controls for the repaired first speech and first text corresponding to the repaired first speech. A third operation is received, which selects a first target character in the first text. In response to the third operation, at least one candidate character related to the first target character is displayed. A fourth operation is received, which selects a second target character from the at least one candidate character. In response to the fourth operation, second text and playback controls for the speech corresponding to the second text are displayed, the second text being obtained by replacing the first target character in the first text with the second target character.
[0363] For example, the second interface can be, for example, the interface shown in Figure 8. The playback control for the repaired first voice can be, for example, playback control 803. The first text corresponding to the first voice can be, for example, "Are you going this afternoon?". The third operation can be a user operation. The first target character can be a character selected by the user. Optionally, the first target character can be, but is not limited to, text, punctuation marks, or emoticons. For example, the first target character can be, for example, "afternoon". At least one candidate character can be displayed in a list. For example, at least one candidate character can be, for example, characters displayed in candidate word list 806, such as "morning", "good at dancing", and "foggy". The fourth operation can be a user operation. The second target character can be a character selected by the user from at least one candidate character. For example, the second target character can be, for example, "morning". For example, the second text can be, for example, "Are you going this morning?". The playback control for the voice corresponding to the second text can also be, for example, playback control 803.
[0364] In this embodiment, a second interface is displayed, including playback controls for the repaired first speech and first text corresponding to the repaired first speech. A third operation is received, whereby a first target character in the first text is selected. In response to the third operation, at least one candidate character related to the first target character is displayed. A fourth operation is received, whereby a second target character is selected from the at least one candidate character. In response to the fourth operation, second text and playback controls for the speech corresponding to the second text are displayed. This allows the user to quickly and manually correct inaccurate speech repair results, thereby improving the accuracy of speech communication.
[0365] In another possible implementation, the user could re-enter the voice when they find the voice repair result to be inaccurate.
[0366] In another possible implementation, the user can generate the second text and the corresponding playback control with one click. In this case, at least one character or word in the first text can be replaced to generate the second text and the corresponding playback control.
[0367] In one possible implementation, the method further includes:
[0368] Cancel the display of the playback controls for the first audio clip after the repair.
[0369] In this embodiment of the application, by canceling the display of the playback control for the repaired first voice, only the latest voice can be played at any given time, which improves the convenience of selecting the appropriate voice for playback.
[0370] In another possible implementation, the display of the playback control for the repaired first voice recording can be left undisplayed; that is, not only the playback control for the latest voice recording, but also the playback controls for historical voice recordings can be retained.
[0371] In one possible implementation, for remote communication scenarios, the method also includes:
[0372] A third interface is displayed, which is the interface for communicating with the second contact. The third interface includes a second control and a third control. The second control is used to instruct the sending of a first voice message to the second contact, and the third control is used to instruct the sending of a repaired first voice message to the second contact. Then, in response to an operation on the third control, the repaired first voice message can be sent to the second contact. Alternatively, in response to an operation on the second control, the first voice message can be sent to the second contact.
[0373] For example, the third interface could be the interface shown in Figures 7(a)-7(f). The second control could be, for example, directly sending option 703, and the third control could be, for example, repairing and sending option 704.
[0374] In this embodiment, the user can selectively send the repaired voice or the original voice to the second contact through the second and third controls. That is, even if the voice repair function is enabled, the user can still choose to send the original voice to the second contact, thereby improving the user's flexibility in remote communication.
[0375] In another possible implementation, there could be only one voice sending control. After the electronic device detects that the voice sending control is being used, it sends the repaired voice, which can improve the efficiency of sending the repaired voice in remote communication scenarios.
[0376] In one possible implementation, for face-to-face communication scenarios, the method also includes:
[0377] A fourth interface is displayed, including a virtual keyboard. In response to an operation on the virtual keyboard, third text is displayed. A fifth operation is received on a fourth control in the fourth interface, which instructs the generation of speech. In response to the fifth operation, speech corresponding to the third text is generated based on first speech features.
[0378] For example, the fourth interface can be, for instance, the interface shown in (c)-(i) of Figure 9. The third text can be text entered by the user via a virtual keyboard, such as "Hello, I am Zhang San" or "I would like to ask what time the store closes today," etc. The fourth control can be, for example, control 809. The fifth operation can be a user operation. In this embodiment, the speech corresponding to the third text is generated based on the first speech feature, so the speech corresponding to the third text can be played according to the first speech feature.
[0379] In this embodiment, a fourth interface is displayed, including a virtual keyboard. In response to an operation on the virtual keyboard, third text is displayed. A fifth operation is received on a fourth control in the fourth interface, the fourth control indicating the generation of speech. In response to the fifth operation, speech corresponding to the third text is generated based on a first speech feature. That is, the user can also input text and then generate speech based on the registered first speech feature, thus achieving text-to-speech conversion. Users can then communicate by choosing to input either speech or text, improving the selectivity and flexibility of user communication.
[0380] In another possible implementation, the speech corresponding to the entered third text could be automatically generated after the user finishes entering the third text via the virtual keyboard, thus reducing user operations. Optionally, the third text entry could be considered complete if no operation on the virtual keyboard is detected within a certain period of time.
[0381] In one possible implementation, the third text includes punctuation marks and / or emojis. When generating the speech corresponding to the third text based on the first speech features, the tone of the speech corresponding to the third text is also controlled based on the punctuation marks and / or emojis.
[0382] In this embodiment, the tone of the speech corresponding to the third text is controlled based on punctuation marks and / or emoticons, so that the tone of the text corresponding to the third speech matches the punctuation marks and / or emoticons. For example, if the last character of the third text is marked with an exclamation mark "!", the tone of the speech corresponding to that third text can be exclamatory; if the last character of the third text is marked with a question mark "?", the tone of the speech corresponding to that third text can be interrogative. For example, if the emoticons in the third text include a smiley face, the tone of the speech corresponding to the third text is cheerful; if the emoticons in the third text include a crying face, the tone of the speech corresponding to the third text is sad.
[0383] In this embodiment, the tone of the speech corresponding to the third text can be controlled by punctuation marks and / or emoticons in the third text input by the user. This allows the tone of the output speech to be adjusted adaptively according to the user input text, thereby improving the flexibility of voice broadcasting and enhancing the user experience.
[0384] In another possible implementation, punctuation marks or emojis in the third-party text can be disregarded. This reduces the computational resources required to determine the tone matching for punctuation marks or emojis. Optionally, a mapping relationship between punctuation marks and tone, as well as between emojis and tone, can be pre-defined to achieve the desired tone matching for punctuation marks or emojis.
[0385] In one possible implementation, the user's initial voice characteristics are obtained through voice registration, including:
[0386] The fifth interface is displayed. This fifth interface is for voice registration and includes a third prompt message and a fifth control. The third prompt message indicates the recording content for voice registration, and the fifth control instructs on recording speech. In response to an action on the fifth control, a fourth voice recording is made. First speech features are extracted from the fourth voice recording.
[0387] For example, the fifth interface can be, for instance, the interface shown in Figure 2. The third prompt message can be, for example, text prompts such as "The giant panda sleeps wherever it goes," or voice prompts; there are no restrictions here. The fifth control can be, for example, a recording control 208. The fourth voice can be voice input by the user, used to register the user's voice so as to extract the user's first voice feature, thereby enabling the playback of the repaired voice according to the user's first voice feature.
[0388] In one possible implementation, the method also includes:
[0389] Speech recognition is performed on the fourth speech to obtain the fourth text. Then, the fourth text is compared with the recorded content to obtain the similarity between the fourth text and the recorded content. Then, a target restoration model can be selected from multiple restoration models based on the similarity. The target restoration model is used to restore the first speech. The speech restoration capabilities of the multiple restoration models are different, and the speech restoration capability of the target restoration model is negatively correlated with the similarity.
[0390] In this embodiment of the application, the text obtained by speech recognition of the user's registered speech is compared with the recorded content to obtain a similarity score. Then, a target repair model is selected from multiple repair models based on the similarity score. This allows for the selection of a suitable repair model based on the user's speech impairment level, thereby balancing the accuracy of speech repair with the computing resources required for speech repair.
[0391] It should be understood that multiple repair models can be deployed on a single server, allowing the corresponding repair model to be retrieved from that server for voice repair via model identifier. Alternatively, they can be deployed on multiple servers, allowing the server corresponding to the target repair model to be called based on the server-repair model mapping. Furthermore, multiple repair models can also be deployed on electronic devices; this is not a limitation.
[0392] In another possible implementation, only one restoration model can be configured, thus eliminating the need for similarity matching with the recorded content and improving the efficiency of voice restoration.
[0393] The above embodiments describe an example of repairing a user's voice. The following embodiments illustrate an example of repairing the voice of a contact person who communicates with the user.
[0394] Please refer to Figure 14, which is a flowchart illustrating another voice processing method provided in an embodiment of this application. The method of this application embodiment can be executed by an electronic device, a server, or a system including both an electronic device and a server. This embodiment can also be used to repair the voice of a contact person in user communication. The method shown in Figure 14 may include:
[0395] 1401. Receive setting input. The setting input is used to indicate the scenario for enabling voice repair, the contact for enabling voice repair, or the application for enabling voice repair. The scenario includes face-to-face communication scenario or remote communication scenario.
[0396] S1401 can be referred to the description of S1301, and will not be repeated here.
[0397] 1402. Receive the fifth voice message from the target contact.
[0398] The target contact can be a contact who communicates with the user; for example, the target contact can be the first contact or the second contact.
[0399] 1403. Obtain the second speech feature, which is a preset speech feature or a speech feature extracted from the fifth speech feature.
[0400] The second speech feature may include at least one of voiceprint features or prosodic features. In another possible implementation, it could be speech features extracted from a user's registered voice. For example, if the user of the electronic device is a first user, and the first user wants to hear a second user's voice during speech restoration, the second user could register their voice, and then speech features could be extracted from the second user's registered voice.
[0401] 1404. Repair the fifth voice based on the input settings and the second voice characteristics.
[0402] In this embodiment, the fifth voice is repaired based on the setting input and the second voice feature. Since the setting input is used to indicate the scenario where voice repair is enabled, the contact person who enables voice repair, or the application that enables voice repair, when the electronic device is detected to be in a scenario where voice repair is enabled, when the contact person being communicated with by the electronic device is the contact person who enables voice repair, or when the application being run by the electronic device is the application that enables voice repair, the fifth voice is repaired based on the second voice feature, so that the intelligibility of the repaired fifth voice is higher than that of the unrepaired fifth voice. Intelligibility can represent the degree of understanding when the voice is heard, or it can be understood as the accuracy with which the voice expresses what the contact person intends to convey.
[0403] In one possible implementation, for remote communication scenarios, the method also includes:
[0404] The sixth interface is displayed, which is the interface for communicating with the target contact. The sixth interface includes the fifth voice message. Then, in response to an action on the fifth voice message, a sixth control and a seventh control can be displayed. The sixth control instructs the conversion of the fifth voice message into text, and the seventh control instructs the conversion of the repaired fifth voice message into text. Then, in response to an action on the seventh control, the text corresponding to the repaired fifth voice message can be displayed. Alternatively, in response to an action on the sixth control, the text corresponding to the fifth voice message can be displayed.
[0405] For example, the sixth interface can be, for example, the interface shown in (g)-(i) of Figure 7, the sixth control can be, for example, the speech-to-text option 708, and the seventh control can be, for example, the speech-to-text option 707 after voice repair.
[0406] In this embodiment of the application, after receiving the voice from the target contact, the fifth voice can be selectively converted into text, which can also be understood as converting the original voice of the target contact into text; or the repaired fifth voice can be converted into text. Users can choose to display the text as needed, thereby improving the flexibility and experience of user communication.
[0407] In another possible implementation, the sixth control can be used to instruct the playback of the fifth voice, and the seventh control can be used to instruct the playback of the repaired fifth voice. In this embodiment, the user can choose to play either the original voice of the target contact or the repaired voice, thereby improving the flexibility and experience of user communication.
[0408] In one possible implementation, for face-to-face communication scenarios, the method also includes:
[0409] The seventh interface is displayed, which is the interface for communicating with the target contact. The seventh interface includes an eighth control. Then, in response to the sixth action on the eighth control, the fifth text is displayed. Then, in response to the seventh action on the eighth control, the display of the fifth text stops. The fifth text includes the text corresponding to the fifth voice within the target time period or the text corresponding to the repaired fifth voice. The target time period includes the time period from the response to the sixth action to the response to the seventh action.
[0410] For example, the seventh interface could be the interface shown in Figures 9(a)-(b), and the eighth control could be, for example, a recording control 802. The sixth operation could be a user input operation. The fifth text could be, for example, "I am a hearing person, I am speaking."
[0411] In this embodiment of the application, the voice of the target contact can also be converted into text, thereby improving the flexibility of communication between the user and the contact.
[0412] In another possible implementation, the eighth control could be omitted to control the voice-to-text conversion of contacts. For example, all voice messages from contacts could be converted to text, thus reducing the number of voice-to-text operations.
[0413] Based on the above embodiments, the following is an exemplary description of how to implement voice restoration.
[0414] Please refer to Figure 15, which is a flowchart illustrating another voice processing method provided in an embodiment of this application. The method of this application embodiment can be executed by an electronic device, a server, or a system including both an electronic device and a server. The method shown in Figure 15 may include:
[0415] S1501. Obtain the target speech and speech features.
[0416] The target speech can be unrepaired speech, which can be understood as the original audio or audio to be repaired. For example, the target speech may include, but is not limited to, the first speech or the fifth speech. Speech features may include, but are not limited to, the first speech feature or the second speech feature.
[0417] S1502. Input the target speech and speech features into the target restoration model. The target restoration model is used to extract the content of the target speech, obtain continuous speech features based on the content and speech features of the target speech, and synthesize the restored target speech based on the continuous speech features.
[0418] Optionally, the content of the target speech can be obtained through speech recognition. In this embodiment, the obtained target speech can be the content of the extracted target speech read aloud according to the speech features.
[0419] S1503. Obtain the repaired target speech output by the target repair model.
[0420] For example, this embodiment can be referred to the relevant description in FIG12, which will not be repeated here.
[0421] In this embodiment, after the target speech is input, the electronic device can acquire the target speech and speech features, and then input the target speech and speech features into the target restoration model. The target restoration model is used to extract the content of the target speech, obtain speech continuity features based on the content and speech features of the target speech, and synthesize the restored target speech based on the speech continuity features. In this way, people with speech impairments can communicate by inputting speech, thereby improving the convenience of communication for people with speech impairments.
[0422] In one possible implementation, the target restoration model includes a first module, a second module, and a third module. The first module is used to extract the content of the target speech, the second module is used to obtain continuous speech features based on the content and speech features of the target speech, and the third module is used to synthesize the restored target speech based on the continuous speech features.
[0423] For example, the first module may include a streaming ASR encoder. The second module may include a streaming audio restoration large model. The third module may include a vocoder.
[0424] In one possible implementation, the target speech and speech features are input into the target inpainting model, including:
[0425] The speech features and a first portion of the target speech are input into the target restoration model. The first module extracts the content of the first portion of speech, the second module obtains continuous features of the target speech based on the speech features and the content of the first portion, and the third module synthesizes the restored first portion of speech based on the continuous features of the target speech. Then, the speech features and a second portion of the target speech are input into the target restoration model. The first module extracts the content of the second portion of speech, the second module obtains continuous features of the second portion of speech based on the speech features and the content of the second portion, and the third module synthesizes the restored second portion of speech based on the continuous features of the second portion.
[0426] For example, the first part of the voice and the second part of the voice in the embodiments of this application can be described with reference to (c) in FIG6. For example, the first part of the voice includes "Thank you, put it downstairs" and the first part of the voice may include "That's fine".
[0427] The repaired target speech includes the repaired first part of the speech and the repaired second part of the speech.
[0428] In this embodiment, by first repairing a portion of the target speech and then repairing another portion, speech repair can begin when only a portion of the speech is acquired. In other words, speech repair can begin without acquiring the complete target speech, thereby improving the efficiency of speech repair.
[0429] In another possible implementation, speech restoration can begin only after the complete target speech has been acquired. This reduces the likelihood of target speech acquisition and restoration occurring simultaneously, thereby reducing the resources required for speech restoration.
[0430] In one possible implementation, the target restoration model further includes a fourth module, which is used to discretize the first part of the speech to obtain discrete features of the target speech. The second module is also used to make predictions based on the discrete features of the target speech to obtain the restored discrete features of the target speech, and to obtain the second continuous features of the speech based on the restored discrete features of the target speech, the speech features, and the content of the second part of the speech.
[0431] For example, the fourth module may include an audio discretization unit.
[0432] In this embodiment, the first part of the speech is discretized by the fourth module to obtain the discrete features of the target speech. Then, the second module predicts based on the discrete features of the target speech to obtain the repaired discrete features of the target speech. Then, the second continuous features of the speech are obtained based on the repaired discrete features of the target speech, the speech features, and the content of the second part of the speech. In this way, the second continuous features of the speech can be obtained by combining the discrete features of the target speech, the speech features, and the content of the second part of the speech, thereby improving the accuracy of the obtained second continuous features of the speech and thus improving the accuracy of speech repair.
[0433] In another possible implementation, the fourth module may not be necessary, which would improve the efficiency of voice restoration.
[0434] In one possible implementation, the fourth module is used to discretize the first part of the speech to obtain discrete features of the target speech, including:
[0435] The fourth module is used to discretize the first part of the speech based on its content to obtain the discrete features of the target speech.
[0436] In this embodiment of the application, the first part of the speech is discretized according to the content of the first part of the speech to obtain the discrete features of the target speech. That is, the content of the first part of the speech is used as a reference for the discretization process, thereby improving the accuracy of the obtained discrete features of the target speech and thus improving the accuracy of speech restoration.
[0437] In another possible implementation, the fourth module can be used to discretize the first part of the speech based on its content to obtain the discrete features of the target speech. This can improve the efficiency of obtaining the discrete features of the target speech, thereby improving the efficiency of speech restoration.
[0438] It should be noted that the solution in this application embodiment can be used not only for sound restoration tasks, but also extended to tasks such as dialect to Mandarin conversion and cross-language translation.
[0439] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation portals are provided for users to choose to authorize or refuse.
[0440] It should also be understood that the steps of the various embodiments described above can be coupled to each other, and this application does not limit this. Furthermore, the sequence number of each process does not imply the order of execution; the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0441] This application also proposes a voice processing device that can execute the steps of the above-described method embodiments. For example, the voice processing device includes a receiving module and an output module, wherein the receiving module is used to receive setting input, obtain the user's first voice feature through voice registration, receive the user's input first voice, etc.; the output module is used to repair the first voice based on the setting input and the first voice feature, etc.
[0442] In another possible implementation, the receiving module is used to receive setting input, receive the fifth voice from the target contact, obtain the second voice features, etc., and the output module is used to repair the fifth voice based on the setting input and the second voice features, etc.
[0443] In another possible implementation, the receiving module is used to acquire the target speech and speech features, the output module is used to input the target speech and speech features into the target restoration model, the target restoration model is used to extract the content of the target speech, obtain speech continuous features based on the content and speech features of the target speech, and synthesize the restored target speech based on the speech continuous features; and the restored target speech output by the target restoration model is acquired.
[0444] The steps performed by the voice processing device can be referred to the description of the method embodiments, and will not be repeated here.
[0445] It should be understood that the steps performed by the apparatus in this application embodiment can be referred to the description of the method embodiment above, and will not be repeated here.
[0446] It should be understood that the device described here is embodied in the form of functional modules. The term "module" here can refer to application-specific integrated circuits (ASICs), electronic circuits, processors (e.g., shared processors, proprietary processors, or group processors, etc.) and memories for executing one or more software or firmware programs, integrated logic circuits, and / or other suitable components supporting the described functions. In an alternative example, those skilled in the art will understand that the device may specifically be the first electronic device, the second electronic device, or the fourth electronic device in the above embodiments; or, the functions described in the above embodiments may be integrated into the device, and the device may be used to execute the various processes and / or steps corresponding to the electronic device or server in the above method embodiments. To avoid repetition, further details will not be provided here.
[0447] The aforementioned device has the function of implementing the corresponding steps of the electronic device or server in the aforementioned method; the aforementioned function can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the aforementioned function.
[0448] In embodiments of this application, the device may also be a chip or a chip system, such as a system on a chip (SoC).
[0449] This application also provides a schematic block diagram of another voice processing device. The device includes a processor 1601, a transceiver 1602, and a memory 1603. The processor 1601, transceiver 1602, and memory 1603 communicate with each other via an internal connection path. The memory 1603 stores instructions, and the processor 1601 executes the instructions stored in the memory 1603 to control the transceiver 1602 to transmit and / or receive signals.
[0450] It should be understood that the apparatus may specifically be the electronic device or server in the above embodiments, and may be used to execute the various steps and / or processes corresponding to the electronic device or server in the above method embodiments. Optionally, the memory 1603 may include a read-only memory 1603 and a random access memory 1603, and provide instructions and data to the processor 1601. A portion of the memory 1603 may also include a non-volatile random access memory 1603. For example, the memory 1603 may also store device type information. The processor 1601 may be used to execute the instructions stored in the memory 1603, and when the processor 1601 executes the instructions stored in the memory 1603, the processor 1601 is used to execute the various steps and / or processes of the above method embodiments. The transceiver 1602 may include a transmitter and a receiver, the transmitter may be used to implement the various steps and / or processes corresponding to the transceiver 1602 for performing a transmitting action, and the receiver may be used to implement the various steps and / or processes corresponding to the transceiver 1602 for performing a receiving action.
[0451] It should be understood that, in the embodiments of this application, the processor 1601 may be a central processing unit (CPU), or it may be other general-purpose processors 1601, digital signal processors 1601 (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor 1601 may be a microprocessor 1601, or it may be any conventional processor 1601, etc.
[0452] In implementation, each step of the above method can be completed by the integrated logic circuitry of the hardware in the processor 1601 or by instructions in software form. The steps of the method disclosed in the embodiments of this application can be directly manifested as execution by the hardware processor 1601, or as a combination of hardware and software modules in the processor 1601. The software modules can reside in mature storage media in the art, such as random access memory 1603, flash memory, read-only memory 1603, programmable read-only memory 1603, electrically erasable programmable memory 1603, registers, etc. This storage medium is located in memory 1603, and the processor 1601 executes the instructions in memory 1603, combining with its hardware to complete the steps of the above method. To avoid repetition, detailed descriptions are not provided here.
[0453] This application also provides a voice processing system, including a terminal device and a server. The steps performed by the terminal device and the server can be referred to the description of the method embodiments above, and will not be repeated here.
[0454] This application also provides a computer-readable storage medium for storing a computer program that implements the methods shown in the above-described method embodiments.
[0455] This application also provides a computer program product, which includes a computer program (also referred to as code or instructions). When the computer program is run on a computer, the computer can execute the methods shown in the above-described method embodiments.
[0456] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0457] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0458] In the embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0459] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of the embodiments of this application, depending on actual needs.
[0460] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0461] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
Claims
1. A speech processing method, characterized in that, include: Receive setting input, which is used to indicate the scenario for enabling voice repair, the contact for enabling voice repair, or the application for enabling voice repair. The scenario includes face-to-face communication scenario or remote communication scenario. Obtain the user's initial voice characteristics through voice registration; Receive the user's first voice input; The first speech is repaired based on the input settings and the first speech features.
2. The method according to claim 1, characterized in that, The remote communication scenario includes a call scenario, and in the call scenario, the method further includes: Send a first notification message to the first contact in the call, the first notification message being used to indicate that the voice repair function has been enabled; And / or, Upon receiving the first voice message, a second notification message is sent to the first contact, indicating that the first voice message is being repaired.
3. The method according to claim 2, characterized in that, The second notification message includes a notification tone. Sending the second notification message to the first contact upon receiving the first voice message includes: Upon receiving the first voice message, the notification tone is continuously sent to the first contact. The method further includes: Once the first restored voice message is sent to the first contact, the notification sound is stopped from being sent to the first contact.
4. The method according to any one of claims 1-3, characterized in that, The remote communication scenario includes a call scenario, and in the call scenario, the method further includes: Send the repaired first voice message to the first contact in the call; The first interface is displayed, which is an interface for making a call with the first contact. The first interface includes a first control, which is used to control the voice repair function to be turned off. In response to a first operation on the first control, the voice repair function is turned off to send a second voice message to the first contact, the second voice message including the user's voice message received after the voice repair function is turned off.
5. The method according to claim 4, characterized in that, In response to the first operation, the first control is also switched from a first state to a second state, wherein the first state indicates that the voice repair function is enabled and the second state indicates that the voice repair function is disabled. The method further includes: In response to a second operation on the first control, the first control is switched from the second state to the first state, and the voice repair function is enabled to send a repaired third voice message to the first contact, the third voice message including the voice message received from the user after the voice repair function is enabled.
6. The method according to any one of claims 1-5, characterized in that, The method further includes: The second interface is displayed, which includes playback controls for the repaired first audio and the first text corresponding to the repaired first audio. A third operation is received, the third operation being used to select a first target character in the first text; In response to the third operation, at least one candidate character associated with the first target character is displayed; A fourth operation is received, the fourth operation being used to select a second target character from the at least one candidate character; In response to the fourth operation, a playback control for the second text and the corresponding audio is displayed, wherein the second text is obtained by replacing the first target character in the first text with the second target character.
7. The method according to claim 6, characterized in that, The method further includes: Cancel the display of the playback control for the repaired first voice message.
8. The method according to any one of claims 1-7, characterized in that, In the aforementioned remote communication scenario, the method further includes: The third interface is displayed, which is an interface for communicating with the second contact. The third interface includes a second control and a third control. The second control is used to instruct the first voice message to be sent to the second contact, and the third control is used to instruct the repaired first voice message to be sent to the second contact. In response to an operation on the third control, the repaired first voice message is sent to the second contact; or... In response to an operation on the second control, the first voice message is sent to the second contact.
9. The method according to any one of claims 1-8, characterized in that, In the face-to-face communication scenario, the method further includes: A fourth interface is displayed, which includes a virtual keyboard; In response to an operation on the virtual keyboard, a third text is displayed; A fifth operation is received for the fourth control in the fourth interface, the fourth control being used to instruct the generation of speech; In response to the fifth operation, the speech corresponding to the third text is generated based on the first speech feature.
10. The method according to claim 9, characterized in that, The third text includes punctuation marks and / or emoticons. When generating the speech corresponding to the third text based on the first speech features, the tone of the speech corresponding to the third text is also controlled based on the punctuation marks and / or emoticons.
11. The method according to any one of claims 1-10, characterized in that, The process of obtaining a user's first voice feature through voice registration includes: The fifth interface is displayed. The fifth interface is the sound registration interface. The fifth interface includes a third prompt message and a fifth control. The third prompt message is used to prompt the recording content of the sound registration, and the fifth control is used to indicate the recording of voice. In response to an operation on the fifth control, a fourth voice message is recorded; Extract the first speech feature from the fourth speech.
12. A speech processing method, characterized in that, include: Receive setting input, which is used to indicate the scenario for enabling voice repair, the contact for enabling voice repair, or the application for enabling voice repair. The scenario includes face-to-face communication scenario or remote communication scenario. Receive a fifth voice message from the target contact; Obtain a second speech feature, which is a preset speech feature or a speech feature extracted from the fifth speech feature; The fifth speech is repaired based on the input settings and the second speech features.
13. The method according to claim 12, characterized in that, In the aforementioned remote communication scenario, the method further includes: The sixth interface is displayed, which is an interface for communicating with the target contact, and the sixth interface includes the fifth voice message; In response to an operation on the fifth speech, a sixth control and a seventh control are displayed, the sixth control indicating that the fifth speech be converted into text and the seventh control indicating that the repaired fifth speech be converted into text; In response to an operation on the seventh control, display the text corresponding to the repaired fifth voice; or, In response to an operation on the sixth control, the text corresponding to the fifth voice is displayed.
14. The method according to claim 12 or 13, characterized in that, In the face-to-face communication scenario, the method further includes: The seventh interface is displayed, which is an interface for communicating with the target contact person, and the seventh interface includes an eighth control; In response to the sixth operation on the eighth control, the fifth text is displayed; In response to the seventh operation on the eighth control, the display of the fifth text is stopped. The fifth text includes the text corresponding to the fifth voice within a target time period or the text corresponding to the repaired fifth voice. The target time period includes the time period from the response to the sixth operation to the response to the seventh operation.
15. A speech processing method, characterized in that, include: Obtain the target speech and speech features; The target speech and the speech features are input into the target restoration model. The target restoration model is used to extract the content of the target speech, obtain speech continuity features based on the content of the target speech and the speech features, and synthesize the restored target speech based on the speech continuity features. Obtain the repaired target speech output by the target repair model.
16. A voice processing device, characterized in that, include: A processor coupled to a memory storing computer-executable instructions, the processor executing the computer-executable instructions stored in the memory, such that the processor performs the method as claimed in any one of claims 1 to 11, or the method as claimed in any one of claims 12 to 14, or the method as claimed in claim 15.
17. A voice processing system, characterized in that, The device includes an electronic device and a server, wherein the electronic device performs the method as claimed in any one of claims 1 to 11, or performs the method as claimed in any one of claims 12 to 14, and the server performs the method as claimed in claim 15.
18. A computer-readable storage medium, characterized in that, For storing a computer program, the computer program including instructions for implementing the method as claimed in any one of claims 1 to 11, or including instructions for implementing the method as claimed in any one of claims 12 to 14, or including instructions for implementing the method as claimed in claim 15.
19. A computer program product, said computer program product comprising computer program code, characterized in that, When the computer program code is run on a computer, it causes the computer to perform the method as described in any one of claims 1 to 11, or the method as described in any one of claims 12 to 14, or the method as described in claim 15.
Citation Information
Patent Citations
Text-to-voice conversion processing method and device and electronic equipment
CN111312209A
End-to-end accent conversion method
CN111462769A
Speech synthesis method, related device, electronic equipment and storage medium
CN114664284A
Speech synthesis method and device, electronic equipment and storage medium
CN115547288A
Pathological voice repairing method based on voice conversion
CN116312469A
Cited By
Conference segment acquisition and summary export method and system for limited H5 environment
CN122339869A