Voice processing method, device and system, storage medium and program product

By acquiring speech features and repairing the speech of people with special speech impairments in specific scenarios, more intelligible speech output is generated, solving the problem of inconvenient communication for people with special speech impairments and improving the convenience and accuracy of communication.

CN121662019APending Publication Date: 2026-03-13HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-08-09
Publication Date
2026-03-13

AI Technical Summary

Technical Problem

People with speech impairments have defects in pronunciation, resulting in low intelligibility of their speech and difficulty in being understood by listeners. Existing technologies that convert text input into speech present inconveniences in communication.

Method used

By receiving settings input, the system obtains the user's voice characteristics and repairs the user's voice when a voice repair scenario or contact is detected, generating a more intelligible repaired voice. It supports voice-to-text conversion and repair, provides a control interface and prompts for the voice repair function, and allows users to flexibly choose between repair or original output.

Benefits of technology

It improves the convenience and experience of communication for users with speech impairments, reduces call conflicts, enhances the accuracy and flexibility of communication, maintains the consistency of user voice timbre, and supports voice restoration in various communication scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121662019A_ABST
    Figure CN121662019A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a voice processing method, device and system, a storage medium and a program product, and relates to the technical field of terminals. According to the voice processing method, in a face-to-face communication scene and a conversation scene, the speech barrier voice of the speech barrier user can be restored, the restored voice is obtained, the intelligibility of the restored voice is higher than that of the speech barrier voice, so that the speech barrier user can communicate by inputting the voice, and the user experience is improved. According to the invention, the technical problem of insufficient communication convenience of the speech-impaired users can be solved, so that the technical effect of improving the communication convenience of the speech-impaired users is realized.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application, with parent application number 202411097752.4 and filing date August 9, 2024, the entire contents of which are incorporated herein by reference. The parent application claims priority to Chinese Patent Application No. 202410808992.4, filed on June 20, 2024, entitled "A Method and Apparatus for Assisting Communication," filed with the China National Intellectual Property Administration, the entire contents of which are incorporated herein by reference. Technical Field

[0002] This application relates to the field of terminal technology, and in particular to a voice processing method, apparatus, system, storage medium, and program product. Background Technology

[0003] With the development of terminal technology, more and more functions are being applied to terminals, such as user communication functions.

[0004] Currently, some users, due to congenital or acquired factors, cannot communicate as freely as others. For example, people with hearing impairments, ALS, vocal cord damage, or other speech impairments have pronunciation defects, resulting in low intelligibility of their speech and difficulty in being understood by listeners, which significantly impacts their daily social interactions. Therefore, for people with speech impairments, in communication scenarios, they need to input text into electronic devices, which can then convert the user's input text into speech and output it.

[0005] However, for people with speech impairments, communicating by inputting text on electronic devices can be inconvenient. Summary of the Invention

[0006] This application provides a voice processing method, apparatus, system, storage medium, and program product, which helps to improve the convenience of communication between users.

[0007] In a first aspect, embodiments of this application provide a voice processing method, which may include:

[0008] The system receives settings input, which indicates the scenario, contact, or application for enabling voice repair. Scenarios include face-to-face or remote communication. Then, it obtains the user's initial voice characteristics through voice registration. The order of voice registration and settings input can be reversed. When the user speaks, the system receives their initial voice input. Since the settings input was received previously, the initial voice can be repaired based on the settings input and the initial voice characteristics.

[0009] In this embodiment, the first speech is repaired based on the setting input and the first speech feature. Since the setting input is used to indicate the scenario where speech repair is enabled, the contact person for enabling speech repair, or the application for enabling speech repair, when the electronic device is detected to be in a scenario where speech repair is enabled, or when the contact being communicated by the electronic device is the contact person for enabling speech repair, or when the application being run by the electronic device is the application for enabling speech repair, the first speech is repaired based on the first speech feature. This allows the repaired first speech to be played using the first speech feature, and the intelligibility of the repaired first speech is higher than that of the unrepaired first speech. Intelligibility can represent the accuracy of expressing what the user wants to say, or it can be understood as the listener's understanding of the speech signal transmitted by the speaker.

[0010] In this embodiment, by receiving setting input—which indicates the scenario, contact, or application for enabling voice restoration—and then obtaining the user's first voice feature through voice registration, the system can restore the first voice based on the setting input and the first voice feature after receiving the user's input. This allows users with speech or hearing impairments to communicate via voice input, improving their communication convenience. Furthermore, by generating the restored first voice using the user's registered first voice feature, the restored first voice can be played according to that feature, making it more closely resemble the user's own voice and further enhancing the user experience.

[0011] In one possible implementation, the remote communication scenario includes a call scenario, and in the call scenario, the method further includes:

[0012] Send a first notification message to the first contact in the call, indicating that voice repair has been enabled. And / or, upon receiving the first voice message, send a second notification message to the first contact, indicating that the first voice message is being repaired.

[0013] In this embodiment, by sending a first notification message to the first contact in the call to inform the user that the voice repair function has been enabled, the first contact can be informed that the user has enabled the voice repair function, thereby improving the call experience. Furthermore, by sending a second notification message to the first contact upon receiving the first voice message to indicate that the first voice message is being repaired, the first contact can be informed that the delay is due to voice repair, reducing the possibility of conflicting speaking times in two-party or multi-party calls, thereby improving the call experience.

[0014] In one possible implementation, if the second notification message includes a notification tone, then sending the second notification message to the first contact upon receiving the first voice message may include:

[0015] Upon receiving the first voice message, continue sending a notification tone to the first contact. Then, upon starting to send the repaired first voice message to the first contact, stop sending notification to the first contact.

[0016] In this embodiment of the application, by continuously sending a prompt tone to the first contact when the first voice is received, the first contact can know that the user on the other end is speaking when they hear the prompt tone. And when the prompt tone is stopped after the repaired first voice is sent to the first contact, the first contact can then hear the repaired first voice. This can reduce the situation of conflict between two or more parties during the call, thereby improving the experience of both parties in the call.

[0017] In one possible implementation, the remote communication scenario includes a call scenario, and in the call scenario, the method further includes:

[0018] A repaired first voice message is sent to the first contact in the call. Then, a first interface can be displayed, which is the interface for communicating with the first contact. The first interface includes a first control for disabling the voice repair function. Then, in response to a first operation on the first control, the voice repair function is disabled, and a second voice message is sent to the first contact. The second voice message includes the user's voice received after the voice repair function was disabled.

[0019] In this embodiment, during a call, the user can also control the voice repair function to be turned off. This allows the user to disable voice repair when it is not needed, improving the flexibility of switching between using and not using voice repair. For example, the user can control the voice repair function to be turned off via the first control when the electronic device's battery is low or the electronic device is running slowly, thereby improving the ability to disable voice repair during calls and enhancing the user experience.

[0020] In one possible implementation, in response to the first operation, the first control is also switched from a first state to a second state, where the first state indicates that the voice repair function is enabled and the second state indicates that the voice repair function is disabled. The method further includes:

[0021] In response to a second operation on the first control, the first control is switched from a second state to a first state, and the voice repair function is enabled, thereby sending a repaired third voice message to the first contact, the third voice message including the voice message received from the user after the voice repair function is enabled.

[0022] In this embodiment, the state of the first control indicates whether the voice repair function is currently enabled, thereby improving the accuracy of the user's choice to use or disable the voice repair function. Furthermore, this embodiment not only allows the voice repair function to be disabled during a call but also to be re-enabled, thus increasing the flexibility of enabling or disabling voice repair. In addition, enabling or disabling the voice repair function through the same control improves the simplicity of the interface during calls.

[0023] It should be understood that WeChat and face-to-face communication applications can also enable or disable the voice repair function within the application interface. Please refer to the relevant descriptions on how to enable or disable the voice repair function in the call interface, which will not be repeated here.

[0024] In one possible implementation, the method further includes:

[0025] A second interface is displayed, including playback controls for the repaired first audio and corresponding first text. Then, if the user selects a first target character from the first text, a third operation is received, selecting that first target character from the first text. In response to the third operation, at least one candidate character associated with the first target character is displayed. Then, if the user selects one of the candidate characters, a fourth operation is received, selecting a second target character from the at least one candidate character. In response to the fourth operation, second text and playback controls for the corresponding audio are displayed; the second text is obtained by replacing the first target character in the first text with the second target character.

[0026] In this embodiment, a second interface is displayed, including playback controls for the repaired first speech and first text corresponding to the repaired first speech. A third operation is received, whereby a first target character in the first text is selected. In response to the third operation, at least one candidate character related to the first target character is displayed. A fourth operation is received, whereby a second target character is selected from the at least one candidate character. In response to the fourth operation, second text and playback controls for the speech corresponding to the second text are displayed. This allows the user to quickly and manually correct inaccurate speech repair results, thereby improving the accuracy of speech communication.

[0027] In one possible implementation, the method further includes:

[0028] Cancel the display of the playback controls for the first audio clip after the repair.

[0029] In this embodiment of the application, by canceling the display of the playback control for the repaired first voice, only the latest voice can be played at any given time, which improves the convenience of selecting the appropriate voice for playback.

[0030] In one possible implementation, for remote communication scenarios, the method also includes:

[0031] A third interface is displayed, which is the interface for communicating with the second contact. The third interface includes a second control and a third control. The second control is used to instruct the sending of a first voice message to the second contact, and the third control is used to instruct the sending of a repaired first voice message to the second contact. Then, in response to an operation on the third control, the repaired first voice message can be sent to the second contact. Alternatively, in response to an operation on the second control, the first voice message can be sent to the second contact.

[0032] In this embodiment, the user can selectively send the repaired voice or the original voice to the second contact through the second and third controls. That is, even if the voice repair function is enabled, the user can still choose to send the original voice to the second contact, thereby improving the user's flexibility in remote communication.

[0033] In one possible implementation, for face-to-face communication scenarios, the method also includes:

[0034] A fourth interface is displayed, including a virtual keyboard. The user can then interact with the virtual keyboard, and in response to these interactions, third text is displayed. The user can then interact with a fourth control on the fourth interface, which instructs the generation of speech. In response to this fifth interaction, speech corresponding to the third text is generated based on first speech characteristics.

[0035] In this embodiment, a fourth interface is displayed, including a virtual keyboard. In response to an operation on the virtual keyboard, third text is displayed. A fifth operation is received on a fourth control in the fourth interface, the fourth control indicating the generation of speech. In response to the fifth operation, speech corresponding to the third text is generated based on a first speech feature. That is, the user can also input text and then generate speech based on the registered first speech feature, thus achieving text-to-speech conversion. Users can then communicate by choosing to input either speech or text, improving the selectivity and flexibility of user communication.

[0036] In one possible implementation, the third text includes punctuation marks and / or emojis. In this case, when generating the speech corresponding to the third text based on the first speech features, the tone of the speech corresponding to the third text can also be controlled based on the punctuation marks and / or emojis.

[0037] In this embodiment, the tone of the speech corresponding to the third text can be controlled by punctuation marks and / or emoticons in the third text input by the user. This allows the tone of the output speech to be adjusted adaptively according to the user input text, thereby improving the flexibility of voice broadcasting and enhancing the user experience.

[0038] In one possible implementation, the user's initial voice characteristics are obtained through voice registration, including:

[0039] The fifth interface is displayed, which is the sound registration interface. This fifth interface includes a third prompt message and a fifth control. The third prompt message indicates the recording content for sound registration, and the fifth control indicates the recording of audio. If the user interacts with the fifth control, a fourth audio recording can be initiated in response to the interaction with the fifth control. The first speech feature is then extracted from the fourth audio recording.

[0040] In this embodiment, the recording of voice is triggered by a fifth control, which can improve the accuracy and effectiveness of voice recording and make the subsequent comparison between the recorded content and the voice recognition result more accurate, thereby improving the accuracy of the repair model selection.

[0041] In one possible implementation, the method also includes:

[0042] Speech recognition is performed on the fourth speech to obtain the fourth text. Then, the fourth text is compared with the recorded content to obtain the similarity between the fourth text and the recorded content. Then, a target restoration model can be selected from multiple restoration models based on the similarity. The target restoration model is used to restore the first speech. The speech restoration capabilities of the multiple restoration models are different, and the speech restoration capability of the target restoration model is negatively correlated with the similarity.

[0043] In this embodiment of the application, the text obtained by speech recognition of the user's registered speech is compared with the recorded content to obtain a similarity score. Then, a target repair model is selected from multiple repair models based on the similarity score. This allows for the selection of a suitable repair model based on the user's speech impairment level, thereby balancing the accuracy of speech repair with the computing resources required for speech repair.

[0044] Secondly, embodiments of this application also provide another speech processing method. This method may include:

[0045] The system receives settings input, which indicates the scenario for enabling voice repair, the contact to enable voice repair, or the application to enable voice repair. Scenarios include face-to-face communication and remote communication. Then, it can receive a fifth voice message from the target contact. When voice repair is needed, a second voice feature can be obtained, which can be a preset voice feature or a voice feature extracted from the fifth voice message. Then, the fifth voice message is repaired based on the settings input and the second voice feature.

[0046] In this embodiment, the fifth voice is repaired based on the setting input and the second voice feature. Since the setting input is used to indicate the scenario where voice repair is enabled, the contact person who enables voice repair, or the application that enables voice repair, when the electronic device is detected to be in a scenario where voice repair is enabled, when the contact person being communicated with by the electronic device is the contact person who enables voice repair, or when the application being run by the electronic device is the application that enables voice repair, the fifth voice is repaired based on the second voice feature, so that the intelligibility of the repaired fifth voice is higher than that of the unrepaired fifth voice. Intelligibility can represent the degree of understanding when the voice is heard, or it can be understood as the accuracy with which the voice expresses what the contact person intends to convey.

[0047] In one possible implementation, for remote communication scenarios, the method also includes:

[0048] The sixth interface is displayed, which is the interface for communicating with the target contact. The sixth interface includes the fifth voice message. Then, if the user interacts with the fifth voice message, a sixth and seventh control can be displayed in response. The sixth control instructs the user to convert the fifth voice message to text, and the seventh control instructs the user to convert the repaired fifth voice message to text. The user can then interact with either the sixth or seventh control. Alternatively, an interaction with the seventh control can display the text corresponding to the repaired fifth voice message.

[0049] In this embodiment of the application, after receiving the voice from the target contact, the fifth voice can be selectively converted into text, which can also be understood as converting the original voice of the target contact into text; or the repaired fifth voice can be converted into text. Users can choose to display the text as needed, thereby improving the flexibility and experience of user communication.

[0050] In one possible implementation, for face-to-face communication scenarios, the method also includes:

[0051] The seventh interface is displayed, which is the interface for communicating with the target contact. The seventh interface includes an eighth control. The user can then interact with the eighth control, which, in response to a sixth action on the eighth control, displays the fifth text. The user can then interact with the eighth control again, which, in response to a seventh action on the eighth control, stops the display of the fifth text. The fifth text includes either the text corresponding to the fifth voice message within a target time period or the text corresponding to the repaired fifth voice message. The target time period includes the period between the responses to the sixth and seventh actions.

[0052] In this embodiment of the application, the voice of the target contact can also be converted into text, thereby improving the flexibility of communication between the user and the contact.

[0053] Thirdly, embodiments of this application also provide another speech processing method, which may include:

[0054] The target speech and its features are obtained. Then, the target speech and features are input into a target restoration model. This model extracts the content of the target speech, obtains continuous speech features based on the content and features, and synthesizes the restored target speech based on these features. Finally, the restored target speech output by the target restoration model can be obtained.

[0055] In this embodiment, after the target speech is input, the electronic device can acquire the target speech and speech features, and then input the target speech and speech features into the target restoration model. The target restoration model is used to extract the content of the target speech, obtain speech continuity features based on the content and speech features of the target speech, and synthesize the restored target speech based on the speech continuity features. In this way, people with speech impairments can communicate by inputting speech, thereby improving the convenience of communication for people with speech impairments.

[0056] In one possible implementation, the target restoration model includes a first module, a second module, and a third module. The first module is used to extract the content of the target speech, the second module is used to obtain continuous speech features based on the content and speech features of the target speech, and the third module is used to synthesize the restored target speech based on the continuous speech features.

[0057] In one possible implementation, the target speech and speech features are input into the target inpainting model, including:

[0058] The speech features and a first portion of the target speech are input into the target restoration model. The first module extracts the content of the first portion of speech, the second module obtains continuous features of the target speech based on the speech features and the content of the first portion, and the third module synthesizes the restored first portion of speech based on the continuous features of the target speech. Then, the speech features and a second portion of the target speech are input into the target restoration model. The first module extracts the content of the second portion of speech, the second module obtains continuous features of the second portion of speech based on the speech features and the content of the second portion, and the third module synthesizes the restored second portion of speech based on the continuous features of the second portion.

[0059] In this embodiment of the application, by first repairing a portion of the target speech and then repairing another portion of the target speech, speech repair can begin when only a portion of the speech is acquired. In other words, speech repair can begin without acquiring the complete target speech, thereby improving the efficiency of speech repair.

[0060] In one possible implementation, the target restoration model further includes a fourth module, which is used to discretize the first part of the speech to obtain discrete features of the target speech. The second module is also used to make predictions based on the discrete features of the target speech to obtain the restored discrete features of the target speech, and to obtain the second continuous features of the speech based on the restored discrete features of the target speech, the speech features, and the content of the second part of the speech.

[0061] In this embodiment, the first part of the speech is discretized by the fourth module to obtain the discrete features of the target speech. Then, the second module predicts based on the discrete features of the target speech to obtain the repaired discrete features of the target speech. Then, the second continuous features of the speech are obtained based on the repaired discrete features of the target speech, the speech features, and the content of the second part of the speech. In this way, the second continuous features of the speech can be obtained by combining the discrete features of the target speech, the speech features, and the content of the second part of the speech, thereby improving the accuracy of the obtained second continuous features of the speech and thus improving the accuracy of speech repair.

[0062] In one possible implementation, the fourth module is used to discretize the first part of the speech to obtain discrete features of the target speech, including:

[0063] The fourth module is used to discretize the first part of the speech based on its content to obtain the discrete features of the target speech.

[0064] In this embodiment of the application, the first part of the speech is discretized according to the content of the first part of the speech to obtain the discrete features of the target speech. That is, the content of the first part of the speech is used as a reference for the discretization process, thereby improving the accuracy of the obtained discrete features of the target speech and thus improving the accuracy of speech restoration.

[0065] It should be noted that the solution in this application embodiment can be used not only for sound restoration tasks, but also extended to tasks such as dialect to Mandarin conversion and cross-language translation.

[0066] Fourthly, another voice processing apparatus is provided, including a processor coupled to a memory for executing instructions in the memory to implement the methods in any of the possible implementations of the first aspect described above. Optionally, the apparatus further includes a memory. Optionally, the apparatus also includes a communication interface to which the processor is coupled.

[0067] Fifthly, a processor is provided, comprising: an input circuit, an output circuit, and a processing circuit. The processing circuit is used to receive signals through the input circuit and transmit signals through the output circuit, causing the processor to execute the method in any possible implementation of the first aspect described above.

[0068] In specific implementation, the processor can be a chip, the input circuit can be an input pin, the output circuit can be an output pin, and the processing circuit can be a transistor, gate circuit, flip-flop, and various logic circuits. The input signal received by the input circuit can be received and input by, for example, but not limited to, a receiver, and the signal output by the output circuit can be output to, for example, but not limited to, a transmitter and transmitted by the transmitter. Furthermore, the input circuit and the output circuit can be the same circuit, which is used as the input circuit and the output circuit at different times. This application does not limit the specific implementation of the processor and various circuits.

[0069] In a sixth aspect, a processing apparatus is provided, including a processor and a memory. The processor is configured to read instructions stored in the memory and to receive signals via a receiver and transmit signals via a transmitter to execute the method in any of the possible implementations of the first aspect described above.

[0070] Optionally, there may be one or more processors and one or more memories.

[0071] Alternatively, the memory can be integrated with the processor, or the memory can be set separately from the processor.

[0072] In the specific implementation process, the memory can be a non-transitory memory, such as read-only memory (ROM), which can be integrated with the processor on the same chip or set on different chips. The embodiments of this application do not limit the type of memory or the way the memory and processor are set.

[0073] It should be understood that the relevant data interaction process, such as sending indication information, can be the process of outputting indication information from the processor, and receiving capability information can be the process of the processor receiving input capability information. Specifically, the processed output data can be output to the transmitter, and the input data received by the processor can come from the receiver. Here, the transmitter and receiver can be collectively referred to as a transceiver.

[0074] The processing device in the fourth aspect above can be a chip. The processor can be implemented in hardware or software. When implemented in hardware, the processor can be a logic circuit, integrated circuit, etc. When implemented in software, the processor can be a general-purpose processor that reads software code stored in memory. The memory can be integrated into the processor or located outside the processor and exist independently.

[0075] In a seventh aspect, a computer program product is provided, comprising: a computer program (also referred to as code or instructions) that, when executed, causes a computer to perform the method in any of the possible implementations of the first aspect described above.

[0076] Eighthly, a computer-readable storage medium is provided that stores a computer program (also referred to as code or instructions) that, when run on a computer, causes the computer to perform the methods in any of the possible implementations of the first aspect above. Attached Figure Description

[0077] Figure 1 A schematic diagram of a system architecture provided for an embodiment of this application;

[0078] Figure 2 A schematic diagram of a sound registration interface provided for an embodiment of this application;

[0079] Figure 3 A schematic diagram of an interface for configuring an application using the sound restoration function, provided as an embodiment of this application;

[0080] Figure 4 A schematic diagram of an interface for configuring a whitelist for sound restoration, provided as an embodiment of this application;

[0081] Figure 5A schematic diagram of an interface for a remote real-time communication scenario provided in an embodiment of this application;

[0082] Figure 6 A schematic diagram of a call scenario provided in an embodiment of this application;

[0083] Figure 7 A schematic diagram of an interface for a remote, non-real-time communication scenario provided in an embodiment of this application;

[0084] Figure 8 A schematic diagram of an interface for a face-to-face communication scenario provided in an embodiment of this application;

[0085] Figure 9 A schematic diagram of an interface for another face-to-face communication scenario provided in an embodiment of this application;

[0086] Figure 10 A schematic flowchart of a speech processing method provided in an embodiment of this application;

[0087] Figure 11 A schematic flowchart illustrating a sound registration method provided in an embodiment of this application;

[0088] Figure 12 A schematic diagram of the architecture of a repair model provided in an embodiment of this application;

[0089] Figure 13 A flowchart illustrating another speech processing method provided in an embodiment of this application;

[0090] Figure 14 A flowchart illustrating another speech processing method provided in an embodiment of this application;

[0091] Figure 15 A flowchart illustrating another speech processing method provided in an embodiment of this application;

[0092] Figure 16 This is a schematic diagram of the structure of a voice processing device provided in an embodiment of this application. Detailed Implementation

[0093] To facilitate a clear description of the technical solutions in the embodiments of this application, some terms and technologies involved in the embodiments of this application will be briefly introduced below:

[0094] 1. Electronic equipment:

[0095] The electronic devices in the embodiments of this application may also be referred to as: terminal equipment, user equipment (UE), mobile station (MS), mobile terminal (MT), access terminal, user unit, user station, mobile station, mobile station, remote station, remote terminal, mobile device, user terminal, terminal, wireless communication equipment, user agent, or user device, etc.

[0096] Electronic devices can be devices that provide users with voice / data connectivity, such as handheld devices with wireless connectivity, in-vehicle devices, etc. Currently, examples of terminals include: mobile phones, tablets, laptops, PDAs, mobile internet devices (MIDs), wearable devices, virtual reality (VR) devices, augmented reality (AR) devices, wireless terminals in industrial control, wireless terminals in self-driving vehicles, wireless terminals in remote medical surgery, wireless terminals in smart grids, wireless terminals in transportation safety, wireless terminals in smart cities, wireless terminals in smart homes, cellular phones, cordless phones, session initiation protocol (SIP) phones, wireless local loop (WLL) stations, personal digital assistants (PDAs), handheld devices with wireless communication capabilities, computing devices or other processing devices connected to wireless modems, in-vehicle devices, wearable devices, electronic devices in 5G networks, or future evolution of public land mobile communication networks. The embodiments of this application do not limit the scope of electronic devices in a network (PLMN).

[0097] By way of example and not limitation, in this embodiment, the electronic device can also be a wearable device. Wearable devices, also known as wearable smart devices, are a general term for devices that utilize wearable technology to intelligently design and develop everyday wearables, such as glasses, gloves, watches, clothing, and shoes. Wearable devices are portable devices that are worn directly on the body or integrated into the user's clothing or accessories. Wearable devices are not merely hardware devices, but also achieve powerful functions through software support, data interaction, and cloud interaction. Broadly speaking, wearable smart devices include those that are feature-rich, large in size, and can achieve complete or partial functions without relying on a smartphone, such as smartwatches or smart glasses, as well as those that focus on a specific type of application function and require the use of other devices such as smartphones, such as various smart bracelets and smart jewelry for vital sign monitoring.

[0098] Furthermore, in this application embodiment, the electronic device can also be an electronic device in an Internet of Things (IoT) system. IoT is an important component of future information technology development, and its main technical feature is connecting objects to networks through communication technology, thereby realizing an intelligent network of human-machine interconnection and object-to-object interconnection. The electronic device in this application can also be an on-board unit, on-board module, on-board component, on-board chip, or on-board unit built into a vehicle as one or more components or units. The vehicle can implement the methods of this application through the built-in on-board unit, on-board module, on-board component, on-board chip, or on-board unit. Therefore, the embodiments of this application can be applied to vehicle networking, such as vehicle-to-everything (V2X), long-term evolution-vehicle (LTE-V) communication technology, and vehicle-to-vehicle (V2V). Optionally, the electronic device in the embodiments of this application can also be simply referred to as a device.

[0099] 2. Access network equipment:

[0100] The access network device in this application embodiment can also be called a wireless access network device. It can be a transmission reception point (TRP), an evolved NodeB (eNB or eNodeB) in an LTE system, a home base station (e.g., home evolved NodeB or home Node B, HNB), a base band unit (BBU), or a wireless controller in a cloud radio access network (CRAN) scenario. Alternatively, the access network device can be a relay station, access point, vehicle-mounted device, wearable device, or access network device in a 5G network or an access network device in a future evolved PLMN network. It can be an access point (AP) in a WLAN, a gNB in ​​a new radio (NR) system, or a satellite base station in a satellite communication system. This application embodiment is not limited to these categories.

[0101] The access network equipment in this embodiment may include centralized unit (CU) nodes, distributed unit (DU) nodes, or access network equipment including CU nodes and DU nodes, or access network equipment including control plane CU nodes (CU-CP nodes), user plane CU nodes (CU-UP nodes), and DU nodes. Access network equipment including CU nodes and DU nodes can separate the protocol layers of the access network equipment, with some protocol layer functions centrally controlled by the CU, and the remaining part or all protocol layer functions distributed in the DU, which is centrally controlled by the CU. As one implementation, the CU deployment protocol stack includes a radio resource control (RRC) layer, a packet data convergence protocol (PDCP) layer, and a service data adaptation protocol (SDAP) layer. The DU deployment protocol stack includes a radio link control (RLC) layer, a media access control (MAC) layer, and a physical layer (PHY) layer. Thus, the CU has RRC, PDCP, and SDAP processing capabilities. The DU has RLC, MAC, and PHY processing capabilities. The above functional division is merely an example and does not constitute a limitation on the CU and DU. That is, there can be other ways to divide functions between the CU and DU, which will not be elaborated upon in this embodiment. The functions of the CU can be implemented by a single entity or by different entities. For example, the functions of the CU can be further divided, such as separating the control plane (CP) and user plane (UP), i.e., the CU control plane (CU-CP) and the CU user plane (CU-UP). For example, CU-CP and CU-UP can be implemented by different functional entities, and CU-CP and CU-UP can be coupled with the DU to jointly complete the functions of the access network device. In one possible approach, CU-CP is responsible for control plane functions, mainly including RRC and PDCP-C, where PDCP-C is mainly responsible for encryption / decryption, integrity protection, and data transmission of control plane data. CU-UP is responsible for user plane functions, mainly including SDAP and PDCP-U, where SDAP is mainly responsible for processing data from the core network device and mapping data flows to bearers. PDCP-U is primarily responsible for data plane encryption / decryption, integrity protection, header compression, sequence number maintenance, and data transmission. CU-CP and CU-UP are connected via an E1 interface. CU-CP represents the access network device connecting to the core network device through the interface between the core network device and the access network device.The CU-UP is connected to the DU via F1-C (control plane). The CU-UP is connected to the DU via F1-U (user plane). Alternatively, the PDCP-C may also be located in the CU-UP, but this application does not limit this implementation. The access network equipment can be, for example, a base station.

[0102] 3. Artificial Intelligence (AI):

[0103] AI is a new technical science that studies and develops theories, methods, technologies, and application systems for simulating, extending, and expanding human intelligence. AI encompasses fields including, but not limited to, at least one of automatic speech recognition (ASR), text-to-speech (TTS), image recognition, or natural language processing. ASR is a technology that converts speech into text, processing and analyzing speech signals to identify the linguistic content and converting it into readable text. TTS is a technology that converts text information into speech output, transforming text information into speech feature vectors, then converting the speech features into audio signals. It can also provide personalized pronunciation services for businesses and individuals through timbre selection, customizable volume, and speech rate. Speech can also be referred to as sound, audio, etc.

[0104] 4. Other terms

[0105] In the embodiments of this application, terms such as "first" and "second" are used to distinguish identical or similar items with essentially the same function and effect. For example, "first speech" and "second speech" are used only to distinguish different speech and do not limit their order. Those skilled in the art will understand that terms such as "first" and "second" do not limit the quantity or execution order, and that terms such as "first" and "second" do not necessarily imply that they are different.

[0106] It should be noted that, in the embodiments of this application, the terms "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design scheme described as "exemplary" or "for example" in this application should not be construed as being more preferred or advantageous than other embodiments or design schemes. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.

[0107] In this application embodiment, "at least one" refers to one or more, and "more than one" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can represent: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple.

[0108] Currently, some users, due to congenital or acquired factors, cannot communicate as freely as others. For example, people with hearing impairments, ALS, or vocal cord damage have speech defects, resulting in low intelligibility and difficulty in being understood, which significantly impacts their daily social interactions. To address these communication challenges, systems that assist with daily communication are needed for these individuals. These people with specific speech impairments can also be referred to as speech-impaired users.

[0109] For individuals with specific speech impairments, a language comprehension system is needed to address daily communication challenges. In related technologies, text can be input, and an electronic device can convert the text into speech for output, enabling these individuals to communicate through sound.

[0110] However, people with speech impairments need to input text to communicate via sound, which is cumbersome, resulting in insufficient convenience for them to communicate via sound. For people with speech impairments, their situation is that their speech has low intelligibility and is difficult for listeners to understand, not that they cannot produce sound.

[0111] In view of this, embodiments of this application propose a speech processing method, apparatus, system, storage medium, and program product that can assist speech-impaired users in communication by repairing the speech input by them, namely, dysarthric speech reconstruction (DSR). DSR can convert the low-intelligibility speech of a special speech-impaired group into higher-intelligibility normal speech. Speech can also be referred to as sound, audio, or sound source, etc. The speech processing method of this application embodiment can convert difficult-to-understand speech into more intelligible speech while maintaining the timbre consistent with the speaker. This voice restoration function has high practical value in both short-range and long-range communication scenarios.

[0112] In one possible implementation, voice restoration can be achieved by an electronic device performing voice restoration processing (referred to as voice restoration or repair) on the voice input by a speech-impaired user to obtain restored voice. The restored voice is more intelligible than the input voice, so the electronic device can output the restored voice, which enables the party communicating with the speech-impaired user to better understand the speech-impaired user's expression.

[0113] In another possible implementation, after obtaining the repaired speech, the repaired speech can be processed by speech recognition to obtain the recognized text. The recognized text converted from the repaired speech can be understood as the repaired recognized text. Then, the electronic device can output the repaired recognized text, which also enables the party communicating with the speech-impaired user to better understand the speech-impaired user's expression.

[0114] In another possible implementation, speech restoration can be achieved by having an electronic device perform speech recognition processing on the input speech to obtain recognized text, and then perform restoration processing on the recognized text to obtain restored recognized text. The restored recognized text is more intelligible than the original recognized text, which also enables the party communicating with the speech-impaired user to better understand the expression of the speech-impaired user.

[0115] Therefore, it can be seen that voice restoration can enable users with speech impairments to communicate by inputting voice, thereby improving the convenience and experience of communication for them.

[0116] Please see Figure 1 , Figure 1 This is a schematic diagram of a system architecture provided for an embodiment of this application.

[0117] like Figure 1 As shown in (a) of this application embodiment, the system architecture may include a first device 101, a second device 102, an access network device 103, and a cloud 104. The first device 101 and the second device 102 can communicate via the access network device 103, and at least one of the first device 101 and the second device 102 is a device used by a user with a speech impairment. The cloud 104 has a sound restoration function, capable of restoring the input speech to obtain the output speech. The cloud 104 may be, for example, a server.

[0118] In one possible implementation, the first device 101 can be a device used by a user with a speech impairment. After receiving the voice input from the user with a speech impairment, the first device 101 sends the input voice to the cloud 104. The cloud 104 can then repair the input voice and send the repaired voice back to the first device 101. The first device 101 then sends the repaired voice to the second device 102 through the access network device 103. After receiving the repaired voice, the second device 102 can then play the repaired voice.

[0119] In another possible implementation, the second device 102 can be a device used by a user with a speech impairment. After receiving the voice input from the user with a speech impairment, the second device 102 sends the input voice to the first device 101. After receiving the input voice, the first device 101 sends the input voice to the cloud 104. The cloud 104 can then repair the input voice and send the repaired voice to the first device 101. After receiving the repaired voice, the first device 101 can play the repaired voice.

[0120] It should be understood that the above-mentioned voice repair processing can also be performed on at least one of the first device 101 or the second device 102, so the cloud 104 may not be necessary.

[0121] It should be noted that the cloud-based 104 can be configured with multiple speech restoration models. These models, also known as restoration models, are used to process the input speech to obtain restored speech. Optionally, the speech restoration capabilities of the multiple restoration models may differ.

[0122] like Figure 1 As shown in (b) of this application, the system architecture of this embodiment may include a first device 101 and a cloud 104.

[0123] In one possible implementation, after receiving the voice input from a user with a speech impairment, the first device 101 sends the input voice to the cloud 104. The cloud 104 can then repair the input voice and send the repaired voice back to the first device 101. The first device 101 can then output the repaired voice or convert the repaired voice into text and output it.

[0124] It should be understood that the above-mentioned voice repair processing can also be performed on the first device 101, so the cloud 104 is not needed.

[0125] It should be noted that the solutions in this application embodiment can be used in scenarios including but not limited to the following:

[0126] Use Case 1: Remote Communication Scenarios. Remote communication scenarios can include, but are not limited to, real-time or non-real-time remote communication scenarios. Real-time remote communication scenarios can include, but are not limited to, real-time call scenarios. Non-real-time remote communication scenarios can include, but are not limited to, non-real-time remote voice communication scenarios.

[0127] Use Case 2: Close-range communication scenarios. Close-range communication scenarios can include, but are not limited to, face-to-face communication scenarios.

[0128] In other words, in this embodiment, sound restoration capabilities can be integrated into the interaction in both call and face-to-face communication scenarios. In general, this embodiment can be applied to users with speech impairments, enabling real-time restoration of input speech and playback of the restored audio to the listener or transmission to the other end of the call. Application scenarios include remote communication scenarios such as phone calls and video calls, as well as short-range communication scenarios such as face-to-face interactions.

[0129] The following examples illustrate remote communication scenarios and face-to-face communication scenarios respectively.

[0130] First, an example illustration will be provided for a remote communication scenario. This embodiment uses a remote real-time communication scenario as an example.

[0131] As understood, the terms "interface" and "user interface" used in this application refer to the medium through which an application or operating system interacts and exchanges information with the user. It facilitates the conversion between the internal form of information and a form acceptable to the user. A common form of user interface is the graphical user interface (GUI), which refers to a user interface related to computer operation displayed graphically. It can be an icon, window, control, or other interface element displayed on the screen of an electronic device. Controls can include visual interface elements such as icons, buttons, menus, tabs, text boxes, dialog boxes, status bars, navigation bars, and widgets.

[0132] Please see Figure 2 , Figure 2 This is a schematic diagram of a sound registration interface provided in an embodiment of this application.

[0133] like Figure 2The interface shown in (a) displays a page with application icons, which may include at least one of the following: settings app icon 201, weather app icon, calendar app icon, photo app icon, email app icon, or app store app icon. Below these application icons, a page indicator may also be displayed to show the positional relationship between the currently displayed page and other pages. Below the page indicator are multiple application icons (e.g., camera app icon, contacts app icon, messaging app icon, dialer app icon), which remain displayed when switching pages.

[0134] It is understood that the application icons in this application embodiment are icons for launching an application or icons for navigating to a certain interface. For example, the camera application icon is the icon of the camera application (i.e., the camera app), meaning that the camera application icon can be used to trigger the launch of the camera application. As another example, the settings application icon 201 can be used to trigger navigation to the settings interface.

[0135] In this embodiment, the electronic device can detect user actions performed on the setting application icon 201, and in response to the user actions, the electronic device can display, as shown below. Figure 2 The user interface shown in (b) is as follows. Figure 2 The interface shown in (b) can be, for example, the settings application interface, on which the user can perform at least one of the following functions: querying information about the electronic device or configuring the user.

[0136] It is understood that the user operations mentioned in this application may include, but are not limited to, touch (e.g., click), voice control, gestures, etc., and this application does not limit them.

[0137] like Figure 2 The interface shown in (b) displays a page with settings options. This page may include one or more settings options. These settings options may include at least one of the following: accessibility settings option 202, battery and performance settings, security and privacy settings, display and brightness settings, or network connectivity management. Accessibility settings option 202 may be designed for users with special needs, helping them use electronic devices more conveniently through a series of assistive functions. The functions that accessibility settings option 202 in this embodiment can set may include, but are not limited to, at least one of the following: sound optimization, screen reading, or color correction. Sound optimization can also be referred to as sound restoration.

[0138] It should be noted that, in Figure 2In the interface shown in (b), for configuration options, such as setting option 2, an indication of whether the option is on or off may also be included. Optionally, if the indication for a setting option is "off," the configuration corresponding to that setting option is not enabled; if the indication for a setting option is "on," the configuration corresponding to that setting option is enabled. It should be understood that the indications for enabling or disabling a setting option can be different, and are not limited to specific text styles such as "off" or "on." By displaying the indication of whether the option is on or off on the interface, users can quickly know whether the setting option is enabled, improving the user experience. For setting options used to query information, the setting option may not include an indication of whether the option is on or off. Furthermore, as... Figure 2 The interface shown in (b) may also include an indication that the option has a next-level interface, such as the ">" icon. By displaying an indication of a next-level interface on the interface, the user can be informed that a next-level interface exists, thus improving the user experience.

[0139] In this embodiment, the electronic device can detect user actions on accessibility setting option 202, and in response to the action, the electronic device can display, as shown below. Figure 2 The user interface shown in (c) is as follows. Figure 2 The interface shown in (c) can be, for example, the interface of accessibility settings option 202, on which users can perform the setting function of accessibility settings option 202.

[0140] like Figure 2 The interface shown in (c) displays a page that includes accessibility features. This page may include one or more accessibility controls, such as a sound restoration control 203. The sound restoration control 203 can be used to turn the sound restoration function on or off; in other words, the sound restoration control 203 can be understood as the master switch for the sound restoration function. Optionally, the interface may also include an indicator to show whether the sound restoration function is on or off, allowing users to determine whether the sound restoration function is enabled. Figure 2 As shown in (c), the sound restoration function is off, meaning it is not enabled. Optional, such as... Figure 2 The interface shown in (c) may also include an indication that the sound repair function has a next-level interface. Optionally, such as... Figure 2 The interface shown in (c) may also include a control for returning to the previous level interface, such as a return function. Figure 2 The controls of the interface shown in (b) are shown in the image.

[0141] In this embodiment, the electronic device can detect user operations on the control 203 that acts on the sound restoration function, and in response to the operation, the electronic device can display as follows: Figure 2 The user interface shown in (d) is as follows. Figure 2 The interface shown in (d) could be, for example, the interface for the sound restoration function, which can be understood as a detailed settings page for sound restoration. Users can configure the sound restoration function on this interface, such as configuring at least one of the following: sound recording for sound restoration or the application or scene supported by the sound restoration function.

[0142] like Figure 2 The interface shown in (d) includes a sound recording control 204 and a sound repair application configuration control 205. The sound repair application configuration control 205 may include, but is not limited to, a call application configuration control 2051, a face-to-face communication application configuration control 2052, and a WeChat application configuration control 2053.

[0143] The sound recording control 204 can be used to record the user's voice. Optionally, such as... Figure 2 The interface shown in (d) may also include a sound recording status indicator. This indicator shows whether sound has been recorded. Optionally, if the sound recording status indicator is "Pending Recording," it means the user has not yet recorded sound; if the indicator is "Recorded," it means the user has already recorded sound. Optionally, even if the user has already recorded sound, additional sound can be added, thus increasing the number of recorded sounds. It should be understood that the indicator indicating recorded sound can be different from the indicator indicating no recorded sound, and is not limited to the example above. Figure 2 The sound recording status indicator (d) indicates that no sound has been recorded. This embodiment of the application uses the sound recording status indicator to indicate to the user whether sound has been recorded, allowing the user to quickly determine whether sound has been recorded, thereby improving the user experience. In another possible implementation, the sound recording status indicator may not need to be displayed, thus reducing the amount of content displayed on the interface and improving the efficiency of the interface display.

[0144] The sound repair application configuration control 205 is used to configure whether the sound repair function is enabled or disabled. This control can also be understood as a switch for applications that support the sound repair function. When the sound repair application configuration control 205 is in the enabled state, the sound repair function in the corresponding application or scene is activated. For example... Figure 2The interface shown in (d) allows configuration of the sound restoration function for applications or scenarios that may include, but are not limited to, at least one of the following: call applications, face-to-face communication applications, or WeChat applications. WeChat can be understood as a non-real-time remote communication application. Figure 2 (d) in the text may include the sound repair application configuration control 205 for each application, such as the sound repair application configuration control 205 for a call application, the sound repair application configuration control 205 for a face-to-face communication application, and the sound repair application configuration control 205 for a WeChat application. Optionally, when the sound repair application configuration control 205 for an application or scenario is in a disabled state, the sound repair function cannot be used in that application or scenario; when the sound repair application configuration control 205 for an application or scenario is in a enabled state, the sound repair function can be used in that application or scenario. This embodiment of the application configures the sound repair application configuration control 205 for each application or scenario, meaning that the sound repair function of each application or scenario can be individually enabled or disabled. This allows users to select which application or scenario to enable or disable the sound repair function as needed, thereby improving the user experience. In another possible implementation, the sound repair application configuration control 205 can also be used to enable or disable the sound repair function for multiple applications or scenarios. This allows for one-click enabling or disabling of the sound repair function for multiple applications or scenarios, thereby improving the efficiency of enabling or disabling the sound repair function for multiple applications or scenarios. In another possible implementation, the application or scene that supports the sound repair function configuration and the corresponding sound repair application configuration control 205 may not be displayed. In other words, the sound repair function may be enabled by default for all applications or scenes. This can reduce the content displayed on the interface and improve the efficiency of the interface display.

[0145] For example, such as Figure 2 The sound repair application configuration control 205 shown in (d) is in the off state, which means that the sound repair function cannot be used by any application.

[0146] It should be noted that sound recording and supported application or scene configuration can be mutually constrained. This constraint can be that one of the sound recording and supported application or scene configurations can only be performed after the other operation is completed. For example, sound recording can be completed before configuring the application or scene supported by the sound repair function. In other words, if sound recording is not completed, no application or scene can enable the sound repair function through the sound repair application configuration control 205. This embodiment of the application allows sound recording and supported application or scene configuration to be mutually constrained, prompting the user to perform sound recording and configure the application or scene for enabling the sound repair function, thereby improving the user experience.

[0147] In this embodiment, the electronic device can detect user operations applied to the sound recording control 204, and in response to the operation, the electronic device can display as follows: Figure 2 The interface shown in (e) is as follows. In this embodiment, since the sound optimization function requires the user to register their voice, after enabling sound repair, the user can be automatically redirected to the voice registration interface, for example... Figure 2 The interface shown in (e) could be a sound registration interface. For example... Figure 2 The interface shown in (e) may include a new voiceprint control 206. The new voiceprint control 206 can be used to trigger the acquisition of voiceprint features (referred to as voiceprints). Voiceprints can be used to indicate the timbre of a registered user's voice. Optionally, Figure 2 The interface shown in (e) may also include a control for returning to the previous level interface, such as a return function. Figure 2 The controls of the interface shown in (d) are shown in the diagram.

[0148] Then, the electronic device can detect user actions performed on the newly created voiceprint control 206, and in response to the action, the electronic device can display, as shown below. Figure 2 The interface shown in (f) is as follows. Figure 2 The interface shown in (f) may include sound recording content 207, recording control 208, and save control. Sound recording content 207 can guide the user to the content they want to record, such as text like "Giant pandas don't have fixed sleeping times; they sleep wherever they go." Recording control 208 can be used to trigger sound acquisition. Optionally, sound acquisition can be triggered by long-pressing the recording control 208, with the electronic device capturing the sound recorded by the user during the long press; or the user can click the recording control 208 to trigger sound acquisition, and the electronic device will start capturing sound, stopping the acquisition when the user clicks the recording control 208 again. The save control is used to include the acquired sound. Optionally, the user can choose to record sound in segments. When the electronic device detects a user action on the save control, it splices the segmented sound recordings in chronological order to obtain a single recording.

[0149] In this embodiment, the electronic device detects a user operation on the recording control 208 and, in response to the operation, begins to collect the user's voice.

[0150] Then, as Figure 2 As shown in (g), the electronic device detects a user operation on the save control, and in response to the operation, saves the collected sound and extracts the voiceprint of the collected sound. It should be understood that the voiceprint extraction process can be performed at any time after the sound is collected, and is not limited to extraction immediately after the sound is saved.

[0151] Optionally, the electronic device can compare the recorded content with the audio recording content 207. If the similarity between the recorded content and the audio recording content 207 is higher than a similarity threshold, the voiceprint of the recorded content is used as the voiceprint for subsequent voice broadcast. If the similarity between the recorded content and the audio recording content 207 is not higher than the similarity threshold, the user is prompted to re-record the audio according to the instructions in the audio recording content 207. Optionally, the electronic device in this embodiment supports recording 1 to n sentences of text, that is, when recording... Figure 2 Following the text content shown in (g), another piece of text can be displayed for the user to record, where n is an integer not less than 2. After recording, the user can save the text using the save control. It should be understood that the user can also choose to record m sentences of text and then save, where 1 ≤ m < n, and m is an integer.

[0152] Then, as Figure 2 As shown in (h), the electronic device detects a user action performed on the control that returns to the previous screen, and in response to this action, can display something like... Figure 2 The interface shown in (i) is shown in the diagram. Figure 2 In the interface shown in (i), the audio recording status indicator is marked as "Recorded". Figure 2 (i) in Figure 2 The similar part of (d) can be found by referring to Figure 2 The description of (d) in the text will not be repeated here.

[0153] In another possible implementation Figure 2 The interface shown in (f) may also exclude the audio recording content 207, thus reducing the content displayed on the interface and improving its efficiency. However, by displaying the audio recording content 207, this embodiment guides the user to record audio according to the content, which helps the electronic device determine the validity of the audio recording and improves the accuracy of voiceprint extraction. Furthermore, it can also be used to determine the validity of the audio recording based on the content of the collected audio. Figure 2 The audio recording content 207 of the interface shown in (f) is compared with the input audio, and then a suitable repair model is selected based on the similarity to repair the subsequent user input audio. The similarity can also be called intelligibility. The higher the similarity, the more accurate the intelligibility of the collected audio content. Figure 2 The more similar the audio recording content 207 of the interface shown in (f) is, the less impairment the user's speaking ability is. In other words, the smaller the similarity of the comparison, the more similar the content of the collected audio is to the audio recording content. Figure 2The less similar the audio recording content 207 of the interface shown in (f) is, the greater the impairment of the user's speaking ability. Generally speaking, the more complex the architecture of the restoration model, the stronger its speech restoration capability and the more accurate the speech restoration. However, the electronic device also requires more computing resources. Therefore, when selecting a suitable restoration model based on the similarity of the comparison, the smaller the similarity, the more complex the architecture of the chosen restoration model, and vice versa. For example, when the similarity is the first similarity, the first restoration model is selected; when the similarity is the second similarity, the second restoration model is selected. Wherein, the first similarity is higher than the second similarity, the model architecture complexity of the first restoration model is lower than that of the second restoration model, and the corresponding capability of the first restoration model is lower than that of the second restoration model.

[0154] In another possible implementation Figure 2 The interface shown in (f) may also exclude the recording control 208, in which case it will enter as shown in the image. Figure 2 The method described in (f) starts collecting sound when the interface is activated, thus reducing the amount of content displayed on the interface and improving display efficiency. However, this embodiment of the application, by displaying the recording control 208 and only starting sound collection after detecting an action on the recording control 208, can reduce the resource utilization of the electronic device and increase its lifespan.

[0155] In another possible implementation Figure 2 The interface shown in (f) can also omit the save control, in which case the recorded sound is saved directly after recording is complete. For example, if no recording is detected within a certain time, the recording is considered complete, or the length of the recorded content is... Figure 2 If the length of the content shown in (f) is consistent, the recording is considered complete. In this embodiment of the application, by setting a save control, users can flexibly record sound.

[0156] It should be understood that the interface of this application embodiment can be added or removed as needed, and is not limited to this. Figure 2 The number and style of the interfaces shown are not limited here.

[0157] In this embodiment of the application, by having users register their voices, the user's voiceprint can be extracted from the collected voice. In this way, during the broadcast, the repaired voice can be broadcast using the user's voiceprint. That is, the user can select a specific voiceprint to broadcast the repaired voice as needed, thereby improving the user experience.

[0158] In another possible implementation Figure 2The audio recording process can also be eliminated, allowing electronic devices to use built-in voiceprints to broadcast the restored speech. This reduction in audio recording also improves the ease of communication.

[0159] Below, we will provide an example of how to configure an application that uses the sound restoration feature.

[0160] Please see Figure 3 , Figure 3 This is a schematic diagram of an interface for configuring and using the sound restoration function in an embodiment of this application.

[0161] like Figure 3 The interface shown in (a) includes a sound recording control 204 and a sound repair application configuration control 205. Figure 3 The interface shown in (a) can be referenced. Figure 2 The description of the interface shown in (i) is omitted here.

[0162] In this embodiment, the electronic device can detect user operations applied to the sound repair application configuration control 205, such as detecting user operations applied to the call application configuration control 2051. In response to this operation, it can display... Figure 3 The interface shown in (b) is shown in the image. Figure 3 In the interface shown in (b), the call application configuration control 2051 in the sound repair application configuration control 205 switches from the off sound repair state to the on sound repair state. For example, Figure 3 In the interface shown in (b), if the call application configuration control 2051 is in the "sound repair enabled" state, the call application can use the sound repair function. However, if both the face-to-face communication application configuration control 2052 and the WeChat application configuration control 2053 are in the "sound repair disabled" state, then neither the face-to-face communication application nor the WeChat application can use the sound repair function. However, when using the face-to-face communication application or the WeChat application, the voice repair function can be enabled or disabled via controls within the interface of the face-to-face communication application or the WeChat application. Optionally, in this embodiment, the face-to-face communication application, the WeChat application, and the call application can control the voice repair function to be enabled or disabled via controls in their respective application interfaces. This embodiment will not explain how to control the voice repair function via controls in the application interface; further explanation will be provided in later embodiments.

[0163] It should be noted that if the electronic device can detect user operations on the sound repair application configuration control 205, the sound repair application configuration control 205 can also switch from the sound repair enabled state to the sound repair disabled state.

[0164] In this embodiment of the application, by configuring a sound repair control for each application, the enabling or disabling of the sound repair function of each application can be controlled independently. Users can choose to enable or disable the sound repair function of the application as needed, thereby improving the user experience.

[0165] exist Figure 3 The illustrated interface shows that the sound restoration function for the call application has been enabled. Therefore, the following example illustrates the call scenario.

[0166] Figure 4 This is a schematic diagram of an interface for configuring a whitelist for voice restoration, provided as an embodiment of this application. In this embodiment, for people with special speech impairments, there may be friends who are already accustomed to understanding their pronunciation; for these friends, when the voice restoration switch is turned on, such as Figure 4 As shown, users with speech impairments can set a whitelist in their phone's contacts so that sound restoration is not enabled by default when calling these friends.

[0167] like Figure 4 The interface shown in (a) displays a page with application icons, which may include multiple application icons (e.g., the Contacts application 401). Figure 4 The interface shown in (a) can be referred to Figure 2 The description of the interface shown in (a) is omitted here.

[0168] In this embodiment, the electronic device can detect user operations applied to the address book application 401, and in response to the operation, the electronic device can display, as shown below. Figure 4 The interface shown in (b) is as follows. Figure 4 The interface shown in (b) can be, for example, the main interface of the address book. Figure 4 The interface shown in (b) may include at least one of the following: a contact configuration control 403, a contact search bar, or a contact list. The contact configuration control 403 is used to configure the same operation on the selected contact, such as adding the selected contact to the sound repair whitelist, or enabling call logging by default for the selected contact during a call. The contact search bar can be used to quickly search for contacts. The contact list includes one or more contacts. Optionally, the contacts in the contact list can be categorized according to certain rules, such as categorizing by the first letter of the contact's name, and so on. Figure 4 The interface shown in (b) also includes controls for quickly locating categories. Optional, such as Figure 4 The interface shown in (b) may also include dial controls and favorites controls. Optional, such as Figure 4 The interface shown in (b) may also include a contact addition control.

[0169] Then, as Figure 4 As shown in (c), the electronic device can detect user actions applied to the interface, and in response to these actions, the electronic device can display a contact selection control 402. The contact selection control 402 can be used to select contacts to add them to a whitelist. In this embodiment, each contact corresponds to a contact selection control 402. Figure 4 The contact selection control 402 shown in (c) is in an unselected state.

[0170] In this embodiment, the electronic device can detect user actions applied to the contact selection control 402, and in response to the action, the electronic device can display as follows: Figure 4 The interface shown in (d) is shown in the image. Figure 4 In the interface shown in (d), the contact selection control 402 is in a selected state, for example, the contact selection control 402 for the contact "B2" is in a selected state.

[0171] Then, the electronic device can detect the user action performed on the contact configuration control 403, and in response to the action, the electronic device can display, as shown below. Figure 4 The interface shown in (e) of this application. The contact configuration control 403 in this embodiment can be used to configure selected contacts. Figure 4 The interface shown in (e) may also include a voice repair whitelist option 404 and a call recording option. The call recording option is used to set the contact selected by the contact selection control 402 as the contact to save call recordings. The voice repair whitelist option 404 is used to add the contact selected by the contact selection control 402 to the voice repair whitelist, that is, to add the selected contact to the voice repair whitelist.

[0172] Then, as Figure 4 As shown in (f), if the electronic device can detect a user operation on the voice repair whitelist option 404, then the electronic device can add the contacts whose contact selection control 402 is selected to the voice repair whitelist. Optionally, in this embodiment, the voice repair whitelist is a list where voice repair is not enabled, that is, the contacts in the voice repair whitelist do not enable the voice repair function when making a call; or, the voice repair whitelist is a list where voice repair is enabled, that is, the contacts in the voice repair whitelist enable the voice repair function when making a call.

[0173] This application embodiment allows users to configure a whitelist for voice repair, enabling them to select contacts for use or disuse of the voice repair function during calls, based on actual circumstances. This flexibility allows users to adjust which contacts require or do not require voice repair during calls, thereby improving the user experience.

[0174] It should be noted that you can also choose not to configure the voice repair whitelist. In this case, you can either perform voice repair on all contacts or not.

[0175] After configuring the voice repair whitelist, you can initiate a call with one of the contacts. For ease of understanding, the following example illustrates how contacts on the voice repair whitelist are not subject to voice repair.

[0176] This application embodiment improves the flexibility of using the voice repair function by setting a voice repair whitelist, allowing different contacts to use or not use the voice repair function.

[0177] Please see Figure 5 , Figure 5 This is a schematic diagram of an interface for a remote real-time communication scenario provided in an embodiment of this application. In this embodiment, contacts in the voice repair whitelist are those for whom voice repair is disabled by default.

[0178] like Figure 5 The interface shown in (a) displays a page with application icons, which may include multiple application icons (e.g., phone application 501). Figure 5 The interface shown in (a) can be referred to Figure 2 The description of the interface shown in (a) is omitted here.

[0179] In this embodiment of the application, the electronic device can detect user operations applied to the telephone application 501, and in response to the operation, can display as follows: Figure 5 The interface shown in (b) is as follows. Figure 5 The interface shown in (b) displays the user's recent call history, which may include the caller and the corresponding call time. Figure 5 The interface shown in (b) can include multiple contacts. Then, if the electronic device can detect a user action performed on one of the contacts, for example, a user action performed on contact "B2", a call can be initiated in response to that action. Figure 4 As shown in (d), since contact "B2" has been added to the voice repair whitelist, meaning contact "B2" is configured not to use the voice repair function, then as follows... Figure 5As shown in (c), when initiating a call with contact "B2", the call is conducted in its original audio format by default, meaning the call is conducted with the audio from before the repair. Optionally, in Figure 5 The interface shown in (c) also includes a sound switching control 502, the avatar of the call contact, a sound mute control, a keyboard access control, a sound speaker control, and an end call control. Figure 5 The audio switching control 502 shown in (c) is in original audio playback mode, meaning the other end of the call is playing the original audio captured by this end. The audio switching control 502 can be used to control the switching between the original audio and the restored audio. For example... Figure 5 As shown in (d), the electronic device can detect user operations applied to the voice switching control 502. In response to this operation, the voice switching control 502 of the electronic device is in the state of playing the repaired voice, that is, the other end of the call plays the repaired voice, and the voice heard by the caller "B2" is the repaired voice. In this embodiment of the application, the repaired voice can be obtained by repairing the original voice. The original voice can also be called the original pronunciation, which can be the voice collected from the user.

[0180] Optionally, if the electronic device detects another user operation on the sound switching control 502, the state of the sound switching control 502 switches to the original sound playback state, and the corresponding user conducts the call using the original sound. In this embodiment, by setting the sound switching control 502 in the call interface, the user can choose to use the original sound or the restored sound through the sound switching control 502, which can improve the user experience.

[0181] It should be noted that if an electronic device initiates a call with a contact not on the voice repair whitelist, the voice repair function will be enabled by default in the call interface.

[0182] It should be noted that if the call is initiated by the electronic device on the other end, the local electronic device can also display the message. Figure 5 (c) or Figure 5 As shown in (d) of the interface, the local end can use the sound switching control 502 to control whether to play the original sound captured by the other end or the sound after the original sound captured by the other end has been repaired.

[0183] In another possible implementation, if the caller is a contact outside the voice repair whitelist, that is, the caller is configured to enable the voice repair function, then the voice switching control 502 is in the repaired voice playback state when the call is connected.

[0184] In another possible implementation, the voice switching control 502 may not be necessary, meaning that the original voice or the repaired voice can be used for the call depending on the configuration of the voice repair whitelist.

[0185] In some cases, when electronic devices repair the sound, it takes a certain amount of time. Therefore, there will be a delay from when the local user speaks to when the remote user hears the local user's speech. This delay is related to the time required for the electronic device to repair the sound. As a result, it will take a long time for the remote user to receive the repaired voice from the local user. The remote user may then think that the local user did not speak or that the signal is poor, which is why the remote user cannot hear it.

[0186] It should be noted that it can also be used when making a call with a contact in the voice repair whitelist, repairing the voice transmitted by the contact in the voice repair whitelist and then playing it. In other words, the local device repairs the voice transmitted from the other end and then plays it, for example, repairing the voice of the contact "B2" and then playing it.

[0187] Please see Figure 6 , Figure 6 This is a schematic diagram of a call scenario provided by an embodiment of this application. In this embodiment, there are local users and remote users, respectively. The electronic device used by the local user is called the local device, and the electronic device used by the remote user is called the remote device. The local device can also be called the first electronic device, and the remote device can also be called the second electronic device. The local user can also be called the first user, and the remote user can also be called the second user. In this embodiment, the call audio will increase the call latency after being processed by the sound restoration model, which may affect the call quality of multiple parties. Therefore, this embodiment describes improving the call quality of multiple parties in a call scenario.

[0188] like Figure 6As shown in (a), if the user inputs "Hello," the local device needs to perform repair processing before sending the repaired "Hello" to the remote device. The remote device then receives and plays the repaired "Hello," allowing the user to hear it. However, there is a significant time delay between the user inputting the voice and the remote device hearing it, due to the processing time required for the local device to repair the voice. Alternatively, the remote device can send the user input "Hello, your takeout, I'm downstairs" to the local device, which will then play the input. When the user on this end wants to reply, they can type "Thank you, just leave it downstairs." The local device needs to process this voice message, but this process takes time, which is called processing delay. If the user on the other end doesn't hear the response from the local user for a long time, they might assume that the local user didn't hear their voice message and repeat "Hello, your takeout, I'm downstairs." However, the local user has already heard it, but because the processing takes time, the user on the other end may not hear the local user's feedback for an extended period.

[0189] Therefore, in one possible implementation, during a call, the local device can send a prompt to the remote device, which then plays the prompt, thereby informing the remote user that the local device needs to process the voice input from the user. This allows the remote user to know that the local device needs to repair the voice input.

[0190] like Figure 6 As shown in (b), after establishing a communication channel with the peer device, the local device can send a prompt to the peer device. Upon receiving the prompt, the peer device can broadcast it, allowing the peer user to hear it. This prompt can indicate that the local device needs to repair the voice input by the local user. For example, the prompt might include phrases such as "The other party has enabled voice optimization; there is a delay in the call. Please wait patiently for the prompt tone to finish before replying," to inform the other party that a call delay may occur.

[0191] It should be understood that the timing of the local device sending a prompt to the remote device can be either immediately after establishing a communication channel with the remote device or after a certain interval; there is no restriction on this.

[0192] In this embodiment of the application, sending a prompt message from the local device to the remote device can improve the experience for both parties in the call.

[0193] like Figure 6As shown in (b), although the local device can send prompts to the remote device, there are still situations during the communication process where the remote user thinks that the local user did not hear their voice.

[0194] In this application embodiment, in call scenarios, such as mobile phone calls and video calls, the voice changing and audio repair functions are enabled by switching on the call interface. After being enabled, the receiving party will receive a corresponding prompt message to inform the other party that the sound they hear has been optimized and that there will be a delay, so as to improve the call experience between the two parties.

[0195] Therefore, in one possible implementation, after receiving each segment of voice input from the local user, the local device sends a prompt tone to the remote device, which can then play the prompt tone to indicate that the local user has input voice.

[0196] like Figure 6 As shown in (c), for example, after establishing a communication channel with the peer device, the local device can send a prompt to the peer device. Then, after detecting the user's voice input of "Hello," the local device sends a prompt tone to the peer device, which is then played back by the peer device. Then, after detecting the user's voice input of "Thank you, just leave it downstairs," the local device sends a prompt tone to the peer device, which is then played back by the peer device. The peer user can then know that the local user has input a voice message through the prompt tone, and after receiving the "Thank you, just leave it downstairs" voice message, can reply with "Okay, goodbye." For example, the prompt tone can be, for instance, a "ding," a "beep," or a "tick-tock." In a call scenario, the AI ​​model's voice conversion can cause call delays. After the local device speaks, a "beep" prompt tone is played back to indicate waiting, thus reducing overlap in the conversation. Optionally, during the call, before the sound repair result is given, the other party can be guided to wait patiently and reply through a call prompt tone.

[0197] In this embodiment of the application, the timing of the prompt tone sent by the local terminal may be as follows: after receiving the user's voice input, or after receiving a segment of voice input. No limitation is imposed here.

[0198] It should be noted that the prompts in this embodiment can also be used to indicate the function of notification sounds during a call. For example, the prompts may include phrases such as "You will hear a notification sound during the call, indicating that the user is currently inputting voice; please wait patiently."

[0199] In this embodiment of the application, by sending a prompt tone to the peer device after receiving each segment of voice input by the local device, the experience of both parties in the call when using the voice repair function can be improved, and the problem of communication inefficiency caused by latency can be reduced.

[0200] In one possible implementation, it is also possible to send only one of the prompts or sounds; this is not a limitation.

[0201] The above examples illustrate remote real-time communication scenarios. The following examples illustrate remote non-real-time communication scenarios.

[0202] Please see Figure 7 , Figure 7 This is a schematic diagram of an interface for a remote, non-real-time communication scenario provided in an embodiment of this application.

[0203] like Figure 7 The interface shown in (a) displays a page with application icons, which may include multiple application icons (e.g., chat application 701). In this embodiment, chat application 701 may be a non-real-time remote chat application 701. Figure 7 The interface shown in (a) can be referred to Figure 2 The description of the interface shown in (a) is omitted here.

[0204] In this embodiment, the electronic device can detect user actions performed on the chat application 701, and in response to the action, the electronic device can display, as shown below. Figure 7 The interface shown in (b) is optional. Figure 7 The interface shown in (b) may include a search bar, chat functionality controls, contact list functionality controls, send functionality controls, and a personal center functionality controls. The search bar can be used to quickly search for information, such as quickly searching for contacts. Figure 7 In the interface shown in (b), the chat function control is selected. Figure 7 The interface shown in (b) can be understood as the interface corresponding to the chat function. This interface includes one or more contacts and their corresponding contact information. Optionally, the contact information may include the contact name and chat content, which may be, for example, the latest chat message. For example, one of the contacts' names may be "Xiaoming", and their latest chat message may be, for example, "What did you do today?"

[0205] In this embodiment of the application, the electronic device can detect a user action performed on one of the contacts, and in response to the action, the electronic device can display as follows: Figure 7The interface shown in (c) can be a chat interface with the selected contact. This chat interface can include each chat message and the user corresponding to that message. Optionally, the chat message can include, but is not limited to, at least one of text, voice, video, or files. Figure 7 The interface shown in (c) may include a chat content input control, which may include a voice input control 702. Optionally, the chat content input control may also include a text input control, a file / video transfer control, and a video recording control, etc. The voice input control 702 is used to trigger voice recording. For details on how the voice input control 702 triggers voice recording, please refer to... Figure 2 The explanation of how the recording control in the embodiment triggers sound recording will not be repeated here.

[0206] In this embodiment, the electronic device can detect user operations applied to the voice input control 702, and in response to these operations, can record the user's voice input. Figure 7 As shown in (d), the electronic device can also display a recording in progress notification during voice recording. Optionally, this recording in progress notification may include text messages such as "Recording in progress," and may also include text content converted from the recorded voice. It should be understood that by displaying the text content converted from the recorded voice during the recording process, users can know whether the electronic device has captured their voice and whether the text content converted by the electronic device matches the recorded voice, thereby improving the user experience.

[0207] Then, as Figure 7 As shown in (e), the electronic device can display a recording completion message after the recording of voice is finished. This completion message may include text prompts such as "Recording Complete," and may also include the text content converted from the recorded voice. Figure 7 The interface shown in (e) includes a direct send option 703 and a repair and send option 704. The direct send option 703 triggers the transmission of the captured user's original audio; that is, if the electronic device detects a user action on the direct send option 703, it sends the unrepaired audio in the chat interface. The repair and send option 704 triggers the transmission of the repaired audio 705; that is, if the electronic device detects a user action on the repair and send option 704, it sends the repaired audio 705 in the chat interface.

[0208] In this embodiment of the application, the electronic device can detect the user operation acting on the repair and sending option 704, and in response to this operation, can send the repaired voice 705 in the chat interface. Figure 7In the interface shown in (f), the repaired voice message 705 is sent. The electronic device can detect the user operation applied to the repaired voice message 705 and play the repaired voice message 705.

[0209] Optionally, if the user chooses to send the original audio, the electronic device can also play the original audio after detecting a user action applied to the original audio.

[0210] This application embodiment allows users to send the repaired voice 705 to their contacts, thereby improving the experience of non-real-time remote communication between users and their contacts.

[0211] Then, as Figure 7 In the interface shown in (g), voice 706 was received from the contact.

[0212] In the embodiments of this application, such as Figure 7 As shown in (h), the electronic device can detect user operation 706 applied to a contact's voice, and in response to this operation, can display a voice-to-text conversion option 707 and a voice-to-text conversion option 708. The voice-to-text conversion option 707 is triggered to display the text converted from the selected voice after repair. The voice-to-text conversion option 708 is triggered to display the text converted from the selected voice.

[0213] It should be noted that, optional, you can click on the voice to play the voice message, or long-press the voice message to display the voice-to-text option 707 and the voice-to-text option 708.

[0214] Then, as Figure 7 As shown in (i), the electronic device can detect the user operation on the voice-to-text option 707, and in response to the operation, can send the selected voice to be converted into text after repair in the chat interface.

[0215] It should be noted that the electronic device may detect a user operation that applies the voice repair to the text conversion option 707, repair the voice, and then convert it into text for display; or, the electronic device may receive the voice, repair the voice to obtain the repaired voice, and then, upon detecting a user operation that applies the voice repair to the text conversion option 707, display the selected voice converted into text.

[0216] In this embodiment of the application, if the received voice message is the audio of a user with a special speech impairment, the user on this end can perform the user operation 706 on the contact's voice to display the voice-to-text option 707 and the voice-to-text option 708. Then, by performing the user operation on the voice-to-text option 707, more accurate text can be obtained.

[0217] In another possible implementation, after the electronic device detects the user operation 706 applied to the contact's voice, it can also display a voice repair option. The voice repair option is used to trigger the playback of the repaired voice 705. Then, after the electronic device detects the user operation applied to the voice repair option, it can play the repaired voice 705, which is to say, play the repaired voice 706 from the contact.

[0218] This application embodiment can repair the voice 706 of a contact after receiving the voice 706 from the contact, so that even if the other party is a speech impaired user, normal communication can be achieved, thereby improving the experience of non-real-time remote communication between the user and the contact.

[0219] In summary, in remote communication scenarios, the other party can hear the restored, more intelligible, and identical voice in real time, making the call smoother. Furthermore, in remote communication scenarios, the other party will receive a notification that "the local device has enabled the voice restoration function." Simultaneously, the local device can switch between the original and restored audio sources at any time via buttons, allowing the other user to access the original audio information, which may include at least one of the original timbre, ambient sound, or original pronunciation.

[0220] In another possible implementation, Figure 7 In section (h), a playback voice option and a playback repaired voice option can also be displayed. If a user operation is detected acting on the playback voice option, the voice from the contact is played; if an operation is detected acting on the playback repaired voice option, the voice from the contact is repaired before playback. This embodiment of the application, by displaying the playback voice option and the playback repaired voice option, can repair the voice from the contact, thereby improving the communication experience for both parties.

[0221] The above describes remote communication scenarios. Below, we will provide examples of near-field communication scenarios. Unlike real-time calls, face-to-face rehabilitation is used by individuals with speech impairments when communicating with others offline.

[0222] Please see Figure 8 , Figure 8 This is a schematic diagram of an interface for a face-to-face communication scenario provided in an embodiment of this application.

[0223] like Figure 8 The interface shown in (a) displays a page with application icons, which may include multiple application icons (e.g., Face-to-Face Repair Application 801). The Face-to-Face Repair Application 801 in this embodiment can be used for voice repair in face-to-face communication scenarios. Figure 8 The interface shown in (a) can be referred to Figure 2 The description of the interface shown in (a) is omitted here.

[0224] In this embodiment, the electronic device can detect user actions applied to the face-to-face repair application 801, and in response to the action, can display as follows: Figure 8 The interface shown in (b) is as follows. Figure 8 The interface shown in (b) may include a recording control 802, which is used to trigger voice acquisition. For example, the recording control may be in the shape of a recording ball. In this embodiment, the electronic device can detect user operation on the recording control 802 and, in response to the operation, start recording voice. Optionally, the user operation in this embodiment may be, for example, a long press operation. During the long press of the recording control 802, a swipe-up cancel prompt and a release repair prompt may be displayed. The swipe-up cancel prompt indicates that voice recording can be canceled by swiping up. The release repair prompt indicates that recording can be completed by releasing the recording control 802.

[0225] like Figure 8 As shown in (c), after the electronic device detects the release operation on the recording control 802, it displays a playback control 803 for a repaired audio recording. The playback control 803 for the repaired audio recording is the audio obtained by repairing the recorded audio. Furthermore, Figure 8 The interface shown in (c) may also include a text display area 804, which is used to display the text content corresponding to the playback control 803 of the repaired audio. For example, wherein... Figure 8 The text displayed on the interface shown in (c) could be, for example, "Are you going this afternoon?". It should be noted that if the electronic device detects an action... Figure 8 User operation of the playback control 803 for the repaired voice in the interface shown in (c) can play the repaired voice, such as playing the voice corresponding to "Are you going this afternoon?". Optionally, after detecting a release operation on the recording control 802, the animation effect of the recording control 802 scrolls for a short period of time, then displays the repaired text and automatically plays the repaired audio. If the user is not satisfied with the recording during the recording process, they can cancel the repair by swiping up while holding down the button.

[0226] However, since voice restoration may not be entirely accurate, this application embodiment also supports rapid correction of the restoration results to improve the effectiveness of face-to-face communication. For example, if the restored text contains misidentified words, the user can click on the word they want to modify, and a candidate word list 806 will appear, from which the user can select the desired word for quick replacement. Users can also directly edit the text using an input method. After completing the correction, the user can choose to play the corrected audio.

[0227] like Figure 8 As shown in (d), the electronic device can detect user actions on a portion of the text content 805, and in response to the action, can display as shown in (d). Figure 8 The interface shown in (e) is as follows. Figure 8 The interface shown in (e) includes a candidate word list 806 corresponding to the word 805 selected by the user. The candidate word list 806 includes one or more candidate words. For example, if the word 805 is "afternoon", the corresponding candidate words may include "morning", "good at dancing", and "foggy".

[0228] Then, as Figure 8 As shown in (f), the electronic device can detect a user action applied to one of the candidate words, and in response to that action, can display as shown in (f). Figure 8 The interface shown in (g) is shown in the image. Figure 8 In the interface shown in (g), the word 805 in the displayed text content will be replaced with the candidate word that the user action was applied to. For example, if the electronic device detects a user action applied to the candidate word "morning", then the word "afternoon" in the displayed text content "Today afternoon or not" will be replaced with "morning", and the corresponding displayed text content will be replaced with "Today morning or not".

[0229] Then, as Figure 8 As shown in (h), the electronic device can detect the action acting on Figure 8 User interaction with the repaired voice playback control 803 in the interface shown in (h) will trigger the playback of the voice prompt "Are you going this morning or not?". This allows the other party communicating with the user to understand the user's voice expression.

[0230] In other words, in this embodiment of the application, the speech can be broadcast according to the text content displayed in the text display area 804.

[0231] In this embodiment, by recording and repairing the user's voice during face-to-face communication, the playback control 803 displays the repaired voice, improving the user's communication experience. Furthermore, after recording the user's input voice, the corresponding text content can be displayed. This allows the user to determine the accuracy of the voice repair result through the text content. If the user deems the repair result inaccurate, they can adjust the text content, thereby adjusting the content of the played voice. In other words, face-to-face communication scenarios can quickly correct semantic errors, making user operation more convenient and faster, and improving the accuracy of voice communication. Moreover, in this embodiment, voice input during face-to-face communication allows for the playback of the repaired audio, eliminating the need for plain text input, making it faster. Furthermore, for words with errors after repair, the system supports quick replacement of the identified words in the text and replays the audio, improving the efficiency of face-to-face communication.

[0232] It should be noted that, in Figure 8 In the interface shown in (c)-(e), one possible implementation is as follows: When performing ASR recognition processing on the speech, multiple ASR recognition results may be obtained, each corresponding to a probability value. The electronic device can then display the ASR recognition result with the highest probability value among the multiple ASR recognition results. For example, the ASR recognition results obtained from the ASR recognition processing may include: "Today, shall we go to Shanwu?", "Today, shall we go to Xiawu?", "Today, shall we go to Shanwu?", and "Today, shall we go to Shanwu?". Since the probability value of the recognition result "Today, shall we go to Shanwu?" is the highest, this ASR recognition result is displayed. The words in the other ASR recognition results besides the displayed ASR recognition result are also saved. Then, the displayed ASR recognition results can be segmented to obtain one or more words. When a user action is detected acting on one of the words (i.e., part of word 805), words different from that part of word 805 are searched from the remaining ASR recognition results and displayed as candidate words, such as "morning," "good at dancing," and "foggy." This embodiment improves the efficiency of obtaining candidate words by saving the remaining ASR recognition results and then searching for words different from the user-selected word (i.e., part of word 805) from the remaining ASR recognition results.

[0233] In another possible implementation, the remaining ASR recognition results may not be saved. In this way, candidate words can be predicted directly from the words selected by the user. Since the remaining ASR recognition results do not need to be stored, the storage resources required to store the ASR recognition results can be reduced.

[0234] It should be noted that, in the embodiments of this application, Figure 8 The interface shown can have a playback control 803 that can display at most one repaired audio recording, such as the playback control 803 for the latest repaired audio recording. This improves the simplicity of the interface display. In another possible implementation, a new audio recording can be displayed for each recorded audio recording, and a new audio recording can also be displayed for each change to the text content. This allows for recording every audio recording from the user and tracking the user's editing process, thereby improving the user experience.

[0235] In another possible implementation, such as Figure 8 In the interface shown in (c), the voice or text display area 804 can also be used, which can improve the simplicity of the interface display. In this embodiment of the application, by displaying both voice and text in the display area 804, users can quickly determine whether the voice repair result of the electronic device is accurate through the text content displayed in the text display area 804, thereby improving the user experience.

[0236] In another possible implementation, the user can select punctuation marks from the text content. In this case, candidate punctuation marks can be displayed, and the user-selected punctuation marks can be used to replace the selected punctuation marks in the text content. The tone of the speech can then be changed based on the updated punctuation. For example, replacing “?” with “!” changes the tone of the speech from a question to an exclamation.

[0237] It should be noted that users with speech impairments often also have some degree of hearing impairment; that is, users with speech impairments may also be hearing impaired. Therefore, the following examples illustrate how to improve the communication experience of users with hearing or speech impairments.

[0238] Please see Figure 9 , Figure 9 This is a schematic diagram of an interface for another face-to-face communication scenario provided in an embodiment of this application.

[0239] like Figure 9 The interface shown in (a) could be, for example, the interface of a face-to-face repair application. In, for example... Figure 9 The interface shown in (a) includes a recording control 802, which can be used to record speech. Optionally, the recording control 802 can be used to record the speech of hearing-impaired / speech-impaired users, or it can be used to record the speech of normal users. Normal users can be users with normal hearing ability or users with normal speaking ability. It should be noted that, although Figure 9 and Figure 8 The recording control 802 in the game has a different style, but its function is the same; both are used to trigger voice recording.

[0240] In this embodiment, the electronic device can detect a click operation on the recording control 802, and in response to the operation, enter a continuous recording mode, during which the other party's voice is converted into text in real time. For example... Figure 9 In the interface shown in (a), during recording, a notification message can be displayed to indicate that recording is in progress, such as "Listening, we can help you transcribe it into text." Then, as... Figure 9 In the interface shown in (a), the electronic device can display text converted from the speech of a normal user, such as "I am a hearing person, I am speaking." Optionally, the content of the conversation can be displayed in a dialog box (chat box). The electronic device can cancel the continuous recording mode by detecting a click operation on the recording control 802 again. Optionally, the electronic device can also enter the continuous recording mode when it detects a long press operation on the recording control 802, and then cancel the continuous recording mode after detecting an operation to leave the recording control 802, that is, it is in the continuous recording mode while the recording control 802 is long-pressed. Optionally, in this embodiment, the speech of a normal user can be repaired and then converted into text for display, thereby improving the face-to-face communication experience.

[0241] Then, as Figure 9 As shown in (b), the electronic device can detect user actions applied to the interface, such as detecting user actions applied to a blank area of ​​the interface, and in response to the action, display as shown in Figure (b). Figure 9 The interface shown in (c) is as follows. Figure 9 The interface shown in (a) could be, for example, the interface for the voice recording function. The interface for the voice recording function can be found in [reference needed]. Figure 8 The description of the embodiments will not be repeated here. Figure 9 The interface shown in (c) could be, for example, a text input interface. Figure 9 The interface shown in (a) switches to Figure 9 The interface shown in (c) can be understood as switching from voice recording to text input. Optionally, this user action could be, for example, a swipe up. Figure 9 The interface shown in (c) includes a virtual keyboard 808 for inputting text, a speech generation control 809 for generating speech, and a text display box 810 for displaying text. Optionally, the recording control 802 can also be used as the speech generation control, that is, after inputting text through the virtual keyboard 808, the speech is displayed in the dialog box through user operations on the recording control 802.

[0242] In another possible implementation, such as Figure 9The interface shown in (b) may also include a keyboard trigger control, which is used to trigger the display of the virtual keyboard 808. After detecting a user operation on the keyboard trigger control, the electronic device can display, as shown in [example code]. Figure 9 The interface shown in (c) is shown in the image.

[0243] In the embodiments of this application, such as Figure 9 As shown in (c), the electronic device can detect user actions on the virtual keyboard 808 and, in response to the actions, display the text entered by the user in the text display box 810.

[0244] Then, as Figure 9 In the interface shown in (d), the electronic device can detect user actions on the speech generation control 809 and, in response to these actions, generate speech based on the text displayed in the text display box 810. For example... Figure 9 The interface shown in (d) can include the voice of the text displayed in the text display box 810. For example, if the text displayed in the text display box 810 is "Hello, I am Zhang San", then the electronic device will generate the voice of "Hello, I am Zhang San" after detecting a user operation on the voice generation control 809. In addition, when generating the voice of "Hello, I am Zhang San", the tone of voice can also be generated based on the punctuation marks or emoticons in the text.

[0245] For example, if the text includes a period (.), the message "Hello, I am Zhang San" can be announced in a neutral tone; if the text includes an exclamation mark (!), the message can be announced in an exclamatory tone; and if the text includes a question mark (?), the message can be announced in a questioning tone. Furthermore, if the text includes a smiley face, the message can be announced in a cheerful tone; and if the text includes a crying face, the message can be announced in a sad tone.

[0246] Then, as Figure 9 In the interface shown in (e), the electronic device can detect user actions performed on the voice and, in response to the action, can play the voice message "Hello, I am Zhang San".

[0247] In this embodiment of the application, the user can continue to input text. Then, as... Figure 9 In the interface shown in (f), the electronic device can detect user operations on the virtual keyboard 808 and, in response to the operation, display the text entered by the user in the text display box 810.

[0248] Then, as Figure 9In the interface shown in (g), the electronic device can detect user operations on the speech generation control 809 and, in response to the operation, generate speech according to the text displayed in the text display box 810, such as displaying... Figure 9 The interface shown in (h) is as follows. Figure 9 The interface shown in (h) can include the voice of the text displayed in the text display box 810. For example, if the text displayed in the text display box 810 is "I want to ask what time it closes today", then the electronic device will generate the voice of "I want to ask what time it closes today" after detecting a user operation on the voice generation control 809.

[0249] Then, as Figure 9 In the interface shown in (i), the electronic device can detect user actions performed on the interface, such as detecting user actions performed on a blank area of ​​the interface. In response to this action, it can switch to the interface for the voice recording function. Optionally, the user action can be, for example, a swipe down operation.

[0250] In this embodiment of the application, the voice input by a normal user can be converted into text, and the hearing / speech impaired user can input text, and then the electronic device will display the corresponding voice, which is beneficial to the communication between the hearing / speech impaired user and the normal user.

[0251] In general, in face-to-face communication scenarios, it supports functions such as speech-to-text conversion for the other party, text-to-speech conversion for the local user, and voice restoration for the local user.

[0252] In another possible implementation, the face-to-face repair application may only include a text input interface, omitting voice input functionality. This simplifies the application's functionality, reduces its package size, and consequently decreases the storage resources required by the electronic device. Furthermore, the embodiments of this application, by configuring both voice recording and text input functions within the face-to-face repair application, allow users to choose their preferred communication method, thereby enhancing communication flexibility and improving the user experience.

[0253] It should be noted that when the Face-to-Face Repair app has both voice recording and text input functions, the default interface upon entering the app can be either the voice recording interface or the text input interface; there is no restriction on this. Optionally, the default interface of the Face-to-Face Repair app can be the interface displayed when exiting the app. For example, if the exit interface was the voice recording interface, the default interface upon re-entry will be the voice recording interface; similarly, if the exit interface was the text input interface, the default interface upon re-entry will be the text input interface.

[0254] It should be noted that a system that can integrate voice restoration capabilities into electronic devices can provide voice restoration capabilities for some or all applications or scenarios involving voice input functions. In other words, it can restore the voice before it is emitted for some or all applications or scenarios involving voice input functions, thus comprehensively improving the user experience.

[0255] The above examples illustrate the interface for communication scenarios. The following examples illustrate the specific implementation of voice processing.

[0256] Please see Figure 10 , Figure 10 This is a flowchart illustrating a speech processing method provided in an embodiment of this application. Figure 10 The method shown can be performed by at least one of an electronic device or in the cloud. For example... Figure 10 The methods shown may include:

[0257] S1001, Obtain sound input.

[0258] In this embodiment of the application, the received input can be user voice input. For example, the voice input method can be referred to... Figure 5 , Figure 7 , Figure 8 or Figure 9 The method shown will not be elaborated upon here.

[0259] S1002. Input the acquired sound into the ASR model.

[0260] The ASR model is used to convert sound into text. In this embodiment, the ASR model can be a pre-trained model. This embodiment inputs the acquired sound into the ASR model, which then converts the input sound into text output.

[0261] S1003. Obtain the text output by the ASR model.

[0262] S1004. Correct the text.

[0263] In this embodiment, if the acquired sound is input by a user with a speech impairment, the converted text may contain errors, thus requiring text correction. Optionally, text correction can be automatic or manual. Automatic text correction can include, but is not limited to, statistical text correction or context-aware text correction. Statistical text correction can utilize large amounts of text data to learn word frequencies and contextual information, using a statistical model to predict the most likely correct text. Context-aware text correction can combine contextual information for text correction; contextual information can be, for example, the preceding and following sentences. Manual correction can be, for example... Figure 8 As shown in the diagram.

[0264] S1005. Obtain the result of the voice registration.

[0265] In this embodiment, the result of voice registration may include the registered voice or voiceprint features extracted from the registered voice. Optionally, when a user uses the voice repair function, if it is detected that the user has not registered a voice, the user is redirected to the voice registration interface to register; alternatively, the user may have already registered a voice before using the voice repair function. For example, the method of registering a voice can be referred to... Figure 2 Related descriptions.

[0266] S1006. Perform TTS broadcasting based on the corrected text and voice registration results.

[0267] In this embodiment, the corrected text can be converted into speech and then played back according to the registered voiceprint. This not only results in higher intelligibility of the output speech compared to the acquired speech, but also improves the user experience by playing back the voiceprint of the registered user. Optionally, TTS can use a neural network model to convert the text into audio with the user's voice and play it back.

[0268] It should be understood that, through the registration of voiceprints in this application embodiment, voiceprint features can be obtained in advance. When performing voice broadcasting, the pre-registered voiceprints can be obtained, which can reduce the time for extracting voiceprints after obtaining the sound, and thus reduce the time between obtaining the sound and broadcasting the speech, thereby improving the efficiency of speech processing.

[0269] In one possible implementation, S1005 may not be necessary. In this case, S1006 can be based on the corrected text to broadcast the speech. In this case, the voiceprint of the broadcast speech can be the voiceprint built into the electronic device or the voiceprint of the acquired sound.

[0270] It should be noted that, since the acquired sound is converted into text in this embodiment, the solution of this embodiment can be applied to scenarios that require text display, such as... Figure 7 , Figure 8 or Figure 9 The scenario is illustrated. Furthermore, since it also needs to be converted into text, the solution in this application embodiment can be applied to scenarios where real-time requirements are not particularly high.

[0271] Another possible implementation is to correct the acquired sound and then convert it into text using an ASR model.

[0272] The following is an example illustrating the specific implementation of sound restoration.

[0273] Please see Figure 11 , Figure 11 This is a flowchart illustrating a sound registration method provided in an embodiment of this application. Figure 11 The method shown can be performed by at least one of an electronic device or in the cloud. For example... Figure 11 The methods shown may include:

[0274] S1101, Voiceprint Registration.

[0275] In this embodiment of the application, the electronic device can trigger voiceprint registration processing in response to a user's voiceprint registration operation. For example, the voiceprint registration operation can be, for instance,... Figure 2 The operation shown in (e) is not restricted here.

[0276] S1102, Sound Input.

[0277] In this embodiment of the application, after the electronic device responds to the user's voiceprint registration operation, it can collect the user's voice input. For example, the method of receiving the user's voice input can be as follows: Figure 2 The method shown in (e) is not limited here.

[0278] S1103, ASR recognition rate grading judgment.

[0279] In this embodiment, after collecting the user's voice input, the electronic device can perform ASR recognition on the voice input to obtain an ASR recognition result. It should be understood that the ASR recognition result can be interpreted as the recognized text corresponding to the user's voice input. Then, the ASR recognition result is compared with pre-recorded audio content to obtain the ASR recognition rate. The ASR recognition rate can also be called the similarity score, used to indicate the degree of similarity between the ASR recognition result and the pre-recorded audio content. For example, the audio recording content can be, for instance,... Figure 2The text content shown in (f) is not limited here. After obtaining the ASR recognition rate, it can be determined which recognition rate level it belongs to. Optionally, multiple recognition rate levels can be pre-defined, and these levels do not share a common recognition rate. For example, there are multiple recognition rate levels such as (0, 70%) and (70%, 100%). If the ASR recognition rate is 50%, it belongs to the (0, 70%) level; if the ASR recognition rate is 80%, it belongs to the (70%, 100%) level.

[0280] This application embodiment improves the efficiency of repair model selection by classifying ASR recognition rates before selecting a repair model.

[0281] S1104, Repair Model Selection.

[0282] The repair model is used to repair the input speech to obtain the repaired speech. In this embodiment, the repair model can be selected based on the ASR recognition rate classification result. Optionally, a mapping relationship between each recognition rate level and the repair model is pre-configured. After obtaining the ASR recognition rate classification result, the repair model corresponding to the ASR recognition rate classification result can be obtained from the mapping relationship. Optionally, the speech repair capabilities of the repair models corresponding to different recognition rate levels are different. Optionally, the higher the recognition rate in a recognition rate level, the weaker the speech repair capability of the corresponding repair model; that is, the lower the recognition rate in a recognition rate level, the stronger the speech repair capability of the corresponding repair model.

[0283] For example, by calculating the recognition rate of user-recorded audio, the intelligibility of the user's speech can be indirectly determined, allowing for the selection of at least one suitable restoration model from a pool of alternative restoration models. For instance, users with an ASR recognition rate below 70% use a more robust restoration model, while users with an ASR recognition rate above 70% use a less robust restoration model that focuses on voice enhancement and accent correction.

[0284] This application embodiment selects a repair model based on the ASR recognition rate classification results. In other words, it can select a repair model based on the degree of impairment of the user's speaking ability. This allows for the selection of a weaker repair model for users with less impairment of speaking ability, resulting in less computational power required for speech repair. Conversely, it allows for the selection of a weaker repair model for users with more impairment of speaking ability, resulting in better speech repair performance. This approach balances the effectiveness of speech repair with the computational resources required for speech repair.

[0285] S1105, Voiceprint Feature Extraction Module.

[0286] The voiceprint feature extraction module is used to extract voiceprint features (referred to as voiceprint). In this embodiment, the voiceprint feature extraction module is used to extract the voiceprint of the user's input voice. Optionally, the voiceprint feature extraction module can convert audio of arbitrary length into a fixed-length voiceprint feature vector. This voiceprint feature extraction module can be an algorithm or a neural network model. Optionally, the voiceprint extraction module can select a voiceprint extraction model. In this embodiment, the voiceprint extraction model can include, but is not limited to, an emphasized channel attention, propagation and aggregation in time delay neural network (ECAPA-TDNN) structure.

[0287] S1106, a fixed-length eigenvector.

[0288] In this embodiment, the fixed-length feature vector is a vector obtained by extracting voiceprint features from the user's recorded voice, used to indicate the user's voiceprint characteristics. By converting the extracted voiceprint features into a fixed-length voiceprint feature vector, this embodiment helps reduce the storage resources required to store the feature vector.

[0289] S1107, Save.

[0290] In this embodiment of the application, the feature vector can be saved so that the voiceprint can be used for subsequent speech playback.

[0291] In another possible implementation, the length of the feature vector saved by S1106 can be related to the length of the input sound, rather than being a fixed-length feature vector. This allows for adaptive saving based on the length of the input sound, improving the accuracy of the saved voiceprint feature vector.

[0292] In another possible implementation, in S1103, after obtaining the ASR recognition rate, gear determination can be skipped, and the repair model can be selected directly based on the ASR recognition rate.

[0293] In another possible implementation, S1103 and S1104 may not be necessary. Instead, a pre-configured repair model can be used to repair the sound, which can reduce the time required for ASR recognition rate classification and repair model selection.

[0294] The architecture of one of the repair models will be described below.

[0295] Please see Figure 12 , Figure 12This is a schematic diagram of the architecture of a repair model provided in an embodiment of this application. The repair model of this embodiment can be deployed on at least one of electronic devices or in the cloud, such as... Figure 12 The restoration model shown may include a prosody encoder 1201, a streaming ASR encoder 1202, an audio discretization unit 1203, a streaming audio restoration big model 1204, and a vocoder 1205.

[0296] The prosody encoder 1201 is used to extract prosodic features (hereinafter referred to as prosody). For example, the prosody encoder 1201 can extract prosodic features from an input audio signal. Prosodic features may include, but are not limited to, at least one of pitch, rhythm, or intonation. Optionally, an audio signal from a normal user can be acquired, and then prosodic features can be extracted from the normal user's audio signal as the prosodic features for playing audio.

[0297] The streaming ASR encoder 1202 is an encoder capable of converting audio signals into text sequences in real-time or near real-time. In this embodiment, the streaming ASR encoder 1202 is used to convert user-input audio into text in real-time or near real-time. The streaming ASR encoder 1202 may include a neural network model.

[0298] The audio discretization unit 1203 is used to convert a continuous audio signal into a discrete sequence of symbols or units. For example, a continuous audio signal (such as a waveform signal) is decomposed into a series of discrete, identifiable units or symbols by some method. These units can be phoneme-based, feature-based, or obtained through other means (such as cluster analysis).

[0299] The streaming audio restoration model 1204 is used to restore or optimize audio features. Audio features may include, but are not limited to, at least one of content features, prosodic features, or voiceprint features. Optionally, the streaming audio restoration model 1204 utilizes streaming language model-based audio conversion (streaming LM-based VC) for restoration or optimization, combining the advantages of streaming processing and language models (LM) to achieve real-time, efficient audio conversion. Optionally, the streaming audio restoration model 1204 may include, but is not limited to, the streaming prosodic restoration model 1204.

[0300] Among them, the vocoder 1205 is a synthesizer that converts audio features into playable audio signals.

[0301] In the embodiments of this application, the repair model can be a pre-trained model.

[0302] In one possible implementation, the restoration model can perform the following processing when performing audio restoration:

[0303] The prosodic encoder 1201 can extract prosodic features from built-in normal user voices. The streaming ASR encoder 1202 performs ASR recognition processing on the input audio to be repaired, obtaining the ASR recognition result (also known as content features or content). Then, the result of concatenating the ASR recognition result, prosodic features, and user-registered voiceprint features is input into the streaming audio restoration model 1204. Furthermore, the streaming audio restoration model 1204 fuses the voiceprint features, prosodic features, and ASR recognition result in the concatenated result with the initial audio discretization features (also known as discretization features), and generates repaired discretization features and repaired continuous audio features through autoregressive inference. The discretization features generated by autoregressive inference will be used as input for subsequent autoregressive inference, fused and inferred by the streaming audio restoration model 1204, while the generated continuous audio features will be used as input for the vocoder 1205. Then, by converting the continuous features of the audio into the repaired audio through the vocoder 1205, the repaired audio can be broadcast through an audio broadcasting device (e.g., a loudspeaker). The voiceprint of the broadcast audio can be the voiceprint registered by the user, and the prosody of the broadcast audio can be the prosody extracted by the prosody encoder 1201.

[0304] The audio to be repaired can be user-input audio, acquired through a sound acquisition device (e.g., a microphone) of an electronic device. For example, the audio to be repaired in this embodiment may include, but is not limited to, audio input by a user. Figure 5 , Figure 7 , Figure 8 or Figure 9 The input speech in the scenario shown. Optionally, the initial audio discretization features may be obtained by the audio discretization unit 1203 based on at least one of the ASR recognition results or the audio to be repaired.

[0305] In this embodiment of the application, since the ASR recognition result is obtained by performing ASR recognition processing on the audio to be repaired, and the audio discretization feature is also obtained by performing discretization processing on the audio to be repaired, the ASR recognition result and the audio discretization feature are fused together. That is, the processing results of the audio to be repaired from different dimensions can be fused together, thereby improving the accuracy of the obtained audio continuous features, and then improving the accuracy of the content of the obtained audio to be repaired.

[0306] Optionally, audio repair can be performed as soon as the user inputs audio, thus improving the real-time performance of audio repair. Alternatively, repair can begin after a certain amount of audio has been acquired. For example, if one second of audio has been acquired while the user is inputting audio, repair processing can begin on that one second, and then on the next second, until a complete segment of audio input by the user has been repaired.

[0307] In another possible implementation, generating the restored discretized features may not be necessary, which simplifies the architecture of the restoration model and improves the efficiency of sound restoration. However, the embodiments of this application generate restored discretized features, which can be used as input for subsequent autoregressive inference and fused and inferred by the streaming audio restoration large model 1204, thereby improving the accuracy of sound restoration.

[0308] Alternatively, each segment of the audio to be repaired can be of the same length, which can improve the accuracy of audio repair; or the different segments of the audio to be repaired can be of different lengths, which can improve the flexibility of audio repair.

[0309] For example, assuming the user inputs the audio "Going or not this afternoon", after obtaining "this afternoon", it is used as the audio to be repaired. Then, the streaming ASR encoder 1202 performs ASR recognition processing on the input "this afternoon" audio segment to obtain the ASR recognition result. Then, the result of concatenating the ASR recognition result, prosodic features, and user-registered voiceprint features is input into the streaming audio repair model 1204. Furthermore, the audio discretization unit 1203 discretizes the "this afternoon" audio segment based on the ASR recognition result to obtain audio discretization features. Then, the audio discretization features are input into the streaming audio repair model 1204. The streaming audio repair model 1204 can output the repaired audio discretization features and repaired audio continuous features corresponding to the "this afternoon" audio segment. The vocoder 1205 can output the repair result of the "this afternoon" audio segment based on the repaired audio continuous features. Furthermore, the streaming audio restoration model 1204 can use the discretized features of the restored audio corresponding to the "this afternoon" audio segment as a reference for whether or not to remove this audio segment during restoration processing, for example, as a reference for whether or not to remove this audio segment during discretization processing.

[0310] The training of the repair model will be illustrated below.

[0311] Obtain a training sample set, which includes one or more training samples. Each training sample includes an audio sample, a text sample corresponding to the audio sample, and a repaired audio sample corresponding to the audio sample.

[0312] During training, audio samples are used as input to the streaming ASR encoder 1202, which outputs predicted text. The predicted text is compared with the text samples corresponding to the audio samples to calculate a first training loss. If the first training loss meets a first training termination condition, the streaming ASR encoder 1202 terminates training. If the first training loss does not meet the first training termination condition, the parameters of the streaming ASR encoder 1202 are updated, and training continues until the first training loss meets the first training termination condition. Optionally, the first training termination condition may include the first training loss being less than a first threshold.

[0313] Then, the restoration model continues to be trained. Audio samples are used as input to the streaming ASR encoder 1202, or the corresponding text samples are used as input to the streaming audio restoration model 1204. The restored audio output by the vocoder 1205 is obtained, and then the restored audio output by the vocoder 1205 is compared with the restored audio samples corresponding to the audio samples to calculate the second training loss. If the second training loss meets the second training termination condition, the restoration model ends training; if the second training loss does not meet the second training termination condition, the parameters of the restoration model are updated and training continues until the second training loss meets the second training termination condition. Optionally, the second training termination condition may include the second training loss being less than a second threshold. Optionally, at least one of the vocoder 1205 or the prosodic encoder 1201 in this embodiment may be pre-trained. In this case, updating the parameters of the restoration model may involve updating at least one of the streaming audio restoration model 1204 or the audio discretization unit 1203.

[0314] Optionally, the audio samples may include a first audio sample and a second audio sample. For example, if an audio sample could include "Today's weather is sunny," then the first audio sample could include "Today's weather," and the second audio sample could include "sunny." Correspondingly, the text samples corresponding to the audio samples include the text of the first audio sample and the text of the second audio sample, and the repaired audio samples corresponding to the audio samples include the repaired audio samples corresponding to the first and second audio samples.

[0315] Then, during training, the repaired audio corresponding to the first audio sample output by the repair model is compared with the repaired audio sample to calculate the second training loss. Furthermore, the discretized features of the repaired audio corresponding to the first audio sample output by the repair model are compared with the discretized features of the second audio sample to calculate the third training loss. If the third training loss satisfies the third training termination condition, the repair model training ends; if the third training loss does not satisfy the third training termination condition, the parameters of the repair model are updated and training continues until the third training loss satisfies the third training termination condition. Optionally, the third training termination condition may include the third training loss being less than a third threshold.

[0316] It should be understood that the processing flow of the repair model in the embodiment of this application during the training process can be referred to the description of the sound repair processing flow in the above embodiment, and will not be repeated here.

[0317] It should be noted that users can access the personalized portal and record a large amount of text as training samples to train a personalized repair model, which can improve the user experience.

[0318] In another possible implementation, the audio discretization unit 1203 may not require ASR recognition results when discretizing the audio to be repaired. This can improve the efficiency of discretization processing and thus improve the efficiency of audio repair. In this embodiment, the audio discretization unit 1203 uses ASR recognition results to discretize the audio to be repaired, which can improve the accuracy of audio repair.

[0319] In another possible implementation, the restoration model may include either the audio discretization unit 1203 or the streaming ASR encoder 1202. This improves the efficiency of discretization processing, while audio restoration using both the audio discretization unit 1203 and the streaming ASR encoder 1202 enhances the accuracy of audio restoration. It should be noted that if the restoration model includes the audio discretization unit 1203 but not the streaming ASR encoder 1202, the result of concatenating the audio discretization features, prosodic features, and user-registered voiceprint features can be input into the large streaming audio restoration model 1204.

[0320] In another possible implementation, the prosody encoder 1201 may not be necessary during audio restoration. Instead, by pre-storing prosodic features, audio restoration can be performed using these pre-stored prosodic features, thus improving the efficiency of audio restoration.

[0321] In another possible implementation, the streaming audio inpainting large model 1204 may not require the prediction of audio discretization features, which can improve the efficiency of audio inpainting, while the prediction of audio discretization features can improve the accuracy of audio inpainting.

[0322] It should be noted that both the prosodic encoder 1201 and the streaming prosodic restoration model 1204 can adopt a generative pre-trained transformer (GPT) structure. The speech discretization unit can use a vector-quantized variational autoencoder (VQVAE) model, and the streaming ASR encoder 1202 can adopt a recurrent neural network transducer (RNNT). The vocoder 1205 can include an efficient and high-fidelity generative adversarial network (HiFi GAN). Specifically, VQVAE can convert audio into frame-level discretized units (discrete features). HiFi GAN can convert audio features into audio sample points (audio signals).

[0323] The following examples illustrate the process of the embodiments of this application based on the above examples.

[0324] Please see Figure 13 , Figure 13 This is a flowchart illustrating another voice processing method provided in an embodiment of this application. The method of this application embodiment can be executed by an electronic device, a server, or a system including both an electronic device and a server. This application embodiment can repair user-inputted voice. For example... Figure 13 The methods shown may include:

[0325] S1301, Receive setting input, the setting input is used to indicate the scenario for enabling voice repair, the contact for enabling voice repair, or the application for enabling voice repair, the scenario includes face-to-face communication scenario or remote communication scenario.

[0326] The face-to-face communication scenario can be a scenario where at least two users are communicating face-to-face. The remote communication scenario can be a scenario where communication is conducted using communication means. In this embodiment, at least one of the following can be selectively enabled: a scenario for enabling voice repair, a contact for enabling voice repair, or an application for enabling voice repair. Optionally, if an input is set to enable a scenario for enabling voice repair, then if the electronic device is detected to be in a scenario where voice repair is enabled, such as a face-to-face communication scenario or a remote communication scenario, the user's input voice is repaired. For example, a face-to-face communication scenario can be, for example, a scenario where... Figure 8 or Figure 9 The scenarios described in the embodiments will not be repeated here. Remote communication scenarios could be, for example,... Figure 5 or Figure 7 The scenarios described in the embodiments are not limited here. If a contact is set as the input to enable voice repair, then when the electronic device communicates with a contact whose voice repair is enabled, the electronic device will repair the user's input voice. For example, the method of enabling voice repair by setting a contact could be as follows: Figure 4 The description of the embodiments will not be repeated here. If the input is set to enable the application for voice repair, then when the application running on the electronic device is the application with voice repair enabled, the user's voice input will be repaired. For example, enabling the application for voice repair could be done in the following ways: Figure 3 The implementation method is not described in detail here.

[0327] S1302. Obtain the user's first voice characteristics through voice registration.

[0328] The first speech feature can be a user's speech feature. In this embodiment, the first speech feature is used to represent the characteristics of the user's pronunciation. Optionally, the first speech feature can include at least one of voiceprint features or prosodic features. Voiceprint features can be a sound wave spectrum carrying speech information displayed by an electroacoustic instrument. Prosodic features can refer to those features in speech that are not directly manifested as changes in timbre, but are reflected through changes in factors such as pitch, duration, and intensity. These features are manifested as suprasegmental components in the speech signal, which, together with timbre components (such as vowels and consonants), constitute a complete speech signal. The first speech feature can be obtained by feature extraction from the user's registered speech. For example, the method of user registration speech can refer to... Figure 2 The description of the embodiments is omitted here. It should be noted that the users in the embodiments of this application can include users with speech impairments or hearing impairments. Users with speech impairments can be, for example, users who have difficulty or defects in pronunciation. Users with hearing impairments can be, for example, users who have difficulty or defects in hearing.

[0329] S1303, Receive the first voice input from the user.

[0330] S1304. Repair the first speech based on the input settings and the first speech characteristics.

[0331] In this embodiment, the first speech is repaired based on the setting input and the first speech feature. Since the setting input is used to indicate the scenario where speech repair is enabled, the contact person for enabling speech repair, or the application for enabling speech repair, when the electronic device is detected to be in a scenario where speech repair is enabled, or when the contact being communicated by the electronic device is the contact person for enabling speech repair, or when the application being run by the electronic device is the application for enabling speech repair, the first speech is repaired based on the first speech feature. This allows the repaired first speech to be played using the first speech feature, and the intelligibility of the repaired first speech is higher than that of the unrepaired first speech. Intelligibility can represent the accuracy of expressing what the user wants to say, or it can be understood as the listener's understanding of the speech signal transmitted by the speaker.

[0332] In this embodiment, by receiving setting input—which indicates the scenario, contact, or application for enabling voice restoration—and then obtaining the user's first voice feature through voice registration, the system can restore the first voice based on the setting input and the first voice feature after receiving the user's input. This allows users with speech or hearing impairments to communicate via voice input, improving their communication convenience. Furthermore, by generating the restored first voice using the user's registered first voice feature, the restored first voice can be played according to that feature, making it more closely resemble the user's own voice and further enhancing the user experience.

[0333] In another possible implementation, voice repair can be performed using pre-defined voice features. For example, before a user registers a voice feature, the voice is repaired according to pre-defined features, allowing the repaired voice to be played according to those features. After a user registers a voice feature, the repair is performed based on the user's initial voice features. This means that the voice repair function can be used even if the user has not registered a voice feature, thus expanding the applicable scenarios for the voice repair function. Furthermore, it reduces the processing of extracting voice features from registered voice features, thereby improving the efficiency of voice repair.

[0334] In one possible implementation, the remote communication scenario includes a call scenario, and in the call scenario, the method further includes:

[0335] Send a first notification message to the first contact in the call. The first notification message is used to indicate that the voice repair function has been enabled.

[0336] And / or,

[0337] Upon receiving the first voice message, a second notification message is sent to the first contact, indicating that the first voice message is being repaired.

[0338] The first notification message may include notification text or notification sound. Optionally, the notification text in the first notification message may be displayed on the screen of the terminal corresponding to the first contact person, and the notification sound in the first notification message may be played through the speaker of the terminal corresponding to the first contact person. For example, the first notification message may refer to... Figure 6 The relevant descriptions of the prompts in the embodiments are not repeated here. The second prompt information may include prompt text or prompt sound. Optionally, the prompt text in the second prompt information may be displayed on the screen of the terminal corresponding to the first contact, and the prompt sound in the first prompt information may be played through the speaker of the terminal corresponding to the first contact. For example, the prompt sound in the second prompt information may be, for example, Figure 6 The "ding" and other notification sounds exemplified in the embodiments will not be described in detail here.

[0339] In this embodiment, by sending a first notification message to the first contact in the call to inform the user that the voice repair function has been enabled, the first contact can be informed that the user has enabled the voice repair function, thereby improving the call experience. Furthermore, by sending a second notification message to the first contact upon receiving the first voice message to indicate that the first voice message is being repaired, the first contact can be informed that the delay is due to voice repair, reducing the possibility of conflicting speaking times in two-party or multi-party calls, thereby improving the call experience.

[0340] In another possible implementation, at least one of the first or second prompt messages may not be sent, which can reduce the consumption of transmission resources during the call.

[0341] In one possible implementation, the second notification message includes a notification tone, which is sent to the first contact upon receiving the first voice message, including:

[0342] Once the first voice message is received, continuously send notification to the first contact.

[0343] The method also includes:

[0344] Once the first restored voice message is sent to the first contact, stop sending notification sounds to the first contact.

[0345] For example, embodiments of this application may refer to Figure 6 The explanation of (c) in the text will not be repeated here.

[0346] In this embodiment of the application, by continuously sending a prompt tone to the first contact when the first voice is received, the first contact can know that the user on the other end is speaking when they hear the prompt tone. And when the prompt tone is stopped after the repaired first voice is sent to the first contact, the first contact can then hear the repaired first voice. This can reduce the situation of conflict between two or more parties during the call, thereby improving the experience of both parties in the call.

[0347] In another possible implementation, a brief notification tone could be sent to the first contact when the first voice message is received, or after a certain interval following the receipt of the first voice message. This would reduce the transmission resources required to transmit the notification tone.

[0348] In one possible implementation, the remote communication scenario includes a call scenario, and in the call scenario, the method further includes:

[0349] A repaired first voice message is sent to the first contact in the call. Then, a first interface can be displayed, which is the interface for communicating with the first contact. The first interface includes a first control for disabling the voice repair function. Then, in response to a first operation on the first control, the voice repair function is disabled, and a second voice message is sent to the first contact. The second voice message includes the user's voice received after the voice repair function was disabled.

[0350] For example, the first interface can be, for instance, Figure 5 The interface shown. The first control can be, for example, a sound switching control 502. The first operation can be a user operation. In this embodiment, before the voice repair function is turned off, that is, when the voice repair function is on, the repaired voice is sent to the first contact, and after the voice repair function is turned off, the unrepaired voice is sent to the first contact.

[0351] In this embodiment, during a call, the user can also control the voice repair function to be turned off. This allows the user to disable voice repair when it is not needed, improving the flexibility of switching between using and not using voice repair. For example, the user can control the voice repair function to be turned off via the first control when the electronic device's battery is low or the electronic device is running slowly, thereby improving the ability to disable voice repair during calls and enhancing the user experience.

[0352] In another possible implementation, the voice repair function can be disabled on the first interface, meaning that the voice repair function is not supported during a call, thereby improving the simplicity of the call interface.

[0353] In another possible implementation, the voice repair function can be turned off via voice commands.

[0354] In one possible implementation, in response to the first operation, the first control is also switched from a first state to a second state, where the first state indicates that the voice repair function is enabled and the second state indicates that the voice repair function is disabled. The method further includes:

[0355] In response to a second operation on the first control, the first control is switched from a second state to a first state, and the voice repair function is enabled to send a repaired third voice message to the first contact, the third voice message including the voice message received from the user after the voice repair function is enabled.

[0356] The second operation can be a user action. The first state can be, for example, a user action. Figure 5 The state of the sound switching control 502 shown in (d) is as follows: the second state can be, for example, Figure 5 The state of the sound switching control 502 shown in (c) is shown.

[0357] In this embodiment, the state of the first control indicates whether the voice repair function is currently enabled, thereby improving the accuracy of the user's choice to use or disable the voice repair function. Furthermore, this embodiment not only allows the voice repair function to be disabled during a call but also to be re-enabled, thus increasing the flexibility of enabling or disabling voice repair. In addition, enabling or disabling the voice repair function through the same control improves the simplicity of the interface during calls.

[0358] It should be understood that WeChat and face-to-face communication applications can also enable or disable the voice repair function within the application interface. Please refer to the relevant descriptions on how to enable or disable the voice repair function in the call interface, which will not be repeated here.

[0359] In another possible implementation, once the voice repair function is turned off during a call, it can no longer be re-enabled.

[0360] In another possible implementation, the voice repair function can be turned on or off through different controls. For example, one control can be used to turn on the voice repair function during a call, and another control can be used to turn it off. This can improve the independence of the control over turning the voice repair function on or off.

[0361] In one possible implementation, the method further includes:

[0362] A second interface is displayed, including playback controls for the repaired first speech and first text corresponding to the repaired first speech. A third operation is received, which selects a first target character in the first text. In response to the third operation, at least one candidate character related to the first target character is displayed. A fourth operation is received, which selects a second target character from the at least one candidate character. In response to the fourth operation, second text and playback controls for the speech corresponding to the second text are displayed, the second text being obtained by replacing the first target character in the first text with the second target character.

[0363] For example, the second interface could be, for instance, Figure 8 In the interface shown, the playback control for the repaired first voice message can be, for example, playback control 803, and the first text corresponding to the first voice message can be, for example, "Are you going this afternoon?". The third operation can be a user operation. The first target character can be a character selected by the user. Optionally, the first target character can be, but is not limited to, text, punctuation marks, or emoticons. For example, the first target character can be, for example, "afternoon". At least one candidate character can be displayed in a list. For example, at least one candidate character can be, for example, characters displayed in candidate word list 806, such as "morning", "good at dancing", and "foggy". The fourth operation can be a user operation. The second target character can be a character selected by the user from at least one candidate character. For example, the second target character can be, for example, "morning". For example, the second text can be, for example, "Are you going this morning?". The playback control for the voice message corresponding to the second text can also be, for example, playback control 803.

[0364] In this embodiment, a second interface is displayed, including playback controls for the repaired first speech and first text corresponding to the repaired first speech. A third operation is received, whereby a first target character in the first text is selected. In response to the third operation, at least one candidate character related to the first target character is displayed. A fourth operation is received, whereby a second target character is selected from the at least one candidate character. In response to the fourth operation, second text and playback controls for the speech corresponding to the second text are displayed. This allows the user to quickly and manually correct inaccurate speech repair results, thereby improving the accuracy of speech communication.

[0365] In another possible implementation, the user could re-enter the voice when they find the voice repair result to be inaccurate.

[0366] In another possible implementation, the user can generate the second text and the corresponding playback control with one click. In this case, at least one character or word in the first text can be replaced to generate the second text and the corresponding playback control.

[0367] In one possible implementation, the method further includes:

[0368] Cancel the display of the playback controls for the first audio clip after the repair.

[0369] In this embodiment of the application, by canceling the display of the playback control for the repaired first voice, only the latest voice can be played at any given time, which improves the convenience of selecting the appropriate voice for playback.

[0370] In another possible implementation, the display of the playback control for the repaired first voice recording can be left undisplayed; that is, not only the playback control for the latest voice recording, but also the playback controls for historical voice recordings can be retained.

[0371] In one possible implementation, for remote communication scenarios, the method also includes:

[0372] A third interface is displayed, which is the interface for communicating with the second contact. The third interface includes a second control and a third control. The second control is used to instruct the sending of a first voice message to the second contact, and the third control is used to instruct the sending of a repaired first voice message to the second contact. Then, in response to an operation on the third control, the repaired first voice message can be sent to the second contact. Alternatively, in response to an operation on the second control, the first voice message can be sent to the second contact.

[0373] For example, the third interface can be, for instance, Figure 7 (a) in Figure 7 The interface shown in (f) is as follows. The second control can be, for example, to directly send option 703, and the third control can be, for example, to repair and send option 704.

[0374] In this embodiment, the user can selectively send the repaired voice or the original voice to the second contact through the second and third controls. That is, even if the voice repair function is enabled, the user can still choose to send the original voice to the second contact, thereby improving the user's flexibility in remote communication.

[0375] In another possible implementation, there could be only one voice sending control. After the electronic device detects that the voice sending control is being used, it sends the repaired voice, which can improve the efficiency of sending the repaired voice in remote communication scenarios.

[0376] In one possible implementation, for face-to-face communication scenarios, the method also includes:

[0377] A fourth interface is displayed, including a virtual keyboard. In response to an operation on the virtual keyboard, third text is displayed. A fifth operation is received on a fourth control in the fourth interface, which instructs the generation of speech. In response to the fifth operation, speech corresponding to the third text is generated based on first speech features.

[0378] For example, the fourth interface could be, for instance, Figure 9 The interface shown in (c)-(i) is as follows. The third text can be text entered by the user via a virtual keyboard, such as "Hello, I am Zhang San" or "I would like to ask what time the store closes today," etc. The fourth control can be, for example, control 809. The fifth operation can be a user operation. In this embodiment, the voice corresponding to the third text is generated based on the first voice feature, so the voice corresponding to the third text can be played according to the first voice feature.

[0379] In this embodiment, a fourth interface is displayed, including a virtual keyboard. In response to an operation on the virtual keyboard, third text is displayed. A fifth operation is received on a fourth control in the fourth interface, the fourth control indicating the generation of speech. In response to the fifth operation, speech corresponding to the third text is generated based on a first speech feature. That is, the user can also input text and then generate speech based on the registered first speech feature, thus achieving text-to-speech conversion. Users can then communicate by choosing to input either speech or text, improving the selectivity and flexibility of user communication.

[0380] In another possible implementation, the speech corresponding to the entered third text could be automatically generated after the user finishes entering the third text via the virtual keyboard, thus reducing user operations. Optionally, the third text entry could be considered complete if no operation on the virtual keyboard is detected within a certain period of time.

[0381] In one possible implementation, the third text includes punctuation marks and / or emojis. When generating the speech corresponding to the third text based on the first speech features, the tone of the speech corresponding to the third text is also controlled based on the punctuation marks and / or emojis.

[0382] In this embodiment, the tone of the speech corresponding to the third text is controlled based on punctuation marks and / or emoticons, so that the tone of the text corresponding to the third speech matches the punctuation marks and / or emoticons. For example, if the last character of the third text is marked with an exclamation mark "!", the tone of the speech corresponding to that third text can be exclamatory; if the last character of the third text is marked with a question mark "?", the tone of the speech corresponding to that third text can be interrogative. For example, if the emoticons in the third text include a smiley face, the tone of the speech corresponding to the third text is cheerful; if the emoticons in the third text include a crying face, the tone of the speech corresponding to the third text is sad.

[0383] In this embodiment, the tone of the speech corresponding to the third text can be controlled by punctuation marks and / or emoticons in the third text input by the user. This allows the tone of the output speech to be adjusted adaptively according to the user input text, thereby improving the flexibility of voice broadcasting and enhancing the user experience.

[0384] In another possible implementation, punctuation marks or emojis in the third-party text can be disregarded. This reduces the computational resources required to determine the tone matching for punctuation marks or emojis. Optionally, a mapping relationship between punctuation marks and tone, as well as between emojis and tone, can be pre-defined to achieve the desired tone matching for punctuation marks or emojis.

[0385] In one possible implementation, the user's initial voice characteristics are obtained through voice registration, including:

[0386] The fifth interface is displayed. This fifth interface is for voice registration and includes a third prompt message and a fifth control. The third prompt message indicates the recording content for voice registration, and the fifth control instructs on recording speech. In response to an action on the fifth control, a fourth voice recording is made. First speech features are extracted from the fourth voice recording.

[0387] For example, the fifth interface could be, for instance, Figure 2 The interface shown. The third prompt can be a text prompt such as "The giant panda sleeps wherever it goes," or a voice prompt; there are no restrictions here. The fifth control can be, for example, a recording control 208. The fourth voice can be the user's recorded voice, used to register the user's voice so that the user's first voice feature can be extracted, and thus, when playing the repaired voice, it can be played according to the user's first voice feature.

[0388] In one possible implementation, the method also includes:

[0389] Speech recognition is performed on the fourth speech to obtain the fourth text. Then, the fourth text is compared with the recorded content to obtain the similarity between the fourth text and the recorded content. Then, a target restoration model can be selected from multiple restoration models based on the similarity. The target restoration model is used to restore the first speech. The speech restoration capabilities of the multiple restoration models are different, and the speech restoration capability of the target restoration model is negatively correlated with the similarity.

[0390] In this embodiment of the application, the text obtained by speech recognition of the user's registered speech is compared with the recorded content to obtain a similarity score. Then, a target repair model is selected from multiple repair models based on the similarity score. This allows for the selection of a suitable repair model based on the user's speech impairment level, thereby balancing the accuracy of speech repair with the computing resources required for speech repair.

[0391] It should be understood that multiple repair models can be deployed on a single server, allowing the corresponding repair model to be retrieved from that server for voice repair via model identifier. Alternatively, they can be deployed on multiple servers, allowing the server corresponding to the target repair model to be called based on the server-repair model mapping. Furthermore, multiple repair models can also be deployed on electronic devices; this is not a limitation.

[0392] In another possible implementation, only one restoration model can be configured, thus eliminating the need for similarity matching with the recorded content and improving the efficiency of voice restoration.

[0393] The above embodiments describe an example of repairing a user's voice. The following embodiments illustrate an example of repairing the voice of a contact person who communicates with the user.

[0394] Please see Figure 14 , Figure 14 This is a flowchart illustrating another voice processing method provided in an embodiment of this application. The method of this embodiment can be executed by an electronic device, a server, or a system including both an electronic device and a server. This embodiment can also be used to repair the voice of a user's contacts. Figure 14 The methods shown may include:

[0395] 1401. Receive setting input. The setting input is used to indicate the scenario for enabling voice repair, the contact for enabling voice repair, or the application for enabling voice repair. The scenario includes face-to-face communication scenario or remote communication scenario.

[0396] S1401 can be referred to the description of S1301, and will not be repeated here.

[0397] 1402. Receive the fifth voice message from the target contact.

[0398] The target contact can be a contact who communicates with the user; for example, the target contact can be the first contact or the second contact.

[0399] 1403. Obtain the second speech feature, which is a preset speech feature or a speech feature extracted from the fifth speech feature.

[0400] The second speech feature may include at least one of voiceprint features or prosodic features. In another possible implementation, it could be speech features extracted from a user's registered voice. For example, if the user of the electronic device is a first user, and the first user wants to hear a second user's voice during speech restoration, the second user could register their voice, and then speech features could be extracted from the second user's registered voice.

[0401] 1404. Repair the fifth voice based on the input settings and the second voice features.

[0402] In this embodiment, the fifth voice is repaired based on the setting input and the second voice feature. Since the setting input is used to indicate the scenario where voice repair is enabled, the contact person who enables voice repair, or the application that enables voice repair, when the electronic device is detected to be in a scenario where voice repair is enabled, when the contact person being communicated with by the electronic device is the contact person who enables voice repair, or when the application being run by the electronic device is the application that enables voice repair, the fifth voice is repaired based on the second voice feature, so that the intelligibility of the repaired fifth voice is higher than that of the unrepaired fifth voice. Intelligibility can represent the degree of understanding when the voice is heard, or it can be understood as the accuracy with which the voice expresses what the contact person intends to convey.

[0403] In one possible implementation, for remote communication scenarios, the method also includes:

[0404] The sixth interface is displayed, which is the interface for communicating with the target contact. The sixth interface includes the fifth voice message. Then, in response to an action on the fifth voice message, a sixth control and a seventh control can be displayed. The sixth control instructs the conversion of the fifth voice message into text, and the seventh control instructs the conversion of the repaired fifth voice message into text. Then, in response to an action on the seventh control, the text corresponding to the repaired fifth voice message can be displayed. Alternatively, in response to an action on the sixth control, the text corresponding to the fifth voice message can be displayed.

[0405] For example, the sixth interface could be, for instance, Figure 7In the interface shown in (g)-(i), the sixth control can be, for example, the speech-to-text option 708, and the seventh control can be, for example, the speech-to-text after repair option 707.

[0406] In this embodiment of the application, after receiving the voice from the target contact, the fifth voice can be selectively converted into text, which can also be understood as converting the original voice of the target contact into text; or the repaired fifth voice can be converted into text. Users can choose to display the text as needed, thereby improving the flexibility and experience of user communication.

[0407] In another possible implementation, the sixth control can be used to instruct the playback of the fifth voice, and the seventh control can be used to instruct the playback of the repaired fifth voice. In this embodiment, the user can choose to play either the original voice of the target contact or the repaired voice, thereby improving the flexibility and experience of user communication.

[0408] In one possible implementation, for face-to-face communication scenarios, the method also includes:

[0409] The seventh interface is displayed, which is the interface for communicating with the target contact. The seventh interface includes an eighth control. Then, in response to the sixth action on the eighth control, the fifth text is displayed. Then, in response to the seventh action on the eighth control, the display of the fifth text stops. The fifth text includes the text corresponding to the fifth voice within the target time period or the text corresponding to the repaired fifth voice. The target time period includes the time period from the response to the sixth action to the response to the seventh action.

[0410] For example, the seventh interface could be, for instance, Figure 9 In the interface shown in (a)-(b), the eighth control can be, for example, a recording control 802. The sixth operation can be a user input operation. The fifth text can be, for example, "I am a hearing person, I am speaking."

[0411] In this embodiment of the application, the voice of the target contact can also be converted into text, thereby improving the flexibility of communication between the user and the contact.

[0412] In another possible implementation, the eighth control could be omitted to control the voice-to-text conversion of contacts. For example, all voice messages from contacts could be converted to text, thus reducing the number of voice-to-text operations.

[0413] Based on the above embodiments, the following is an exemplary description of how to implement voice restoration.

[0414] Please see Figure 15 , Figure 15This is a flowchart illustrating another voice processing method provided in an embodiment of this application. The method of this application embodiment can be executed by an electronic device, a server, or a system including both an electronic device and a server. Figure 15 The methods shown may include:

[0415] S1501. Obtain the target speech and speech features.

[0416] The target speech can be unrepaired speech, which can be understood as the original sound or audio to be repaired. For example, the target speech may include, but is not limited to, the first speech or the fifth speech. Speech features may include, but are not limited to, the first speech feature or the second speech feature.

[0417] S1502. Input the target speech and speech features into the target restoration model. The target restoration model is used to extract the content of the target speech, obtain continuous speech features based on the content and speech features of the target speech, and synthesize the restored target speech based on the continuous speech features.

[0418] Optionally, the content of the target speech can be obtained through speech recognition. In this embodiment, the obtained target speech can be the content of the extracted target speech read aloud according to the speech features.

[0419] S1503. Obtain the repaired target speech output by the target repair model.

[0420] For example, this embodiment can be referred to Figure 12 The relevant descriptions will not be repeated here.

[0421] In this embodiment, after the target speech is input, the electronic device can acquire the target speech and speech features, and then input the target speech and speech features into the target restoration model. The target restoration model is used to extract the content of the target speech, obtain speech continuity features based on the content and speech features of the target speech, and synthesize the restored target speech based on the speech continuity features. In this way, people with speech impairments can communicate by inputting speech, thereby improving the convenience of communication for people with speech impairments.

[0422] In one possible implementation, the target restoration model includes a first module, a second module, and a third module. The first module is used to extract the content of the target speech, the second module is used to obtain continuous speech features based on the content and speech features of the target speech, and the third module is used to synthesize the restored target speech based on the continuous speech features.

[0423] For example, the first module may include a streaming ASR encoder. The second module may include a streaming audio restoration large model. The third module may include a vocoder.

[0424] In one possible implementation, the target speech and speech features are input into the target inpainting model, including:

[0425] The speech features and a first portion of the target speech are input into the target restoration model. The first module extracts the content of the first portion of speech, the second module obtains continuous features of the target speech based on the speech features and the content of the first portion, and the third module synthesizes the restored first portion of speech based on the continuous features of the target speech. Then, the speech features and a second portion of the target speech are input into the target restoration model. The first module extracts the content of the second portion of speech, the second module obtains continuous features of the second portion of speech based on the speech features and the content of the second portion, and the third module synthesizes the restored second portion of speech based on the continuous features of the second portion.

[0426] For example, the first part of the speech and the second part of the speech in the embodiments of this application can be referred to Figure 6 The description of (c) in the text, for example, the first part of the voice includes "Thank you, put it downstairs", the first part of the voice may include "That's fine".

[0427] The repaired target speech includes the repaired first part of the speech and the repaired second part of the speech.

[0428] In this embodiment of the application, by first repairing a portion of the target speech and then repairing another portion of the target speech, speech repair can begin when only a portion of the speech is acquired. In other words, speech repair can begin without acquiring the complete target speech, thereby improving the efficiency of speech repair.

[0429] In another possible implementation, speech restoration can begin only after the complete target speech has been acquired. This reduces the likelihood of target speech acquisition and restoration occurring simultaneously, thereby reducing the resources required for speech restoration.

[0430] In one possible implementation, the target restoration model further includes a fourth module, which is used to discretize the first part of the speech to obtain discrete features of the target speech. The second module is also used to make predictions based on the discrete features of the target speech to obtain the restored discrete features of the target speech, and to obtain the second continuous features of the speech based on the restored discrete features of the target speech, the speech features, and the content of the second part of the speech.

[0431] For example, the fourth module may include an audio discretization unit.

[0432] In this embodiment, the first part of the speech is discretized by the fourth module to obtain the discrete features of the target speech. Then, the second module predicts based on the discrete features of the target speech to obtain the repaired discrete features of the target speech. Then, the second continuous features of the speech are obtained based on the repaired discrete features of the target speech, the speech features, and the content of the second part of the speech. In this way, the second continuous features of the speech can be obtained by combining the discrete features of the target speech, the speech features, and the content of the second part of the speech, thereby improving the accuracy of the obtained second continuous features of the speech and thus improving the accuracy of speech repair.

[0433] In another possible implementation, the fourth module may not be necessary, which would improve the efficiency of voice restoration.

[0434] In one possible implementation, the fourth module is used to discretize the first part of the speech to obtain discrete features of the target speech, including:

[0435] The fourth module is used to discretize the first part of the speech based on its content to obtain the discrete features of the target speech.

[0436] In this embodiment of the application, the first part of the speech is discretized according to the content of the first part of the speech to obtain the discrete features of the target speech. That is, the content of the first part of the speech is used as a reference for the discretization process, thereby improving the accuracy of the obtained discrete features of the target speech and thus improving the accuracy of speech restoration.

[0437] In another possible implementation, the fourth module can be used to discretize the first part of the speech based on its content to obtain the discrete features of the target speech. This can improve the efficiency of obtaining the discrete features of the target speech, thereby improving the efficiency of speech restoration.

[0438] It should be noted that the solution in this application embodiment can be used not only for sound restoration tasks, but also extended to tasks such as dialect to Mandarin conversion and cross-language translation.

[0439] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.

[0440] It should also be understood that the steps of the various embodiments described above can be coupled to each other, and this application does not limit this. Furthermore, the sequence number of each process does not imply the order of execution; the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0441] This application also proposes a voice processing device that can execute the steps of the above-described method embodiments. For example, the voice processing device includes a receiving module and an output module, wherein the receiving module is used to receive setting input, obtain the user's first voice feature through voice registration, receive the user's input first voice, etc.; the output module is used to repair the first voice based on the setting input and the first voice feature, etc.

[0442] In another possible implementation, the receiving module is used to receive setting input, receive the fifth voice from the target contact, obtain the second voice features, etc., and the output module is used to repair the fifth voice based on the setting input and the second voice features, etc.

[0443] In another possible implementation, the receiving module is used to acquire the target speech and speech features, the output module is used to input the target speech and speech features into the target restoration model, the target restoration model is used to extract the content of the target speech, obtain speech continuous features based on the content and speech features of the target speech, and synthesize the restored target speech based on the speech continuous features; and the restored target speech output by the target restoration model is acquired.

[0444] The steps performed by the voice processing device can be referred to the description of the method embodiments, and will not be repeated here.

[0445] It should be understood that the steps performed by the apparatus in this application embodiment can be referred to the description of the method embodiment above, and will not be repeated here.

[0446] It should be understood that the device described here is embodied in the form of functional modules. The term "module" here can refer to application-specific integrated circuits (ASICs), electronic circuits, processors (e.g., shared processors, proprietary processors, or group processors) and memories for executing one or more software or firmware programs, integrated logic circuits, and / or other suitable components supporting the described functions. In an alternative example, those skilled in the art will understand that the device may specifically be the first electronic device, the second electronic device, or the fourth electronic device in the above embodiments; or, the functions described in the above embodiments may be integrated into the device, which may be used to execute the various processes and / or steps corresponding to the electronic device or server in the above method embodiments. To avoid repetition, further details are omitted here.

[0447] The aforementioned device has the function of implementing the corresponding steps of the electronic device or server in the aforementioned method; the aforementioned function can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the aforementioned function.

[0448] In embodiments of this application, the device may also be a chip or a chip system, such as a system on chip (SoC).

[0449] This application also provides a schematic block diagram of another voice processing device. The device includes a processor 1601, a transceiver 1602, and a memory 1603. The processor 1601, transceiver 1602, and memory 1603 communicate with each other via an internal connection path. The memory 1603 stores instructions, and the processor 1601 executes the instructions stored in the memory 1603 to control the transceiver 1602 to transmit and / or receive signals.

[0450] It should be understood that the apparatus may specifically be the electronic device or server in the above embodiments, and may be used to execute the various steps and / or processes corresponding to the electronic device or server in the above method embodiments. Optionally, the memory 1603 may include a read-only memory 1603 and a random access memory 1603, and provide instructions and data to the processor 1601. A portion of the memory 1603 may also include a non-volatile random access memory 1603. For example, the memory 1603 may also store device type information. The processor 1601 may be used to execute the instructions stored in the memory 1603, and when the processor 1601 executes the instructions stored in the memory 1603, the processor 1601 is used to execute the various steps and / or processes of the above method embodiments. The transceiver 1602 may include a transmitter and a receiver, the transmitter may be used to implement the various steps and / or processes corresponding to the transceiver 1602 for performing a transmitting action, and the receiver may be used to implement the various steps and / or processes corresponding to the transceiver 1602 for performing a receiving action.

[0451] It should be understood that, in the embodiments of this application, the processor 1601 may be a central processing unit (CPU), or it may be other general-purpose processors 1601, digital signal processors 1601 (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor 1601 may be a microprocessor 1601, or it may be any conventional processor 1601, etc.

[0452] In implementation, each step of the above method can be completed by the integrated logic circuitry of the hardware in the processor 1601 or by instructions in software form. The steps of the method disclosed in the embodiments of this application can be directly manifested as execution by the hardware processor 1601, or as a combination of hardware and software modules in the processor 1601. The software modules can reside in mature storage media in the art, such as random access memory 1603, flash memory, read-only memory 1603, programmable read-only memory 1603, electrically erasable programmable memory 1603, registers, etc. This storage medium is located in memory 1603, and the processor 1601 executes the instructions in memory 1603, combining with its hardware to complete the steps of the above method. To avoid repetition, detailed descriptions are not provided here.

[0453] This application also provides a voice processing system, including a terminal device and a server. The steps performed by the terminal device and the server can be referred to the description of the method embodiments above, and will not be repeated here.

[0454] This application also provides a computer-readable storage medium for storing a computer program that implements the methods shown in the above-described method embodiments.

[0455] This application also provides a computer program product, which includes a computer program (also referred to as code or instructions) that, when run on a computer, allows the computer to execute the methods shown in the above-described method embodiments.

[0456] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0457] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0458] In the embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0459] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of the embodiments of this application, depending on actual needs.

[0460] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0461] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

Claims

1. A speech processing method, characterized in that, include: Receive the user's first voice input; Repair the first audio recording; The second interface is displayed, which includes playback controls for the repaired first audio and the first text corresponding to the repaired first audio. A third operation is received, the third operation being used to select a first target character in the first text; In response to the third operation, at least one candidate character associated with the first target character is displayed; A fourth operation is received, the fourth operation being used to select a second target character from the at least one candidate character; In response to the fourth operation, a playback control for the second text and the corresponding audio is displayed, wherein the second text is obtained by replacing the first target character in the first text with the second target character.

2. The method according to claim 1, characterized in that, In a call scenario, the method further includes: Send a first notification message to the first contact in the call, the first notification message being used to indicate that the voice repair function has been enabled; And / or, Upon receiving the first voice message, a second notification message is sent to the first contact, indicating that the first voice message is being repaired.

3. The method according to claim 2, characterized in that, The second notification message includes a notification tone. Sending the second notification message to the first contact upon receiving the first voice message includes: Upon receiving the first voice message, the notification tone is continuously sent to the first contact. The method further includes: Once the first restored voice message is sent to the first contact, the notification sound is stopped from being sent to the first contact.

4. The method according to any one of claims 1-3, characterized in that, In a call scenario, the method further includes: Send the repaired first voice message to the first contact in the call; The first interface is displayed, which is an interface for making a call with the first contact. The first interface includes a first control, which is used to control the voice repair function to be turned off. In response to a first operation on the first control, the voice repair function is turned off to send a second voice message to the first contact, the second voice message including the user's voice message received after the voice repair function is turned off.

5. The method according to claim 4, characterized in that, In response to the first operation, the first control is also switched from a first state to a second state, wherein the first state indicates that the voice repair function is enabled and the second state indicates that the voice repair function is disabled. The method further includes: In response to a second operation on the first control, the first control is switched from the second state to the first state, and the voice repair function is enabled to send a repaired third voice message to the first contact, the third voice message including the voice message received from the user after the voice repair function is enabled.

6. The method according to any one of claims 1-5, characterized in that, The method further includes: Cancel the display of the playback control for the repaired first voice message.

7. The method according to any one of claims 1-6, characterized in that, In remote communication scenarios, the method further includes: The third interface is displayed, which is an interface for communicating with the second contact. The third interface includes a second control and a third control. The second control is used to instruct the first voice message to be sent to the second contact, and the third control is used to instruct the repaired first voice message to be sent to the second contact. In response to an operation on the third control, the repaired first voice message is sent to the second contact; or... In response to an operation on the second control, the first voice message is sent to the second contact.

8. The method according to any one of claims 1-7, characterized in that, In face-to-face communication scenarios, the method further includes: A fourth interface is displayed, which includes a virtual keyboard; In response to an operation on the virtual keyboard, a third text is displayed; A fifth operation is received for the fourth control in the fourth interface, the fourth control being used to instruct the generation of speech; In response to the fifth operation, the speech corresponding to the third text is generated based on the first speech feature.

9. The method according to claim 8, characterized in that, The third text includes punctuation marks and / or emoticons. When generating the speech corresponding to the third text based on the first speech features, the tone of the speech corresponding to the third text is also controlled based on the punctuation marks and / or emoticons.

10. The method according to any one of claims 1-9, characterized in that, The method further includes: The fifth interface is displayed. The fifth interface is the sound registration interface. The fifth interface includes a third prompt message and a fifth control. The third prompt message is used to prompt the recording content of the sound registration, and the fifth control is used to indicate the recording of voice. In response to an operation on the fifth control, a fourth voice message is recorded; Extract the first speech feature from the fourth speech.

11. A voice processing device, characterized in that, include: A processor coupled to a memory storing computer-executable instructions, the processor executing the computer-executable instructions stored in the memory, causing the processor to perform the method as described in any one of claims 1 to 10.

12. A voice processing system, characterized in that, It includes an electronic device and a server, wherein the electronic device performs the method as described in any one of claims 1 to 10.

13. A computer-readable storage medium, characterized in that, For storing a computer program, the computer program including instructions for implementing the method as claimed in any one of claims 1 to 11, or including instructions for implementing the method as claimed in any one of claims 12 to 14, or including instructions for implementing the method as claimed in claim 15.

14. A computer program product, said computer program product comprising computer program code, characterized in that, When the computer program code is run on a computer, it causes the computer to perform the method as described in any one of claims 1 to 10.