Voice session method, system, device, equipment, storage medium and program product

By receiving and converting the original speech into reference text and combining it with a speech style feature extraction model to generate the target speech, the problem of low speech completion accuracy caused by insufficient base station wireless signal coverage is solved, thereby improving call quality.

CN120808744APending Publication Date: 2025-10-17CHINA TELECOM CORP LTD TECHNOLOGY INNOVATION CENTER +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511064058.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-31
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

In the case of insufficient base station wireless signal coverage, the accuracy of voice completion in the existing technology is low, resulting in poor call quality.

Method used

By receiving the original voice and reference text sent by the second terminal, the target voice is generated, the original voice is converted into the reference text using the automatic speech recognition model, and the target voice is generated based on the comparison results of the proofread text and the reference text. The voice style feature extraction model is combined to ensure the fidelity of the voice style and the integrity of the content.

Benefits of technology

The accuracy of voice completion is improved, ensuring the integrity of call content and the fidelity of voice style, thereby improving the call quality between terminals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120808744A_ABST
    Figure CN120808744A_ABST
Patent Text Reader

Abstract

The invention relates to a voice session method, system and device, equipment, a storage medium and a program product. The method comprises the steps that a first terminal receives an original voice and a reference text sent by a second terminal, generates a target voice according to the original voice and the reference text, and plays the target voice. Wherein the original voice is obtained by triggering a voice session request on a second terminal by a user, and the reference text is generated by converting the original voice. By adopting the method, the accuracy of voice completion can be improved, and the call quality is further improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of communication, and in particular to a voice conversation method, system, device, equipment, storage medium and program product. BACKGROUND

[0002] In the case of insufficient base station wireless signal coverage, the voice between two terminals will be discontinuous and missing. Therefore, it is necessary to complete the voice to ensure the call quality between terminals.

[0003] In related technologies, when completing the voice, a multi-modal model or a time series model is usually used to combine historical chat text, predict the missing voice part through model reasoning, and then combine the original voice to obtain the complete voice.

[0004] However, the accuracy of voice completion in related technologies is low, resulting in poor call quality. SUMMARY

[0005] Therefore, it is necessary to provide a voice conversation method, system, device, equipment, storage medium and program product capable of improving the accuracy of voice completion and improving the call quality.

[0006] In a first aspect, the present application provides a voice conversation method applied to a first terminal, comprising:

[0007] receiving original voice and reference text sent by a second terminal; the original voice is obtained by a user triggering a voice conversation request on the second terminal, and the reference text is generated by converting the original voice;

[0008] generating target voice according to the original voice and the reference text, and playing the target voice.

[0009] In one of the embodiments, generating the target voice according to the original voice and the reference text comprises:

[0010] converting the original voice into proofreading text;

[0011] obtaining the target voice according to the comparison result of the proofreading text and the reference text.

[0012] In one of the embodiments, obtaining the target voice according to the comparison result of the proofreading text and the reference text comprises:

[0013] if the comparison result indicates that the proofreading text and the reference text are inconsistent, performing voice conversion on the reference text based on the voice style of the original voice to obtain the target voice;

[0014] If the comparison result indicates that the proofread text is consistent with the reference text, the original speech corresponding to the proofread text is taken as the target speech.

[0015] In one of the embodiments, the speech conversion is performed on the reference text based on the speech style of the original speech to obtain the target speech, including:

[0016] The speech style of the original speech is determined by performing speech style feature extraction on the original speech based on a speech feature extraction model deployed on the first terminal.

[0017] The speech conversion is performed on the reference text according to the speech style to obtain the target speech.

[0018] In one of the embodiments, the method further includes:

[0019] The channel establishment request sent by the second terminal is received, and a multimedia channel and an application data channel between the first terminal and the second terminal are established based on the channel establishment request.

[0020] The original speech and the reference text sent by the second terminal are received, including:

[0021] The original speech sent by the second terminal is received based on the multimedia channel, and the reference text sent by the second terminal is received based on the application data channel.

[0022] In one of the embodiments, the method further includes:

[0023] In response to the voice session request initiated to the second terminal, the original speech corresponding to the voice session request is obtained, and the original speech is converted into the reference text.

[0024] The original speech and the reference text are sent to the second terminal to instruct the second terminal to generate the target speech according to the received original speech and reference text.

[0025] In a second aspect, the present application further provides a voice session system, including a first terminal and a second terminal.

[0026] The second terminal obtains the original speech in response to the voice session request triggered by the user, and converts the original speech into the reference text.

[0027] The second terminal sends the original speech and the reference text to the first terminal.

[0028] The first terminal generates the target speech according to the original speech and the reference text, and plays the target speech.

[0029] In a third aspect, the present application further provides a voice session device, including:

[0030] The information receiving module is configured to receive original speech and reference text sent by the second terminal, wherein the original speech is obtained by triggering a voice conversation request by a user on the second terminal, and the reference text is generated by converting the original speech.

[0031] The voice playing module is configured to generate target speech according to the original speech and the reference text, and play the target speech.

[0032] In a fourth aspect, the present application further provides a computer device, comprising a memory and a processor, the memory stores a computer program, and the processor implements the method steps of any one of the embodiments of the first aspect when executing the computer program.

[0033] In a fifth aspect, the present application further provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the method steps of any one of the embodiments of the first aspect.

[0034] In a sixth aspect, the present application further provides a computer program product, comprising a computer program, and the computer program is executed by a processor to implement the method steps of any one of the embodiments of the first aspect.

[0035] The voice conversation method, system, device, equipment, storage medium and program product, the first terminal receives the original speech and the reference text sent by the second terminal, then generates the target speech according to the original speech and the reference text, and plays the target speech. Wherein, the original speech is obtained by triggering a voice conversation request by a user on the second terminal, and the reference text is generated by converting the original speech. In this method, the first terminal receives the original speech and the reference text generated by converting the original speech sent by the second terminal, and since the data amount of the reference text is less than that of the original speech, it means that the reference text is less likely to be distorted and damaged in the transmission process, so as to ensure that the first terminal receives complete and accurate voice text information. Then, the first terminal generates the target speech according to the original speech and the reference text, integrates the content integrity of the reference text and the voice style of the original speech, so that the target speech is more accurate and complete in voice content while preserving the voice style, and plays the target speech, which improves the call quality between terminals from two dimensions of content and style. BRIEF DESCRIPTION OF DRAWINGS

[0036] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the related art, the following will briefly introduce the drawings needed to be used in the description of the embodiments of the present application or the related art. Obviously, the drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other related drawings can also be obtained without creative labor on the basis of these drawings.

[0037] Figure 1 It is an application environment diagram of the voice conversation method in an embodiment.

[0038] Figure 2 Flowchart of a voice conversation method in an embodiment;

[0039] Figure 3 Flowchart of a target voice determination step in an embodiment;

[0040] Figure 4 Flowchart of a voice text sending step in an embodiment;

[0041] Figure 5 Signaling interaction diagram of a voice conversation method in an embodiment;

[0042] Figure 6 Structural block diagram of a voice conversation device in an embodiment;

[0043] Figure 7 Internal structural diagram of a computer device in an embodiment. DETAILED DESCRIPTION

[0044] In order to make the purposes, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and not to limit the present application.

[0045] It should be noted that the terms "first", "second", etc. used in the present application can be used to describe various elements, but these elements are not limited by these terms. These terms are only used to distinguish the first element from the second element. The terms "include" and "have" and any variations thereof used in the present application are intended to cover non-exclusive inclusion. The term "multiple" used in the present application refers to two and more than two. The term "and / or" used in the present application refers to one of the options or any combination of multiple options.

[0046] In the field of communication, two terminal devices communicate through network signals, such as voice communication, video communication, etc. However, in the case where the two terminal devices belong to a base station with insufficient wireless signal coverage, it may cause the voice between the two terminals to be discontinuous and missing. Based on this, it is necessary to complete the voice of the terminal.

[0047] In the related art, when completing the voice, the received audio data is usually input into a multi-modal model or a time series neural network model, and the missing voice part is inferred from the multi-modal information, knowledge graph or historical voice text through the model, and then combined with the original voice to restore the complete voice data.

[0048] However, the inference result of the model in the related art deviates greatly from the actual call content, resulting in low accuracy of voice completion, distorted call content between terminals, and poor call quality.

[0049] Based on this, the present application provides a voice conversation method, system, device, equipment, storage medium and program product. The first terminal generates target voice and plays the target voice by using the original voice and the reference text sent by the second terminal, so as to improve the accuracy of voice completion, guarantee the content integrity of the target voice, and further improve the call quality between terminals.

[0050] The voice conversation method provided by the embodiments of the present application can be applied to an application environment as shown in Figure 1 . Figure 1 In the application environment, two terminals realize communication, such as voice call, through a mobile communication network. The terminal can be, but is not limited to, various personal computers, notebook computers, smart phones, tablet computers, unmanned aerial vehicles, low-altitude aircraft, Internet of Things devices, and portable wearable devices. The Internet of Things device can be a smart speaker, a smart television, a smart air conditioner, a smart vehicle device, a projection device, etc. The portable wearable device can be a smart watch, a smart bracelet, a head-mounted device, etc. The head-mounted device can be a virtual reality (VR) device, an augmented reality (AR) device, smart glasses, etc.

[0051] In an exemplary embodiment, as shown in Figure 2 , a voice conversation method is provided. Taking the application of the method to a first terminal as an example, the method includes the following steps:

[0052] S201, receiving original voice and reference text sent by a second terminal; the original voice is obtained by a user triggering a voice conversation request on the second terminal, and the reference text is generated by converting the original voice.

[0053] The first terminal and the second terminal refer to two devices that can perform voice call services with each other. The two terminals can perform high-definition calls based on a 4G network, that is, the terminal supports a voice over long-term evolution (VoLTE) function. The two terminals can also perform high-definition calls based on a 5G network, that is, the terminal supports a voice over new radio (VoNR) function.

[0054] In the embodiments of the present application, the first terminal and the second terminal both have voice receiving and voice sending capabilities, and both support IP Multimedia Subsystem Data Channel (IMS DC) network architecture. The IMS DC session adds a data channel on the basis of a traditional audio / video channel, and can transmit additional data in the process of a call to improve user experience.

[0055] Taking the second terminal sending voice to the first terminal as an example, the information interaction steps of the first terminal and the second terminal are described: a user initiates voice on the second terminal, triggers a voice session request, the second terminal obtains original voice carried in the voice session request, and performs text conversion on the original voice to generate reference text; the second terminal sends the original voice and the reference text to the first terminal; the first terminal receives the original voice and the reference text sent by the second terminal.

[0056] It should be noted that in the case of terminal network quality decline, for example, insufficient base station wireless signal coverage, the original voice and the reference text may be lost in the transmission process, but since the data volume of the reference text is much smaller than that of the original voice, and the transmission bandwidth occupied by the reference text is much smaller than that occupied by the original voice, the present application considers that the reference text does not exist distortion.

[0057] Optionally, the second terminal inputs the original voice into a local Automatic Speech Recognition (ASR) model, converts the original voice into text through the ASR model, and outputs the reference text; the first terminal receives the original voice sent by the second terminal and the reference text output by the ASR model.

[0058] The working principle of the ASR model is: converting the original voice into a digital signal, then pre-processing the converted digital signal, such as noise reduction, endpoint monitoring, etc., then extracting the spectral features of the pre-processed digital signal, and mapping the extracted features to a text sequence to obtain the reference text.

[0059] In actual application scenarios, the second terminal can send the original voice and the reference text to the first terminal synchronously, in which case the first terminal synchronously receives the original voice and the reference text; the second terminal can also send the original voice first, and then send the original voice association identifier and the reference text, in which case the first terminal device receives the original voice first, and then determines the reference text corresponding to the original voice according to the received original voice association identifier.

[0060] S202, generating target voice according to the original voice and the reference text, and playing the target voice.

[0061] The original voice is in the form of audio, including a voice style and literal content. In the embodiment of the present application, the original voice in the form of audio and the reference text are fused to obtain the target voice, and the target voice is played through a playing interface of the first terminal.

[0062] Optionally, the original voice and the reference text are input into a voice generation model, the literal content in the original voice is replaced by the reference text through the voice generation model, and the target voice is output.

[0063] Optionally, the integrity of the original voice is verified according to the reference text, if the verification is passed, the original voice is determined as the target voice, and the target voice is played in the first terminal, if the verification is not passed, the reference text is converted into voice based on the voice style of the original voice to obtain the target voice, and the target voice is played.

[0064] In the embodiment of the present application, the first terminal receives the original voice and the reference text sent by the second terminal, then generates the target voice according to the original voice and the reference text, and plays the target voice. The original voice is obtained by triggering a voice conversation request by a user on the second terminal, and the reference text is generated by conversion of the original voice. In this method, the first terminal receives the original voice sent by the second terminal and the reference text generated by conversion of the original voice, since the data amount of the reference text is less than that of the original voice, it means that the reference text is less likely to be distorted and damaged in the transmission process, so as to ensure that the first terminal receives complete and accurate voice text information. Then, the first terminal generates the target voice according to the original voice and the reference text, integrates the content integrity of the reference text and the voice style of the original voice, so that the target voice is more accurate and complete in voice content while the voice style is faithful, and the target voice is played, thereby improving the call quality between terminals from two dimensions of content and style.

[0065] As known from the foregoing embodiments, after the first terminal receives the original voice and the reference text, the original voice and the reference text can be directly fused to generate the target voice, or the received original voice can be verified for integrity, and the target voice is generated according to the verification result. Based on this, the steps of verifying the received original voice by the first terminal are described through an embodiment.

[0066] In an exemplary embodiment, as shown in Figure 3 the target voice is generated according to the original voice and the reference text, including:

[0067] S301, text conversion is performed on the original voice to obtain a proofreading text.

[0068] The first terminal device inputs the original voice into an automatic speech recognition model, performs text conversion on the original voice through the automatic speech recognition model for preprocessing, and outputs the proofreading text.

[0069] Optionally, the original speech is converted into a digital signal through an automatic speech recognition model, and then the converted digital signal is preprocessed, such as noise reduction, endpoint monitoring, etc., then the spectral features of the preprocessed digital signal are extracted, and the extracted features are mapped into a text sequence to obtain the proofread text.

[0070] S302, according to the comparison result of the proofread text and the reference text, the target speech is obtained.

[0071] The proofread text and the reference text are compared word by word, if the proofread text and the reference text are completely matched, it is determined that the proofread text and the reference text are consistent, and the original speech corresponding to the proofread text is determined as the target speech, otherwise, it is determined that the proofread text and the reference text are inconsistent, and then the reference text and the original speech need to be fused to generate the target speech.

[0072] In another comparison scenario, the proofread features of the proofread text and the reference features of the reference text can also be extracted through a semantic analysis model, and then the feature similarity of the proofread features and the reference features is obtained, if the feature similarity is greater than a preset similarity threshold, it is determined that the proofread text and the reference text are consistent, and the original speech corresponding to the proofread text is determined as the target speech, otherwise, it is determined that the proofread text and the reference text are inconsistent, and then the reference text and the original speech need to be fused to generate the target speech.

[0073] In the embodiments of the present application, the original speech is converted into a text to obtain a proofread text, and according to the comparison result of the proofread text and the reference text, a target speech is obtained, and the reference text is used as a reference to verify the integrity of the original speech received by the first terminal, so as to obtain an accurate comparison result and provide a reliable basis for the target speech generation strategy.

[0074] In an exemplary embodiment, the aforementioned S302 "according to the comparison result of the proofread text and the reference text, the target speech is obtained", includes:

[0075] If the comparison result indicates that the proofread text and the reference text are inconsistent, the reference text is speech converted based on the speech style of the original speech to obtain the target speech.

[0076] The comparison result indicates that the proofread text and the reference text are inconsistent, which means that there is content loss in the process of transmitting the original speech from the second terminal to the first terminal, then the reference text is determined as the text content of the target speech, and the reference text is speech converted based on the speech style of the original speech to obtain the target speech.

[0077] If the comparison result indicates that the proofread text and the reference text are consistent, the original speech corresponding to the proofread text is taken as the target speech.

[0078] The comparison result indicates that the proofread text is consistent with the benchmark text, meaning that no content loss occurs in the process of transmitting the original speech from the second terminal to the first terminal, and thus the original speech corresponding to the proofread text, i.e., the original speech received by the first terminal, is directly determined as the target speech.

[0079] In the case where the proofread text is inconsistent with the benchmark text, the speech style of the original speech is used to perform speech conversion on the benchmark text, so that the content of the target speech is completely matched with the benchmark text, and the integrity and accuracy of the target speech are improved. In the case where the proofread text is consistent with the benchmark text, the original speech corresponding to the proofread text is used as the target speech, so that the authenticity of the target speech is ensured, and the computational burden of speech fusion is reduced.

[0080] In an exemplary embodiment, the speech style of the original speech is used to perform speech conversion on the benchmark text to obtain the target speech, including:

[0081] The speech feature extraction model deployed on the first terminal is used to perform speech style feature extraction on the original speech to determine the speech style of the original speech, and the benchmark text is converted according to the speech style to obtain the target speech.

[0082] The input of the speech feature extraction model is multi-modal data (original speech and benchmark text), and the output is the target speech. The working principle of the speech feature extraction model is as follows: the speech style features of the original speech are extracted, which are used to describe the speaking mode of the original speech, such as volume, pitch, intensity, speed, emotion, timbre, etc.; and then the speech style features are migrated to the benchmark text to output the target speech.

[0083] Optionally, the speech feature extraction model includes a style extraction module and a style conversion module. The input of the style extraction module is the original speech, and the output is the speech style. The input of the style conversion module is the speech style (output of the style extraction module) and the benchmark text, and the output is the target speech.

[0084] In the embodiment of the present application, the speech feature extraction model deployed on the first terminal is used to perform speech style feature extraction on the original speech to determine the speech style of the original speech, and the benchmark text is converted according to the speech style to generate a target speech that conforms to the content of the benchmark text and retains the original speech style. In addition, compared with the traditional pipeline model, the end-to-end speech generation method of the speech feature extraction model in the embodiment of the present application avoids artificial intermediate design steps and reduces system complexity.

[0085] The target speech is generated according to the original speech and the benchmark text. Next, an implementable manner of receiving the original speech and the benchmark text by the first terminal is described.

[0086] In an example embodiment, the method further comprises:

[0087] receiving a channel establishment request sent by the second terminal, and establishing the multimedia channel and the application data channel with the second terminal based on the channel establishment request.

[0088] The channel establishment request is sent by the second terminal to the first terminal before the second terminal and the first terminal perform the call service.

[0089] For example, the second terminal initiatively sends the voice to the first terminal, the second terminal and the first terminal establish the bootstrap data channel (BDC) of the IMS respectively, and the second terminal and the first terminal download the IMS DC applet list. The second terminal starts the "voice completion" applet first, and then initiates the IMS DC peer-to-peer type application data channel (ADC) with the second terminal. After the second terminal and the first terminal establish the application data channel, the "voice completion" applet is also started.

[0090] On the basis of the first terminal and the second terminal establishing the multimedia channel and the application data channel, the original voice and the reference text sent by the second terminal are received, comprising:

[0091] The original voice sent by the second terminal is received based on the multimedia channel, and the reference text sent by the second terminal is received based on the application data channel.

[0092] The first terminal and the second terminal both include a multimedia interface and an application data interface. The first terminal and the second terminal transmit multimedia data such as voice through the respective multimedia interfaces, and transmit data other than the call such as the reference text through the respective application data interfaces.

[0093] The second terminal sends the original voice to the first terminal through the multimedia channel, and the first terminal acquires the original voice transmitted by the multimedia channel through the multimedia interface. The second terminal sends the reference text to the first terminal through the application data channel, and the first terminal acquires the reference text sent by the application data channel through the application data interface.

[0094] In the embodiment of the application, based on the channel establishment request sent by the second terminal, the multimedia channel and the application data channel with the second terminal are established, the original voice is received through the multimedia channel, the reference text is received through the application data channel, the voice and the text are decoupled and transmitted, and the flexibility of the transmission mode is improved.

[0095] It is emphasized again that the first terminal and the second terminal have both the ability to receive voice and the ability to send voice. The foregoing embodiment details the step of generating target voice by the first terminal in the scenario that the first terminal receives voice and text data sent by the second terminal. Next, the voice sending step of the first terminal in the scenario that the first terminal sends voice to the second terminal is described.

[0096] In one exemplary embodiment, as shown in Figure 4 the method further comprises:

[0097] S401, in response to the voice session request initiated to the second terminal, obtaining original voice corresponding to the voice session request, and converting the original voice into reference text.

[0098] The user initiates voice at the first terminal and triggers a voice session request. The first terminal obtains original voice carried in the voice session request and performs text conversion on the original voice carried in the voice session request to generate reference text.

[0099] In an actual application scenario, the first terminal can locally deploy an automatic speech recognition model to perform text conversion on the original voice by using the automatic speech recognition model to output reference text. For example, the original voice is converted into a digital signal, and then the converted digital signal is preprocessed, such as noise reduction, endpoint monitoring, and the like. Then, the spectral features of the preprocessed digital signal are extracted, and the extracted features are mapped into a text sequence to obtain the reference text.

[0100] S402, sending the original voice and the reference text to the second terminal to instruct the second terminal to generate target voice according to the received original voice and reference text.

[0101] The first terminal establishes a multimedia channel and a data transmission channel with the second terminal. The first terminal sends the original voice to the second terminal through the multimedia channel, and the first terminal sends the reference text through the application data channel. In this way, the second terminal generates target voice according to the received original voice and reference text. The implementation of the second terminal to generate the original voice can refer to the description of the first terminal to generate target voice in the foregoing embodiment, and will not be described here.

[0102] It is emphasized that the original voice, the reference text, and the target voice in the embodiments of the present application are obtained or generated by taking the first terminal as a voice sender, which is different from the original voice, the reference text, and the target voice obtained by taking the first terminal as a voice receiver in the foregoing Figure 2 、 Figure 3 embodiments. For ease of understanding, the original voice, the reference text, and the target voice in the foregoing Figure 2 、 Figure 3The original voice, the reference text and the target voice shown are respectively understood as the first original voice, the first reference text and the first target voice, and the original voice, the reference text and the target voice in the embodiment of the application are respectively understood as the second original voice, the second reference text and the second target voice.

[0103] In the embodiment of the application, in response to the voice conversation request initiated to the second terminal, the original voice corresponding to the voice conversation request is acquired, the original voice is converted into the reference text, and the original voice and the reference text are sent to the second terminal to instruct the second terminal to generate the target voice according to the received original voice and reference text, which is equivalent to considering the case that the original voice is missing due to signal discontinuity, instructing the second terminal to restore and complete the original voice based on the reference text to improve user experience and also reduce the network optimization cost of the operator.

[0104] In one exemplary embodiment, a voice conversation system is provided, which includes a first terminal and a second terminal; wherein the second terminal acquires an original voice in response to a voice conversation request triggered by a user, and converts the original voice into a reference text; the second terminal sends the original voice and the reference text to the first terminal; the first terminal generates a target voice according to the original voice and the reference text, and plays the target voice.

[0105] Taking a voice conversation between a terminal A and a terminal B as an example, please refer to Figure 5 , Figure 5 The signaling interaction diagram for the voice conversation method includes the following steps:

[0106] S501, the terminal A and the terminal B establish a voice call.

[0107] The terminal A, the terminal B and the IMS network all support the IMS DC feature. The terminal A and the terminal B establish a VoLTE / VoNR voice call of the operator mobile communication network.

[0108] S502, a guided data transmission channel is established.

[0109] The terminal A and the IMS network, and the terminal B and the IMS network respectively establish a BDC channel of the IMS DC, and download an IMS DC applet list.

[0110] S503, the terminal A starts a voice completion applet.

[0111] S504, an application data channel is established.

[0112] The terminal A initiates a request for establishing an application data channel of the IMS DC end-to-end type with the terminal B.

[0113] S505, the terminal B starts a voice completion applet.

[0114] After terminal B establishes an ADC channel with terminal A, terminal B also starts the voice completion applet.

[0115] S506, the voice completion applet of terminal A starts the automatic speech recognition model of the local terminal to recognize the local voice in real time and convert it into a reference text.

[0116] S507, terminal A sends the reference text and the local voice to terminal B.

[0117] S508, the voice completion applet of terminal B starts the automatic speech recognition model of the local terminal to recognize the opposite voice and convert it into a proofreading text.

[0118] S509, terminal B compares the proofreading text of the opposite party with the reference text.

[0119] S510, in the case of missing in the proofreading text, the voice completion applet of terminal B extracts the missing text from the reference text of the opposite party, and starts the text-to-speech conversion model of the local terminal to convert the missing text into voice and play it on the local terminal, so that the user of terminal B can hear it, thereby realizing voice completion.

[0120] In the embodiments of the present application, with the channel capability of the operator IMS DC, in the poor network condition of low bandwidth and intermittent voice, a small amount of call text is transmitted to the opposite end as the proofreading reference of the call content, so that the opposite end can identify the missing call content, and the local small model of ASR and TTS is combined to complete the voice completion.

[0121] On the one hand, the complete call text of both parties is exchanged as the proofreading reference, the opposite voice is converted into text by the local ASR, the two texts are compared, the missing part is accurately found out, and then the missing voice is restored by TTS. Compared with the traditional scheme of completing by AI reasoning, the present scheme is more accurate.

[0122] On the other hand, when the voice is intermittent, it indicates that the network quality is declining. At this time, the voice completion scheme based on the cloud-side AI model in the traditional technology is not ideal, but the data amount of the text is much smaller than that of the voice, so it can still be transmitted with high probability. The embodiments of the present application ingeniously use this feature as a breakthrough, transmit the short text to the opposite party through the IMS DC channel as the proofreading reference, and thus realize the complete voice completion scheme.

[0123] In summary, compared with the existing inference scheme using cloud-side artificial intelligence large model, the voice completion strategy of the voice conversation scheme provided by the embodiments of the present application is more accurate, has smaller time delay, and has lower cost.

[0124] It should be understood that although the steps in the flowcharts involved in the embodiments described above are shown in sequence according to the arrows, the steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, the execution of the steps is not strictly limited in sequence, and the steps can be executed in other orders. Moreover, at least some of the steps in the flowcharts involved in the embodiments described above can include multiple steps or multiple stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution order of the steps or stages is not necessarily sequential, but can be alternately or alternately executed with at least part of other steps or steps or stages in other steps. It can be understood that the steps in different embodiments can be freely combined as needed, and various non-contradictory schemes formed by the combination are within the scope of protection of the present application.

[0125] Based on the same inventive concept, the embodiments of the present application also provide a voice conversation device for implementing the voice conversation method described above. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme described in the above method, so the specific limitations in one or more voice conversation device embodiments provided below can refer to the limitations of the voice conversation method described above, and will not be repeated here.

[0126] In an exemplary embodiment, as shown in Figure 6 a voice conversation device is provided, comprising: an information receiving module 601 and a voice playing module 602, wherein:

[0127] The information receiving module 601 is configured to receive the original voice and the reference text sent by the second terminal, the original voice is obtained by triggering the voice conversation request by the user on the second terminal, and the reference text is generated by converting the original voice;

[0128] The voice playing module 602 is configured to generate target voice according to the original voice and the reference text, and play the target voice.

[0129] In an exemplary embodiment, the voice playing module 602 comprises a text conversion unit and a voice generation unit, wherein:

[0130] The text conversion unit is configured to perform text conversion on the original voice to obtain the proofreading text;

[0131] The voice generation unit is configured to obtain the target voice according to the comparison result of the proofreading text and the reference text.

[0132] In an exemplary embodiment, the voice generation unit comprises a first generation subunit and a second generation subunit, wherein:

[0133] The first generating sub-unit is configured to perform speech conversion on the reference text based on a speech style of the original speech to obtain target speech, in a case where the comparison result indicates that the proofread text is inconsistent with the reference text.

[0134] The second generating sub-unit is configured to take the original speech corresponding to the proofread text as the target speech, in a case where the comparison result indicates that the proofread text is consistent with the reference text.

[0135] In an exemplary embodiment, the first generating sub-unit is further configured to perform speech feature extraction on the original speech according to a speech feature extraction model deployed on the first terminal to determine the speech style of the original speech, and perform speech conversion on the reference text according to the speech style to obtain the target speech.

[0136] In an exemplary embodiment, the voice session apparatus further includes a channel establishing module configured to receive a channel establishing request sent by the second terminal and establish a multimedia channel and an application data channel between the second terminal based on the channel establishing request.

[0137] The information receiving module 601 is further configured to receive the original speech sent by the second terminal based on the multimedia channel and receive the reference text sent by the second terminal based on the application data channel.

[0138] In an exemplary embodiment, the voice session apparatus further includes a request responding module and an information sending module, wherein:

[0139] The request responding module is configured to, in response to a voice session request initiated to the second terminal, acquire original speech corresponding to the voice session request and convert the original speech into reference text.

[0140] The information sending module is configured to send the original speech and the reference text to the second terminal to instruct the second terminal to generate target speech according to the received original speech and reference text.

[0141] The above voice session apparatus includes various modules, which can be all or partially implemented by software, hardware, and combinations thereof. The above modules can be embedded in or independent of a processor in a computer device in a hardware form, or stored in a memory in a computer device in a software form, so as to be called and executed by a processor to perform operations corresponding to the above modules.

[0142] In an exemplary embodiment, a computer device is provided, which can be a terminal, and an internal structure diagram of the computer device can be as shown in Figure 7The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit and an input device. The processor, the memory and the input / output interface are connected through a system bus, and the communication interface, the display unit and the input device are connected to the system bus through the input / output interface. The processor of the computer device is configured to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for running the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is configured to exchange information between the processor and external devices. The communication interface of the computer device is configured to perform wired or wireless communication with external terminals. The wireless communication can be achieved through WIFI, mobile cellular network, near field communication (NFC) or other technologies. The computer program is executed by the processor to implement a voice conversation method. The display unit of the computer device is configured to form a visually visible picture, which can be a display screen, a projection device or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer overlaid on the display screen, or a key, trackball or touchpad arranged on the shell of the computer device, or an external keyboard, touchpad or mouse, etc.

[0143] Those skilled in the art can understand that Figure 7 The structure shown in the figure is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device to which the scheme of the present application is applied. The specific computer device can include more or fewer components than those shown in the figure, or combine certain components, or have a different component arrangement.

[0144] In one exemplary embodiment, a computer device is provided, including a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the following steps:

[0145] Receiving original speech and reference text sent by the second terminal, the original speech being obtained by a user triggering a voice conversation request on the second terminal, and the reference text being generated by converting the original speech;

[0146] Generating target speech according to the original speech and the reference text, and playing the target speech.

[0147] In one exemplary embodiment, the processor executing the computer program further implements the following steps:

[0148] Converting the original speech into text to obtain proofreading text;

[0149] According to a comparison result of the proofread text and the benchmark text, the target speech is obtained.

[0150] In an exemplary embodiment, the processor, when executing the computer program, further implements the following steps:

[0151] If the comparison result indicates that the proofread text and the benchmark text are inconsistent, the benchmark text is converted into speech based on the speech style of the original speech, and the target speech is obtained;

[0152] If the comparison result indicates that the proofread text and the benchmark text are consistent, the original speech corresponding to the proofread text is taken as the target speech.

[0153] In an exemplary embodiment, the processor, when executing the computer program, further implements the following steps:

[0154] According to a speech feature extraction model deployed on the first terminal, speech style features of the original speech are extracted to determine the speech style of the original speech;

[0155] The benchmark text is converted into speech according to the speech style, and the target speech is obtained.

[0156] In an exemplary embodiment, the processor, when executing the computer program, further implements the following steps:

[0157] A channel establishment request sent by the second terminal is received, and a multimedia channel and an application data channel between the first terminal and the second terminal are established based on the channel establishment request;

[0158] The original speech sent by the second terminal is received based on the multimedia channel, and the benchmark text sent by the second terminal is received based on the application data channel.

[0159] In an exemplary embodiment, the processor, when executing the computer program, further implements the following steps:

[0160] In response to a voice session request initiated to the second terminal, the original speech corresponding to the voice session request is obtained, and the original speech is converted into the benchmark text;

[0161] The original speech and the benchmark text are sent to the second terminal to instruct the second terminal to generate the target speech according to the received original speech and the benchmark text.

[0162] In an exemplary embodiment, the processor, when executing the computer program, further implements the following steps:

[0163] The original speech and the benchmark text sent by the second terminal are received; the original speech is obtained by triggering a voice session request on the second terminal by a user, and the benchmark text is generated by conversion from the original speech;

[0164] The target speech is generated according to the original speech and the reference text, and the target speech is played.

[0165] In one example embodiment, the computer program, when executed by the processor, further implements the following steps:

[0166] The original speech is text-converted to obtain a proofread text;

[0167] The target speech is obtained according to a comparison result of the proofread text and the reference text.

[0168] In one example embodiment, the computer program, when executed by the processor, further implements the following steps: if the comparison result indicates that the proofread text and the reference text are inconsistent, performing speech conversion on the reference text based on a speech style of the original speech to obtain the target speech;

[0169] If the comparison result indicates that the proofread text and the reference text are consistent, the original speech corresponding to the proofread text is taken as the target speech.

[0170] In one example embodiment, the computer program, when executed by the processor, further implements the following steps:

[0171] The speech style of the original speech is determined by performing speech style feature extraction on the original speech according to a speech feature extraction model deployed on the first terminal;

[0172] The target speech is obtained by performing speech conversion on the reference text according to the speech style.

[0173] In one example embodiment, the computer program, when executed by the processor, further implements the following steps:

[0174] A multimedia channel and an application data channel between the second terminal are established based on the channel establishment request sent by the second terminal;

[0175] The original speech sent by the second terminal is received based on the multimedia channel, and the reference text sent by the second terminal is received based on the application data channel.

[0176] In one example embodiment, the computer program, when executed by the processor, further implements the following steps:

[0177] In response to a voice session request initiated to the second terminal, the original speech corresponding to the voice session request is obtained, and the original speech is converted into the reference text;

[0178] The original speech and the reference text are sent to the second terminal to instruct the second terminal to generate the target speech according to the received original speech and the reference text.

[0179] In one example embodiment, a computer program product is provided, comprising a computer program which, when executed by a processor, implements the following steps:

[0180] receiving original speech and reference text sent by the second terminal; the original speech is obtained by a user triggering a speech session request on the second terminal, and the reference text is generated by conversion from the original speech;

[0181] generating target speech according to the original speech and the reference text, and playing the target speech.

[0182] In one example embodiment, the computer program, when executed by the processor, further implements the following steps:

[0183] performing text conversion on the original speech to obtain proofread text;

[0184] obtaining target speech according to a comparison result of the proofread text and the reference text.

[0185] In one example embodiment, the computer program, when executed by the processor, further implements the following steps: if the comparison result indicates that the proofread text and the reference text are inconsistent, performing speech conversion on the reference text based on a speech style of the original speech to obtain the target speech;

[0186] if the comparison result indicates that the proofread text and the reference text are consistent, taking original speech corresponding to the proofread text as the target speech.

[0187] In one example embodiment, the computer program, when executed by the processor, further implements the following steps:

[0188] performing speech style feature extraction on the original speech according to a speech feature extraction model deployed on the first terminal to determine a speech style of the original speech;

[0189] performing speech conversion on the reference text according to the speech style to obtain the target speech.

[0190] In one example embodiment, the computer program, when executed by the processor, further implements the following steps:

[0191] receiving a channel establishment request sent by the second terminal, and establishing a multimedia channel and an application data channel with the second terminal based on the channel establishment request;

[0192] receiving original speech sent by the second terminal based on the multimedia channel, and receiving reference text sent by the second terminal based on the application data channel.

[0193] In one example embodiment, the computer program, when executed by the processor, further implements the following steps:

[0194] In response to the voice session request initiated to the second terminal, original voice corresponding to the voice session request is acquired, and the original voice is converted into reference text;

[0195] The original voice and the reference text are sent to the second terminal to instruct the second terminal to generate target voice according to the received original voice and the reference text.

[0196] It should be noted that the user information (including but not limited to user equipment information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or authorized by all parties, and the collection, use and processing of related data need to comply with relevant regulations.

[0197] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when executed, can include the processes of the above-mentioned embodiment methods. Any reference to memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. The non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical storage, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. The volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration but not limitation, the RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The database involved in the embodiments provided in the present application can include at least one of a relational database and a non-relational database. The non-relational database can include a distributed database based on a block chain, etc., without being limited thereto. The processor involved in the embodiments provided in the present application can be a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, an artificial intelligence (AI) processor, etc., without being limited thereto.

[0198] The technical features of the above embodiments can be combined in any manner. To make the description concise, not all possible combinations of the technical features in the above embodiments are described, but as long as the combinations of the technical features do not exist contradictions, they should be considered as the scope of the present application.

[0199] The above-described embodiments are merely illustrative of several embodiments of the present application, and the description is relatively specific and detailed, but should not be understood as a limitation on the scope of the patent. It should be noted that for those skilled in the art, without departing from the concept of the present application, a number of modifications and improvements can be made, which are all within the scope of the present application. Therefore, the scope of protection of the present application should be subject to the appended claims.

Claims

1. A voice conversation method, characterized in that: Applied to a first terminal, the method includes: Receiving an original voice and a reference text sent by a second terminal; the original voice is obtained by a user triggering a voice session request on the second terminal, and the reference text is generated by converting the original voice; A target voice is generated according to the original voice and the reference text, and the target voice is played.

2. The method according to claim 1, characterized in that Generating the target speech according to the original speech and the reference text includes: Converting the original speech into text to obtain a proofread text; The target speech is obtained according to the comparison result between the proofread text and the reference text.

3. The method according to claim 2, characterized in that Obtaining the target speech according to the comparison result between the proofread text and the reference text includes: If the comparison result indicates that the proofread text and the reference text are inconsistent, performing voice conversion on the reference text based on the voice style of the original voice to obtain the target voice; If the comparison result indicates that the proofread text is consistent with the reference text, the original speech corresponding to the proofread text is used as the target speech.

4. The method according to claim 3, characterized in that The performing speech conversion on the reference text based on the speech style of the original speech to obtain the target speech includes: extracting speech style features of the original speech according to a speech feature extraction model deployed on the first terminal to determine the speech style of the original speech; The reference text is voice-converted according to the voice style to obtain the target voice.

5. The method according to any one of claims 1 to 4, characterized in that The method further comprises: receiving a channel establishment request sent by the second terminal, and establishing a multimedia channel and an application data channel with the second terminal based on the channel establishment request; The receiving the original voice and the reference text sent by the second terminal includes: The original voice sent by the second terminal is received based on the multimedia channel, and the reference text sent by the second terminal is received based on the application data channel.

6. The method according to any one of claims 1 to 4, characterized in that The method further comprises: In response to a voice conversation request initiated to the second terminal, obtaining original speech corresponding to the voice conversation request, and converting the original speech into a reference text; The original speech and the reference text are sent to the second terminal to instruct the second terminal to generate a target speech according to the received original speech and the reference text.

7. A voice conversation system, characterized in that: The system includes a first terminal and a second terminal; The second terminal obtains original speech in response to a voice conversation request triggered by the user, and converts the original speech into a reference text; The second terminal sends the original voice and the reference text to the first terminal; The first terminal generates a target voice according to the original voice and the reference text, and plays the target voice.

8. A voice conversation device, characterized in that: The device comprises: An information receiving module, configured to receive an original voice and a reference text sent by a second terminal; the original voice is obtained by a user triggering a voice conversation request on the second terminal, and the reference text is generated by converting the original voice; The speech playing module is used to generate a target speech according to the original speech and the reference text, and play the target speech.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

11. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.