Voice message processing method, device and electronic equipment

By correcting key voice messages using a trained voice correction model, the target voice message can be obtained, solving the problem of users having to click through long voice messages one by one, thus improving the efficiency and accuracy of information acquisition.

CN115132206BActive Publication Date: 2026-05-05VIVO MOBILE COMM CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
VIVO MOBILE COMM CO LTD
Filing Date
2022-06-28
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

When users receive multiple long voice messages, they need to open them one by one to listen to their content, resulting in low information acquisition efficiency.

Method used

By using a trained speech correction model, speech correction is performed on the key speech message based on the similarity between the key speech message and the speech message to be processed, and the target speech message is obtained. The target speech message has the speech characteristics of the speech message to be processed and the voiceprint characteristics of the target object.

Benefits of technology

It improves the efficiency of users obtaining voice message content, allows users to perceive information such as the tone, intonation, and speaking speed of the voice message sender, and makes the target voice message closer to the original voice, thus improving the accuracy of voice message processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115132206B_ABST
    Figure CN115132206B_ABST
Patent Text Reader

Abstract

The application discloses a voice message processing method and device and electronic equipment, and belongs to the technical field of computers. The method comprises the following steps: obtaining a voice message to be processed; determining a key voice message corresponding to the voice message to be processed; performing voice correction on the key voice message based on the similarity between the key voice message and the voice message to be processed by using a trained voice correction model, so as to obtain a target voice message; wherein the trained voice correction model is obtained by fine-tuning training of sample voice of a target object; the target object is a message sending object corresponding to the voice message to be processed; and the target voice message has voice characteristics of the voice message to be processed and voiceprint characteristics of the target object.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of computer technology, and specifically relates to a voice message processing method, apparatus and electronic device. Background Technology

[0002] With the continuous advancement of technology, people are using electronic devices more and more frequently. When people communicate, they often use the voice messaging function in some applications. Voice messaging brings great convenience, and it also has the characteristics of being vivid and having strong user features, with less information loss during transmission.

[0003] However, during use, when users receive multiple long voice messages, they need to open them one by one to listen to their content. Since there is little useful information in the voice messages, the efficiency of obtaining information from the voice messages is low. Summary of the Invention

[0004] This application provides a voice message processing method, apparatus, and electronic device, which can solve the problem that existing methods require listening to a large number of long voice messages one by one, resulting in a long time consumption and low efficiency in obtaining information from voice messages.

[0005] In a first aspect, embodiments of this application provide a voice message processing method, the method comprising:

[0006] Retrieve voice messages to be processed;

[0007] Identify the key voice message corresponding to the voice message to be processed;

[0008] By using a trained speech correction model, the key speech message is corrected based on the similarity between the key speech message and the speech message to be processed, and the target speech message is obtained.

[0009] The trained speech correction model is obtained by fine-tuning the training using sample speech of the target object; the target object is the message sending object corresponding to the speech message to be processed; the target speech message has the speech characteristics of the speech message to be processed and the voiceprint characteristics of the target object.

[0010] Secondly, embodiments of this application provide a voice message processing apparatus, the apparatus comprising:

[0011] The acquisition module is used to acquire voice messages to be processed.

[0012] The determination module is used to determine the key voice message corresponding to the voice message to be processed;

[0013] The correction module is used to correct the key speech message based on the similarity between the key speech message and the speech message to be processed using a trained speech correction model, so as to obtain the target speech message.

[0014] The trained speech correction model is obtained by fine-tuning the training using sample speech of the target object; the target object is the message sending object corresponding to the speech message to be processed; the target speech message has the speech characteristics of the speech message to be processed and the voiceprint characteristics of the target object.

[0015] Thirdly, embodiments of this application provide an electronic device including a processor, a memory, and a program or instructions stored in the memory and executable on the processor, wherein the program or instructions, when executed by the processor, implement the steps of the method described in the first aspect.

[0016] Fourthly, embodiments of this application provide a readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method described in the first aspect.

[0017] Fifthly, embodiments of this application provide a chip, the chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the method as described in the first aspect.

[0018] In this embodiment, the process first acquires the voice message to be processed, then determines the key voice message corresponding to the voice message to be processed, and finally, by fine-tuning the trained voice correction model obtained from the sample voice of the target object, the key voice message is corrected based on the similarity between the key voice message and the voice message to be processed, thereby obtaining a target voice message with the voice characteristics of the voice message to be processed and the voiceprint characteristics of the target object. This embodiment obtains the key voice message by processing at least one voice message, and then corrects the key voice message using the trained voice correction model. This results in a target voice message with the voice characteristics of the voice message to be processed and the voiceprint characteristics of the target object. This improves the efficiency of users acquiring voice message content and also allows users to perceive the emotions of the voice message sender, making the target voice message heard by the user closer to the original voice message, more like a simplified and corrected version of the original voice message by the voice message sender, thus improving the accuracy of voice message processing. Attached Figure Description

[0019] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0020] Figure 1 This is a flowchart illustrating a voice message processing method provided in one embodiment of this application;

[0021] Figure 2 This is a schematic diagram illustrating the acquisition of a voice message to be processed, provided in one embodiment of this application.

[0022] Figure 3 This is a schematic diagram illustrating another method for acquiring voice messages to be processed, provided in one embodiment of this application.

[0023] Figure 4 This is an operational diagram illustrating the process of entering the voice message processing interface, provided in one embodiment of this application.

[0024] Figure 5 This is a schematic diagram of the structure of a speech correction model provided in one embodiment of this application;

[0025] Figure 6 This is a schematic diagram of the structure of a voice network provided in one embodiment of this application;

[0026] Figure 7 This is a simplified structural diagram of the input and output during the training of a speech correction model provided in one embodiment of this application;

[0027] Figure 8 This is a schematic diagram of the conversation interface after voice message processing is completed, provided in one embodiment of this application;

[0028] Figure 9 This is a schematic diagram of the overall flow of a voice message processing method provided in one embodiment of this application;

[0029] Figure 10 This is a schematic diagram of the structure of a voice message processing device provided in one embodiment of this application;

[0030] Figure 11 This is a schematic diagram of the structure of an electronic device provided in one embodiment of this application;

[0031] Figure 12 This is a schematic diagram of the hardware structure of an electronic device provided in one embodiment of this application. Detailed Implementation

[0032] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.

[0033] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0034] In some embodiments, when a user receives multiple long voice messages from a message sender, they need to open them one by one to listen to the content, which takes a long time. Moreover, since there may be few useful messages in these long voice messages, the efficiency of the user in obtaining information from these voice messages is low. To solve the above problems, this solution proposes a voice message processing method, apparatus, and electronic device. By converting at least one voice message sent by a message sender into a target voice message, the target voice message has the voiceprint characteristics of the message sender and has the same voice characteristics as the at least one voice message. For example, if the at least one voice message sent by the message sender is in a happy tone, the resulting target voice message is also in a happy tone. Finally, the target voice message is displayed in the conversation interface, so that the user hears a concise voice message with the voiceprint characteristics of the message sender and the same voice characteristics as the at least one voice message, reducing the time spent by the user and improving the efficiency of voice information acquisition.

[0035] The following description, in conjunction with the accompanying drawings, details a voice message processing method, apparatus, and electronic device provided in this application through specific embodiments and application scenarios.

[0036] like Figure 1 The diagram shown is a flowchart illustrating a voice message processing method according to an embodiment of this application. The voice message processing method may include the contents shown in S101 to S103.

[0037] In S101, the voice message to be processed is obtained.

[0038] Here, the voice message to be processed refers to the voice message sent by the message sender, and includes at least one voice message. If the voice message to be processed includes multiple voice messages, the voice messages to be processed can be obtained by selecting them one by one; or by selecting multiple voice messages by swiping, such as... Figure 2 As shown, you can select multiple voice messages sent by user 1 by sliding; you can also select multiple voice messages sent to that recipient by selecting the message recipient, such as... Figure 3 As shown, multiple voice messages sent by user 1 can be selected by selecting user 1's avatar. Selecting user 1's avatar can be done by long-pressing user 1's avatar or by double-clicking user 1's avatar, etc. The specific method is not limited in this embodiment and can be determined according to the actual application.

[0039] The above method can be used to obtain the voice messages to be processed. Users can then swipe left to enter the voice message processing interface and convert the selected voice message into the target voice message, such as... Figure 4 The diagram shows how to slide to the left to enter the voice message processing interface. The specific voice message processing flow is as follows.

[0040] In S102, the key voice message corresponding to the voice message to be processed is determined.

[0041] Among them, the key voice message refers to the voice message obtained after the voice message to be processed has been processed by voice editing. This voice message can correct the errors in the voice message to be processed. When there are many voice messages (such as voice messages with long duration and many voice messages), the key content can be extracted, that is, a concise voice message is obtained. The specifics will be described in detail in subsequent embodiments, and will not be repeated in this embodiment.

[0042] In S103, the trained speech correction model corrects the speech of the key speech message based on the similarity between the key speech message and the speech message to be processed, thereby obtaining the target speech message.

[0043] The trained speech correction model is obtained by fine-tuning the training using sample speech of the target object; the target object is the message sender corresponding to the speech message to be processed; the target speech message has the speech characteristics of the speech message to be processed and the voiceprint characteristics of the target object.

[0044] It's worth noting that the sample voice of the target object refers to the target user's local voice, that is, the historical voice messages sent by the target object to the user. The target user's local voice can be stored in a local database and can be directly retrieved when needed. Voice characteristics can include at least one of tone, intonation, speech rate, and volume.

[0045] Among them, voice correction refers to modifying key voice messages into target voice messages that have the voice characteristics of the voice message to be processed and the voiceprint characteristics of the target object. This makes the target voice message that the user hears in the end similar to the voice message to be processed in terms of tone, intonation, and voiceprint. This can improve the efficiency of users in obtaining voice information, and at the same time, accurately receive information such as the tone, intonation, and speaking speed of the message sender when sending the voice message to be processed, making the voice message that the user hears more like the voice of the message sender.

[0046] In this embodiment, the process first acquires the voice message to be processed, then determines the key voice message corresponding to the voice message to be processed, and finally, by fine-tuning the trained voice correction model obtained from the sample voice of the target object, the key voice message is corrected based on the similarity between the key voice message and the voice message to be processed, thereby obtaining a target voice message with the voice characteristics of the voice message to be processed and the voiceprint characteristics of the target object. This embodiment obtains the key voice message by processing at least one voice message, and then corrects the key voice message using the trained voice correction model. This results in a target voice message with the voice characteristics of the voice message to be processed and the voiceprint characteristics of the target object. This improves the efficiency of users acquiring voice message content and also allows users to perceive the emotions of the voice message sender, making the target voice message heard by the user closer to the original voice message, more like a simplified and corrected version of the original voice message by the voice message sender, thus improving the accuracy of voice message processing.

[0047] In one possible implementation of this application, determining the key voice message corresponding to the voice message to be processed may include: converting the voice message to be processed into message text using a voice conversion model; extracting key content from the message text using a text extraction model to obtain key text; and converting the key text into key voice message using a text conversion model.

[0048] In this embodiment, the speech conversion model is a model for converting speech messages into text. It can convert the speech messages to be processed into message text, enabling faster extraction of key content. The text extraction model is used to streamline speech messages that are long or have many segments, extracting key content to reduce the time users spend listening to the speech messages and improve the efficiency of obtaining content from them. The text conversion model is a model for converting text into speech messages. It can convert the extracted key text into key speech messages, allowing users to easily obtain effective content and emotional information about the message recipient.

[0049] The text extraction model can be the BERTSUM model or other models, as long as it can extract the key content from the message text. This application does not impose any specific limitations on the implementation.

[0050] Optionally, by using a text extraction model to extract key content from the message text and obtain key texts, the voice message processing method may also include: performing error correction processing on the message text using a text error correction model.

[0051] Since the voice messages to be processed may contain pronunciation errors or slips of the tongue, text correction models can be used to correct these errors in the converted text. These text correction models can utilize existing model structures, such as traditional natural speech correction models like the Chinese Language Model (N-Gram), without requiring a complete model reconstruction, thus saving resources.

[0052] Correspondingly, extracting key content from message text using a text extraction model to obtain key text can include: extracting key content from error-corrected message text using a text extraction model to obtain key text.

[0053] In this embodiment of the application, after the text information is corrected, the key content is extracted from the corrected message text using a text extraction model, which makes the obtained key text more accurate and more accurately expresses the original meaning that the information sender intended to express.

[0054] The speech conversion model and text conversion model in the above embodiments can both be converted according to the text-speech mapping relationship corresponding to the target object, so that the converted text and speech have more personal characteristics of the information sending object, solve the problems of accent and personalized expression when converting speech to text, and make text-to-speech more common and vivid, as detailed in the following embodiments.

[0055] In one possible implementation of this application, converting a voice message to be processed into message text using a speech conversion model may include: converting the voice message to be processed into message text based on the text-speech mapping relationship corresponding to the target object using a speech conversion model.

[0056] The text-to-speech mapping relationship is obtained by fine-tuning the speech correction model using sample speech from the target object. The text-to-speech mapping relationship includes the mapping relationship between text and at least one speech segment, where each speech segment possesses the voiceprint characteristics of the target object, and the speech characteristics of each speech segment are different.

[0057] It is worth noting that the speech correction model is trained using common sample speech. The common sample speech can be public speech obtained from the Internet or database. The speech correction model to be trained can utilize the existing model structure, such as the joint task learning model, without having to rebuild the model structure, which can save resources.

[0058] In this embodiment, the speech correction model is fine-tuned using sample speech of the target object, resulting in a speech correction model that is more characteristic of the target object. The text-speech mapping relationship obtained from this target object-specific speech correction model is also more characteristic of the target object. Through the text-speech mapping relationship corresponding to the target object, the speech message to be processed can be quickly segmented, multiple speech candidate segments can be trained, and then the weights of the corresponding text segments can be determined, thereby identifying the corresponding text.

[0059] In one possible implementation of this application, converting key text into key speech messages using a text conversion model may include: converting key text into key speech messages based on the text-speech mapping relationship corresponding to the target object using a text conversion model.

[0060] The text-to-speech mapping relationship is obtained by fine-tuning the speech correction model using sample speech from the target object. The text-to-speech mapping relationship includes the mapping relationship between text and at least one speech segment, where each speech segment possesses the voiceprint characteristics of the target object, and the speech characteristics of each speech segment are different.

[0061] The training of the speech correction model is the same as in the above embodiments, and the fine-tuning training process is also the same, so it will not be described again in this embodiment.

[0062] In this embodiment, the speech correction model is fine-tuned and trained using sample speech of the target object, resulting in a speech correction model that is more characteristic of the target object. The text-speech mapping relationship obtained based on this target object-specific speech correction model is also more characteristic of the target object. Through the text-speech mapping relationship corresponding to the target object, the edited speech can be generated quickly and accurately.

[0063] The text-to-speech mapping relationship in the above embodiments is obtained during the fine-tuning training of the speech correction model using sample speech of the target object. It refers to the mapping relationship between a text and multiple speech segments. The text-to-speech mapping relationship table x u,i,j →y u,i,j It can be as follows:

[0064]

[0065] Where I represents the speech dimension, J represents the text dimension, and ω i,j The filtering weight for text j corresponding to user u's voice segment i.

[0066] In one possible implementation of this application, a trained speech correction model is used to correct the speech of a key speech message based on the similarity between the key speech message and the speech message to be processed, thereby obtaining the target speech message. This may include: determining the similarity between the key speech message and the speech message to be processed using the speech correction model; if the similarity is less than a preset threshold, passing the similarity to at least one of a speech conversion model and a text conversion model, so that at least one of the speech conversion model and the text conversion model adjusts its respective output until the similarity between the key speech message and the speech message to be processed is greater than or equal to the preset threshold, thereby obtaining the target speech message.

[0067] The preset threshold can be set by the user or determined based on the similarity of sample speech during the training of the speech correction model. The higher the value of the preset threshold, the closer the final target speech message is to the speech message restated by the target user.

[0068] In other words, this application uses the PID (proportional-integral-derivative) control principle. The PID control principle is based on the control deviation formed by the given value and the actual output value. This deviation is then linearly combined proportionally, integrally, and derivatively to form the control quantity, which controls the controlled object. In this application, the given value refers to a preset threshold, the actual output value refers to the similarity level, the deviation refers to the difference between the similarity level and the preset threshold, and the controlled object refers to the speech correction model and / or text conversion model. The difference between the similarity level and the preset threshold is used to correct the text and speech weight distribution in the speech conversion, making the text converted by the speech correction model more accurate and the speech converted by the text conversion model more similar to the original speech in terms of speech characteristics.

[0069] The PID control principle in this embodiment is mainly used to actively feed back the difference between the similarity level and the preset threshold to the speech conversion model and the text conversion model in real time during the use of the model. The speech conversion model and the text conversion model correct the text and speech weight distribution of the speech conversion based on the difference, so that the text converted by the speech correction model is more accurate and the speech converted by the text conversion model is more similar to the speech characteristics of the original speech.

[0070] Based on the above text-speech mapping relationship, it can be seen that the text corresponding to a speech segment has different filtering weights. In speech-to-text or text-to-speech conversion, there may be errors in weight allocation, resulting in input-output speech errors. When the similarity level is calculated to be less than the preset threshold, the error value is fed back to at least one of the speech conversion module and the text conversion model. The final similarity level is adjusted by adjusting the output results of the speech conversion module and / or the text conversion model to obtain the target speech message.

[0071] In one possible implementation of this application, determining the similarity between a key speech message and a speech message to be processed using a speech correction model may include: obtaining speech features of the key speech message and the speech features of the speech message to be processed using the speech network of the speech correction model; obtaining text-speech combination features of the key speech message based on the speech features of the key speech message, and obtaining text-speech combination features of the speech message to be processed based on the speech features of the speech message to be processed using the text network of the speech correction model; and determining the similarity between the key speech message and the speech message to be processed using the similarity evaluation network of the speech correction model based on the text-speech combination features of the key speech message and the text-speech combination features of the speech message to be processed.

[0072] The speech correction model comprises three networks: a speech network, a text network, and a similarity evaluation network. The speech network divides the input speech message into multiple speech segment latent vectors and outputs speech feature vectors. These feature vectors express the speech characteristics of the input speech message, such as tone, pitch, speech rate, and volume. In other words, the speech network can reconstruct the input speech message in terms of tone, pitch, speech rate, and volume, deriving weights for each speech segment on various speech characteristics, and then reconstructing the speech message based on these weights. The reconstructed speech message is then input into the text network, which converts it into a text-speech combination feature that includes both textual and speech characteristics. That is, the text obtained through the text network possesses both textual and speech features. Finally, the text-speech combination features obtained from processing two input speech messages through the speech and text networks are input into the similarity evaluation network to determine the degree of similarity between the two input speech messages.

[0073] In this embodiment of the application, the key speech message and the speech message to be processed are processed separately through three networks of the speech correction model, and then the similarity is determined. Among them, the speech network can more accurately determine the speech characteristics such as tone, pitch, speech rate, and volume of the input speech message, so that after the speech message passes through the text network, the converted text can have text-level features and speech-level features, making the converted text more accurate, so that the evaluation result is more accurate when evaluated in the similarity evaluation network.

[0074] like Figure 5 The image shown is a schematic diagram of the speech correction model. Figure 5 As can be seen, the speech correction model consists of three parts: a speech network, a text network, and a similarity evaluation network. The speech message is input into the speech correction model, and the speech network of the model divides it into multiple speech segment latent vectors, such as... Figure 6 As shown, the speech feature vector can express the speech characteristics of the input speech message, such as tone, intonation, speech rate, and volume. In other words, the speech network can reconstruct the input speech message in terms of tone, intonation, speech rate, and volume, obtaining a weight for each speech segment of the speech message on various speech characteristics, and then reconstructing the speech message based on these weights. This differs from existing technologies that directly convert speech messages into text. In this application, after the speech message passes through a text network, the converted text possesses both textual and speech characteristics, making the converted text more accurate. The reconstructed speech message is then input into the text network, which converts it into a text-speech combination feature that includes both textual and speech characteristics. That is, the text obtained through the text network has both textual and speech characteristics, unlike existing technologies that only convert text messages into speech messages. This application converts the speech message into a speech message with its original speech characteristics. Finally, the text-speech combination features obtained from the speech network and text network of two input speech messages can be input into a similarity evaluation network to determine the degree of similarity between the two input speech messages.

[0075] In one possible implementation of this application, the training steps of the speech correction model include: pre-training the speech correction model to be trained using general sample speech; fine-tuning the pre-trained speech correction model using sample speech of the target object until the training is completed, and obtaining the trained speech correction model.

[0076] In this context, "general sample speech" refers to public speech, not the speech of a specific person or group. General sample speech can be obtained from the internet or from a speech database. "Target object sample speech" refers to the speech of the target object, which can be a voice message sent by the target object before the speech to be processed, or a voice message of the target object from a user-authorized device's local speech database. The speech correction model to be trained can utilize the existing model structure, such as a joint task learning model, without needing to rebuild the model structure, thus saving resources.

[0077] In this embodiment, the speech correction model is first pre-trained using common sample speech, i.e., public speech, to obtain a general speech correction model. Then, the general model is fine-tuned using sample speech of the target object to obtain a speech correction model with the characteristics of the target object. This makes it possible to use the model to determine the similarity between the speech message to be processed and the key speech message, and to correct the speech message of the key speech message. The corrected speech message has more characteristics of the target object and is more similar to the target object's own re-expression.

[0078] Optionally, pre-training the speech correction model to be trained using general sample speech can include: obtaining speech features of at least two sample speech messages from the general sample speech through the speech network of the speech correction model to be trained, wherein any two sample speech messages have a preset similarity level; obtaining text-speech combination features of each sample speech message based on the speech features of the at least two sample speech messages through the text network of the speech correction model to be trained; determining the similarity level between any two sample speech messages based on the text-speech combination features of any two sample speech messages through the similarity evaluation network of the speech correction model to be trained; and training the speech network and text network based on the difference between the similarity level between any two sample speech messages and the preset similarity level between any two sample speech messages until the similarity level between any two sample speech messages is greater than or equal to the preset similarity level, thereby obtaining the pre-trained speech correction model.

[0079] The speech correction model comprises three networks: a speech network, a text network, and a similarity evaluation network. The speech network divides the input speech message into multiple speech segment latent vectors and outputs speech feature vectors. These feature vectors express the speech characteristics of the input speech message, such as tone, pitch, speech rate, and volume. In other words, the speech network can reconstruct the input speech message in terms of tone, pitch, speech rate, and volume, deriving weights for each speech segment on various speech characteristics, and then reconstructing the speech message based on these weights. The reconstructed speech message is then input into the text network, which converts it into a text-speech combination feature that includes both textual and speech characteristics. That is, the text obtained through the text network possesses both textual and speech features. Finally, the text-speech combination features obtained from processing two input speech messages through the speech and text networks are input into the similarity evaluation network to determine the degree of similarity between the two input speech messages.

[0080] In this embodiment, any two sample speech messages with a preset similarity level are processed through the aforementioned speech network and text network to obtain their respective text-speech combination features. These features are then input into a similarity evaluation network to determine the similarity between the two sample speech messages. If the determined similarity between any two sample speech messages is greater than or equal to the preset similarity, the speech correction model is considered to have been trained successfully. Otherwise, the speech network and text network are trained continuously until the above conditions are met. Through this training process, a general speech correction model can be trained, making the simplified speech message corrected by this model closer to the original speech message.

[0081] After pre-training the speech correction model to obtain a general speech correction model, the general speech correction model can be fine-tuned using sample speech from the target object to obtain a speech correction model that conforms to the speech characteristics of the target object. Specifically, the fine-tuning process can include: obtaining the speech features of at least two sample speech messages from the target object's sample speech through the speech network of the trained speech correction model, wherein any two sample speech messages have a preset similarity level; obtaining the text-speech combination features of each sample speech message based on the speech features of at least two sample speech messages through the text network of the trained speech correction model; determining the similarity level between any two sample speech messages based on the text-speech combination features of any two sample speech messages through the similarity evaluation network of the trained speech correction model; and training the speech network and text network based on the difference between the similarity level between any two sample speech messages and the preset similarity level between any two sample speech messages until the similarity level between any two sample speech messages is greater than or equal to the preset similarity level, thus obtaining the trained speech correction model.

[0082] The specific descriptions of the speech network, text network, and similarity evaluation network of the trained speech correction model have been described in detail in the above embodiments, and will not be repeated in this embodiment.

[0083] In this embodiment, any two sample speech messages of the target object with a preset similarity level are processed through the aforementioned speech network and text network to obtain their respective text-speech combination features. These two text-speech combination features are then input into a similarity evaluation network to determine the similarity between the two sample speech messages. If the determined similarity between any two sample speech messages is greater than or equal to the preset similarity level, it indicates that the speech correction model has been trained successfully; otherwise, the speech network and text network are trained continuously until the above conditions are met. Through this training process, a speech correction model with the characteristics of the target object can be trained, making the target speech message corrected by the speech correction model closer to the speech message to be processed.

[0084] like Figure 7 The diagram shows a simplified input and output structure for training a speech correction model. The inputs are general sample speech and the target object's sample speech. During model training, the similarity between the two speech messages, the implicit quantity of speech segments, the text-speech mapping relationship, and the preset similarity level can be obtained. These details have been described in detail in the above embodiments and will not be repeated in this embodiment.

[0085] In one possible implementation of this application, the step of obtaining the text-speech mapping relationship may include: obtaining multiple speech segments from the sample speech of the target object, obtaining the text corresponding to each speech segment through the text network of the speech correction model, and determining the text-speech mapping relationship based on each speech segment and the text corresponding to each speech segment.

[0086] In other words, multiple speech segments can be obtained from sample speech of the target object, input into the text network of the speech correction model, and the text corresponding to each speech segment can be obtained. Based on each speech segment and the text corresponding to each speech segment, the mapping relationship between the text and multiple speech segments can be determined, that is, the text-speech mapping relationship.

[0087] In one possible implementation of this application, acquiring the voice message to be processed may include: receiving at least one voice message sent by a target object in a conversation interface. Correspondingly, the voice message processing method may further include: displaying the target voice message in the conversation interface.

[0088] In other words, the voice message to be processed is at least one voice message sent by the target object received in the user's conversation interface. This at least one voice message can be obtained by selecting messages one by one, by directly selecting multiple messages by swiping, or by selecting the target object to obtain at least one voice message sent by the target object. Specifically, this embodiment does not impose limitations and is determined according to the actual application. The final obtained target voice message can be displayed in the conversation interface for the user to listen to, such as... Figure 8 As shown, after the voice message is processed, a read control can be displayed in the conversation interface. After the user clicks the control, they can hear a concise voice message with the user's voiceprint characteristics and the same voice characteristics as the original voice message, which reduces the time spent by the user and improves the efficiency of voice information acquisition.

[0089] like Figure 9The diagram shows the overall flow of the voice message processing method of this application. Specifically, upon obtaining the voice message to be processed, the key voice message is obtained after passing through a voice conversion model, a text correction model, a text extraction model, and a text conversion model. The voice message to be processed and the key voice message are then input into a voice correction model to obtain the target voice message. During this process, the voice correction model can, in real time, feed back the difference between the similarity level of the voice message to be processed and the key voice message and a preset similarity level to the voice conversion model and the text conversion model. This allows the voice conversion model and the text conversion model to correct the text and voice weight distribution of the voice conversion based on the difference, making the text converted by the voice correction model more accurate and the voice converted by the text conversion model more similar to the original voice in terms of voice characteristics. These details have been described in detail in the above embodiments and will not be repeated in this embodiment.

[0090] It should be noted that the voice message processing method provided in this application embodiment can be executed by a voice message processing device or a control module within that voice message processing device for executing the voice message processing method. This application embodiment uses the execution of the voice message processing method by a voice message processing device as an example to illustrate the voice message processing device provided in this application embodiment.

[0091] like Figure 10 The diagram shown is a schematic representation of a voice message processing device according to an embodiment of this application. The voice message processing device may include: an acquisition module 1001, a determination module 1002, and a correction module 1003.

[0092] The system includes an acquisition module 1001 for acquiring the voice message to be processed; a determination module 1002 for determining the key voice message corresponding to the voice message to be processed; and a correction module 1003 for correcting the key voice message based on the similarity between the key voice message and the voice message to be processed using a trained voice correction model to obtain the target voice message. The trained voice correction model is fine-tuned using sample voices of the target object. The target object is the message sending object corresponding to the voice message to be processed. The target voice message has the voice characteristics of the voice message to be processed and the voiceprint characteristics of the target object.

[0093] In this embodiment, the acquisition module 1001 first acquires the voice message to be processed, then the determination module 1002 determines the key voice message corresponding to the voice message to be processed, and finally the correction module 1003, by fine-tuning the trained voice correction model obtained from the sample voice of the target object, corrects the key voice message based on the similarity between the key voice message and the voice message to be processed, thereby obtaining a target voice message with the voice characteristics of the voice message to be processed and the voiceprint characteristics of the target object. This embodiment, by processing at least one voice message to obtain the key voice message and correcting it using the trained voice correction model, can obtain a target voice message with the voice characteristics of the voice message to be processed and the voiceprint characteristics of the target object. This improves the efficiency of users acquiring voice message content and also allows users to perceive the emotions of the voice message sender, making the target voice message heard by the user closer to the original voice message, more like a simplified and corrected version of the original voice message by the voice message sender, thus improving the accuracy of voice message processing.

[0094] Optionally, the determining module 1002 can be used to: convert the voice message to be processed into message text through a speech conversion model; extract key content from the message text through a text extraction model to obtain key text; and convert the key text into key voice messages through a text conversion model.

[0095] Optionally, the determining module 1002 can be used to: perform error correction processing on the message text using a text error correction model; and extract key content from the error-corrected message text using a text extraction model to obtain key text.

[0096] Optionally, the determining module 1002 can be used to: convert the speech message to be processed into message text based on the text-speech mapping relationship corresponding to the target object through a speech conversion model; wherein the text-speech mapping relationship is obtained by fine-tuning and training the speech correction model using sample speech of the target object.

[0097] Optionally, the determining module 1002 can be used to: convert key text into key speech messages based on the text-speech mapping relationship corresponding to the target object through a text conversion model; wherein the text-speech mapping relationship is obtained by fine-tuning and training the speech correction model using sample speech of the target object.

[0098] Optionally, the determining module 1002 can be used to: the text-speech mapping relationship includes the mapping relationship between text and at least one speech segment, at least one speech segment has the voiceprint characteristics of the target object, and each speech segment has different speech characteristics.

[0099] Optionally, the determining module 1002 can be used to: determine the similarity between the key speech message and the speech message to be processed through the speech correction model; if the similarity is less than a preset threshold, pass the similarity to at least one of the speech conversion model and the text conversion model so that at least one of the speech conversion model and the text conversion model adjusts its respective output until the similarity between the key speech message and the speech message to be processed is greater than or equal to the preset threshold, thereby obtaining the target speech message.

[0100] Optionally, the determining module 1002 can be used to: obtain the speech features of the key speech message and the speech features of the speech message to be processed through the speech network of the speech correction model; obtain the text-speech combination features of the key speech message based on the speech features of the key speech message through the text network of the speech correction model, and obtain the text-speech combination features of the speech message to be processed based on the speech features of the speech message to be processed; and determine the degree of similarity between the key speech message and the speech message to be processed through the similarity evaluation network of the speech correction model based on the text-speech combination features of the key speech message and the text-speech combination features of the speech message to be processed.

[0101] Optionally, the correction module 1003 can be used to: pre-train the speech correction model to be trained using general sample speech; and fine-tune the pre-trained speech correction model using sample speech of the target object until the training ends, thereby obtaining the trained speech correction model.

[0102] Optionally, the correction module 1003 can be used to: obtain the speech features of at least two sample speech messages in a general sample speech through the speech network of the speech correction model to be trained, wherein any two sample speech messages have a preset similarity level; obtain the text-speech combination features of each sample speech message based on the speech features of the at least two sample speech messages through the text network of the speech correction model to be trained; determine the similarity level between any two sample speech messages based on the text-speech combination features of any two sample speech messages through the similarity evaluation network of the speech correction model to be trained; and train the speech network and text network based on the difference between the similarity level between any two sample speech messages and the preset similarity level between any two sample speech messages until the similarity level between any two sample speech messages is greater than or equal to the preset similarity level, thereby obtaining a pre-trained speech correction model.

[0103] Optionally, the determining module 1002 can be used to: obtain multiple speech segments from the sample speech of the target object; obtain the text corresponding to each speech segment through the text network of the speech correction model; and determine the text-speech mapping relationship based on each speech segment and the text corresponding to each speech segment.

[0104] Optionally, the acquisition module 1001 can be used to: receive at least one voice message sent by the target object in the conversation interface; correspondingly, the voice message processing device further includes: a display module.

[0105] The display module is used to display the target voice message in the conversation interface.

[0106] The voice message processing device in this application embodiment can be a device, or a component, integrated circuit, or chip in a terminal. The device can be a mobile electronic device or a non-mobile electronic device. For example, mobile electronic devices can be mobile phones, tablets, laptops, PDAs, in-vehicle electronic devices, wearable devices, ultra-mobile personal computers (UMPCs), netbooks, or personal digital assistants (PDAs), etc., while non-mobile electronic devices can be servers, network attached storage (NAS), personal computers (PCs), televisions (TVs), ATMs, or self-service machines, etc. This application embodiment does not impose specific limitations.

[0107] The voice message processing device in this application embodiment can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit the specific operating system used.

[0108] The voice message processing device provided in this application embodiment can achieve... Figures 1-9 To avoid repetition, the various processes implemented in the method embodiments shown will not be described again here.

[0109] Optionally, such as Figure 11 As shown, this application embodiment also provides an electronic device 1100, including a processor 1101, a memory 1102, and a program or instructions stored in the memory 1102 and executable on the processor 1101. When the program or instructions are executed by the processor 1101, they implement the various processes of the above-described voice message processing method embodiment and achieve the same technical effect. To avoid repetition, they will not be described again here.

[0110] It should be noted that the electronic devices in the embodiments of this application include the mobile electronic devices and non-mobile electronic devices described above.

[0111] Figure 12A schematic diagram of the hardware structure of an electronic device to implement an embodiment of this application.

[0112] The electronic device 1200 includes, but is not limited to, components such as: radio frequency unit 1201, network module 1202, audio output unit 1203, input unit 1204, sensor 1205, display unit 1206, user input unit 1207, interface unit 1208, memory 1209, and processor 1210.

[0113] Those skilled in the art will understand that the electronic device 1200 may also include a power supply (such as a battery) for supplying power to various components. The power supply may be logically connected to the processor 1210 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. Figure 12 The electronic device structure shown does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.

[0114] The processor 1210 is used to acquire the speech message to be processed; determine the key speech message corresponding to the speech message to be processed; and perform speech correction on the key speech message based on the similarity between the key speech message and the speech message to be processed using a trained speech correction model to obtain the target speech message. The trained speech correction model is obtained by fine-tuning the training using sample speech of the target object. The target object is the message sending object corresponding to the speech message to be processed. The target speech message has the speech characteristics of the speech message to be processed and the voiceprint characteristics of the target object.

[0115] In this embodiment, the process first acquires the voice message to be processed, then determines the key voice message corresponding to the voice message to be processed, and finally, by fine-tuning the trained voice correction model obtained from the sample voice of the target object, the key voice message is corrected based on the similarity between the key voice message and the voice message to be processed, thereby obtaining a target voice message with the voice characteristics of the voice message to be processed and the voiceprint characteristics of the target object. This embodiment obtains the key voice message by processing at least one voice message, and then corrects the key voice message using the trained voice correction model. This results in a target voice message with the voice characteristics of the voice message to be processed and the voiceprint characteristics of the target object. This improves the efficiency of users acquiring voice message content and also allows users to perceive the emotions of the voice message sender, making the target voice message heard by the user closer to the original voice message, more like a simplified and corrected version of the original voice message by the voice message sender, thus improving the accuracy of voice message processing.

[0116] It should be understood that, in this embodiment, the input unit 1204 may include a graphics processing unit (GPU) 12041 and a microphone 12042. The GPU 12041 processes image data of still images or videos obtained by an image capture device (such as a camera) in video capture mode or image capture mode. The display unit 1206 may include a display panel 12061, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, etc. The user input unit 1207 includes a touch panel 12071 and other input devices 12072. The touch panel 12071 is also called a touch screen. The touch panel 12071 may include a touch detection device and a touch controller. Other input devices 12072 may include, but are not limited to, physical keyboards, function keys (such as volume control buttons, power buttons, etc.), trackballs, mice, joysticks, etc., which will not be described in detail here. The memory 1209 can be used to store software programs and various data, including but not limited to applications and operating systems. Processor 1210 may integrate an application processor and a modem processor. The application processor mainly handles the operating system, user interface, and applications, while the modem processor mainly handles wireless communication. It is understood that the modem processor may also not be integrated into processor 1210.

[0117] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described voice message processing method embodiments and achieve the same technical effect. To avoid repetition, they will not be described again here.

[0118] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0119] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-described voice message processing method embodiments and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0120] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.

[0121] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0122] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0123] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. A voice message processing method, characterized in that, include: Retrieve voice messages to be processed; The voice message to be processed is converted into message text using a speech conversion model; Using a text extraction model, key content is extracted from the message text to obtain key text. Using a text conversion model, the key text is converted into key voice messages based on the text-speech mapping relationship corresponding to the target object; By using a trained speech correction model, the key speech message is corrected based on the similarity between the key speech message and the speech message to be processed, and the target speech message is obtained. The text-speech mapping relationship includes a mapping relationship between text and at least one speech segment, wherein the at least one speech segment has the voiceprint characteristics of the target object, and each speech segment has different voice characteristics; the trained speech correction model is obtained by fine-tuning and training using sample speech of the target object; the target object is the message sending object corresponding to the speech message to be processed; the target speech message has the voice characteristics of the speech message to be processed and the voiceprint characteristics of the target object; The training steps for the speech correction model include: The speech network of the speech correction model to be trained is used to obtain the speech features of at least two sample speech messages in the general sample speech, wherein any two sample speech messages have a preset similarity. By using the text network of the speech correction model to be trained, and based on the speech features of the at least two sample speech messages, the text-speech combination features of each sample speech message are obtained. The similarity evaluation network of the speech correction model to be trained determines the degree of similarity between any two sample speech messages based on the text-speech combination features of any two sample speech messages. Based on the difference between the similarity between any two sample speech messages and the preset similarity between any two sample speech messages, train the speech network and the text network until the similarity between any two sample speech messages is greater than or equal to the preset similarity, and obtain the pre-trained speech correction model. Using sample speech from the target object, the pre-trained speech correction model is fine-tuned until the training ends, thus obtaining the trained speech correction model.

2. The method according to claim 1, characterized in that, Before extracting key content from the message text using a text extraction model to obtain the key text, the method further includes: The message text is corrected using a text correction model. The step of extracting key content from the message text using a text extraction model to obtain key text includes: By using a text extraction model, key content is extracted from the message text after error correction, and the key text is obtained.

3. The method according to claim 1, characterized in that, The step of converting the voice message to be processed into message text using a speech conversion model includes: Using a speech conversion model, the speech message to be processed is converted into the message text based on the text-speech mapping relationship corresponding to the target object; The text-to-speech mapping relationship is obtained by fine-tuning the speech correction model using sample speech from the target object.

4. The method according to claim 1, characterized in that, The text-to-speech mapping relationship is obtained by fine-tuning and training the speech correction model using sample speech from the target object.

5. The method according to claim 1, characterized in that, The step of using a trained speech correction model to correct the key speech message based on the similarity between the key speech message and the speech message to be processed, thereby obtaining the target speech message, includes: The similarity between the key voice message and the voice message to be processed is determined by a voice correction model. If the similarity is less than a preset threshold, the similarity is passed to at least one of the speech conversion model and the text conversion model, so that at least one of the speech conversion model and the text conversion model adjusts its respective output until the similarity between the key speech message and the speech message to be processed is greater than or equal to the preset threshold, thereby obtaining the target speech message.

6. The method according to claim 5, characterized in that, The step of determining the similarity between the key speech message and the speech message to be processed using a speech correction model includes: The speech features of the key speech message and the speech features of the speech message to be processed are obtained through the speech network of the speech correction model. The text network of the speech correction model obtains the text-speech combination features of the key speech message based on the speech features of the key speech message, and obtains the text-speech combination features of the speech message to be processed based on the speech features of the speech message to be processed. The similarity evaluation network of the speech correction model determines the degree of similarity between the key speech message and the speech message to be processed based on the text-speech combination features of the key speech message and the text-speech combination features of the speech message to be processed.

7. The method according to any one of claims 3 or 4, characterized in that, The steps for obtaining the text-to-speech mapping relationship include: Obtain multiple speech segments from sample speech of the target object; The text network of the speech correction model is used to obtain the text corresponding to each speech segment; The text-speech mapping relationship is determined based on each speech segment and the corresponding text.

8. The method according to claim 1, characterized in that, The acquisition of the voice message to be processed includes: Receive at least one voice message sent by the target object in the conversation interface; The method further includes: The target voice message is displayed in the conversation interface.

9. A voice message processing device, characterized in that, include: The acquisition module is used to acquire voice messages to be processed. The determination module is used to convert the voice message to be processed into message text through a voice conversion model; Using a text extraction model, key content is extracted from the message text to obtain key text. Using a text conversion model, the key text is converted into key voice messages based on the text-speech mapping relationship corresponding to the target object; The correction module is used to correct the key speech message based on the similarity between the key speech message and the speech message to be processed using a trained speech correction model, so as to obtain the target speech message. The text-speech mapping relationship includes a mapping relationship between text and at least one speech segment, wherein the at least one speech segment has the voiceprint characteristics of the target object, and each speech segment has different voice characteristics; the trained speech correction model is obtained by fine-tuning and training using sample speech of the target object; the target object is the message sending object corresponding to the speech message to be processed; the target speech message has the voice characteristics of the speech message to be processed and the voiceprint characteristics of the target object; The correction module is used for: The speech network of the speech correction model to be trained is used to obtain the speech features of at least two sample speech messages in the general sample speech, wherein any two sample speech messages have a preset similarity. By using the text network of the speech correction model to be trained, and based on the speech features of the at least two sample speech messages, the text-speech combination features of each sample speech message are obtained. The similarity evaluation network of the speech correction model to be trained determines the degree of similarity between any two sample speech messages based on the text-speech combination features of any two sample speech messages. Based on the difference between the similarity between any two sample speech messages and the preset similarity between any two sample speech messages, train the speech network and the text network until the similarity between any two sample speech messages is greater than or equal to the preset similarity, and obtain the pre-trained speech correction model. Using sample speech from the target object, the pre-trained speech correction model is fine-tuned until the training ends, thus obtaining the trained speech correction model.

10. An electronic device, characterized in that, It includes a processor, a memory, and a program or instructions stored in the memory and executable on the processor, wherein the program or instructions, when executed by the processor, implement the steps of the voice message processing method as described in any one of claims 1-8.

11. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a processor, implement the steps of the voice message processing method as described in any one of claims 1-8.

Citation Information

Patent Citations

  • Processing method of voice message and mobile terminal

    CN107123418A

  • Model training and voice interaction method and device

    CN113314092A