Voice message processing method, device and electronic equipment
By obtaining the target correction degree and similarity threshold of the pending voice message, the target voice message is generated using the voice correction model, which solves the repeated recording problem caused by errors in the recording and improves the efficiency and accuracy of voice information acquisition.
Patent Information
- Application Number
- CN202210744413.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-28
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2042-06-28
AI Technical Summary
When sending multiple voice messages with a long duration, if an error occurs in the middle of recording, it is necessary to record a new message repeatedly, resulting in a long time for voice messages and the inefficient efficiency of obtaining effective information from voice messages.
By obtaining the target correction degree of the pending voice message, determining the similarity threshold, using the voice correction model to personalize the key voice message, generating the target voice message, ensuring that it has voice characteristics and voiceprint characteristics, and avoiding repeated recording.
It improves the efficiency of obtaining voice information, makes the final generated voice message more like what the message is said by the message recording himself, and improves the accuracy and efficiency of voice message processing.
Smart Images

Figure CN115132207B_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of computer technology, and specifically relates to a voice message processing method, device and electronic device. Background Art
[0002] With the continuous advancement of technology, people use electronic devices more and more frequently. When people communicate, they often use the voice message function in some applications. While voice messages bring great convenience, they are also vivid and have strong user characteristics, and there is less information loss during transmission.
[0003] However, when a user sends multiple long voice messages, if an error occurs during recording, new messages need to be recorded repeatedly. This results in a long voice message sending time and low efficiency in obtaining effective information from the voice messages. Summary of the Invention
[0004] The embodiments of the present application provide a voice message processing method, device and electronic device, which can solve the problems in the prior art of sending voice messages, such as the need to repeatedly record new messages if an error occurs during recording, the long voice message sending time and the low efficiency of obtaining effective information from the voice message.
[0005] In a first aspect, an embodiment of the present application provides a method for processing a voice message, the method comprising:
[0006] Obtaining a voice message to be processed and a target correction degree corresponding to the voice message to be processed;
[0007] Determining a similarity threshold based on the target correction degree;
[0008] determining, by a voice correction model, a degree of similarity between the voice message to be processed and a key voice message, and performing personalized voice correction on the key voice message based on the degree of similarity and a similarity threshold to obtain a target voice message; the key voice message corresponds to the voice message to be processed;
[0009] The speech correction model is trained using sample speech of a target object; the target object is a message recording object corresponding to the voice message to be processed; and the target voice message has the speech characteristics of the voice message to be processed and the voiceprint characteristics of the target object.
[0010] In a second aspect, an embodiment of the present application provides a voice message processing device, the device comprising:
[0011] An acquisition module, configured to acquire a voice message to be processed and a target correction degree corresponding to the voice message to be processed;
[0012] a determination module, configured to determine a similarity threshold value according to the target correction degree;
[0013] a correction module, configured to determine, using a voice correction model, a degree of similarity between the voice message to be processed and a key voice message, and perform personalized voice correction on the key voice message based on the degree of similarity and a similarity threshold to obtain a target voice message; the key voice message corresponds to the voice message to be processed;
[0014] The speech correction model is trained using sample speech of a target object; the target object is a message recording object corresponding to the voice message to be processed; and the target voice message has the speech characteristics of the voice message to be processed and the voiceprint characteristics of the target object.
[0015] In a third aspect, an embodiment of the present application provides an electronic device comprising a processor, a memory, and a program or instruction stored in the memory and executable on the processor, wherein the program or instruction, when executed by the processor, implements the steps of the method described in the first aspect.
[0016] In a fourth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the method described in the first aspect are implemented.
[0017] In a fifth aspect, an embodiment of the present application provides a chip, which includes a processor and a communication interface, the communication interface and the processor are coupled, and the processor is used to run programs or instructions to implement the method described in the first aspect.
[0018] In an embodiment of the present application, first, a voice message to be processed and a target correction degree corresponding to the voice message to be processed are obtained, and then a similarity threshold is determined based on the target correction degree. Finally, a voice correction model obtained by training the sample voice of the target object is used to determine the similarity between the voice message to be processed and the key voice message. Based on the similarity and the similarity threshold, personalized voice correction is performed on the key voice message to obtain the target voice message. In an embodiment of the present application, a similarity threshold is determined by obtaining the target correction degree, and personalized voice correction is performed on the key voice message using the voice correction model. A target voice message having the voice characteristics of the voice message to be processed and the voiceprint characteristics of the target object can be obtained. This can avoid the user from re-recording when an error occurs in recording a voice message, thereby improving the efficiency of obtaining voice information. At the same time, the tone, intonation, speed, and other information of the target voice message obtained in the end can be made the same as those when the user recorded the voice message, so that the final voice message is more like what the message recorder said, thereby improving the accuracy of voice message processing. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0020] Figure 1 This is a flow chart of a voice message processing method provided by one embodiment of the present application;
[0021] Figure 2 This is a schematic diagram of obtaining a voice message to be processed provided by an embodiment of the present application;
[0022] Figure 3 is a schematic diagram of selecting a target correction degree provided by an embodiment of the present application;
[0023] Figure 4 is a schematic diagram of another method for selecting a target correction degree provided by an embodiment of the present application;
[0024] Figure 5 is a structural diagram of a speech correction model provided by an embodiment of the present application;
[0025] Figure 6 This is a schematic diagram of the structure of a voice network provided by an embodiment of the present application;
[0026] Figure 7 This is a simple structural diagram of the input and output during the training of the speech correction model provided by one embodiment of the present application;
[0027] Figure 8 This is a schematic diagram of the overall flow of a voice message processing method provided by an embodiment of the present application;
[0028] Figure 9 This is a structural diagram of a voice message processing device provided by an embodiment of the present application;
[0029] Figure 10 This is a schematic structural diagram of an electronic device provided by an embodiment of the present application;
[0030] Figure 11 This is a schematic diagram of the hardware structure of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION
[0031] The following will be combined with the accompanying drawings in the embodiments of the present application to clearly describe the technical solutions in the embodiments of the present application. Obviously, the embodiments described are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field are within the scope of protection of this application.
[0032] The terms "first," "second," and the like in the specification and claims of this application are used to distinguish similar objects and are not used to describe a particular order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate so that the embodiments of this application can be implemented in an order other than those illustrated or described herein. In addition, the term "and / or" in the specification and claims refers to at least one of the connected objects, and the character " / " generally indicates that the objects associated with each other are in an "or" relationship.
[0033] In some embodiments, when a user sends a long voice message, some errors may occur during the recording, requiring the user to record repeatedly to generate the voice message to be sent. This makes it take a long time for the user to send the voice message, but the effective information contained in the voice message is less, which leads to low efficiency in obtaining information from the voice message. To solve the above problem, the embodiments of the present application provide a voice message processing method, device and electronic device. In the scenario where the user sends a voice message, when the user makes a recording error, the user can continue recording without re-recording. After the recording is completed, the voice message processing method provided by the present application is used to convert at least one recorded voice message into a target voice message. The target voice message has the voiceprint characteristics of the user who recorded the message, and the target voice message has the same voice characteristics as the at least one recorded voice message. For example, if the at least one voice message recorded by the user is in a happy tone, the target voice message obtained is also in a happy tone. Finally, the target voice message is displayed on the conversation interface, so that the user obtains a voice message with concise content, the voiceprint characteristics of the message recording object, and the same voice characteristics as the at least one voice message, thereby reducing the time spent by the user and improving the efficiency of voice information acquisition.
[0034] In the following, in conjunction with the accompanying drawings, a voice message processing method, device and electronic device provided by the embodiments of the present application are described in detail through specific embodiments and their application scenarios.
[0035] like Figure 1 FIG. 1 is a flow chart of a method for processing a voice message according to an embodiment of the present invention. The method for processing a voice message may include the contents shown in S101 to S103.
[0036] In S101, a voice message to be processed and a target correction degree corresponding to the voice message to be processed are obtained.
[0037] Among them, the voice message to be processed is the voice message recorded by the message recording object, and the voice message to be processed includes at least one voice message. If the voice message to be processed includes multiple voice messages, the method of obtaining the voice message to be processed can be to select them one by one; or to select multiple voice messages by sliding, and multiple recorded voice messages can be selected by sliding. Specifically, this is not limited in the embodiment of the present application and is determined according to actual application. The voice message to be processed can be obtained by the above method, and then the user can enter the voice message processing interface by sliding to the left and convert the selected voice message into the target voice message, specifically as follows Figure 2 shown.
[0038] It is worth noting that the target correction degree can be selected by the user, such as Figure 3 As shown, after entering the voice message processing interface, the user can manually select the target correction level through the voice correction control on the interface. Specifically, the user can slide the voice correction control up and down. Sliding up means that the user needs less correction, sliding to the top means that only error correction is needed without text simplification; sliding down means that the user needs more correction, simplifies the text, and outputs a summary. The system can also automatically confirm based on the conversation scenario, in which case the user can click the automatic button on the interface, such as Figure 4 As shown, the system can automatically confirm the degree of target correction based on the conversation scenario, etc.
[0039] The target correction degree can be any value between 0 and 1, where 0 represents a high degree of simplification and only a summary needs to be output; and 1 represents no simplification and only error correction is required.
[0040] In S102 , a similarity threshold is determined according to the target correction degree.
[0041] It is worth noting that the target correction degree can be expressed as ω = a*S thd +b*S type Determine, where ω is the target correction degree, a and b are normalized weight coefficients set according to needs, S thd is the similarity threshold, S type It is the degree of content conciseness of the text, ranging from 0 to 1, corresponding to the value of ω.
[0042] For example, when ω=0, S type =0; ω = (0,0.3], S type =0.3;ω=(0.3,0.6], S type =0.6; ω = (0.6, 0.9], S type =0.9; ω = (0.9, 1], S type =1.
[0043] And Sthd The formula Calculated.
[0044] The larger the value of the similarity threshold is, the closer the final target voice message is to the voice message reformulated by the user who recorded the voice.
[0045] In S103, the similarity between the voice message to be processed and the key voice message is determined by the voice correction model, and based on the similarity and the similarity threshold, personalized voice correction is performed on the key voice message to obtain a target voice message; the key voice message corresponds to the voice message to be processed.
[0046] The speech correction model is trained using sample speech of a target object; the target object is a message recording object of a voice message to be processed; and the target voice message has the speech characteristics of the voice message to be processed and the voiceprint characteristics of the target object.
[0047] A key voice message refers to a voice message obtained after the voice message to be processed has been voice edited. The key voice message can correct the erroneous content in the voice message to be processed. When there are many voice messages (such as long voice messages or a large number of voice messages), the key content can be extracted, that is, a streamlined voice message is obtained. The details will be described in detail in subsequent embodiments and will not be repeated in this embodiment.
[0048] It's worth noting that the target user's sample voice refers to the target user's local voice, i.e., historical voice messages recorded by the target user. The target user's local voice can be stored in a local database and directly accessed when needed. Voice characteristics can include at least one of tone, intonation, speaking speed, and volume.
[0049] Among them, voice personalized correction refers to correcting the key voice message into a target voice message with the voice characteristics of the voice message to be processed and the voiceprint characteristics of the target object, so that the target voice message finally obtained by the user is a voice message with similar tone, intonation, voiceprint, etc. to the voice message to be processed. This can avoid the user from re-recording when an error occurs in recording a voice message, thereby improving the efficiency of obtaining voice information. At the same time, it can also make the final target voice message the same as the tone, intonation, speaking speed and other information when the user recorded the voice message, so that the final voice message is more like what the person who recorded the message said.
[0050] In an embodiment of the present application, first, a voice message to be processed and a target correction degree corresponding to the voice message to be processed are obtained, and then a similarity threshold is determined based on the target correction degree. Finally, a voice correction model obtained by training the sample voice of the target object is used to determine the similarity between the voice message to be processed and the key voice message. Based on the similarity and the similarity threshold, personalized voice correction is performed on the key voice message to obtain the target voice message. In an embodiment of the present application, a similarity threshold is determined by obtaining the target correction degree, and personalized voice correction is performed on the key voice message using the voice correction model. A target voice message having the voice characteristics of the voice message to be processed and the voiceprint characteristics of the target object can be obtained. This can avoid the user from re-recording when an error occurs in recording a voice message, thereby improving the efficiency of obtaining voice information. At the same time, the tone, intonation, speed, and other information of the target voice message obtained in the end can be made the same as those when the user recorded the voice message, so that the final voice message is more like what the message recorder said, thereby improving the accuracy of voice message processing.
[0051] In one possible implementation of the present application, the step of obtaining the key voice message includes: converting the voice message to be processed into a message text through a voice conversion model; extracting key content from the message text to obtain the key text through a text extraction model; and converting the key text into a key voice message through a text conversion model.
[0052] In the embodiment of the present application, the voice conversion model is a model that converts voice messages into text, and can convert the pending voice messages into message text so that the key content can be extracted more quickly later. The text extraction model is used to streamline pending voice messages that are longer or have a large number of voice messages, extracting the key content from the pending voice messages, thereby reducing the time spent by the voice message recipient listening to the pending voice messages and improving the efficiency of users in obtaining the content of the pending voice messages. The text conversion model is a model that converts text into voice messages, and can convert the extracted key text into key voice messages, making it easier for the voice message recipient to obtain effective content and emotional information about the message recording object.
[0053] Among them, the text extraction model can be a BERTSUM model or other models, as long as it can extract the key content in the message text. The embodiments of this application do not make specific limitations.
[0054] Optionally, before extracting key content from the message text using a text extraction model and obtaining the key text, the voice message processing method may further include: determining a content simplification degree according to a target correction degree.
[0055] It can be seen from the above embodiments that the content simplification degree corresponds to the target correction degree. After the target correction degree is determined, the content simplification degree can be determined.
[0056] The content simplification level can take values of 0, 0.3, 0.6, 0.9, or 1. This level can be applied to the text extraction model. When the content simplification level is 1, the text extraction model is not used, and the output of the text extraction model is equal to the input. In this case, only error correction is required for the text content, and simplification is not required, so that the text remains consistent with the original voice message as much as possible. When the content simplification level is 0, the text extraction model is used to maximally simplify the text to obtain the key text. The content simplification level can be used to determine the desired simplification level of the target voice message for the user.
[0057] Accordingly, extracting key content from the message text by using a text extraction model to obtain the key text may include: extracting key content from the message text by using a text extraction model that matches the degree of content simplification to obtain the key text.
[0058] That is to say, each content simplification level has a text extraction model that matches it. According to the text extraction model, the key content required by the user can be extracted from the message text to obtain the key text.
[0059] Optionally, by extracting key content from the message text through a text extraction model to obtain key text, the voice message processing method may further include: performing error correction processing on the message text through a text error correction model.
[0060] Since the processed voice message may contain mispronunciations or slips of the tongue, a text error correction model can be used to correct errors in the converted text. This model can use an existing model structure, such as a traditional natural speech error correction model like the Chinese language model (N-Gram), eliminating the need to rebuild the model structure and saving resources.
[0061] Correspondingly, extracting key content from the message text by using a text extraction model to obtain key text may include: extracting key content from the message text after error correction processing by using a text extraction model to obtain key text.
[0062] In an embodiment of the present application, after the text information is corrected, a text extraction model is used to extract key content from the corrected message text, so that the obtained key text can be more accurate and more accurately express the original meaning of the information sender.
[0063] The speech conversion model and text conversion model in the above embodiments can be converted according to the text-to-speech mapping relationship corresponding to the target object, so that the converted text and speech have more personal characteristics of the information recording object, solve the problems of accent and personalized expression when converting speech to text, and make the text-to-speech conversion more common and vivid. See the following embodiments for details.
[0064] In a possible implementation of the present application, converting the voice message to be processed into message text through a voice conversion model may include: converting the voice message to be processed into message text based on the text-to-speech mapping relationship corresponding to the target object through the voice conversion model.
[0065] The text-to-speech mapping is obtained by fine-tuning the initial speech correction model using sample speech from the target subject. The text-to-speech mapping includes a mapping between text and at least one speech segment, each of which has unique voice characteristics.
[0066] It is worth noting that the initial speech correction model is obtained by training the speech correction model to be trained using general sample speech, wherein the general sample speech can be a public speech obtained from the Internet or a database, and the speech correction model to be trained can utilize the model structure of the existing model, such as the joint task learning training model, without the need to rebuild the model structure, which can save resources.
[0067] In the embodiments of the present application, the initial speech correction model is fine-tuned using sample speech of the target subject, resulting in a speech correction model that is more characteristic of the target subject. The text-to-speech mapping relationship derived from this speech correction model that is more characteristic of the target subject is also more characteristic of the target subject. Using the text-to-speech mapping relationship corresponding to the target subject, the voice message to be processed can be quickly segmented and processed, multiple speech candidate segments can be trained, and the text segment weights corresponding to the multiple speech candidate segments can be determined, thereby corresponding to the corresponding text.
[0068] In a possible implementation of the present application, converting the key text into a key voice message through a text conversion model may include: converting the key text into a key voice message based on a text-to-speech mapping relationship corresponding to a target object through the text conversion model.
[0069] The text-to-speech mapping is obtained by fine-tuning the initial speech correction model using sample speech from the target subject. The text-to-speech mapping includes a mapping between text and at least one speech segment, each of which has unique voice characteristics.
[0070] The training of the initial speech correction model is the same as that in the above embodiment, and the process of fine-tuning training is also the same, which will not be repeated in this embodiment.
[0071] In the embodiments of the present application, fine-tuning the initial speech correction model using sample speech of the target subject can produce a speech correction model that is more characteristic of the target subject. The text-to-speech mapping relationship generated based on this speech correction model that is more characteristic of the target subject is also more characteristic of the target subject. Using the text-to-speech mapping relationship corresponding to the target subject, the edited speech can be generated quickly and accurately.
[0072] The text-to-speech mapping relationship in the above embodiment is obtained during the fine-tuning training of the speech correction model using the sample speech of the target object, and refers to the mapping relationship between a text and multiple speech segments. The text-to-speech mapping relationship table x u,i,j →y u,i,j It can be as follows:
[0073]
[0074] Among them, I is the voice dimension, J is the text dimension, ω i,j is the filtering weight of text j corresponding to the voice segment i of user u.
[0075] In a possible embodiment of the present application, the step of obtaining the text-to-speech mapping relationship may include: obtaining multiple speech segments from the sample speech of the target object, obtaining the text corresponding to each speech segment through the text network of the speech correction model, and determining the text-to-speech mapping relationship based on each speech segment and the text corresponding to each speech segment.
[0076] In other words, we can obtain multiple speech segments from the sample speech of the target object and input them into the text network of the speech correction model to obtain the text corresponding to each speech segment. Based on each speech segment and the text corresponding to each speech segment, we can determine the mapping relationship between the text and multiple speech segments, that is, the text-to-speech mapping relationship.
[0077] In a possible implementation of the present application, determining the degree of similarity between the voice message to be processed and the key voice message through a voice correction model may include: obtaining the voice features of the key voice message and the voice features of the voice message to be processed through the voice network of the voice correction model; obtaining the text-voice combination features of the key voice message based on the voice features of the key voice message through the text network of the voice correction model, and obtaining the text-voice combination features of the voice message to be processed based on the voice features of the voice message to be processed; determining the degree of similarity between the key voice message and the voice message to be processed based on the text-voice combination features of the key voice message and the text-voice combination features of the voice message to be processed through the similarity evaluation network of the voice correction model.
[0078] The speech correction model consists of three networks: a speech network, a text network, and a similarity assessment network. The speech network divides the input speech message into multiple speech segment latent vectors and outputs speech feature vectors that represent the speech characteristics of the input speech message, such as tone, intonation, speaking rate, and volume. In other words, the speech network reconstructs the input speech message in terms of tone, intonation, speaking rate, and volume, deriving weights for each of the speech segments in the message. The speech message is then reconstructed based on these weights. The reconstructed speech message is then fed into the text network, which converts it into a text-speech combination feature that includes both textual and speech characteristics. This means that the text obtained by the text network has both textual and speech characteristics. The text-speech combination features obtained by the speech network and the text network are then fed into the similarity assessment network to determine the degree of similarity between the two input speech messages.
[0079] In an embodiment of the present application, the key voice message and the voice message to be processed are processed separately by the three networks of the voice correction model, and then the degree of similarity is determined. Among them, the voice network can more accurately determine the voice characteristics of the input voice message, such as tone, intonation, speaking speed, volume, etc., so that after the voice message passes through the text network, the converted text can have text-level features and voice characteristics, making the converted text more accurate, so that when evaluated in the similarity evaluation network, the evaluation result can be more accurate.
[0080] like Figure 5 As shown in the figure, it is a structural diagram of the speech correction model. Figure 5 It can be seen that the speech correction model consists of three parts, namely the speech network, the text network and the similarity evaluation network. The speech message is input into the speech correction model, and the speech network of the speech correction model is divided into multiple speech segment latent vectors, such as Figure 6As shown, the speech feature vector can express the speech characteristics of the input voice message, such as tone, intonation, speaking rate, and volume. This means that the voice network can reconstruct the input voice message in terms of tone, intonation, speaking rate, and volume, deriving weights for various speech characteristics for the multiple voice segments of the voice message. The voice message is then reconstructed based on these weights. This differs from the prior art method of directly converting voice messages into text. In this application, after the voice message passes through the text network, the converted text can have both text-level and speech-related characteristics, making the converted text more accurate. The reconstructed voice message is then input into the text network, which converts the reconstructed voice message into a text-speech combination that includes both text-level and speech-related characteristics. This means that the text obtained through the text network has both text-level and speech-related characteristics. This differs from the prior art method of converting only text messages into voice messages. In this application, the converted voice message has the original speech characteristics. Finally, the text-speech combination features obtained by passing the two input voice messages through the voice and text networks can be input into a similarity assessment network to determine the degree of similarity between the two input voice messages.
[0081] In one possible implementation of the present application, based on the similarity level and a similarity level threshold, personalized voice correction is performed on the key voice message to obtain a target voice message, which may include: when the similarity level is less than the similarity level threshold, transmitting the similarity level to at least one of the voice conversion model and the text conversion model, so that at least one of the voice conversion model and the text conversion model adjusts their respective output results until the similarity level between the key voice message and the voice message to be processed is greater than or equal to the similarity level threshold, thereby obtaining the target voice message.
[0082] The larger the value of the similarity threshold is, the closer the final target voice message is to the voice message reformulated by the user who recorded the voice.
[0083] That is to say, the PID control (proportional-integral-derivative control) principle is used in the embodiment of the present application. The PID control principle is to form a control deviation based on a given value and an actual output value, and to form a control quantity by linearly combining the deviation in proportion, integration and differentiation to control the controlled object. In the present application, the given value refers to the similarity threshold, the actual output value refers to the similarity, the deviation refers to the difference between the similarity and the similarity threshold, and the controlled object refers to the speech correction model and / or text conversion model. The difference between the similarity and the similarity threshold is used to correct the text converted by the speech and the speech weight distribution, so that the text converted by the speech correction model is more accurate and the speech converted by the text conversion model is more similar to the speech characteristics of the original speech.
[0084] The PID control principle in the embodiment of the present application is mainly used for the model to actively provide real-time feedback to the speech conversion model and the text conversion model based on the difference between the similarity level and the similarity level threshold during use. The speech conversion model and the text conversion model correct the speech conversion text and the speech weight distribution based on the difference, so that the text converted by the speech correction model is more accurate, and the speech converted by the text conversion model is more similar to the speech characteristics of the original speech.
[0085] According to the above text-to-speech mapping relationship, the screening weights of the text corresponding to a voice segment are different. There may be errors in the weight distribution when converting voice to text or text to voice, so there will be errors in the output and input voice. When the similarity calculated later does not meet the similarity threshold, the error value is fed back to at least one of the voice conversion module and the text conversion model. The final similarity is adjusted by adjusting the output results of the voice conversion module and / or the text conversion model to obtain the target voice message.
[0086] In one possible implementation of the present application, obtaining a voice message to be processed and a target correction degree corresponding to the voice message to be processed may include: determining the voice message to be processed in response to a first input of a user to a voice message in a conversation interface; and determining the target correction degree corresponding to the voice message to be processed in response to a second input of a user to a voice correction control in the conversation interface.
[0087] The first input may be a selection input by item or a sliding selection input, which is not limited in the embodiments of this application and is determined according to actual applications. The second input may be an up and down sliding input or a click input, which is not limited in the embodiments of this application and is determined according to actual applications.
[0088] Specifically, by sliding to select multiple voice messages, you can slide to select multiple recorded voice messages, such as Figure 2 After entering the voice message processing interface, the user can manually select the target correction level through the voice correction control on the interface. Specifically, the user can slide the voice correction control up and down. Sliding up means that the user needs less correction, and sliding to the top means that only error correction is needed without text simplification. Sliding down means that the user needs more correction, simplifies the text, and outputs a summary. The system can also automatically confirm based on the conversation scenario, etc. At this time, the user can click the automatic button on the interface, such as Figure 4 As shown, the system can automatically confirm the degree of target correction based on the conversation scenario, etc.
[0089] In a possible implementation of the present application, the training steps of the speech correction model may include: using general sample speech to pre-train the speech correction model to be trained; using the sample speech of the target object to fine-tune the pre-trained speech correction model until the training is completed to obtain the speech correction model.
[0090] The general voice sample refers to public speech, not the speech of a specific person or group. It can be obtained from the internet or a voice database. The target subject's voice sample refers to the target subject's speech. This can be a voice message recorded by the target subject before the speech to be processed, or a voice message from the target subject stored in a local voice database authorized by the user's device. The speech correction model to be trained can leverage the model structure of an existing model, such as a joint task learning training model, eliminating the need to rebuild the model structure and conserving resources.
[0091] In an embodiment of the present application, the speech correction model is first pre-trained with a general sample speech, that is, a public speech, to obtain a general speech correction model, and then the sample speech of the target object is used to fine-tune the above-mentioned general model to obtain a speech correction model with the characteristics of the target object, so that the model is used to determine the similarity between the voice message to be processed and the key voice message, and when the key voice message is speech-corrected, the corrected voice message has more characteristics of the target object and is more similar to the target object's own re-expression.
[0092] Optionally, pre-training the speech correction model to be trained using a general sample speech can include: obtaining speech features of at least two sample speech messages in the general sample speech through the speech network of the speech correction model to be trained, wherein any two sample speech messages have a preset similarity; obtaining text-speech combination features of each sample speech message based on the speech features of at least two sample speech messages through the text network of the speech correction model to be trained; determining the similarity between any two sample speech messages based on the text-speech combination features of any two sample speech messages through the similarity evaluation network of the speech correction model to be trained; training the speech network and the text network based on the difference between the similarity between any two sample speech messages and the preset similarity between any two sample voice messages, until the similarity between any two sample voice messages is greater than or equal to the preset similarity, thereby obtaining a pre-trained speech correction model.
[0093] The speech correction model consists of three networks: a speech network, a text network, and a similarity assessment network. The speech network divides the input speech message into multiple speech segment latent vectors and outputs speech feature vectors that represent the speech characteristics of the input speech message, such as tone, intonation, speaking rate, and volume. In other words, the speech network reconstructs the input speech message in terms of tone, intonation, speaking rate, and volume, deriving weights for each of the speech segments in the message. The speech message is then reconstructed based on these weights. The reconstructed speech message is then fed into the text network, which converts it into a text-speech combination feature that includes both textual and speech characteristics. This means that the text obtained by the text network has both textual and speech characteristics. The text-speech combination features obtained by the speech network and the text network are then fed into the similarity assessment network to determine the degree of similarity between the two input speech messages.
[0094] In an embodiment of the present application, any two speech samples with a preset degree of similarity are passed through the above-mentioned speech network and text network to obtain their respective text-speech combined features. The two text-speech combined features are then input into a similarity assessment network to determine the similarity between the two speech samples. If the determined similarity between the two speech samples is greater than or equal to the preset similarity between the two speech samples, the speech correction model has been trained. Otherwise, the speech network and text network are continuously trained until the above conditions are met. Through the above-mentioned training process, a universal speech correction model can be trained so that the simplified speech message corrected by the speech correction model is closer to the original speech message.
[0095] After pre-training the speech correction model to be trained to obtain a general speech correction model, the general speech correction model can also be fine-tuned using the sample speech of the target object to obtain a speech correction model that conforms to the speech characteristics of the target object. Specifically, the fine-tuning training process may include: obtaining the speech features of at least two sample speech messages in the sample speech of the target object through the speech network of the trained speech correction model, wherein any two sample speech messages have a preset similarity; obtaining the text-speech combination features of each sample speech message based on the speech features of at least two sample speech messages through the text network of the trained speech correction model; determining the similarity between any two sample speech messages based on the text-speech combination features of any two sample speech messages through the similarity evaluation network of the trained speech correction model; training the speech network and the text network based on the difference between the similarity between any two sample speech messages and the preset similarity between any two sample speech messages until the similarity between any two sample speech messages is greater than or equal to the preset similarity, thereby obtaining a trained speech correction model.
[0096] Among them, the specific introduction of the speech network, text network and similarity evaluation network of the trained speech correction model has been described in detail in the above embodiment and will not be repeated in this embodiment.
[0097] In this embodiment of the present application, any two speech samples of the target object with a preset degree of similarity are passed through the above-mentioned speech network and text network to obtain their respective text-speech combination features. The two text-speech combination features are then input into the similarity assessment network to determine the similarity between the two speech samples. If the similarity between the two speech samples is greater than or equal to the preset similarity, the speech correction model has been trained. Otherwise, the speech network and text network are continuously trained until the above conditions are met. Through the above-mentioned training process, a speech correction model with the characteristics of the target object can be trained, so that the target speech message corrected by the speech correction model is closer to the speech message to be processed.
[0098] like Figure 7 The figure shows a simplified schematic diagram of the input and output structure of the speech correction model training. The input in the figure is a general sample speech and a sample speech of the target subject. During the model training process, the similarity between the two speech messages, the speech segment latent vector, the text-to-speech mapping relationship, the similarity threshold, and other information are obtained. These details have been described in detail in the above embodiments and will not be repeated in this embodiment.
[0099] like Figure 8The figure shows the overall flow chart of the voice message processing method of the present application. Specifically, when the voice message to be processed is obtained, the key voice message is obtained after the voice conversion model, the text error correction model, the text extraction model and the text conversion model. Then, the voice message to be processed and the key voice message are input into the voice correction model to obtain the target voice message. In this process, the voice correction model can feed back the difference between the similarity and the similarity threshold to the voice conversion model and the text conversion model in real time based on the similarity between the voice message to be processed and the key voice message, so that the voice conversion model and the text conversion model correct the voice-converted text and the voice weight distribution according to the difference, so that the text converted by the voice correction model is more accurate, and the voice converted by the text conversion model is more similar to the voice characteristics of the original voice. Specifically, it has been described in detail in the above embodiments and will not be repeated in this embodiment.
[0100] It should be noted that the voice message processing method provided in the embodiments of the present application can be executed by a voice message processing device or a control module in the voice message processing device for executing the voice message processing method. The embodiments of the present application use a voice message processing device executing the voice message processing method as an example to illustrate the voice message processing device provided in the embodiments of the present application.
[0101] like Figure 9 FIG. 1 is a schematic diagram of a voice message processing device provided by an embodiment of the present application. The voice message processing device may include: an acquisition module 901 , a determination module 902 , and a correction module 903 .
[0102] Among them, the acquisition module 901 is used to obtain the voice message to be processed and the target correction degree corresponding to the voice message to be processed; the determination module 902 is used to determine the similarity threshold according to the target correction degree; the correction module 903 is used to determine the similarity between the voice message to be processed and the key voice message through the voice correction model, and based on the similarity and the similarity threshold, perform personalized voice correction on the key voice message to obtain the target voice message; the key voice message corresponds to the voice message to be processed; wherein the voice correction model is obtained by training the sample voice of the target object; the target object is the message recording object corresponding to the voice message to be processed; the target voice message has the voice characteristics of the voice message to be processed and the voiceprint characteristics of the target object.
[0103] In an embodiment of the present application, first, the acquisition module 901 obtains the voice message to be processed and the target correction degree corresponding to the voice message to be processed, then the determination module 902 determines the similarity threshold value based on the target correction degree, and finally the correction module 903 determines the similarity between the voice message to be processed and the key voice message by using the voice correction model obtained by training the sample voice of the target object, and performs personalized voice correction on the key voice message based on the similarity and the similarity threshold value to obtain the target voice message. In an embodiment of the present application, the similarity threshold value is determined by the obtained target correction degree, and personalized voice correction is performed on the key voice message using the voice correction model, so that a target voice message with the voice characteristics of the voice message to be processed and the voiceprint characteristics of the target object can be obtained. This can avoid the user from re-recording when an error occurs in recording a voice message, thereby improving the efficiency of obtaining voice information. At the same time, the tone, intonation, speed and other information of the target voice message obtained in the end can be made the same as that of the user when recording the voice message, so that the final voice message is more like what the message recorder said, thereby improving the accuracy of voice message processing.
[0104] Optionally, the determination module 902 can be used to: convert the voice message to be processed into message text through a voice conversion model; extract key content from the message text to obtain key text through a text extraction model; and convert the key text into a key voice message through a text conversion model.
[0105] Optionally, the determination module 902 may be configured to: determine a content simplification degree according to a target correction degree; and extract key content from the message text using a text extraction model that matches the content simplification degree to obtain key text.
[0106] Optionally, the correction module 903 can be used to: obtain the voice features of the key voice message and the voice features of the voice message to be processed through the voice network of the voice correction model; obtain the text-voice combination features of the key voice message based on the voice features of the key voice message through the text network of the voice correction model, and obtain the text-voice combination features of the voice message to be processed based on the voice features of the voice message to be processed; determine the degree of similarity between the key voice message and the voice message to be processed based on the text-voice combination features of the key voice message and the text-voice combination features of the voice message to be processed through the similarity evaluation network of the voice correction model.
[0107] Optionally, the correction module 903 can be used to: when the similarity is less than a similarity threshold, pass the similarity to at least one of the speech conversion model and the text conversion model, so that at least one of the speech conversion model and the text conversion model adjusts their respective output results until the similarity between the key voice message and the voice message to be processed is greater than or equal to the similarity threshold, thereby obtaining the target voice message.
[0108] Optionally, the acquisition module 901 can be used to: determine the voice message to be processed in response to the user's first input of the voice message in the conversation interface; and determine the target correction degree corresponding to the voice message to be processed in response to the user's second input of the voice correction control in the conversation interface.
[0109] Optionally, the correction module 903 can be used to: pre-train the speech correction model to be trained using general sample speech; fine-tune the pre-trained speech correction model using the sample speech of the target object until the training is completed to obtain the speech correction model.
[0110] The voice message processing device in the embodiments of the present application can be a device, or a component, integrated circuit, or chip in a terminal. The device can be a mobile electronic device or a non-mobile electronic device. For example, the mobile electronic device can be a mobile phone, a tablet computer, a laptop computer, a PDA, an in-vehicle electronic device, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc. The non-mobile electronic device can be a server, a network attached storage (NAS), a personal computer (PC), a television (TV), a teller machine, or a self-service machine, etc., and the embodiments of the present application do not specifically limit this.
[0111] The voice message processing device in the embodiment of the present application may be a device having an operating system. The operating system may be an Android operating system, an iOS operating system, or other possible operating systems, which are not specifically limited in the embodiment of the present application.
[0112] The voice message processing device provided in the embodiment of the present application can achieve Figures 1-8 To avoid repetition, the various processes implemented in the illustrated method embodiment will not be described again here.
[0113] Alternatively, as Figure 10As shown, an embodiment of the present application also provides an electronic device 1000, including a processor 1001, a memory 1002, and a program or instruction stored in the memory 1002 and executable on the processor 1001. When the program or instruction is executed by the processor 1001, each process of the above-mentioned voice message processing method embodiment is implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.
[0114] It should be noted that the electronic devices in the embodiments of the present application include the mobile electronic devices and non-mobile electronic devices mentioned above.
[0115] Figure 11 A schematic diagram of the hardware structure of an electronic device implementing an embodiment of the present application.
[0116] The electronic device 1100 includes but is not limited to components such as a radio frequency unit 1101 , a network module 1102 , an audio output unit 1103 , an input unit 1104 , a sensor 1105 , a display unit 1106 , a user input unit 1107 , an interface unit 1108 , a memory 1109 , and a processor 1110 .
[0117] Those skilled in the art will understand that the electronic device 1100 may also include a power source (such as a battery) to power each component, and the power source may be logically connected to the processor 1110 through a power management system, thereby implementing functions such as charging, discharging, and power consumption management through the power management system. Figure 11 The electronic device structure shown in the figure does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently, which will not be repeated here.
[0118] Among them, the processor 1110 is used to obtain a voice message to be processed and a target correction degree corresponding to the voice message to be processed; determine a similarity threshold based on the target correction degree; determine the similarity between the voice message to be processed and the key voice message through a voice correction model, and perform personalized voice correction on the key voice message based on the similarity and the similarity threshold to obtain a target voice message; the key voice message corresponds to the voice message to be processed; wherein the voice correction model is obtained by training using a sample voice of a target object; the target object is a message recording object corresponding to the voice message to be processed; the target voice message has the voice characteristics of the voice message to be processed and the voiceprint characteristics of the target object.
[0119] In an embodiment of the present application, first, a voice message to be processed and a target correction degree corresponding to the voice message to be processed are obtained, and then a similarity threshold is determined based on the target correction degree. Finally, a voice correction model obtained by training the sample voice of the target object is used to determine the similarity between the voice message to be processed and the key voice message. Based on the similarity and the similarity threshold, personalized voice correction is performed on the key voice message to obtain the target voice message. In an embodiment of the present application, a similarity threshold is determined by obtaining the target correction degree, and personalized voice correction is performed on the key voice message using the voice correction model. A target voice message having the voice characteristics of the voice message to be processed and the voiceprint characteristics of the target object can be obtained. This can avoid the user from re-recording when an error occurs in recording a voice message, thereby improving the efficiency of obtaining voice information. At the same time, the tone, intonation, speed, and other information of the target voice message obtained in the end can be made the same as those when the user recorded the voice message, so that the final voice message is more like what the message recorder said, thereby improving the accuracy of voice message processing.
[0120] It should be understood that in an embodiment of the present application, the input unit 1104 may include a graphics processing unit (GPU) 11041 and a microphone 11042, and the graphics processor 11041 processes the image data of a static picture or video obtained by an image capture device (such as a camera) in a video capture mode or an image capture mode. The display unit 1106 may include a display panel 11061, and the display panel 11061 may be configured in the form of a liquid crystal display, an organic light emitting diode, etc. The user input unit 1107 includes a touch panel 11071 and other input devices 11072. The touch panel 11071 is also called a touch screen. The touch panel 11071 may include two parts: a touch detection device and a touch controller. Other input devices 11072 may include but are not limited to a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, and an operating stick, which will not be repeated here. The memory 1109 can be used to store software programs and various data, including but not limited to applications and operating systems. The processor 1110 may integrate an application processor and a modem processor, wherein the application processor mainly processes the operating system, user interface, and application programs, and the modem processor mainly processes wireless communications. It is understood that the modem processor may not be integrated into the processor 1110.
[0121] An embodiment of the present application also provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the various processes of the above-mentioned voice message processing method embodiment are implemented and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.
[0122] The processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes a computer-readable storage medium, such as a computer read-only memory (ROM), random access memory (RAM), a magnetic disk, or an optical disk.
[0123] An embodiment of the present application further provides a chip, which includes a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the various processes of the above-mentioned voice message processing method embodiment, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.
[0124] It should be understood that the chip mentioned in the embodiments of the present application can also be called a system-level chip, a system chip, a chip system or a system-on-chip chip, etc.
[0125] It should be noted that, in this article, the terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the statement "comprises a ..." does not exclude the presence of other identical elements in the process, method, article or device comprising the element. In addition, it should be noted that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in the opposite order according to the functions involved. For example, the described method may be performed in an order different from that described, and various steps may also be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.
[0126] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a computer software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), including a number of instructions for enabling a terminal (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in each embodiment of the present application.
[0127] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of this application, ordinary technicians in this field can also make many forms without departing from the purpose of this application and the scope of protection of the claims, all of which are within the protection of this application.
Claims
1. A method for processing a voice message, characterized in that: include: Obtaining a voice message to be processed and a target correction degree corresponding to the voice message to be processed; Determining a similarity threshold based on the target correction degree; Determining the similarity between the to-be-processed voice message and the key voice message through a voice correction model, and performing personalized voice correction on the key voice message based on the similarity and the similarity threshold to obtain a target voice message; The speech correction model is trained using sample speech of a target object; the target object is a message recording object corresponding to the voice message to be processed; the target voice message has the speech characteristics of the voice message to be processed and the voiceprint characteristics of the target object; and the step of obtaining the key voice message includes: Converting the voice message to be processed into a message text through a voice conversion model; Determine the degree of content simplification based on the degree of target revision; Extracting key content from the message text using a text extraction model that matches the degree of content simplification to obtain key text; The key text is converted into the key voice message through a text conversion model.
2. The method according to claim 1, characterized in that Determining the similarity between the to-be-processed voice message and the key voice message by using the voice correction model includes: Acquiring the voice features of the key voice message and the voice features of the voice message to be processed through the voice network of the voice correction model; Obtaining, through a text network of a speech correction model, text-speech combination features of the key speech message based on the speech features of the key speech message, and obtaining text-speech combination features of the speech message to be processed based on the speech features of the speech message to be processed; The similarity evaluation network of the speech correction model is used to determine the similarity between the key speech message and the speech message to be processed based on the text-speech combination features of the key speech message and the text-speech combination features of the speech message to be processed.
3. The method according to claim 1, characterized in that The performing personalized voice correction on the key voice message based on the similarity level and the similarity level threshold to obtain a target voice message includes: If the degree of similarity is less than the similarity threshold, the degree of similarity is transmitted to at least one of the speech conversion model and the text conversion model, so that at least one of the speech conversion model and the text conversion model adjusts their respective output results until the degree of similarity between the key voice message and the voice message to be processed is greater than or equal to the similarity threshold, thereby obtaining the target voice message.
4. The method according to claim 1, wherein Obtaining a voice message to be processed and a target correction degree corresponding to the voice message to be processed, including: In response to a first input of a voice message by a user in the conversation interface, determining the voice message to be processed; In response to a second input of the user to the voice correction control in the conversation interface, a target correction degree corresponding to the voice message to be processed is determined.
5. The method according to claim 1, characterized in that The training steps of the speech correction model include: Use common sample speech to pre-train the speech correction model to be trained; The pre-trained speech correction model is fine-tuned using the sample speech of the target object until the training is completed to obtain the speech correction model.
6. A voice message processing device, characterized in that: include: An acquisition module, configured to acquire a voice message to be processed and a target correction degree corresponding to the voice message to be processed; a determination module, configured to determine a similarity threshold value according to the target correction degree; a correction module, configured to determine the degree of similarity between the voice message to be processed and the key voice message through a voice correction model, and perform personalized voice correction on the key voice message based on the degree of similarity and a similarity threshold to obtain a target voice message; The speech correction model is trained using a sample speech of a target object; the target object is a message recording object corresponding to the voice message to be processed; the target voice message has the speech characteristics of the voice message to be processed and the voiceprint characteristics of the target object; The determining module is further configured to: Converting the voice message to be processed into a message text through a voice conversion model; Determine the degree of content simplification based on the degree of target revision; Extracting key content from the message text using a text extraction model that matches the degree of content simplification to obtain key text; The key text is converted into the key voice message through a text conversion model.
7. An electronic device, characterized in that: The method comprises a processor, a memory, and a program or instruction stored in the memory and executable on the processor, wherein the program or instruction, when executed by the processor, implements the steps of the voice message processing method according to any one of claims 1 to 5.
8. A readable storage medium, characterized in that: The readable storage medium stores a program or instruction, and when the program or instruction is executed by the processor, the steps of the voice message processing method according to any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Intelligent voice signature method and device based on text semantic similarity and medium
CN110502610A
Voice data processing method and device, and storage medium
CN110517689A