Voice message processing method and apparatus

CN116708342BActive Publication Date: 2026-09-11VIVO MOBILE COMM CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310753532.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-21
Publication Date
2026-09-11
Estimated Expiration
2043-06-21

AI Technical Summary

Technical Problem

[0003]但是,由于各种原因,消息接收用户有时会错误理解语音消息的内容,导致双方的沟通效率低

Benefits of technology

[0011] In a seventh aspect, embodiments of this application provide a chip including a processor and a communication interface coupled to the processor, the processor being used to run programs or instructions to implement the steps of the methods as described in the first and/or second aspects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116708342B_ABST
    Figure CN116708342B_ABST
Patent Text Reader

Abstract

The application discloses a voice message processing method and device, and belongs to the technical field of communication. The voice message processing method is applied to a first electronic device corresponding to a message sending user, and comprises the following steps: in the case that a voice conversation message input by the message sending user is obtained, a first input is received; in response to the first input, key information associated with the voice conversation message is obtained; the voice conversation message and the key information are sent to a message receiving user, so that the voice conversation message and the key information are displayed on a second electronic device corresponding to the message receiving user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of communication technology, specifically relating to a voice message processing method and apparatus. Background Technology

[0002] With the development of chat software and voice recognition technology, voice, as an information carrier, has become increasingly familiar and used, especially in chat scenarios. Compared to text input, voice input is more convenient and can also preserve timbre information;

[0003] However, for various reasons, message recipients sometimes misunderstand the content of voice messages, leading to low communication efficiency for both parties. Summary of the Invention

[0004] The purpose of this application is to provide a voice message processing method, apparatus, electronic device, and storage medium that improves the communication efficiency between the two parties in a conversation.

[0005] In a first aspect, embodiments of this application provide a voice message processing method applied to a first electronic device corresponding to a message sending user. The method includes: receiving a first input when a voice conversation message is received from the message sending user; in response to the first input, obtaining key information associated with the voice conversation message; and sending the voice conversation message and the key information to a message receiving user to display the voice conversation message and the key information in a second electronic device corresponding to the message receiving user.

[0006] Secondly, embodiments of this application provide a voice message processing method applied to a second electronic device corresponding to a message receiving user. The method includes: receiving a seventh input to the voice conversation message when a voice conversation message sent by a message sending user is received; and displaying a second voice conversation text corresponding to the voice conversation message in response to the seventh input. The second voice conversation text is obtained by the message sending user replacing a first word in a first voice conversation text with a second word, and the first voice conversation text is obtained by text conversion of the voice conversation message.

[0007] Thirdly, embodiments of this application provide a voice message processing device applied to a first electronic device corresponding to a message sending user. The device includes: a first receiving module, configured to receive a first input when a voice conversation message input by the message sending user is obtained; an acquisition module, configured to acquire key information associated with the voice conversation message in response to the first input; and a sending module, configured to send the voice conversation message and the key information to a message receiving user, so as to display the voice conversation message and the key information in a second electronic device corresponding to the message receiving user.

[0008] Fourthly, embodiments of this application provide a voice message processing device applied to a second electronic device corresponding to a message receiving user. The device includes: a second receiving module, configured to receive a seventh input to the voice conversation message when a voice conversation message sent by a message sending user is received; and a display module, configured to display a second voice conversation text corresponding to the voice conversation message in response to the seventh input; wherein the second voice conversation text is obtained by the message sending user replacing a first word in a first voice conversation text with a second word, and the first voice conversation text is obtained by text conversion of the voice conversation message.

[0009] Fifthly, embodiments of this application provide an electronic device including a processor and a memory, the memory storing a program or instructions executable on the processor, the program or instructions, when executed by the processor, implementing the steps of the methods as described in the first and / or second aspects.

[0010] In a sixth aspect, embodiments of this application provide a readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the methods as described in the first and / or second aspects.

[0011] In a seventh aspect, embodiments of this application provide a chip including a processor and a communication interface coupled to the processor, the processor being used to run programs or instructions to implement the steps of the methods as described in the first and / or second aspects.

[0012] Eighthly, embodiments of this application provide a computer program product stored in a storage medium, which is executed by at least one processor to implement the methods of the first aspect and / or the second aspect.

[0013] In this embodiment of the application, when a voice conversation message input by the message sending user is obtained, a first input is received, and in response to the first input, key information associated with the voice conversation message is obtained. The voice conversation message and the key information are then sent to the message receiving user so that the voice conversation message and the key information can be displayed on the second electronic device corresponding to the message receiving user. This avoids the message receiving user from misunderstanding the content of the voice message and improves the communication efficiency between the two parties in the conversation. Attached Figure Description

[0014] Figure 1 One of the flowcharts of a voice message processing method in some embodiments of this application is shown;

[0015] Figure 2 A schematic diagram of the display interface of a first electronic device in some embodiments of this application is shown;

[0016] Figure 3This is a second schematic diagram of the display interface of the first electronic device in some embodiments of this application;

[0017] Figure 4 The third schematic diagram shows the display interface of the first electronic device in some embodiments of this application;

[0018] Figure 5 The fourth schematic diagram shows the display interface of the first electronic device in some embodiments of this application;

[0019] Figure 6 Fifth of several schematic diagrams showing the display interface of the first electronic device in some embodiments of this application;

[0020] Figure 7 Sixth schematic diagram of the display interface of the first electronic device in some embodiments of this application is shown;

[0021] Figure 8 The seventh illustration shows a display interface diagram of a first electronic device in some embodiments of this application;

[0022] Figure 9 Eighth schematic diagram of the display interface of the first electronic device in some embodiments of this application is shown;

[0023] Figure 10 The second schematic flowchart of a voice message processing method in some embodiments of this application is shown;

[0024] Figure 11 One of the schematic diagrams of the display interface of the second electronic device in some embodiments of this application is shown;

[0025] Figure 12 This is a second schematic diagram of the display interface of a second electronic device in some embodiments of this application;

[0026] Figure 13 One of the schematic block diagrams of a voice message processing apparatus according to some embodiments of this application is shown;

[0027] Figure 14 A second schematic block diagram of a voice message processing apparatus according to some embodiments of this application is shown;

[0028] Figure 15 Structural block diagrams of electronic devices in some embodiments of this application are shown;

[0029] Figure 16 The diagram shows a hardware structure schematic of an electronic device in some embodiments of this application. Detailed Implementation

[0030] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.

[0031] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0032] The following is in conjunction with the appendix Figures 1 to 16 The present application provides a detailed description of the voice message processing method, voice message processing device, electronic device, and storage medium provided in the embodiments of this application through specific implementation methods and application scenarios.

[0033] In some embodiments of this application, a voice message processing method is provided, applied to a first electronic device corresponding to a message sending user. Figure 1 This illustration shows one of the flowcharts of a voice message processing method provided in some embodiments of this application. For example... Figure 1 As shown, the voice message processing method includes:

[0034] Step 102: If a voice conversation message with user input is received, the first input is received;

[0035] In this embodiment, the user inputs voice conversation messages to a first electronic device via voice input. These voice conversation messages are then sent to a second electronic device belonging to another user for conversational communication.

[0036] Step 104: In response to the first input, obtain key information associated with the voice conversation message;

[0037] In this embodiment of the application, the first input is used to trigger the first electronic device to extract key information from the voice conversation message. After the first electronic device receives the first input, it can obtain key information related to the voice conversation message.

[0038] For example, the pronunciation of the key information matches the pronunciation of at least part of the speech in the voice conversation message. For instance, if the voice conversation message is "I'm driving to Hengshan tomorrow," then the key information is "Hengshan."

[0039] For example, after obtaining key information, the first electronic device displays the voice conversation message and the key information.

[0040] Step 106: Send the voice conversation message and key information to the message receiving user so that the voice conversation message and key information can be displayed on the second electronic device corresponding to the message receiving user.

[0041] In this embodiment of the application, the key information can be text information, and the semantics of the text information matches the semantics of the voice conversation message. The second electronic device can display the voice conversation message and the key information on the same screen.

[0042] In this embodiment of the application, after obtaining the key information, the first electronic device sends both the key information and the voice conversation message to the message receiving user. After receiving the voice conversation message and the key information, the message receiving user can display the voice conversation message and the key information on the second electronic device. The key information can serve as a prompt to the message receiving user, indicating the true meaning of the voice conversation message.

[0043] Figure 2 This illustration shows one of the schematic diagrams of the display interface of a first electronic device provided in some embodiments of this application. Figure 3 This is a second schematic diagram of the display interface of the first electronic device provided in some embodiments of this application, such as... Figure 2 and Figure 3 As shown, after the electronic device receives voice input, it displays the voice conversation message 202 on the display interface. The user can trigger the first electronic device to obtain the note information 302 by long-pressing the display interface and display the note information 302 on the display interface.

[0044] In this embodiment, the first electronic device can respond to the first input, obtain key information in the voice conversation message, and transmit the key information along with the voice conversation message to the receiving user's second electronic device for display. This can serve as a prompt to the receiving user and prevent the receiving user of the second electronic device from misunderstanding the semantics of the voice conversation message.

[0045] In this embodiment of the application, when a voice conversation message input by the message sending user is obtained, a first input is received, and in response to the first input, key information associated with the voice conversation message is obtained. The voice conversation message and the key information are then sent to the message receiving user so that the voice conversation message and the key information can be displayed on the second electronic device corresponding to the message receiving user. This avoids the message receiving user from misunderstanding the content of the voice message and improves the communication efficiency between the two parties in the conversation.

[0046] In some embodiments of this application, obtaining key information associated with a voice conversation message in response to a first input includes: displaying at least two information categories in response to the first input; receiving a second input for a target information category among the at least two information categories; determining target message content in the voice conversation message that matches the target information category in response to the second input; and determining key information associated with the voice conversation message based on the target message content. In embodiments of this application, the information category is a pre-preset category of key information. By executing the second input, the user can select a target information category among the at least two displayed information categories and extract the target message content from the voice conversation message through the target information category. The target message content is at least a portion of the message content in the voice conversation message. The first electronic device can determine the corresponding key information through the target message content.

[0047] Specifically, when the first electronic device displays at least two information categories, after the user selects the target information category from the at least two information categories through the second input, the first electronic device can generate corresponding target message content according to the target information category selected by the user. The target message content matches the target information category, enabling the user to select the target information category according to actual needs. The first electronic device is controlled to extract the target message content from the voice conversation message according to the target information category and generate key information accordingly, ensuring the accuracy of the generated key information.

[0048] For example, information categories include, but are not limited to: location categories, people's name categories, and item categories. Figure 4 The third schematic diagram shows the display interface of the first electronic device in some embodiments of this application. Figure 5 The fourth illustration shows a schematic diagram of the display interface of the first electronic device in some embodiments of this application, such as... Figure 2 , Figure 4 and Figure 5As shown, the voice session message 202 is displayed on the display interface of the first electronic device, and the content of the voice session message 202 is "Drive to Hengshan Mountain for fun tomorrow". When the user long-presses the voice session message 202, a "location category" button 402 and a "person name category" button 404 are displayed on the display interface. When the user clicks the "location category" button 402, the first electronic device can extract the corresponding target message content, where the target message content is "Hengshan Mountain", and the mobile phone can generate key information 502 based on the target message content. Herein, the user's long-pressing of the voice session message 202 is a first input, and the user's clicking of the "location category" button 402 is a second input.

[0049] In the embodiments of the present application, the user can select a target information category from at least two information categories through a second input, enabling the first electronic device to extract target message content that matches the target information category from the voice session message, so as to ensure the matching degree between the key information obtained based on the target message content and the target information category selected by the user.

[0050] In some embodiments of the present application, determining key information associated with the voice session message based on the target message content includes: displaying at least two candidate information items based on the target message content; receiving a third input for target information among the at least two candidate information items; and in response to the third input, taking the target information as the key information associated with the voice session message.

[0051] In the embodiments of the present application, the at least two candidate information items all match the target message content. The target message content is at least a part of the voice session message, and the target message content matches the target information category. The pronunciations of the at least two candidate information items determined based on the target message content are the same or similar, and the semantics of the at least two candidate information items may be different.

[0052] For example, at least two candidate information items have the same pronunciation, such as "Hengshan Mountain (a mountain in Hunan)" and "Hengshan Mountain (a mountain in Shanxi)", "muscle" and "chicken", "attack" and "rooster", "white dew" and "egret".

[0053] For example, at least two candidate information items have similar pronunciations, such as "affectedly artificial" and "tender and delicate", "insurance" and "fresh-keeping", "route" and "pass through", "Friday" and "noon".

[0054] In the embodiments of the present application, the target message content is the message content identified by the first electronic device through text conversion, and the category of the target message content matches the target information category. Due to the similar or identical pronunciations of different words or sentences, multiple candidate information items can be identified. The user performs a third input on the target information among the plurality of candidate information items to select the key information from the plurality of candidate information items.

[0055] For example, if the content of the voice conversation message is "I'm driving to Hengshan tomorrow", and the target information category is "location information", then the generated target message content is "Hengshan". The displayed candidate information has two options: "Hengshan" and "Hengshan". The user can then use "Hengshan" as the key information by performing a third input on "Hengshan".

[0056] In this embodiment, the first electronic device can display at least two candidate information with similar or related pronunciations based on the pronunciation information in the target message content. The user selects the target information from the at least two candidate information by performing a third input, and uses it as the corresponding key information. This allows the user to select the key information to be sent to the second electronic device for display according to actual needs, further improving the semantic matching degree between the key information transmitted to the second electronic device and the voice conversation message.

[0057] In some embodiments of this application, the target information category is the location category; the key information is the location information.

[0058] In this embodiment of the application, since some location names have similar or identical pronunciations, the user can extract the target message content of the location category in the voice conversation message by using the location category as the target information category, and generate at least two corresponding candidate information. The user can determine the accurate location information as the key information from the at least two candidate information.

[0059] Figure 6 The fifth illustration shows a schematic diagram of the display interface of the first electronic device in some embodiments of this application, such as... Figure 2 , Figure 4 , Figure 5 and Figure 6 As shown, the voice conversation message 202 is displayed on the display interface of the first electronic device. The content of the voice conversation message 202 is "Driving to Hengshan tomorrow". The user long-presses the voice conversation message 202, and the display interface shows the "Location Category" button 402 and the "Name Category" button 404. By clicking the "Location Category" button 402, the first electronic device can extract the corresponding "Hengshan" identifier 602 and "Hengshan" identifier 604. When the user clicks the "Hengshan" identifier 602, the phone identifies "Hengshan" as key information 502. Among them, the user long-pressing the voice conversation message 202 is the first input, the user clicking the "Location Category" button 402 is the second input, and the user clicking the "Hengshan" identifier 602 is the third input. "Hengshan" and "Hengshan" are candidate information.

[0060] In some embodiments of this application, obtaining key information associated with a voice conversation message in response to a first input includes: obtaining a first voice conversation text corresponding to the voice conversation message in response to the first input; receiving a fourth input on a first word in the first voice conversation text; replacing the first word with a second word in response to the fourth input to obtain a second voice conversation text; and using the second voice conversation text as key information.

[0061] In this embodiment of the application, the first input is used to trigger the first electronic device to perform text conversion on the voice conversation message to generate the first voice conversation text, and the pronunciation of the first voice conversation text matches the voice conversation message.

[0062] In this embodiment of the application, the fourth input is the input for editing the first word in the first voice conversation text. By editing the first word to the second word, an updated second voice conversation text is generated, and the second electronic device can use the second voice conversation text edited by the user as key information.

[0063] In this embodiment, the first electronic device responds to the first input by performing text conversion on the voice conversation message to obtain a first voice text whose pronunciation matches the voice conversation message. The first electronic device responds to the fourth input by replacing the first word in the first voice conversation text with a second word, thereby completing the editing of the second voice conversation text. After editing, the edited second voice text is sent to the second electronic device as key information. The second voice text can play the role of explaining the voice conversation message and improving the accuracy of the semantics of the voice conversation message obtained by the receiving user of the second electronic device.

[0064] For example, the display interface shows a text editing box and displays the first voice conversation text in the text editing box. The user can edit the first voice conversation text to obtain the second voice conversation text.

[0065] Figure 7 Sixth schematic diagram of the display interface of the first electronic device in some embodiments of this application is shown. Figure 8 The seventh illustration shows a schematic diagram of the display interface of the first electronic device in some embodiments of this application, such as... Figure 1 , Figure 7 and Figure 8As shown, the display interface shows a voice conversation message 202 and a text editing box 702. The user long-presses the voice conversation message 202, and the text editing box 702 displays "Want to go out for lunch together on Friday?". The user edits the word "Friday" in the text editing box 702, replacing it with "noon." The key information 804 then becomes "Want to go out for lunch together?". Here, "Want to go out for lunch together on Friday?" is the first voice conversation text, "Friday" is the first word, "noon" is the second word obtained after editing, and the key information is "Want to go out for lunch together?" in the text editing box 702. The user's long-press of the voice conversation message 202 is the first input, and the user's editing of "Friday" is the fourth input.

[0066] In this embodiment, the first electronic device can convert voice conversation messages into text to obtain first voice conversation text. The user replaces the first word in the first voice conversation text with the second word by executing the fourth input, and sends the generated second voice text as key information that conforms to the true semantics to the second electronic device. This can serve as a prompt to the receiving user of the second electronic device, further reducing the possibility that the message receiving user may have ambiguity about words with similar or similar pronunciations in the voice conversation message.

[0067] In some embodiments of this application, the voice message processing method further includes: receiving a fifth input of a third word in a second voice conversation text; and in response to the fifth input, treating the third word as key information.

[0068] In this embodiment, the fifth input is used by the user to select a third word from the second speech text. When the edited second speech text is displayed, the user selects the third word from the second speech text as key information through the fifth input, enabling the user to select a portion of the edited second speech text as key information according to actual needs, further improving the flexibility of generating speech text.

[0069] Figure 9 This is illustrated in diagram eight of some embodiments of the first electronic device provided in this application, such as... Figure 9 As shown, the display interface shows the second voice conversation text 902, which is "Shall we go out for lunch together?". The user can drag and select "lunch" to display "lunch" as key information 904 on the display interface. Among them, "Shall we go out for lunch together?" is the second voice conversation text, and the user drags and selects "lunch" as the fifth input.

[0070] In this embodiment, the first electronic device can display the edited second voice text. The user can select the third word in the second voice text as key information through the fifth input. This allows the user to select some words in the edited second voice conversation text as key information according to actual needs. The third word as key information is sent to the second electronic device along with the voice conversation message. This can serve as a prompt to the message receiving user and further reduce the possibility of the message receiving user having ambiguity about words with similar pronunciations in the voice conversation message.

[0071] In some embodiments of this application, obtaining key information associated with a voice conversation message in response to a first input includes: obtaining a first voice conversation text corresponding to the voice conversation message in response to the first input; receiving a sixth input on a fourth word in the first voice conversation text; and using the fourth word as key information in response to the sixth input.

[0072] In this embodiment, the first input is used to trigger the first electronic device to perform text conversion on the voice conversation message, generating the first voice conversation text. The sixth input is used to select a fourth word from the first voice conversation text as key information.

[0073] In this embodiment, the user triggers a first electronic device to obtain the first voice conversation text matching the voice of the voice conversation message through text conversion via a first input. When the first electronic device displays the first voice conversation text, the user selects the fourth word in the first voice conversation text as key information via a sixth input, enabling the user to select key information in the text converted from the original text according to actual needs.

[0074] In some embodiments of this application, a voice message processing method is provided, applied to a second electronic device corresponding to a message receiving user. Figure 10 This is a second schematic flowchart illustrating a voice message processing method provided in some embodiments of this application. For example... Figure 10 As shown, the voice message processing method includes:

[0075] Step 1002: Upon receiving a voice conversation message sent by the message sending user, receive the seventh input to the voice conversation message;

[0076] In this embodiment, the seventh input is used to trigger the second electronic device to display the voice conversation message and perform text conversion on the voice conversation message.

[0077] For example, the seventh input can be a gesture input such as long press input, short press input, drag input, or physical button trigger input, or voice input.

[0078] Step 1004: In response to the seventh input, display the second voice conversation text corresponding to the voice conversation message;

[0079] The second voice conversation text is obtained by replacing the first word in the first voice conversation text with the second word by the message sending user. The first voice conversation text is obtained by converting the voice conversation message into text.

[0080] In this embodiment, after receiving a voice conversation message sent by a user, the second electronic device can display the voice conversation message, and the first electronic device can transmit the edited second voice conversation text along with the voice conversation message to the second electronic device. The user can display the second voice conversation text sent by the first electronic device by performing a seventh input on the voice conversation message.

[0081] In this embodiment of the application, when the message receiving user controls the second electronic device to perform text conversion on the received voice conversation message, the second electronic device can display the second voice conversation text processed by the first electronic device as the text conversion result, which avoids the message receiving user from misunderstanding the content of the voice message and improves the communication efficiency between the two parties in the conversation.

[0082] Figure 11 This illustration shows one of the display interface diagrams of the second electronic device provided in some embodiments of this application, such as... Figure 11 As shown, the display interface shows voice conversation message 1102. When the user long-presses the voice conversation message 1102, the display interface shows the second voice conversation text 1104.

[0083] In this embodiment of the application, the message receiving user can perform a seventh input on the voice conversation message in the second electronic device to display the second voice conversation text sent by the first electronic device. The second voice conversation text is the text content edited by the message sending user, so that the message receiving user can understand the true intention of the voice conversation message based on the second voice conversation text, reducing the misinterpretation of the semantics of the voice conversation message caused by words with similar pronunciations.

[0084] In some embodiments of this application, the voice message processing method further includes: receiving an eighth input to the voice session message when a voice session message sent by a message sending user is received;

[0085] In response to the eighth input, display the fourth word corresponding to the voice conversation message;

[0086] The fourth word is extracted from the first voice conversation text by the message sender through the sixth input. The first voice conversation text is obtained by converting the voice conversation message into text.

[0087] In this embodiment, after receiving a voice conversation message sent by the user, the second electronic device can display the voice conversation message, and the first electronic device can transmit the fourth word along with the voice conversation message to the second electronic device. The user can display the fourth word sent by the first electronic device by performing an eighth input on the voice conversation message.

[0088] In this embodiment of the application, the first voice conversation text is the text content obtained by the first electronic device through text conversion, and the fourth vocabulary is the vocabulary content in the first voice conversation text.

[0089] Figure 12 This is a second schematic diagram of the display interface of the second electronic device provided in some embodiments of this application, such as... Figure 12 As shown, the display interface shows voice conversation message 1202. When the user long-presses the voice conversation message 1202, a note information prompt box 1204 is displayed in the display interface, and the fourth word "noon" and "meal" are displayed in the note information prompt box.

[0090] In this embodiment of the application, the message receiving user can perform an eighth input on the voice conversation message in the second electronic device to display the fourth word sent by the first electronic device. The fourth word is a word extracted by the message sending user in the first voice conversation text, so that the message receiving user can understand the true intention of the voice conversation message based on the second voice conversation text, and reduce the misinterpretation of the semantics of the voice conversation message caused by words with similar or similar pronunciations.

[0091] In some embodiments of this application, a voice message processing device is provided, applied to a first electronic device corresponding to a message sending user. Figure 13 A schematic block diagram of one of the voice message processing apparatuses provided in some embodiments of this application is shown, such as Figure 13 As shown, the voice message processing device 1300 includes:

[0092] The first receiving module 1302 is used to receive a first input when it receives a voice conversation message input by the user who sent the message;

[0093] The acquisition module 1304 is used to obtain key information associated with the voice conversation message in response to the first input;

[0094] The sending module 1306 is used to send voice conversation messages and key information to the message receiving user so that the voice conversation messages and key information can be displayed on the second electronic device corresponding to the message receiving user.

[0095] In this embodiment of the application, when a voice conversation message input by the message sending user is obtained, a first input is received, and in response to the first input, key information associated with the voice conversation message is obtained. The voice conversation message and the key information are then sent to the message receiving user so that the voice conversation message and the key information can be displayed on the second electronic device corresponding to the message receiving user. This avoids the message receiving user from misunderstanding the content of the voice message and improves the communication efficiency between the two parties in the conversation.

[0096] In some embodiments of this application, the acquisition module 1304 is further configured to: display at least two information categories in response to a first input; receive a second input for a target information category among the at least two information categories; determine, in response to the second input, target message content in the voice conversation message that matches the target information category; and determine key information associated with the voice conversation message based on the target message content.

[0097] In this embodiment of the application, the user can select a target information category from at least two information categories through a second input, enabling the first electronic device to extract the target message content in the voice conversation message that matches the target information category, thus ensuring the degree of matching between the key information obtained from the target message content and the target information category selected by the user.

[0098] In some embodiments of this application, the acquisition module 1304 is further configured to display at least two candidate information based on the target message content; receive a third input on the target information among the at least two candidate information; and, in response to the third input, use the target information as key information associated with the voice conversation message.

[0099] In this embodiment, the first electronic device can display at least two candidate information with similar or related pronunciations based on the pronunciation information in the target message content. The user selects the target information from the at least two candidate information by performing a third input, and uses it as the corresponding key information. This allows the user to select the key information to be sent to the second electronic device for display according to actual needs, further improving the semantic matching degree between the key information transmitted to the second electronic device and the voice conversation message.

[0100] In some embodiments of this application, the target information category is the location category; the key information is the location information.

[0101] In this embodiment of the application, since some location names have similar or identical pronunciations, the user can extract the target message content of the location category in the voice conversation message by using the location category as the target information category, and generate at least two corresponding candidate information. The user can determine the accurate location information as the key information from the at least two candidate information.

[0102] In some embodiments of this application, the acquisition module 1304 is configured to, in response to a first input, obtain a first voice conversation text corresponding to a voice conversation message; receive a fourth input on a first word in the first voice conversation text; in response to the fourth input, replace the first word with a second word to obtain a second voice conversation text; and use the second voice conversation text as key information.

[0103] In this embodiment, the first electronic device can convert voice conversation messages into text to obtain first voice conversation text. The user replaces the first word in the first voice conversation text with the second word by executing the fourth input, and sends the generated second voice text as key information that conforms to the true semantics to the second electronic device. This can serve as a prompt to the receiving user of the second electronic device, further reducing the possibility that the message receiving user may have ambiguity about words with similar or similar pronunciations in the voice conversation message.

[0104] In some embodiments of this application, the acquisition module 1304 is configured to receive a fifth input of a third word in the second voice conversation text; in response to the fifth input, the third word is used as key information.

[0105] In this embodiment, the first electronic device can display the edited second voice text. The user can select the third word in the second voice text as key information through the fifth input. This allows the user to select some words in the edited second voice conversation text as key information according to actual needs. The third word as key information is sent to the second electronic device along with the voice conversation message. This can serve as a prompt to the message receiving user and further reduce the possibility of the message receiving user having ambiguity about words with similar pronunciations in the voice conversation message.

[0106] In some embodiments of this application, the acquisition module 1304 is configured to, in response to the first input, acquire the first voice conversation text corresponding to the voice conversation message; receive a sixth input on the fourth word in the first voice conversation text; and, in response to the sixth input, use the fourth word as key information.

[0107] In this embodiment, the user triggers a first electronic device to obtain the first voice conversation text matching the voice of the voice conversation message through text conversion via a first input. When the first electronic device displays the first voice conversation text, the user selects the fourth word in the first voice conversation text as key information via a sixth input, enabling the user to select key information in the text converted from the original text according to actual needs.

[0108] In some embodiments of this application, a voice message processing device is provided, applied to a second electronic device corresponding to a message receiving user. Figure 14A second schematic block diagram of a voice message processing apparatus according to some embodiments of this application is shown, such as... Figure 14 As shown, the voice message processing device 1400 includes:

[0109] The second receiving module 1402 is used to receive a seventh input to the voice session message when it receives a voice session message sent by the message sending user.

[0110] Display module 1404 is used to display the second voice conversation text corresponding to the voice conversation message in response to the seventh input; wherein the second voice conversation text is obtained by the message sending user replacing the first word in the first voice conversation text with the second word, and the first voice conversation text is obtained by text conversion of the voice conversation message.

[0111] In this embodiment of the application, when the message receiving user controls the second electronic device to perform text conversion on the received voice conversation message, the second electronic device can display the second voice conversation text processed by the first electronic device as the text conversion result, thereby improving the matching degree between the text content displayed by the second electronic device and the voice conversation message.

[0112] In this embodiment of the application, when the message receiving user controls the second electronic device to perform text conversion on the received voice conversation message, the second electronic device can display the second voice conversation text processed by the first electronic device as the text conversion result, which avoids the message receiving user from misunderstanding the content of the voice message and improves the communication efficiency between the two parties in the conversation.

[0113] The voice message processing device in this application embodiment can be an electronic device or a component within an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices besides a terminal. For example, the electronic device can be a mobile phone, tablet computer, laptop computer, PDA, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc. It can also be a server, network attached storage (NAS), personal computer (PC), television set (TV), ATM, or self-service machine, etc. This application embodiment does not specifically limit the device.

[0114] The voice message processing device in this application embodiment can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit the specific operating system used.

[0115] The voice message processing device provided in this application embodiment can implement the various processes implemented in the above method embodiments, and will not be described again here to avoid repetition.

[0116] Optionally, embodiments of this application also provide an electronic device, which includes the voice message processing device as described in any of the above embodiments, and thus has all the beneficial effects of the voice message processing device in any of the embodiments, which will not be elaborated further here.

[0117] Optionally, embodiments of this application also provide an electronic device. Figure 15 Structural block diagrams of electronic devices provided in some embodiments of this application are shown, such as... Figure 15 As shown, the electronic device 1500 includes a processor 1502, a memory 1504, and a program or instructions stored in the memory 1504 and executable on the processor 1502. When the program or instructions are executed by the processor 1502, they implement the various processes of the above-described voice message processing method embodiments and achieve the same technical effects. To avoid repetition, they will not be described again here.

[0118] It should be noted that the electronic devices in the embodiments of this application include the aforementioned mobile electronic devices and non-mobile electronic devices.

[0119] Figure 16 The diagram shows a schematic representation of the hardware structure of an electronic device provided in some embodiments of this application.

[0120] The electronic device 1600 includes, but is not limited to, components such as: radio frequency unit 1601, network module 1602, audio output unit 1603, input unit 1604, sensor 1605, display unit 1606, user input unit 1607, interface unit 1608, memory 1609, and processor 1610.

[0121] Those skilled in the art will understand that the electronic device 1600 may also include a power supply (such as a battery) for supplying power to various components. The power supply may be logically connected to the processor 1610 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. Figure 16 The electronic device structure shown does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.

[0122] In the case where the electronic device is the first electronic device of the message sending user, the user input unit 1607 is used to receive the first input when a voice conversation message input by the user is obtained;

[0123] Processor 1610 is configured to, in response to a first input, obtain key information associated with a voice session message;

[0124] Processor 1610 is used to send voice conversation messages and key information to a message receiving user so as to display the voice conversation messages and key information on a second electronic device used by the message receiving user.

[0125] In this embodiment of the application, when a user sends a voice conversation message, the user can extract key information from the voice conversation message by executing a first input, and send the key information together with the voice conversation message to the receiving user. This enables the receiving user to display the semantically correct key information when receiving the voice conversation message, reducing the possibility of the receiving user having misunderstandings about the voice message.

[0126] Furthermore, the display unit 1606 is configured to display at least two information categories in response to the first input;

[0127] User input unit 1607 is configured to receive a second input for a target information category among at least two information categories;

[0128] The determination module is used to determine, in response to the second input, the target message content in the voice conversation message that matches the target information category;

[0129] Processor 1610 is used to determine key information associated with a voice session message based on the content of the target message.

[0130] In this embodiment of the application, the user can select a target information category from at least two information categories through a second input, enabling the first electronic device to extract the target message content in the voice conversation message that matches the target information category, thus ensuring the degree of matching between the key information obtained from the target message content and the target information category selected by the user.

[0131] Furthermore, the display unit 1606 is used to display at least two candidate messages based on the target message content;

[0132] User input unit 1607 is used to receive a third input for target information among at least two candidate information;

[0133] The processor 1610 is used to respond to a third input by treating the target information as key information associated with the voice session message.

[0134] In this embodiment, the first electronic device can display at least two candidate information with similar or related pronunciations based on the pronunciation information in the target message content. The user selects the target information from the at least two candidate information by performing a third input, and uses it as the corresponding key information. This allows the user to select the key information to be sent to the second electronic device for display according to actual needs, further improving the semantic matching degree between the key information transmitted to the second electronic device and the voice conversation message.

[0135] Furthermore, the target information category is the location category; the key information is the location information.

[0136] In this embodiment of the application, since some location names have similar or identical pronunciations, the user can extract the target message content of the location category in the voice conversation message by using the location category as the target information category, and generate at least two corresponding candidate information. The user can determine the accurate location information as the key information from the at least two candidate information.

[0137] Furthermore, the processor 1610 is configured to, in response to the first input, obtain the first voice conversation text corresponding to the voice conversation message;

[0138] User input unit 1607 is used to receive a fourth input of a first word in the first voice conversation text;

[0139] Processor 1610 is configured to, in response to a fourth input, replace the first word with the second word to obtain a second speech conversation text;

[0140] Processor 1610 is used to process the text of the second voice conversation as key information.

[0141] In this embodiment, the first electronic device can convert voice conversation messages into text to obtain first voice conversation text. The user replaces the first word in the first voice conversation text with the second word by executing the fourth input, and sends the generated second voice text as key information that conforms to the true semantics to the second electronic device. This can serve as a prompt to the receiving user of the second electronic device, further reducing the possibility that the message receiving user may have ambiguity about words with similar or similar pronunciations in the voice conversation message.

[0142] Furthermore, the user input unit 1607 is used to receive a fifth input of a third word in the second voice conversation text;

[0143] Processor 1610 is used to respond to the fifth input by taking the third word as key information.

[0144] In this embodiment, the first electronic device can display the edited second voice text. The user can select the third word in the second voice text as key information through the fifth input. This allows the user to select some words in the edited second voice conversation text as key information according to actual needs. The third word as key information is sent to the second electronic device along with the voice conversation message. This can serve as a prompt to the message receiving user and further reduce the possibility of the message receiving user having ambiguity about words with similar pronunciations in the voice conversation message.

[0145] Furthermore, the processor 1610 is configured to obtain the first voice conversation text corresponding to the voice conversation message in response to the first input; the user input unit 1607 is configured to receive a sixth input of the fourth word in the first voice conversation text;

[0146] Processor 1610 is used to respond to the sixth input by taking the fourth word as key information.

[0147] In this embodiment, the user triggers a first electronic device to obtain the first voice conversation text matching the voice of the voice conversation message through text conversion via a first input. When the first electronic device displays the first voice conversation text, the user selects the fourth word in the first voice conversation text as key information via a sixth input, enabling the user to select key information in the text converted from the original text according to actual needs.

[0148] In the case where the electronic device is the first electronic device of the message sending user, the user input unit 1607 is used to receive a seventh input to the voice conversation message when the voice conversation message sent by the message sending user is received.

[0149] Display unit 1606 is used to display the second voice conversation text corresponding to the voice conversation message in response to the seventh input;

[0150] The second voice conversation text is obtained by replacing the first word in the first voice conversation text with the second word by the message sending user. The first voice conversation text is obtained by converting the voice conversation message into text.

[0151] In this embodiment of the application, when the message receiving user controls the second electronic device to perform text conversion on the received voice conversation message, the second electronic device can display the second voice conversation text processed by the first electronic device as the text conversion result, thereby improving the matching degree between the text content displayed by the second electronic device and the voice conversation message.

[0152] It should be understood that, in this embodiment, the input unit 1604 may include a graphics processing unit (GPU) 16041 and a microphone 16042. The GPU 16041 processes image data of still images or videos obtained by an image capture device (such as a camera) in video capture mode or image capture mode. The display unit 1606 may include a display panel 16061, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, or the like. The user input unit 1607 includes at least one of a touch panel 16071 and other input devices 16072. The touch panel 16071 is also called a touch screen. The touch panel 16071 may include a touch detection device and a touch controller. Other input devices 16072 may include, but are not limited to, physical keyboards, function keys (such as volume control buttons, power buttons, etc.), trackballs, mice, and joysticks, which will not be described in detail here.

[0153] The memory 1609 can be used to store software programs and various data. The memory 1609 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, the memory 1609 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM). The memory 1609 in this embodiment includes, but is not limited to, these and any other suitable types of memory.

[0154] Processor 1610 may include one or more processing units; optionally, processor 1610 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into processor 1610.

[0155] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above method embodiments and achieve the same technical effect. To avoid repetition, they will not be described again here.

[0156] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0157] This application also provides a chip, which includes a processor and a communication interface. The communication interface and the processor are coupled. The processor is used to run programs or instructions to implement the various processes of the above-described voice message processing method embodiments and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0158] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.

[0159] This application provides a computer program product, which is stored in a storage medium and executed by at least one processor to implement the various processes of the above-described voice message processing method embodiments, and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0160] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0161] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of this application.

[0162] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. A voice message processing method, applied to a first electronic device corresponding to a message sending user, characterized in that, The method includes: If the user inputs a voice conversation message, the system receives the first input; In response to the first input, at least two information categories are displayed; Receive a second input for a target information category among the at least two information categories; In response to the second input, determine the target message content in the voice conversation message that matches the target information category; Determine key information associated with the voice session message based on the target message content; The voice conversation message and the key information are sent to the message receiving user so that the voice conversation message and the key information are displayed on the second electronic device corresponding to the message receiving user.

2. The voice message processing method according to claim 1, characterized in that, The step of determining the key information associated with the voice session message based on the target message content includes: Display at least two candidate messages based on the target message content; Receive a third input for the target information among the at least two candidate information; In response to the third input, the target information is used as key information associated with the voice session message.

3. The voice message processing method according to claim 1, characterized in that, The target information category is the location category; the key information is the location information.

4. A voice message processing method, applied to a second electronic device corresponding to a message receiving user, characterized in that, The method includes: Upon receiving a voice conversation message sent by the message sending user, a seventh input to the voice conversation message is received; In response to the seventh input, key information associated with the voice conversation message is displayed; The key information is obtained by the message sending user processing the first voice conversation text, which is obtained by converting the voice conversation message into text. The key information is obtained by the message sending user through the following methods: displaying at least two information categories; receiving a second input for a target information category among the at least two information categories; in response to the second input, determining the target message content in the voice conversation message that matches the target information category; and determining the key information based on the target message content.

5. A voice message processing device, applied to a first electronic device corresponding to a message sending user, characterized in that, The device includes: The first receiving module is configured to receive a first input when it receives a voice conversation message input by the user who sent the message; An acquisition module is configured to, in response to the first input, display at least two information categories; receive a second input for a target information category among the at least two information categories; in response to the second input, determine target message content in the voice conversation message that matches the target information category; and determine key information associated with the voice conversation message based on the target message content. The sending module is used to send the voice conversation message and the key information to the message receiving user, so as to display the voice conversation message and the key information on the second electronic device corresponding to the message receiving user.

6. A voice message processing device, applied to a second electronic device corresponding to a message receiving user, characterized in that, The device includes: The second receiving module is used to receive a seventh input to the voice conversation message when it receives the voice conversation message sent by the message sending user; A display module is configured to display key information associated with the voice conversation message in response to the seventh input. The key information is obtained by the message sending user processing the first voice conversation text, which is obtained by converting the voice conversation message into text. The key information is obtained by the message sending user through the following methods: displaying at least two information categories; receiving a second input for a target information category among the at least two information categories; in response to the second input, determining the target message content in the voice conversation message that matches the target information category; and determining the key information based on the target message content.

Citation Information

Patent Citations

  • Voice message processing method and device

    CN110099360A

  • Message processing method and device and terminal equipment

    CN110392158A