Information processing method and apparatus, and electronic device

By generating voice broadcasts and processing information that resemble the voice of the target contact, the problem of electronic devices being unable to distinguish the voice of a contact has been solved, thus improving the efficiency and timeliness of information acquisition.

WO2026012240A1PCT designated stage Publication Date: 2026-01-15HUAWEI TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/106282
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-11-01
Filing Date
2025-06-30
Publication Date
2026-01-15

AI Technical Summary

Technical Problem

In existing technologies, electronic devices cannot effectively distinguish the voices of different contacts when broadcasting information via voice, making it difficult for users to establish an auditory connection between information and contacts. Furthermore, the diverse formats of the broadcast content lead to low efficiency.

Method used

By retrieving voice messages from the target contact, a target voice similar to their real voice is generated. Based on the information type and user familiarity, the target information is then simplified or merged and broadcast.

Benefits of technology

It improves the efficiency and timeliness of users obtaining information, establishes familiarity with contacts through auditory means, and reduces redundant broadcast time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025106282_15012026_PF_FP_ABST
    Figure CN2025106282_15012026_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of the present application are applicable to the technical field of information, and provide an information processing method and apparatus, and an electronic device. By applying the method, when receiving information sent by a contact, an electronic device can acquire contact information on the basis of a user instruction or by means of analyzing the information, and determine a target contact who sent the information. In addition, by comprehensively considering factors such as the sources of a previous piece of information and a current piece of information, the familiarity between a user and the target contact, and the type and complexity of the information, a corresponding playback control policy can be determined. On this basis, the authorization to use sound features of the target contact is obtained from the target contact by means of permission control, and sounds identical or similar to sounds of the target contact can be synthesized for voice playback of information processed according to the playback control policy. By applying the method, it can not only be convenient for the user to develop a concrete sense of familiarity with the contact from the auditory level, but also improve the efficiency and timeliness with which the user acquires information content.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing methods, devices and electronic equipment

[0001] This application claims priority to Chinese invention patent application No. 202410926760.9, filed with the State Intellectual Property Office on July 10, 2024, entitled "Method, Apparatus and Electronic Device for Intelligent Transmission of Information Content", and to Chinese invention patent application No. 202411562369.1, filed with the State Intellectual Property Office on November 1, 2024, entitled "Information Processing Method, Apparatus and Electronic Device". Technical Field

[0002] This application relates to the field of information technology, and in particular to an information processing method, apparatus, and electronic device. Background Technology

[0003] Electronic devices can broadcast received information to users via voice, enriching the ways users access information. In some scenarios, when it's inconvenient for users to directly view information on their electronic devices, voice broadcasts improve the timeliness of information access. For example, while driving, running, or cycling, users may not be able to visually view system messages or information sent by various applications (apps) received on their electronic devices. For instance, it might be inconvenient for users to view messages from contacts received through social media apps. In such cases, electronic devices can transcribe and broadcast the received information to the user via voice, making it easier for the user to understand the content.

[0004] In existing technologies, electronic devices typically use system default or custom audio sources to broadcast information via voice. Whether it's system messages or information transmitted by various apps, electronic devices use the same audio data for broadcasting. In one example, for messages from different contacts in communication or social apps, the electronic device uses the same voice to broadcast them, which can be confusing for users. Users cannot establish a direct auditory connection between the broadcast information and the sender, reducing the efficiency of information retrieval. Furthermore, electronic devices receive information from a wide range of sources, and the content to be broadcast is diverse. For example, the information received by the electronic device may be long text or contain links, special characters, etc. Existing technologies will broadcast these elements during voice broadcasting, resulting in longer broadcast times and hindering users from quickly understanding the specific content of the information. Summary of the Invention

[0005] This application provides an information processing method, apparatus, and electronic device that can use a voice identical or similar to the real voice of the target contact person to broadcast received information content, allowing users to develop a concrete sense of familiarity with the target contact person through auditory perception. Simultaneously, this application can also process the received information, such as simplifying the information, retaining key or important content, supplementing the information source and / or contact person source based on the differences between the current and previous information, and converting non-text information into corresponding broadcastable text information according to different information types. Through these various processing methods, the efficiency and timeliness of users obtaining information content can be improved.

[0006] To achieve the above objectives, the embodiments of this application adopt the following technical solutions:

[0007] The first aspect of this application provides an information processing method, including:

[0008] In response to the received first information, determine the target contact to whom the first information was sent;

[0009] Retrieve voice messages of the target contact and generate a target voice similar to the voice of the target contact based on the voice messages. The voice messages may include voice messages on the first page.

[0010] Generate target information corresponding to the first information;

[0011] The target information is broadcast using the target's voice.

[0012] It should be understood that the first information can be currently received information, and the target contact is the contact who sent the first information. The target contact can send the first information through an application. By applying the information processing method provided in the embodiments of this application, the electronic device can use the same or similar voice as the target contact who sent the first information to broadcast the received first information, thereby facilitating the user to establish a sense of familiarity between the information and the target contact from an auditory perspective, which helps the user to obtain information content efficiently and in a timely manner.

[0013] Simultaneously, after receiving the first information, the electronic device can process it according to a corresponding broadcasting strategy. For example, it can convert non-text information into text information, simplify the text information, and supplement or omit the information source and / or contact person source based on the differences between adjacent information. The target information broadcast by the electronic device using a target voice that is the same as or similar to the target contact person's real voice can be information processed according to the above broadcasting strategy. This can further improve the efficiency of users obtaining information.

[0014] In one possible implementation of the first aspect of this application, the method further includes: retrieving the voice message of the target contact; and generating a target voice that is the same as or similar to the real voice of the target contact based on the voice message of the target contact.

[0015] The retrieval of voice messages from the target contact can be performed on the first page, which may include the conversation page between the current user and the target contact. For example, the conversation page may be a chat page provided by an instant messaging application.

[0016] It should be understood that the target voice is a voice that is the same as or similar to the real voice of the target contact sending the information, and the target voice can be generated based on the voice characteristics of the target contact. This application embodiment retrieves the voice messages of the target contact, extracts the voice characteristics of the target contact from the retrieved voice messages, and synthesizes a target voice that is the same as or similar to the real voice of the target contact for subsequent information broadcasting. This application embodiment does not require the development of third-party application interfaces and can achieve automated and seamless voice retrieval through the system.

[0017] As an example of a first aspect of this application, retrieving the voice messages of the target contact can be achieved through a tagging retrieval method. During tagging retrieval, the electronic device can determine retrieval elements for retrieving the voice messages of the target contact, the retrieval elements including at least the contact information of the target contact; based on the retrieval elements, a search is performed in a tagging information database to obtain the voice messages of the target contact. The tagging information database can be constructed by storing corresponding tagging information after tagging the received voice messages of each contact.

[0018] As another example of the first aspect of this application, retrieving voice messages from the target contact can also be achieved through a page-turning search. When performing a page-turning search, the electronic device can open a first page associated with the target contact in the application, such as the aforementioned conversation page. Then, voice messages sent by the target contact can be retrieved from the first page.

[0019] The first page may include the current conversation page of the target contact, which is displayed directly after the application is opened; or, the first page may also include the historical conversation page between the current user and the target contact. The historical conversation page can be displayed by swiping the current conversation page.

[0020] For example, the first page may include the conversation page between the current user and the target contact. When performing a pagination search, the first page associated with the target contact, i.e., the conversation page between the current user and the target contact, can be displayed, and voice messages sent by the target contact can be retrieved in the conversation page.

[0021] Alternatively, the first page may also include a history page of conversations between the current user and the target contact. In response to the first action, the history page of conversations between the current user and the target contact can be displayed, and voice messages sent by the target contact can be retrieved from the history page. The first action can be swiping the conversation page or other actions capable of achieving the above functionality. For example, swiping the conversation page: in response to the swiping action, the electronic device can display the history page of conversations between the current user and the target contact, and then retrieve voice messages from the history page.

[0022] In one possible implementation of the first aspect of this application, the first page further includes a personal conversation page for direct interaction with the target contact and / or a group conversation page containing the target contact. Therefore, retrieving voice messages sent by the target contact from the first page includes retrieving voice messages sent by the target contact from the personal conversation page and / or the group conversation page. This increases the likelihood of accurately retrieving voice messages from the target contact by performing searches separately on the personal conversation page and the group conversation page.

[0023] In another possible implementation of the first aspect of the embodiments of this application, the voice message further includes voice messages on a second page. The second page may include a historical message record retrieval page. Therefore, when retrieving voice messages of a target contact, in response to the second operation, the second page associated with the target contact can be displayed, and voice messages sent by the target contact can be retrieved on the second page.

[0024] In one example, the second operation could be performed on a historical message entry. For instance, an electronic device could display a historical message record retrieval page by performing an operation on the historical message entry, and then retrieve voice messages sent by the target contact within that page. The operation performed by the electronic device on the historical message entry could be implemented by calling a corresponding interface or simulating user behavior. The aforementioned method of displaying a historical message record retrieval page associated with the target contact through a historical message entry provided by the application, and then performing a retrieval within that page, is a search retrieval.

[0025] The embodiments of this application can retrieve voice messages of target contacts through multiple retrieval methods, expanding the sources of voice messages of target contacts and helping to obtain more useful voice messages.

[0026] In one possible implementation of the first aspect of this application, when an electronic device generates a target voice that is the same as or similar to the real voice of the target contact based on the voice message of the target contact, it can extract the voice features of the target contact from the voice message of the target contact; and perform voice synthesis based on the voice features and the received first information to obtain a target voice that is the same as or similar to the real voice of the target contact.

[0027] As an example of a first aspect of this application, extracting the voice features of the target contact from the target contact's voice messages can be achieved by simulating the playback of the retrieved voice messages and capturing audio data during the simulated playback for voice feature extraction. Specifically, this can be done by simulating clicking and playing the retrieved voice messages of the target contact; capturing audio data during the simulated playback; and extracting the voice features of the target contact from the audio data.

[0028] In this way, the voice features of the target contact can be quickly extracted without the user's awareness, and used for subsequent voice synthesis or cloning.

[0029] It should be understood that in certain contexts or scenarios, "extract" and "fetch" have the same meaning.

[0030] In one possible implementation of the first aspect of this application, the voice message may further include voice messages pre-recorded by the target contact or recorded and stored during a call with the target contact. Therefore, retrieving the voice message of the target contact further includes: retrieving the voice message of the target contact from a database storing the voice message of the target contact based on the target contact's contact information.

[0031] This application embodiment allows contacts to record reference audio for use by others; alternatively, with the contact's authorization, reference audio can be recorded during a call between the user and the contact. The reference audio can be uploaded to a cloud database or stored locally.

[0032] In one possible implementation of the first aspect of this application, if the retrieved voice message contains the voices of multiple contacts, factors for voice filtering can be determined; based on these factors, the voice belonging to the target contact can be extracted from the voice message. These factors may include, but are not limited to, at least one of the following: spectral information, sound intensity information, or duration information of the contact's voice. For example, the electronic device can filter the voice of a target contact based on the spectral information of the contact's voice, or based on the sound intensity information of the contact's voice, or simultaneously based on both spectral information and sound intensity information.

[0033] Thus, when the acquired voice message contains multiple human voices, this embodiment of the application can ensure that high-quality voice messages are obtained for sound feature extraction by identifying and filtering the target human voice. Then, the target voice is synthesized or cloned based on the extracted sound features, thereby improving the quality of the target voice used when broadcasting target information.

[0034] In one possible implementation of the first aspect of this application, generating target information corresponding to the first information includes: determining the information source; and generating the target information based on the information source and the content of the received first information. That is, the electronic device can process the received first information according to a corresponding broadcasting strategy to obtain the target information to be broadcast. The aforementioned information source can be the source that needs to be broadcast when the target information is broadcast via voice; it can be obtained after processing according to the broadcasting strategy. For example, the aforementioned information source can include the complete source of the first information, or it can include source content obtained by omitting or simplifying some of its content.

[0035] The determination of the information source includes: determining the difference between the first information and the adjacent previous information; the difference includes the receiving time interval, the platform, the session type, the group, and the contact person; and determining the specific content of the information source to be broadcast based on the difference.

[0036] It should be understood that the aforementioned receiving time interval can refer to the time interval between receiving two messages. For example, if the time to receive the previous message is t1 and the time to receive this message is t2, then the receiving time interval = t2 - t1. The platform can refer to the application platform on which the message is received, such as an SMS platform or a social application platform. If both messages are SMS messages sent and received through an SMS platform, then they belong to the same platform; if one message is received through an SMS platform and the other is an instant message received through a social application, then they belong to different platforms. The conversation type can refer to whether the current conversation is a personal conversation or a group conversation. A personal conversation can be a one-to-one private chat between a user and a contact, while a group conversation can refer to a conversation within a group containing the user and multiple contacts, such as a group chat. The difference in the group affiliation can refer to whether the messages belong to the same group when it is determined that the messages belong to a group conversation, i.e., whether the messages come from the same group chat. The difference in the contact affiliation refers to whether the contact sending the messages through the application is the same contact.

[0037] The step of determining the specific content of the information source to be broadcast based on the difference includes: if the reception time interval between the first information and the adjacent previous information is greater than a preset interval, then the content of the information source to be broadcast includes the complete information source;

[0038] If the reception time interval between the first message and the adjacent previous message is less than or equal to the preset interval, then the changes in the platform, session type, group, and contact person between the first message and the adjacent previous message are determined sequentially; and the specific content of the information source to be broadcast is determined based on the changes.

[0039] In one possible implementation of the first aspect of this application, determining the specific content of the information source to be broadcast based on the changes includes: if any one of the platform, session type, group, or contact to which the first information belongs changes with the adjacent previous information, then determining the content of the information source to be broadcast includes the corresponding content that has changed.

[0040] This application embodiment determines the differences in the source of two messages by sequentially based on the interval between adjacent messages, the platform to which they belong, the group to which they belong, and the contact to which they belong. This allows for the omission of the broadcast of the information source when the information sources are the same, reducing the increase in time caused by redundant information broadcasting, improving information broadcasting efficiency, and helping users quickly understand the information content.

[0041] In one possible implementation of the first aspect of this application, determining the specific content of the information source to be broadcast based on the difference further includes: determining the familiarity between the current user and the target contact; and determining the specific content of the name of the target contact included in the information source to be broadcast based on the familiarity.

[0042] It should be understood that for familiar contacts, users can identify who the contact is by their voice characteristics. Therefore, when using a voice that is the same as or similar to the contact's voice to broadcast the message, the name of the familiar contact can be omitted, thus improving the efficiency of message broadcasting.

[0043] In one possible implementation of the first aspect of this application, it is possible to determine whether a contact is a familiar contact of the user based on the interaction behavior characteristics between the user and the contact. Therefore, determining the familiarity between the user and the target contact includes: acquiring the interaction behavior characteristics between the current user and the target contact; and determining the familiarity between the user and the target contact based on the interaction behavior characteristics. The aforementioned interaction behavior characteristics may include chat behavior characteristics.

[0044] As an example of the first aspect of the embodiments of this application, if it is determined that the target contact is a familiar contact of the current user based on the familiarity level, it can be determined that the name of the target contact can be omitted from the source of the information to be broadcast.

[0045] If the target contact is determined to be an unfamiliar contact of the current user based on the level of familiarity, the name of the target contact can be simplified, and the name of the target contact included in the information source to be broadcast can be the simplified name of the target contact.

[0046] In one possible implementation of the first aspect of this application, generating target information to be broadcast based on the information source and the content of the received first information includes: estimating the duration of the voice broadcast of the first information; if the duration exceeds a preset value, simplifying the first information; and generating target information to be broadcast based on the information source and the content of the simplified first information.

[0047] It should be understood that information that is lengthy or redundant will consume a significant amount of broadcast time, making it difficult for users to quickly understand the information. Therefore, in the process of generating the target information, this embodiment of the application can estimate the broadcast time of the first information. When the estimated broadcast time exceeds a preset value, the electronic device can simplify the first information, reduce the content of the information to be broadcast, and facilitate users to quickly understand the information.

[0048] In this context, the simplification of the first information by electronic devices can be achieved by reducing the number of words. The simplified information has fewer words than the original information, but it still contains the more critical and core content of the original information. By simplifying the information, users will not miss any key information.

[0049] As an example of an embodiment of this application, the simplification of the first information by the electronic device may include summarizing the original content, retaining key or important information in the original content, deleting connecting words or repetitive or meaningless words, etc.

[0050] In one possible implementation of the first aspect of this application, the first information may further include non-text information. The step of generating target information to be broadcast based on the information source and the content of the received first information further includes: determining the information type of the non-text information; performing text conversion on the non-text information according to the information type; and generating target information to be broadcast based on the information source and the text-converted first information.

[0051] As an example of a first aspect of this application, the non-text information may include all information that is not plain text, such as link information, images, files, emoticons, article pushes, mini-programs, and card information. The electronic device can convert the non-text information into text based on its specific type to obtain the converted text information. Taking link information as an example, link information may include URLs and other forms of information generated based on the content corresponding to the URL, such as cards containing text. Therefore, the step of converting the non-text information into text based on the information type includes: determining the link content corresponding to the link information and summarizing the link content to obtain a summarized text. That is, the electronic device can simulate opening the link, obtain the content corresponding to the link, and obtain the text information through summarization and other processing methods. For example, if the link information is a URL corresponding to a news article, when processing the link information, the electronic device can simulate opening the URL to read the news content and summarize the main content of the news as the converted text information.

[0052] As another example of the first aspect of the embodiments of this application, after converting the non-text information into text according to the information type, the method further includes: determining the session type of receiving the non-text information; adding a transition phrase to the text-converted information according to the session type, wherein the sentence structure of the transition phrase can be a subject-verb-object sentence structure or a verb-object sentence structure or other sentence structures, such as a subject-verb sentence structure, a subject-verb-object complement sentence structure, etc.

[0053] The embodiments of this application can process the received non-text information according to the information type, and convert it into text information that can be used for voice broadcasting, thereby improving the efficiency and accuracy of information transmission.

[0054] In one possible implementation of the first aspect of this application, the first information may include multiple messages sent within a preset time period. Generating target information to be broadcast corresponding to the first information further includes: merging the multiple messages sent within the preset time period; and generating target information to be broadcast corresponding to the merged multiple messages. The electronic device can merge multiple messages using different principles depending on the actual situation. For example, multiple messages sent by the same contact can be spliced ​​together, or, for messages from multiple contacts containing similar content, similar content can be filtered out and merged into a single message while maintaining semantic integrity, and so on.

[0055] The step of merging multiple messages sent within a preset time period includes: determining the session type of the multiple messages sent by the target contact within the preset time period; and merging the multiple messages sent by the target contact within the preset time period that belong to the same session type.

[0056] In another possible implementation of the first aspect of this application, the first information may further include multiple group conversation messages sent by multiple associated contacts within a preset time period. The merging of the multiple messages sent within the preset time period includes: determining the content of the multiple messages sent by the multiple associated contacts within the preset time period; and merging the multiple messages sent by the multiple associated contacts within the preset time period that have similar content.

[0057] This application embodiment can merge multiple messages sent continuously by a target contact, or merge multiple messages with similar content sent by multiple associated contacts, further simplifying the content of the information to be broadcast, avoiding broadcasting each received message individually, reducing broadcast time, and improving the efficiency of information transmission.

[0058] In one possible implementation of the first aspect of this application, the first information may further include non-real-time information, and the method further includes: in response to a user instruction, retrieving one or more pieces of non-real-time information corresponding to the user instruction; generating target information to be broadcast corresponding to one or more pieces of non-real-time information, and broadcasting the target information by voice.

[0059] It should be understood that a user instruction can be an instruction issued by a user to request relevant information. This application embodiment retrieves relevant information that meets the user's needs by responding to user instructions. This enables human-computer interaction during information retrieval and broadcasting, helps to filter information based on the user's actual needs, avoids the cumbersome operation of manual information retrieval, and improves the efficiency of information acquisition for users.

[0060] A second aspect of this application provides an information processing apparatus, including:

[0061] A contact identification module is used to identify the target contact who sent the first information in response to the received first information.

[0062] A voice message retrieval module is used to retrieve voice messages from the target contact.

[0063] A target voice generation module is used to generate a target voice similar to the voice of the target contact based on the voice message, wherein the voice message includes the voice message in the first page;

[0064] The target information generation module is used to generate target information corresponding to the first information;

[0065] The voice broadcasting module is used to broadcast the target information using the target voice.

[0066] A third aspect of this application provides an electronic device that may include a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it can implement the information processing method described in the first aspect above.

[0067] A fourth aspect of this application provides a computer-readable storage medium storing computer instructions that, when executed on an electronic device, cause the electronic device to perform the aforementioned related method steps to implement the information processing method described in the first aspect.

[0068] The fifth aspect of this application provides a computer program product that, when run on a computer, causes the computer to perform the aforementioned related steps to implement the information processing method described in the first aspect.

[0069] It is understood that the beneficial effects of the second to fifth aspects mentioned above can be found in the relevant descriptions in the first aspect mentioned above, and will not be repeated here. Attached Figure Description

[0070] Figure 1 is a schematic flowchart of an information processing method provided in an embodiment of this application;

[0071] Figure 2 is a schematic diagram of the structure of an electronic device to which an information processing method provided in an embodiment of this application is applicable;

[0072] Figure 3 is a schematic diagram of target contact identification and voice feature detection provided in an embodiment of this application;

[0073] Figure 4 is a schematic diagram of a sound data authorization management process provided in an embodiment of this application;

[0074] Figure 5 is a schematic diagram of an electronic device operation page during the process of voice data authorization, provided in an embodiment of this application;

[0075] Figure 6 is a schematic diagram of another voice data authorization management process provided in an embodiment of this application;

[0076] Figure 7 is a schematic diagram of another electronic device operation page during the process of voice data authorization provided in an embodiment of this application;

[0077] Figure 8 is a schematic diagram of another voice data authorization management process provided in an embodiment of this application;

[0078] Figure 9 is a schematic diagram of the operation page of an electronic device during the process of authorizing voice data, provided in another embodiment of this application.

[0079] Figure 10 is a schematic diagram of a sound data processing flow provided in an embodiment of this application;

[0080] Figure 11 is a schematic diagram of a tagging retrieval method provided in an embodiment of this application;

[0081] Figure 12 is a schematic diagram of a page-turning retrieval provided in an embodiment of this application;

[0082] Figure 13 is a schematic diagram of a search retrieval method provided in an embodiment of this application;

[0083] Figure 14 is a schematic diagram of a voice audio data stream capture method provided in an embodiment of this application;

[0084] Figure 15 is a schematic diagram of another voice audio data stream capture provided in an embodiment of this application;

[0085] Figure 16 is a schematic diagram of a target human voice recognition and screening method provided in an embodiment of this application;

[0086] Figure 17 is a schematic diagram of sound feature extraction and storage provided in an embodiment of this application;

[0087] Figure 18 is a schematic diagram of an information processing method provided in an embodiment of this application;

[0088] Figure 19 is a schematic diagram of another information processing method provided in an embodiment of this application;

[0089] Figure 20 is a schematic diagram of a broadcast generation control strategy provided in an embodiment of this application;

[0090] Figure 21 is a schematic diagram of an information source determination process provided in an embodiment of this application;

[0091] Figure 22 is a schematic diagram of a contact source determination process provided in an embodiment of this application;

[0092] Figure 23 is a schematic diagram of an information simplification judgment process provided in an embodiment of this application;

[0093] Figure 24 is a schematic diagram of a non-text information processing flow provided in an embodiment of this application;

[0094] Figure 25 is a schematic diagram of a non-text information processing method provided in an embodiment of this application;

[0095] Figure 26 is a schematic diagram of a real-time information merging process provided in an embodiment of this application;

[0096] Figure 27 is a schematic diagram of a real-time information merging process provided in an embodiment of this application;

[0097] Figure 28 is a schematic diagram of another real-time information merging process provided in an embodiment of this application;

[0098] Figure 29 is a schematic diagram of an information retrieval and summary processing flow provided in an embodiment of this application;

[0099] Figure 30 is a schematic diagram of an information retrieval and summary processing provided in an embodiment of this application;

[0100] Figure 31 is a schematic diagram of an information processing device provided in an embodiment of this application. Detailed Implementation

[0101] It should be noted that in the embodiments of this application, the words "exemplarily" or "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design scheme described as "exemplarily" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design schemes. Specifically, the use of the words "exemplarily" or "for example" is intended to present the relevant concepts in a specific manner.

[0102] The business scenarios described in the embodiments of this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided in the embodiments of this application. As those skilled in the art will know, with the emergence of new business scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.

[0103] In this application embodiment, "at least one" refers to one or more, and "more than one" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can represent: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple.

[0104] The steps involved in the information processing method, apparatus, and electronic device provided in this application are merely examples. Not all steps are mandatory, nor are all contents within each step required. They can be added or removed as needed during use. The same step or steps or contents with the same function in the embodiments of this application can be referenced and learned from each other in different embodiments.

[0105] Typically, the information received via voice broadcast on electronic devices can include incoming call information and other notification messages. For incoming call information broadcasting, this function is usually enabled by default on electronic devices. Furthermore, when the electronic device is connected to headphones or a car (e.g., while the user is driving), it will directly broadcast the incoming call information via voice. For example, upon receiving an incoming call, the electronic device can use the system audio source to announce, "A call from Tom, do you want to answer?" The user can interact with the electronic device to decide whether to answer or hang up. For other notification messages, users can selectively enable the function to broadcast specific types or specific app notification messages through the settings on their electronic device. For example, users can configure their electronic device to enable the function to broadcast messages received from various social applications. For system messages, the function to broadcast this type of message can be disabled through settings. This way, the electronic device will not broadcast system messages; however, for messages received from social applications, the electronic device can directly broadcast them via the system audio source upon receipt. During broadcasting, the electronic device can also indicate the source of the information. For example, when an electronic device receives a message from Tom, a contact in social application APP1, its announcement might be, "Message from Tom in APP1: How about we have dinner together this weekend?" If the message contains links, the announcement might be, "From APP1, Tom - First Company - Application Engineer, I found some relevant information for you, here's the URL: https: / / happy.valley..." Furthermore, for longer messages, the electronic device typically only announces the beginning of the long text, failing to provide a summary of the content. When multiple messages are long, the device also only announces the beginning of each long text, resulting in low efficiency for the user in obtaining the message content.

[0106] To address the aforementioned problems, embodiments of this application provide an information processing method, apparatus, and electronic device. On one hand, when the electronic device receives information, it can first determine the target contact for sending the information. In this way, the electronic device can use a voice identical or similar to the target contact's real voice to broadcast the received information, allowing the user to quickly identify the sender based on the voice used in the broadcast, thus establishing familiarity with the target contact through auditory perception. On the other hand, the electronic device can process the received information according to a specific broadcast control strategy. For example, the electronic device can process received non-text information into text information; or, for text information with substantial content, the electronic device can summarize its content, simplifying the information to be broadcast subsequently. In this way, the electronic device can broadcast only the information processed according to the corresponding strategy, without directly broadcasting redundant text or non-text information, improving the efficiency and timeliness of the user's information acquisition.

[0107] Figure 1 shows a schematic diagram of the overall flow of an information processing method provided in this application embodiment. According to the flow shown in Figure 1, the electronic device can execute the information processing method provided in this application embodiment based on user instructions or when automatic broadcasting conditions are met. By determining the target contact, the device uses a voice that is the same as or similar to the target contact's real voice to broadcast the information according to a certain broadcasting control strategy. As shown in Figure 1, when a user instruction is received or the automatic broadcasting conditions are met, the electronic device can first determine the target contact, which is the contact who sent the information. Specifically, the electronic device can obtain contact information based on user instructions or by analyzing the received information. For example, contact ID, nickname, avatar, etc. By processing the above-mentioned various types of contact information, the electronic device can determine the target contact to send the information. Then, as shown in Figure 1, the electronic device can detect whether the target contact's voice characteristics exist in the system, and whether the target contact has been authorized to use their voice characteristics. If the target contact's voice characteristics do not exist in the system, the electronic device can, with the authorization of the target contact, retrieve the target contact's voice messages and extract the target contact's voice characteristics from the retrieved voice messages for voice synthesis. Alternatively, if the system contains the target contact's voice data but the target contact has not authorized the use of their voice, the electronic device can execute an access control step to request authorization from the target contact to use their voice. After obtaining authorization, the electronic device can process the target contact's voice data stored in the system, such as retrieving the target contact's voice messages and extracting voice features based on the retrieved messages. This extracted voice features are then used to synthesize a target voice that is identical or similar to the target contact's real voice. Based on this, as shown in Figure 1, the electronic device can execute a step to determine a broadcast control strategy. By considering the source of the previous message and the current message, the user's familiarity with the target contact, the type and complexity of the information, and the user's broadcast requirements, a broadcast control strategy is generated, and voice broadcast is performed according to the corresponding strategy. During voice broadcast, the electronic device can use a target voice that is identical or similar to the target contact's real voice to broadcast the information processed according to the corresponding strategy. It should be noted that voice features include features in multiple dimensions, such as timbre, pitch, tone, and rhythm. Therefore, a target voice that is the same as or similar to the real voice of the target contact can refer to a voice that is the same as or similar to one or more of the features such as timbre, pitch, tone, and rhythm.

[0108] As an example of the information processing method applied in this application, a user receives a message from a contact in a social application while driving / doing housework / running. The user can actively ask the electronic device what message they received and request the device to announce it. Alternatively, if the user has pre-set automatic announcement of received information, the electronic device, upon receiving a message from a contact, processes the information using the method provided in this application and automatically announces the received information. The information announced by the electronic device can be the information after processing the original information using the announcement control strategy provided in this application. During announcement, the electronic device can use a target voice that is the same as or similar to the real voice of the target contact who sent the message. The target voice can refer to a voice synthesized based on the voice features of the target contact, and the synthesized target voice has a certain similarity to the real voice of the target contact. During the synthesis of the target voice, the electronic device can determine the similarity between the target voice and the real voice of the target contact according to actual needs. For example, the similarity between the target voice and the real voice of the target contact can be 70% or 90%, etc., which is not limited in this application.

[0109] The information processing method provided in this application can be applied to electronic devices. These electronic devices can be mobile phones, tablets, smart wearable devices, in-vehicle mobile devices, etc. This application does not limit the type of electronic device.

[0110] For example, Figure 2 shows a schematic diagram of the structure of an electronic device 200. The structure of the above-mentioned electronic device can be referred to the structure of the electronic device 200 in Figure 2.

[0111] As shown in Figure 2, the electronic device 200 may include a processor 210, an external memory interface 220, an internal memory 221, a universal serial bus (USB) interface 230, a charging management module 240, a power management module 241, a battery 242, an antenna 1, an antenna 2, a mobile communication module 250, a wireless communication module 260, an audio module 270, a speaker 270A, a receiver 270B, a microphone 270C, a headphone jack 270D, a sensor module 280, buttons 290, a motor 291, an indicator 292, a camera 293, a display screen 294, and a subscriber identification module (SIM) card interface 295, etc. The sensor module 280 may include a pressure sensor 280A, a gyroscope sensor 280B, a barometric pressure sensor 280C, a magnetic sensor 280D, an accelerometer sensor 280E, a distance sensor 280F, a proximity light sensor 280G, a fingerprint sensor 280H, a temperature sensor 280J, a touch sensor 280K, an ambient light sensor 280L, a bone conduction sensor 280M, etc.

[0112] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the electronic device 200. In some embodiments of this application, the electronic device 200 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0113] Processor 210 may include one or more processing units. For example, processor 210 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural network processing unit (NPU). Different processing units may be independent devices or integrated into one or more processors.

[0114] The controller can generate operation control signals based on the instruction opcode and timing signals to control the fetching and execution of instructions. For example, when the electronic device 200 receives a message from a contact through an application, such as a social networking application, the controller can process the message and determine the target contact who sent it.

[0115] The processor 210 may also include a memory for storing instructions and data. For example, the memory may be used to locally store the voice data of contacts within the electronic device 200. Thus, when information needs to be broadcast using a voice identical or similar to that of a particular contact, the electronic device 200 can directly retrieve the voice characteristics of the target contact from its local memory.

[0116] Electronic device 200 can implement audio functions through audio module 270, speaker 270A, receiver 270B, microphone 270C, headphone jack 270D, and application processor. Examples include music playback, recording, and voice broadcasting of information.

[0117] The speaker 270A, also known as a "loudspeaker," is used to convert audio electrical signals into sound signals. The electronic device 200 can listen to music or make hands-free calls through the speaker 270A. In this embodiment, the electronic device 200 can broadcast processed information to the user in voice form through the speaker 270A.

[0118] The receiver 270B and / or microphone 270C can be used to receive user voice. For example, a user can actively ask the electronic device 200 what information has been received and request the electronic device 200 to broadcast it. During the interaction between the user and the electronic device 200, the user's instructions can be transmitted to the electronic device 200 in the form of voice. The electronic device 200 can receive the user's voice through the receiver 270B and / or microphone 270C and convert it into a signal that can be processed by the processor 210.

[0119] The headphone jack 270D can be used to connect headphones, which can be wired or wireless. After connecting headphones to the headphone jack 270D, the functions of the speaker 270A, receiver 270B, and microphone 270C can all be achieved through the headphones.

[0120] The information processing method provided in this application embodiment can be implemented on an electronic device having the above-described hardware structure.

[0121] In this embodiment, when an electronic device receives information, it can determine the target contact who sent the information based on the contact information carried in the information. For example, the electronic device can determine the target contact based on the contact ID, nickname, and / or profile picture information carried in the received information. The information received by the electronic device can be first information. In response to the received first information, the electronic device can determine the target contact who sent the first information.

[0122] To deliver received information via voice using a voice identical or similar to the target contact's real voice, the electronic device needs to identify the target contact after determining their identity. This target voice is then used for voice synthesis to produce the desired voice, which is identical or similar to the target contact's real voice. During this process, the electronic device needs to identify stored voice features based on various types of contact information to confirm whether the relevant voice features belong to the target contact.

[0123] Figure 3 is a schematic diagram of target contact identification and voice feature detection provided in an embodiment of this application. Figure 3 shows the process of detecting voice features stored in the system based on various types of contact information.

[0124] In the process shown in Figure 3, the contact information may include contact ID, nickname, and avatar information. Electronic devices can determine the consistency between the voice characteristics in the system and the voice characteristics of the target contact by sequentially judging various information that can mark and identify the target contact, such as contact ID, nickname, and avatar information.

[0125] Specifically, after receiving information, the electronic device can obtain the contact information carried in the information, such as the contact ID, nickname, and profile picture information as shown in Figure 3. The electronic device first detects whether the voice signature corresponding to the target contact ID exists in the system. If it does, the electronic device can determine whether it has obtained authorization from the target contact to use their voice signature. The electronic device can only use the target contact's voice signature for voice broadcasting of the received information if it has obtained authorization from the target contact. Otherwise, the electronic device should attempt to request authorization from the target contact.

[0126] If the system does not have the voice signature corresponding to the target contact ID, the electronic device can continue to detect whether the system has the target contact's nickname, the target contact's avatar information, and whether the system has the voice signature corresponding to the above-mentioned nickname, the target contact's avatar information, and the target contact's avatar information.

[0127] As shown in Figure 3, if a nickname for the target contact exists in the system, the electronic device can detect whether a corresponding voice feature exists in the system. If the voice feature exists, the electronic device also needs to confirm whether the nickname has been modified and whether the current nickname can uniquely identify the target contact. If so, the electronic device can determine whether it has obtained authorization from the target contact to use their voice feature. Based on the determination result, the electronic device can decide whether to use the voice feature to synthesize a target voice that is the same as or similar to the target contact's real voice and use the target voice to broadcast the received information; or, after obtaining authorization from the target contact, synthesize the target voice and use the target voice to broadcast the information.

[0128] If the target contact's nickname does not exist in the system, or if the nickname exists but the corresponding voice signature does not exist, the electronic device can detect if the target contact's nickname and its corresponding voice signature exist in the system. If the target contact's nickname also does not exist in the system, or if the nickname exists but the corresponding voice signature does not exist, the electronic device can further detect if the target contact's profile picture and its corresponding voice signature exist in the system.

[0129] As shown in Figure 3, after confirming the existence of the target contact's voice characteristics in the system, the electronic device can use the target contact's voice characteristics to synthesize a target voice that is the same as or similar to the target contact's real voice, with the authorization of the target contact.

[0130] Figures 4 and 5 illustrate a schematic diagram of voice data authorization management provided in an embodiment of this application. Figure 4 shows an example of the voice data authorization management process, and Figure 5 shows an example of the electronic device's operation page during the voice data authorization process. Following the authorization management process shown in Figures 4 and 5, users can authorize the use of this solution for a specific application through relevant settings. After authorization, the information received by the relevant application can be processed according to the steps provided in this solution and broadcast using the target voice. For example, following the authorization management process shown in Figures 4 and 5, a user can authorize the use of voice data by an application, such as authorizing the use of social application A shown in Figure 5, or authorizing the use of the SMS application. Thus, when the application receives a message sent by the user, it can synthesize a target voice that is the same as or similar to the user's voice features and broadcast the received message using the target voice. The aforementioned social application A can be an instant messaging application, a chat application, or other applications with communication functions.

[0131] Specifically, referring to Figure 5, in one example, the function of broadcasting information using a voice that is the same as or similar to the real voice of the contact sending the information can be called a voice messenger function. When using the voice messenger function, the user can operate the electronic device to access the corresponding settings page, such as the voice messenger settings page shown in Figure 5(a), and activate the function by turning on the voice messenger function switch 511. After the voice messenger function is activated, as shown in Figure 5(a), the user can select to use this function to automatically broadcast received information when the electronic device is connected to headphones by turning on switch 512. By clicking the voice data authorization management switch 513, the electronic device can jump to the authorization management page shown in Figure 5(b), which displays various applications with authorization records, such as social application A and the phone application shown in Figure 5(b). When the user clicks the control 521 corresponding to social application A, the electronic device jumps to the page shown in Figure 5(c), which displays the authorization records for the user to authorize social application A to use voice data, such as authorization record 1, authorization record 2, and authorization record 3. Users can operate on the page shown in Figure 5(c), for example, by clicking control 531 corresponding to authorization record 2 to view the details of this authorization record. As shown in Figure 5(d), the details of authorization record 2 can be seen. In this authorization, the user has turned on the switch 541 for "Allow the use of my voice for one year," which means that after completing this authorization, the user will allow social application A to use their voice data for the voice messenger function for one year. The voice data authorized for use can be the user's voice characteristics, also known as voice feature data.

[0132] Figures 6 and 7 illustrate another method of voice data authorization management provided in this application embodiment. Figures 6 and 7 show the specific process of a user actively requesting authorization from a target contact to use their voice data. Specifically, referring to Figures 6 and 7, when a user wishes to obtain authorization to use the voice data of a certain contact, the user can operate on the electronic device, open the relevant settings page, and confirm the application containing that contact. The user can then send a copied authorization request to the contact through the conversation page with that contact, requesting authorization to use their voice.

[0133] Specifically, referring to Figure 7, in one example, when a user requests authorization for voice data from a contact, the user can operate on the settings page shown in Figure 7(a). For example, by clicking the authorization management switch 711 corresponding to social application A, the electronic device can jump to the authorization management page shown in Figure 7(b). This page displays operation information on how to make a voice data authorization request. For example, the page displays information 721 that allows clicking to copy the authorization link and jump to the corresponding APP. The user can operate according to the information displayed on the page, click the information 721, and copy the corresponding link. In this way, the electronic device can automatically call the corresponding APP, for example, as shown in Figure 7(d), calling social application A, and automatically sending the authorization request information 741 to the other user.

[0134] Figures 8 and 9 illustrate another type of voice data authorization management provided in this application embodiment. The authorization management process shown in Figures 8 and 9 corresponds to the authorization management process shown in Figures 6 and 7. Figures 6 and 7 show the specific process of a user actively requesting authorization from a target contact to use their voice data; that is, Figures 6 and 7 show the relevant operation process on the user's end. Figures 8 and 9 show the process of the contact's end processing the authorization request sent by the user; that is, Figures 8 and 9 show the relevant operation process on the contact's end after receiving the authorization request information sent by the user. According to the process shown in Figures 8 and 9, the contact can confirm the authorization request sent by the user and select the type of authorization to be granted, thereby authorizing the user's end to use their voice data.

[0135] Specifically, Figure 9 shows the authorization settings page that the electronic device automatically redirects to after the user clicks on the received authorization request information. For example, after the other user Tom in Figure 7(d) clicks on the authorization request information 741 sent by the user, the local electronic device can automatically open the authorization settings page shown in Figure 9. User Tom can authorize the other user to use the voice messenger function by selecting one of the authorization types, such as "Allow me to always use my voice" as shown in Figure 9, and turning on the corresponding switch 911. In this way, when Tom uses the authorized application to send a message to the other user, when the application on the other user's device receives the message, the application can automatically obtain the user Tom's voice characteristics, synthesize a target voice that is the same as or similar to Tom's real voice, and then use the target voice to broadcast the message sent by Tom to the other user.

[0136] After identifying the target contact and obtaining the relevant permissions to use their voice features, the electronic device can use the target contact's voice features to synthesize a target voice that is the same as or similar to the user's real voice, and then use the target voice to broadcast the information sent by the target contact. As shown in Figure 1, in achieving the above objective, the electronic device also needs to obtain the target contact's voice features.

[0137] In one possible implementation of this application, the electronic device can retrieve the voice messages of the target contact, extract the voice features of the target contact, and obtain a target voice that is the same as or similar to the real voice of the target contact through voice synthesis or cloning, so that the received information can be broadcast using the target voice.

[0138] Figure 10 illustrates a sound data processing flow provided in an embodiment of this application. Figure 10 shows the process of retrieving voice messages from a target contact and extracting the target contact's voice features. The flow shown in Figure 10 includes steps such as retrieving voice messages from the target contact, capturing voice audio data streams, extracting and storing the target contact's voice features. Furthermore, during the process of retrieving voice messages from the target contact, the duration of the retrieved voice messages can be determined to ensure that subsequent voice audio data streams are captured from voice messages that meet certain duration requirements. Before extracting the target contact's voice features, the electronic device can also perform quality checks on the captured voice audio data streams to ensure that only voice audio data streams meeting quality requirements can be used for subsequent voice feature extraction, ensuring that the synthesized target voice is as close as possible to the target contact's real voice.

[0139] In one possible implementation of this application, the electronic device can use an automated retrieval method to obtain the voice message of the target contact. The automated retrieval method used by the electronic device may include one or more of the following: tagging retrieval, page-turning retrieval, search retrieval, and direct retrieval.

[0140] Tagging retrieval is a method where electronic devices retrieve tagging information from a tagging information database based on search elements, and then quickly locate and retrieve voice messages from a target contact based on the tagging information. Each time an electronic device receives a voice message sent by a contact through an application, it tags the voice message in the background, marking information such as contact information, message length, date, and application name, and stores the corresponding tagging information in a database. When using tagging retrieval, the electronic device can determine the corresponding tagging information from the tagging information database based on search elements, thereby quickly retrieving the required voice message. In one example, the search elements used in tagging retrieval can be the aforementioned punctuation information, and these search elements should at least include the target contact's contact ID and application name.

[0141] Figure 11 is a schematic diagram of a tagging retrieval method provided in an embodiment of this application. Figure 11 shows the process of an electronic device retrieving voice messages of a target contact using a tagging retrieval method, as well as the process of constructing a tagging information database.

[0142] As shown in Figure 11, during the construction of the tag information database, when an electronic device receives each voice message sent by a target contact through an application, it can identify the contact who sent the voice message, for example, by using the contact ID. Typically, voice messages that are too short have limited use in constructing the tag information database. Therefore, the electronic device can determine the duration of received voice messages, only processing those exceeding a certain duration, such as 10 seconds. It should be noted that the step of determining the duration of received voice messages can be performed before or after identifying the contact who sent the voice message. That is, the electronic device can first identify the contact who sent the voice message and then determine the duration of the voice message; or, after receiving a voice message, the electronic device can first determine the duration of the voice message, and if the duration meets the corresponding requirements, the electronic device will then perform contact identification for that voice message.

[0143] As shown in Figure 11, the electronic device can obtain the time when a voice message is received, which can be used as one of the tagging information for that voice message, namely the reception time. In addition, the tagging information may also include contact information, message duration, and application name, etc. After determining the relevant information, the electronic device can tag the voice message and store the tagged message in a tagging information database.

[0144] In one possible implementation of this application embodiment, as shown in FIG11, the process of an electronic device retrieving tag information from a tag information database using a tagging retrieval method to obtain a voice message from a target contact can be performed even when the voice characteristics of the target contact do not exist in the electronic device. When the voice characteristics of the target contact already exist in the electronic device, the electronic device can directly use the voice characteristics to synthesize the target voice and then use the target voice to broadcast the received information, without repeating the voice feature extraction process and the tagging retrieval process for various voice messages. When the voice characteristics of the target contact do not exist in the electronic device, the electronic device can determine the corresponding retrieval elements. For example, retrieval elements may include contact information, message duration, reception time, and application name. The retrieval elements shown in FIG11 include contact information (Tom), message duration (greater than 10 seconds), date (last month), and application name (social application A). The above retrieval elements indicate that the purpose of this tagging retrieval is to retrieve the corresponding tag information from the tag information database, and then use the tag information as navigation to quickly locate and find a voice message that matches the above tag information, that is, a message received in the last month sent by contact Tom through social application A with a message duration greater than 10 seconds. After retrieving voice information that meets the above requirements using the tagging retrieval method, the electronic device can use it for subsequent feature extraction processes; otherwise, the tagging retrieval fails.

[0145] In this embodiment, page-turning retrieval can be a method where an electronic device opens a first page associated with a target contact and retrieves the target contact's voice messages from that first page. The first page associated with the target contact can be a conversation page or message page of an application on the electronic device that receives the information. For example, the first page could be a chat page when a user chats with a target contact using social media; this chat page can also be called a chat box. In one example, the first page could also be a history conversation page. Therefore, during page-turning retrieval, the electronic device can slide the conversation page and retrieve the target contact's voice messages from the history conversation page displayed after sliding.

[0146] Figure 12 is a schematic diagram of a page-turning retrieval method provided in an embodiment of this application. Figure 12 shows the process by which an electronic device retrieves the voice messages of a target contact from a message page using a page-turning retrieval method.

[0147] Similar to tag-based retrieval, pagination retrieval can also be performed when the target contact's voice signature is not present on the electronic device. When the target contact's voice signature is not present on the electronic device, the device can open the message page associated with the user and the target contact. For example, the electronic device can open the chat window between the user and the target contact. This process of the electronic device opening the message page can be achieved through simulation.

[0148] After opening the message page between the user and the target contact, the electronic device can search the current message page to determine if a voice message sent by the target contact exists. If a voice message from the target contact exists in the current message page, the electronic device can also determine the duration of the voice message, for example, as shown in Figure 12, whether the voice message is longer than 10 seconds. If the duration of the voice message is longer than 10 seconds, the electronic device can use this voice message as the retrieved voice message from the target contact for subsequent voice feature extraction. If no voice message from the target contact exists in the current message page, or although a voice message from the target contact exists, its duration does not meet the relevant requirements, such as being less than 10 seconds, the electronic device can flip through the message page, for example, by swiping up, to search for the target contact's voice message again. The message page obtained after flipping through the page can be regarded as the current message page at this moment, and the electronic device can search for the target contact's voice message in the flipped message page in the same way as described above.

[0149] In one possible implementation of this application embodiment, as shown in Figure 12, if the current message page cannot be scrolled up, it indicates that all messages on the current message page have been retrieved. The electronic device can then retrieve the voice messages of the target contact from a conversation group containing the target contact. The aforementioned conversation group can be a group chat containing the target contact.

[0150] As shown in Figure 12, the electronic device can open a chat group containing the target contact (i.e., a group chat window) and search the current message page of the chat group to determine if a voice message sent by the target contact exists. Similarly, if a voice message sent by the target contact exists in the current message page of the chat group, the electronic device can determine the duration of the voice message and use voice messages whose duration meets the relevant requirements as messages for subsequent voice feature extraction. Otherwise, if a voice message sent by the target contact exists in the current message page of the chat group, or if a voice message sent by the target contact exists but its duration does not meet the relevant requirements, the electronic device can continue searching in the newly presented message page by paging. The way the electronic device paging in the chat group is the same as the way it paging in the page where the user is having a conversation with the target contact, as described above. If no voice message sent by the target contact is found in the page where the user is having a conversation with the target contact or in the chat group containing the target contact, the current paging search fails.

[0151] In this embodiment, tag-based retrieval and page-turning retrieval can be methods for directly retrieving voice messages from the message page of an electronic device. These retrieval methods target historical conversation records between the user and their contacts, or contacts within a conversation group. In one example, historical conversation records can refer to historical chat logs, i.e., chat logs between the user and the target contact, or chat logs within a conversation group including the target contact, i.e., group chat logs. In another example, the electronic device can also access the historical conversation record retrieval page through an entry point provided by the application and use appropriate retrieval methods to achieve the above objectives. Specifically, search retrieval is a method that allows access to the historical conversation record retrieval page through an entry point provided by the application, and direct searching for voice messages from the target contact within the historical conversation record retrieval page.

[0152] Figure 13 illustrates a search and retrieval method provided in this embodiment of the application. The retrieval of voice messages using this method can be performed on a second page, which can be a historical conversation record retrieval page. Figure 13 shows the process by which an electronic device directly searches for voice messages of a target contact from historical conversation records using a search and retrieval method.

[0153] Similar to page-based search, when the target contact's voice characteristics are not present on the electronic device, the device can open the message page associated with the user and the target contact, and access the historical conversation record retrieval page through the historical conversation record entry provided on that message page. The device can then use a search function on the historical conversation record retrieval page to search for the target contact's voice messages. As shown in Figure 13, after opening the message page between the user and the target contact, the electronic device can continue to access historical conversation records, such as those stored in the historical chat history conversation box. The electronic device can search within these historical conversation records to determine if voice chat records exist. If the search on the historical conversation record retrieval page yields voice chat records for the target contact, the electronic device can filter these records until voice messages that meet the relevant duration requirements are found, such as voice messages longer than 10 seconds, for subsequent voice feature extraction.

[0154] Similar to the process of searching within a conversation group during page-turning retrieval, when no voice messages meeting the relevant requirements exist in the user's historical conversation records with the target contact, the electronic device can search within the historical conversation records of the conversation group containing the target contact. These historical conversation records can be group chat history. As shown in Figure 13, the electronic device can open the message page of the group containing the target contact, then open the historical conversation record search page for that group, such as the group's historical chat history search page, and search within it. If a voice message is found in the group's historical chat history, the electronic device can filter out voice messages belonging to the target contact using contact information, and use voice messages with durations meeting the relevant requirements as the messages to be used for subsequent feature extraction. If no voice message belonging to the target contact is found in the historical chat history of the conversation group, the search fails.

[0155] In one possible implementation of this application, the retrieval of voice messages from a target contact can also be performed using a direct retrieval method. This direct retrieval can be based on the target contact's contact information, such as their ID, account information, and phone number, and can be performed directly on the electronic device's local storage or in the cloud to obtain the target contact's voice information.

[0156] In one possible implementation of this application embodiment, the above-mentioned tagging retrieval, page-turning retrieval, search retrieval, and direct retrieval can be performed selectively or simultaneously. For example, the electronic device can use one of the retrieval methods to retrieve the voice message of the target contact. When the voice message of the target contact can be retrieved using that retrieval method, the electronic device can stop the retrieval. When the electronic device cannot retrieve the voice message of the target contact using a certain retrieval method, the electronic device can continue the retrieval using another retrieval method until the voice message of the target contact is obtained. In one example, the electronic device can also use multiple retrieval methods simultaneously to retrieve the voice message of the target contact. When the voice message of the target contact is obtained using at least two retrieval methods, the electronic device can use any one of the retrieved voice messages as the voice message to be used for subsequent voice feature extraction.

[0157] As shown in Figure 10, when an electronic device retrieves a voice message from a target contact and the message meets the required duration, it can capture the voice audio data stream. This voice audio data stream capture process can be automated, capturing the corresponding voice data stream from the electronic device's backend while simulating playback of the retrieved voice message.

[0158] Figure 14 is a schematic diagram of a voice audio data stream capture provided in an embodiment of this application. Figure 14 shows the process of capturing voice audio data streams from the background of an electronic device by simulating the playback of voice messages.

[0159] After the electronic device retrieves the voice message of the target contact, it can simulate playing the message, thereby capturing the relevant audio data stream during the simulated playback. As shown in Figure 14, the electronic device can first mute the speaker or virtual speaker. Mute the speaker or virtual speaker to avoid disturbing the user during the audio data stream capture process. Once the speaker or virtual speaker is muted, the electronic device can simulate clicking to play the retrieved voice message from the target contact, capturing the audio data stream during playback. When the voice message finishes playing, the audio data stream capture process ends. At this point, the electronic device can restore its speaker function. The captured audio data stream, after passing quality checks, can be used for sound feature extraction.

[0160] In one possible implementation of this application, the voice message of the target contact used to capture the voice audio data stream can also be retrieved by the electronic device from a local or cloud database. The voice message of the target contact stored in the cloud database can be pre-recorded by the target contact and uploaded to the cloud database, or obtained by recording voice messages during a call between the user and the target contact with the target contact's authorization. Voice messages recorded during a call between the user and the target contact can also be stored locally on the electronic device.

[0161] Figure 15 is a schematic diagram of another voice audio data stream capture provided in an embodiment of this application. Figure 15 shows the process of obtaining the voice audio data stream of the target contact from the cloud.

[0162] In one example, as shown in Figure 15, the target contact can use their own electronic device, such as their mobile phone, to record an audio clip and bind that audio to their account information. The account information can be uniquely identifying the target contact, such as contact details. The audio data bound to the contact information can then be uploaded to a cloud database by the target contact, making it available to other users for use in voice synthesis or cloning during communication.

[0163] In another example, as shown in Figure 15, during a call between a user and a target contact, the electronic device can determine whether the target contact's voice characteristics have been stored. If the electronic device does not store the target contact's voice characteristics, it can, with the target contact's authorization and consent, extract a certain duration of call recording. For example, the electronic device can extract a 10-second recording during the call and bind the extracted audio data with the target contact's account information, thereby uploading it to a cloud database or storing it directly on the electronic device for subsequent voice synthesis or cloning during communication.

[0164] When an electronic device needs to capture the voice audio data stream of a target contact, it can search locally on the electronic device or from the audio data uploaded to the cloud database. The retrieved audio data can then be used as the voice audio data stream of the target contact captured by the electronic device.

[0165] In another possible implementation of this application, the audio data stored in the local database of the electronic device or in the cloud can also be obtained through screen recording. The user can capture the audio data of the target contact through screen recording and extract the target contact's voice characteristics. With the target contact's authorization to use their voice, the electronic device can synthesize a target voice that is the same as or similar to the target contact's real voice based on the target contact's voice characteristics.

[0166] In one example, when a user makes a video call with a contact using an electronic device, the user, with the contact's authorization, can record their screen on the electronic device, resulting in a screen recording file. This screen recording file can be a video file containing the contact's audio data. The electronic device can extract the contact's voice features from the audio data and bind the voice features to the corresponding contact, storing them locally on the electronic device, or bind the audio data to the contact and store it in a cloud database. The authorization process described above can include authorizing the user to record the video call and extracting the contact's voice features from the screen recording file, to implement the relevant solutions provided in the embodiments of this application.

[0167] In another example, a user can operate an electronic device to find voice messages sent by a contact. While the user is playing the voice message, they can record a screen to obtain a recording file. This recording file contains the contact's audio data. With the contact's authorization, the electronic device can execute the steps of the relevant solutions provided in this application embodiment, extracting the contact's voice features from the audio data and binding the voice features to the corresponding contact, storing them locally on the electronic device, or binding the audio data to the contact and storing it in a cloud database. The audio files stored in the cloud database in both examples can be used by the electronic device to extract the contact's voice features during subsequent implementation of the relevant solutions provided in this application embodiment, for synthesizing a target voice that is the same as or similar to the corresponding contact's real voice.

[0168] In one possible implementation of this application embodiment, as shown in FIG10, in order to ensure that the target voice subsequently synthesized or cloned is as similar as possible to the actual voice data of the target contact, the electronic device can perform audio quality detection after capturing the voice audio data stream of the target contact to determine whether the captured voice audio data stream meets the quality requirements. For audio data that meets the quality requirements, the electronic device can continue to perform the voice feature extraction step; otherwise, the electronic device can re-retrieve the voice message of the target contact.

[0169] In this embodiment, the quality detection of audio data may include detecting the audio quality of the captured voice audio data stream itself, and identifying and filtering the target voice when multiple contacts' voices exist in the voice audio data stream. The aforementioned target voice is the voice of the target contact.

[0170] Figure 16 is a schematic diagram of a target human voice recognition and screening method provided in an embodiment of this application. Figure 16 shows the entire process of quality detection and target human voice recognition and screening of the captured voice audio data stream.

[0171] When the electronic device begins quality inspection of the voice audio data stream, as shown in Figure 16, it first determines whether the signal-to-noise ratio (SNR) of the current voice audio data stream meets the requirements for subsequent sound feature extraction, i.e., whether the voice audio data stream meets the standard. If the voice audio data stream does not meet the requirements, the electronic device can perform noise reduction processing on the voice audio data stream until the SNR meets the standard. Once the SNR of the voice audio data stream meets the standard, the electronic device can determine whether there is human voice in the voice audio data stream. If, after multiple noise reduction processes, the SNR of the voice audio data stream still fails to meet the standard, or if there is no human voice in the voice audio data stream after it meets the standard, the electronic device can re-retrieve the voice messages of the target contact based on the contact information.

[0172] If the qualified voice audio data stream contains human voices, in order to identify the target contact's voice, the electronic device can first determine whether there are different human voices in the voice audio data stream, that is, whether the human voice in the voice audio data stream belongs only to the target contact or includes multiple contacts. If the human voice in the voice audio data stream belongs to multiple contacts, as shown in Figure 16, the electronic device can separate and splice the different human voices, and evaluate the voice belonging to the target contact, i.e., the target human voice, based on information such as spectrum, sound intensity, and duration. After human voice separation and splicing, the electronic device can determine whether the voice audio data stream belonging to the target human voice exceeds a certain duration, such as whether it exceeds 5 seconds. If the target human voice exceeds the above-mentioned requirement of 5 seconds, the electronic device can use it for sound feature extraction; otherwise, if the target human voice is shorter than the required 5 seconds, feature extraction may not obtain effective sound features. In this case, the electronic device can re-retrieve the voice message based on the contact information.

[0173] In this embodiment of the application, as shown in Figure 10, for a voice audio data stream that has undergone quality inspection, the electronic device can extract and store sound features. The stored sound features can be subsequently synthesized or cloned to obtain a target sound that is the same as or similar to the real voice of the target contact, thereby enabling the electronic device to use the target sound to broadcast the information sent by the target contact in a voice manner that is the same as or similar to the target contact.

[0174] Figure 17 illustrates a schematic diagram of sound feature extraction and storage provided in an embodiment of this application. For a quality-checked voice audio data stream, the electronic device can extract the sound features of the target contact person. These sound features can include audio features and text features. Audio features can include HuBERT (hidden-unit BERT) features, spectral features, etc.; text features can include BERT (bidirectional encoder representations from transformers) features, etc. The aforementioned text features can be obtained by processing the qualified voice audio data stream using automatic speech recognition (ASR) technology, and extracting features based on the text data contained within it. The electronic device can store the aforementioned audio and text features based on the target contact person's contact information for subsequent cloning or synthesis of the target sound. Storing the extracted sound features based on contact information also helps to quickly find the required target contact person's sound features based on the contact information when retrieving relevant features later, thus accelerating the speed of target sound cloning or synthesis.

[0175] To facilitate understanding, the information processing method provided in this application embodiment will be described below with detailed process and specific examples. Figures 18 and 19 are schematic diagrams of two different processing methods provided in this application embodiment. Figure 18 is a schematic diagram of processing and voice-reading information received by a social application on an electronic device, using a social application as an example. Figure 19 is a schematic diagram of processing and voice-reading information received by an electronic device, using a messaging app / SMS as an example.

[0176] As shown in Figure 18, the electronic device can be a mobile phone used by the user, with a social networking application installed on it. When the application receives a message sent by a contact, the electronic device can process the information and generate corresponding voice information by executing the steps of the method provided in this application embodiment. Then, it uses a voice that is the same as or similar to the target contact's voice, obtained through voice synthesis or cloning, to broadcast the aforementioned voice information. By using a voice that is the same as or similar to the contact's real voice to broadcast the information sent by the corresponding contact through voice synthesis or cloning, the contact who sent the information can be quickly identified without the user viewing the source of the information, thereby establishing an association between the broadcast voice information and the contact at the user's auditory level. The above process will be described in detail below.

[0177] As shown in Figure 18, when an application receives a message from a target contact, it can first determine whether the system stores the target contact's voice features. This process can be implemented through the aforementioned steps of this application, including judging, comparing, and / or marking elements in the system that can identify and mark contacts, such as contact ID, nickname, avatar, etc., to determine whether the pre-stored voice features in the system belong to the target contact who sent the current message. If the system stores the target contact's voice features, the phone can directly use the target contact's voice features to synthesize a target voice that is the same as or similar to the target contact's real voice, and use the target voice to broadcast the received information. If the phone does not have the target contact's voice features, the phone needs to retrieve the target contact's voice message based on authorization, and obtain the target contact's voice features through feature extraction and other processing, then use the extracted voice features to generate a target voice that is the same as or similar to the target contact's real voice, and use the target voice to broadcast the voice information. The aforementioned authorization includes obtaining authorization from the user and authorization from the target contact. The authorization process can be found in the relevant descriptions of the aforementioned steps of this application.

[0178] In one possible implementation of this application, after extracting the voice features corresponding to a contact, the electronic device can bind the voice features to the same contact on various platforms in the system. For example, based on voice messages retrieved in social media software and extracting their voice features, the electronic device can bind these voice features to contacts belonging to the same contact on other platforms or applications within the electronic device. For instance, if it is determined that the contact "Zhang San" in a social media application, "Zhang JACK" in a contact list application, "Zhang JACK" in a text messaging application, and "Zhang San from the R&D team" in a work application belong to the same contact, after the electronic device extracts the voice features of that contact based on voice messages retrieved in a certain application, it can establish a binding relationship between the voice features and the same contact on the aforementioned platforms or applications. Therefore, when subsequent platforms or applications receive information sent by that contact, they can use the extracted voice features to synthesize a target voice that is the same as or similar to the target contact's real voice, and then use the target voice to broadcast the voice information.

[0179] When retrieving a contact's voice messages, one or more of the aforementioned tag-based search, page-turning search, search search, or direct search methods can be used. For example, a mobile phone can retrieve a contact's voice messages from the user's historical chat history with that contact and establish a correspondence between the contact and the retrieved voice messages.

[0180] As shown in Figure 18, when retrieving a contact's voice message, the mobile phone can simulate opening the chat dialog box between the user and that contact in the application and search through their historical chat history one by one. To improve the accuracy of subsequent voice synthesis or cloning, the retrieved voice messages should meet certain duration requirements. For example, a voice message sent by that contact that is longer than 10 seconds can be retrieved from the chat history.

[0181] In one possible implementation of this application, the mobile phone can retrieve historical chat records in reverse chronological order based on the time the message was received. When a historical voice message whose duration does not meet the required duration is retrieved, the mobile phone may not process that message. For example, if a voice message with a duration of 5 seconds is retrieved, the mobile phone may not process that message and may continue searching until a voice message with a duration of more than 10 seconds is retrieved.

[0182] On the other hand, mobile phones can extract features from retrieved voice messages, and the extracted features can be used for subsequent voice synthesis or cloning processing.

[0183] In one possible implementation of this application, the mobile phone can extract features from a voice message by simulating the playback of the voice message. Specifically, the mobile phone can simulate clicking on the voice message and capture the audio data stream while the voice message is playing in the background for subsequent feature extraction and sound synthesis or cloning. In this way, the entire process is seamless for the user, does not affect other operations on the phone, and does not actually play the retrieved voice message.

[0184] The extracted voice features can be bound to contact information, such as the contact's ID, and stored. This allows the phone to use the voice features to synthesize a target voice that is the same as or similar to the contact's real voice when receiving a message from that contact. This target voice is then used to broadcast the voice information obtained by converting the message.

[0185] Figure 19 illustrates another broadcasting method provided in this application embodiment. In the example shown in Figure 19, the target voice can be synthesized using the voice features of the target contact. The voice features of the target contact can be extracted from reference audio stored in a cloud database. As shown in Figure 19, contact Tom can record a segment of his own audio data using his mobile phone as reference audio. The reference audio can be uploaded to the cloud database and bound to contact Tom's account information for subsequent feature extraction and voice synthesis. Alternatively, the reference audio can also come from a call between the user and the contact. For example, during a call between the user and contact Tom, with contact Tom's authorization, a segment of audio from the call can be recorded as reference audio. This reference audio can be uploaded to the cloud database and bound to contact Tom's account information for subsequent feature extraction and voice synthesis. The reference audio recorded during the call can also be stored locally on the electronic device, and the electronic device can directly retrieve the reference audio from the local storage when needed.

[0186] As shown in Figure 19, when the voice characteristics of contact Tom do not exist in the system, the electronic device can retrieve a reference audio file bound to the account information of contact Tom from the cloud database. The electronic device can perform quality checks and processing on the retrieved reference audio file according to the methods described in the preceding embodiments. The processed audio data can be used for voice feature extraction. The extracted voice features can be bound to and stored with the account information of contact Tom. When it is subsequently necessary to use a target voice that is the same as or similar to contact Tom's voice to broadcast the information sent by him, the electronic device can use the extracted and stored voice features to synthesize or clone the voice to obtain the target voice for broadcasting the information sent by contact Tom.

[0187] In this embodiment, after the electronic device identifies the target contact to whom the information is to be sent and obtains a target voice that is the same as or similar to the real voice of the target contact through voice cloning or synthesis, it can use the target voice to broadcast the received information. In this way, the user can intuitively establish a sense of familiarity with the target contact through hearing.

[0188] Referring back to Figure 1, after acquiring the voice characteristics of the target contact and obtaining authorization from that contact to use those characteristics, the electronic device can process the received information according to a certain broadcast control strategy to obtain the target information to be broadcast. The process by which the electronic device generates the broadcast control strategy can comprehensively consider factors such as the source of the previous and current information, the familiarity between the user and the target contact, and the type and complexity of the received information. That is, the electronic device can process the received information based on one or more of these factors to generate the target information. The electronic device can then use the target voice to broadcast the target information. Below, with corresponding examples, the process by which the electronic device processes the received information according to the broadcast control strategy will be described in detail.

[0189] Figure 20 is a schematic diagram of a broadcast generation control strategy provided in an embodiment of this application. The process shown in Figure 20 is the process by which an electronic device processes received information to generate target information. In this process, the electronic device can process real-time information and non-real-time information separately. The processing of real-time information by the electronic device may include steps such as information source determination, contact source determination, real-time information merging processing, non-text information processing, and information simplification determination.

[0190] In the information source determination step, the electronic device can determine the source and broadcast method of the next message based on the differences between adjacent messages. Adjacent messages can be those that are received sequentially in time. For example, if the electronic device receives a message from a contact and then receives another message, these two messages are adjacent messages. The differences between adjacent messages can include the interval between their receipt, the platform they belong to, the chat group they belong to, and the contact they belong to.

[0191] Figure 21 illustrates a flow chart for determining the source of information according to an embodiment of this application. Following the flow chart, the electronic device can determine the source broadcasting method for the next piece of information based on the differences between adjacent information, such as whether to broadcast the information source or simplify the broadcast. When group information, contact changes, or information source changes occur in the information stream, determining the information source can reduce the amount of information broadcast while avoiding ambiguity, thus improving the efficiency of information transmission.

[0192] As an example of an embodiment of this application, the information source determination process can determine whether the broadcast carries the information source by considering factors such as the time interval between adjacent information, whether they belong to the same platform, whether they come from a conversation group, and whether they belong to the same group or the same contact.

[0193] For example, when an electronic device receives a message and broadcasts it according to the method provided in this application embodiment, it can carry the source of the message, that is, broadcast the source of the message before broadcasting the content of the message. For example, it may be from Tom in a work communication group in the social application APP1, or from a contact named Tom. The former example indicates that the message is a conversation group message, that is, a group chat message, while the latter example indicates that the message is a personal message, that is, a private chat message.

[0194] As shown in Figure 21, when an electronic device receives a new message, it can determine the difference between the latest message and the previous message. For example, it first determines whether the time interval between the two messages is less than a preset interval, such as 1 minute. If the time interval is greater than the preset interval, the electronic device can announce the source of the message according to the processing method for newly received messages. If the time interval is less than the preset interval, the electronic device can determine whether the two messages come from the same platform, the same group chat, or the same contact within the group chat. If these determinations are all the same, i.e., the current message and the previous message come from the same platform and the same contact within the same group chat (e.g., both from Tom in the work communication group of social application APP1 in the example above), the electronic device can omit the source of the message and directly announce the content when announcing the latest message, provided that ambiguity is avoided. When the above determinations are inconsistent, for example, although both messages come from contact Tom, the source of the first message is the work communication group of social application APP1, while the other message is a private conversation, the electronic device should announce the source of the new message to avoid ambiguity.

[0195] In one possible implementation of this application embodiment, the process of determining the information source of the electronic device can be implemented by the information source determination module.

[0196] In the contact source determination step, the electronic device can infer the user's familiarity with the contact based on the interaction characteristics between the contact and the user, and decide whether to broadcast the contact's tag name; or, according to the rules, to broadcast a simplified version of the contact's tag name.

[0197] Figure 22 illustrates a contact source determination process provided in an embodiment of this application. In this process, the electronic device can query the contact to which the information belongs and their interaction characteristics with the user, inferring the user's familiarity with them, and thus confirming whether the target contact of the currently sent information is a contact the user is familiar with. If the target contact is a familiar contact, the electronic device may not announce the target contact's identifier, for example, it may not announce the contact's name; if the target contact is not a familiar contact, the electronic device should announce the target contact's identifier, for example, it should announce the contact's name. In one example, the target contact's identifier may be long, for example, the contact name may contain multiple characters. In this case, the electronic device can simplify the contact's name according to rules to improve the efficiency of information announcement as much as possible.

[0198] As an example of an embodiment of this application, if the contact name is "Tom-Human-Computer Interaction Project-0522", where "Tom" is the contact person's name, "Human-Computer Interaction Project" is the contact person's work group, and 0522 may be the contact person's employee number, then the aforementioned long contact name can be simplified to retain only the contact person's name, or retain both the contact person's name and specific work group. For example, the simplified contact name could be "Tom" or "Tom-Human-Computer Interaction". The employee number "0522", which has limited use in daily communication, can be simplified. The electronic device can only broadcast the simplified contact name in subsequent announcements.

[0199] In one possible implementation of this application, the electronic device can determine whether the target contact is a familiar contact of the current user based on the interaction behavior characteristics between the current user and the target contact. The aforementioned interaction behavior characteristics may include chat behavior characteristics, that is, behavioral characteristics formed based on the chat behavior between the current user and the target user.

[0200] As an example of an embodiment of this application, determining whether a target contact is a familiar contact of the current user can be done by assigning scores to features of different dimensions in the following steps S1-S3.

[0201] S1: Obtain chat behavior characteristics.

[0202] In this application embodiment, various chat behavior characteristics can be acquired. For example, the chat behavior characteristics acquired by the electronic device may include one or more of the following features: duration of being a contact, chat frequency, topic of conversation, response delay, and initiative. The scores assigned to each of the above features can be illustrated in the following example.

[0203] F1: Duration as a contact

[0204] <1 month: 0 points;

[0205] 1-6 months: 3 points;

[0206] 6-12 months: 6 points;

[0207] 12 months and older: 10 points.

[0208] F2: Chat frequency

[0209] Daily: 10 minutes;

[0210] How many times a week: 8 points;

[0211] Several times per month: 5 points;

[0212] Every few months: 2 points;

[0213] Fewer: 0 points.

[0214] F3: Conversation Topic

[0215] Primarily personal life and emotions: 10 points;

[0216] A balance between life and work: 7 points;

[0217] Primarily work-related and etiquette-based conversations: 4 points;

[0218] Rarely involves personal topics: 0 points.

[0219] F4: Response Delay

[0220] Instant response (within minutes): 10 points;

[0221] Fastest response (within 1 hour): 7 points;

[0222] Slower response time (within a few hours): 4 points;

[0223] Rarely responds (more than a day): 0 points.

[0224] F5: Initiative

[0225] Both sides are proactive: 10 points;

[0226] Both sides were in control for most of the time: 7 points;

[0227] Unilateral initiative: 4 points;

[0228] One side almost never takes the initiative: 0 points.

[0229] S2: Rate and summarize all contacts with voice characteristics and sort them in descending order.

[0230] In this embodiment of the application, each chat behavior feature can have different weights. Based on the score of each chat behavior feature and its corresponding weight, the familiarity score of the contact can be calculated.

[0231] In one example, the weight of each chat behavior feature can be the same, that is, the weight of each chat behavior feature F1-F5 above is the same, all being 20%. Therefore, the familiarity score of the contact can be calculated as 0.2×F1+0.2×F2+0.2×F3+0.2×F4+0.2×F5.

[0232] Given the calculated familiarity scores for all contacts, the electronic device can sort the contacts in descending order of their scores.

[0233] S3: Select the top-rated contacts as "familiar contacts".

[0234] In this embodiment, based on the contacts sorted in descending order of familiarity score, the electronic device can select the top-scoring contacts as the user's familiar contacts. For example, the top 7 contacts with the highest familiarity scores can be selected as the user's familiar contacts. The other contacts are considered the user's unfamiliar contacts.

[0235] In one possible implementation of this application embodiment, when an electronic device receives information, it can determine whether the contact is a familiar contact to the user. For example, upon receiving information, the electronic device can query the target contact corresponding to the information and query the user's chat behavior characteristics with the target contact. For instance, following the steps in the example above, it can determine whether the target contact is a familiar contact to the user. For familiar contacts, when broadcasting the converted target information, the electronic device can directly use a voice that is the same as or similar to the contact's voice to broadcast the information content, without broadcasting the contact information. For unfamiliar contacts, the electronic device can first broadcast the contact information and then broadcast the corresponding message content.

[0236] By identifying the source of a contact, users can determine the source of information from familiar contacts by their voices. Omitting the contact's name can shorten the overall broadcast time and improve information retrieval efficiency. For unfamiliar contacts, electronic devices should first announce the contact's source information when broadcasting target information. This avoids misunderstandings caused by users' inability to determine the source by voice and ensures the accuracy of information transmission. Furthermore, simplifying the contact's tagging information (such as nicknames) can improve information retrieval efficiency for unfamiliar contacts.

[0237] In one possible implementation of this application embodiment, the process of determining the source of a contact in an electronic device can be implemented by a contact source determination module.

[0238] In the information simplification judgment step, the electronic device can decide whether to further simplify the information based on the expected broadcast duration of the text content or the information converted into text content.

[0239] Figure 23 illustrates a flowchart of an information simplification judgment process provided in an embodiment of this application. In this process, the electronic device can estimate the playback duration of the text content or information converted into text content and decide whether to further simplify the information. For example, if the electronic device estimates that the playback duration of the currently received information, if broadcast via voice, may exceed 20 seconds, then the electronic device can further simplify the information. For instance, it can simplify the text content of the information by reducing prepositions, subject-verb-object sentence structures, replacing phrases with concise ones, deleting introductory text, and weakening content based on context.

[0240] As an example of an embodiment of this application, the raw information received by the electronic device may be the following:

[0241] "Hey, old friend, long time no see! How have you been? I recently went to Yunnan, it was absolutely beautiful! We visited Lijiang, Dali, and Shangri-La, each place with its own unique charm. The ancient city of Lijiang was truly unforgettable, and Erhai Lake in Dali was also stunning. We rented bicycles and cycled around the lake. Although it was a bit tiring, seeing the azure water and the surrounding pastoral scenery made it all worthwhile."

[0242] Oh, and there's something else important I need to tell you: I'm planning to change jobs. My previous job was good, but I felt there wasn't much room for growth. I recently interviewed at an internet company for a product manager position, and the job sounds quite challenging and interesting.

[0243] Also, do you remember our class reunion? It's scheduled for the first weekend of next month, and everyone is really looking forward to seeing you! You absolutely have to come this time!

[0244] Electronic devices, through information simplification judgment, determine that the aforementioned information exceeds the preset broadcast duration and should therefore simplify its content. Simplification can involve summarizing the content, retaining key or important information from the original content, and removing connecting words, repetitions, or meaningless words. Therefore, the information obtained after simplifying the original information according to certain rules can be as follows:

[0245] "Hey, old friend, long time no see! How have you been? I just went to Yunnan. Lijiang, Dali, and Shangri-La are all beautiful, especially Erhai Lake in Dali. Cycling around the lake is amazing."

[0246] By the way, I'm planning to change jobs. I recently interviewed with an internet company for a product manager position. I hope to get good news.

[0247] Also, our class reunion is next month's first weekend, and everyone is looking forward to seeing you. You absolutely have to come!

[0248] In one possible implementation of this application, when the information broadcast by the electronic device is not the original information, to help the user understand that the broadcast information is a simplified version, the electronic device can add a prompt sound effect before broadcasting the simplified information. That is, by playing the prompt sound effect first, the user is informed in advance that the information to be broadcast is not the content of the received original information, but rather information obtained after certain processing. This information simplification and processing avoids the problem of excessively long broadcast times, which can cause user impatience and occupy their attention for an extended period. Replacing non-self-evident information within the message avoids directly broadcasting meaningless information that could annoy the user. Adding a prompt sound effect helps the user understand that the received information has been modified.

[0249] In one possible implementation of this application embodiment, the information simplification judgment process of the electronic device can be implemented by an information simplification judgment module.

[0250] In the non-text information processing step, the electronic device can understand the content based on the type of non-text information and decide how to convert the non-text information into text.

[0251] Figure 24 illustrates a non-text information processing flow provided in an embodiment of this application. In this flow, the electronic device can understand the content based on the type of non-text information and determine the method for text conversion.

[0252] The non-text information shown in Figure 24 can include links, images, files, emojis, article pushes, mini-programs, card information, and all other information that is not plain text. For received non-text information, the electronic device can select an appropriate transition phrase based on the information type. This transition phrase can be a statement adapted to a specific type of non-text information. For example, when the received non-text information is an article push, the transition phrase adapted to this type of non-text information can be a subject-verb-object sentence, i.e., a subject + verb + object format, such as "I shared an article push with you," or a verb-object sentence, i.e., a verb + object format, such as "I shared an article push with you." Compared to subject-verb-object transition phrases, verb-object sentences omit the subject. In one possible implementation, using a verb-object transition phrase with an omitted subject can be used in scenarios where the subject is explicit. For example, in a one-on-one private chat between a user and a contact. In multi-person chat scenarios, such as group chats, a subject-verb-object transition phrase can be used to help users quickly identify the source of the information. For example, in a group chat scenario, if contact Zhang San shares an article in the group chat, after processing this type of non-text information, the electronic device can use a subject-verb-object transition phrase to form the voice message "Tom shared an article with you." Besides using the aforementioned subject-verb-object or verb-object transition phrases, other sentence structures can be used as needed, such as subject-verb sentences (subject + verb format) or subject-verb-object-complement formats (subject + verb + object + complement format), etc. This application embodiment does not limit these variations.

[0253] As shown in Figure 24, when processing received non-text information, electronic devices can add appropriate transition phrases based on the context of the information. For example, the transition phrase "you" or "you all" can be added depending on whether the information is in a group chat. Specifically, "you all" can be used when the information is in a group chat, while "you" can be used in a private chat. This results in voice messages like "I shared an article with you" or "Tom shared an article with you all" in the example above.

[0254] In one possible implementation of this application, the types of non-text information may include shareable links, files, mini-programs, and transaction information, etc. The transition phrases and corresponding conversion strategies for these various types of non-text information are shown in Table 1 below.

[0255] Table 1. Examples of transition phrases and corresponding conversion strategies for various types of non-textual information.

[0256] Non-text information can be files or links. For information in the form of files or links, electronic devices can push a summary of the content corresponding to the link to form information that can be used for voice broadcast; or, based on the rules corresponding to the file extension, simulate opening the file and extracting the content from the file, summarizing it into a summary to form information that can be used for voice broadcast; or, electronic devices can also simulate opening the file or link, summarizing the content in combination with the context and file name to form information that can be used for voice broadcast.

[0257] As an example of an embodiment of this application, if the non-text information is a link, the electronic device can simulate opening the webpage corresponding to the link, extract or summarize the content of the webpage, generate corresponding broadcastable information, and broadcast the information using a voice that is the same as or similar to the contact's real voice. For example, when receiving non-text information in the form of a link, the voice information broadcast after processing according to this method could be: "Jack pushed a voting link for the National Postgraduate Innovation Competition, which includes the top ten innovative works in the country. The deadline for voting is 24:00 on July 10th."

[0258] Referring to the example shown in Table 1, when the received information contains non-text information of the link type, the electronic device can obtain a summary of the content corresponding to the link when processing the information. For example, the information received by the electronic device is: "I found a relevant paper: doi:10.xxx / 3491102.xxxx, which I think is quite relevant to your research. See if it's helpful to you." This information contains non-text information of the link type, namely the "doi:10.xxx / 3491102.xxxx" part. When processing this information, the electronic device can push a summary of the content corresponding to the link. Therefore, the target information obtained by processing the information containing the link could be: "I found a relevant paper: the title is 'Comparison of Voice Assistant Response Behaviors,' which I think is quite relevant to your research. See if it's helpful to you." That is, the electronic device replaces the link part in the aforementioned information with "the title is 'Comparison of Voice Assistant Response Behaviors.'"

[0259] In another example, if the non-text information is a file, such as a Word document, the electronic device can extract the filename and announce it, such as "Tom sent a Word document named 'Attendance Instructions'." Alternatively, the electronic device can simulate opening the file, extract a summary or main content, and announce the extracted summary or main content, such as "Tom sent a Word document containing the latest revised company attendance instructions."

[0260] In one possible implementation of this application, to help users understand that the content being broadcast has been rewritten and is not the original content of the message sent by the contact, the electronic device can add a corresponding prompt sound effect before broadcasting the rewritten content. For example, when broadcasting the rewritten message, "I found a relevant article: the title is 'Comparison of Voice Assistant Response Behaviors,' which I think is quite relevant to your research. See if it can help you," a prompt sound effect can be added before broadcasting "The title is 'Comparison of Voice Assistant Response Behaviors'" to inform the user that the content to be broadcast next is rewritten.

[0261] In another possible implementation of this application, the prompt sound effects added by the electronic device when broadcasting different types of non-text information can be different. For example, the prompt sound effects added when rewriting and broadcasting received non-text information such as share links are different from the prompt sound effects added when rewriting and broadcasting received non-text information such as files or mini-programs.

[0262] When an electronic device broadcasts the processed target information, it can use a voice that is the same as or similar to the real voice of the contact who sent the message. For example, it can use a voice that is the same as or similar to the voice of the contact Tom to broadcast the voice message "Tom shared a push notification with you".

[0263] Figure 25 illustrates an example of non-text information processing provided in this application embodiment. In the example shown in Figure 25, the information sent by contact Tom is a push notification article, namely tweet 2501 shown in Figure 25(a), with the title "Highlights of the 2024 Speaker and Headphone Exhibition". After processing by the electronic device, information corresponding to the push notification content can be generated and can be used for voice broadcast. Furthermore, the content of the voice broadcast can vary depending on the information scenario. For example, in a one-on-one chat scenario, i.e., a private chat scenario, after processing the content pushed by the contact, the voice broadcast content could be "I shared a push notification with you"; while in a multi-person chat scenario, i.e., a group chat scenario, the voice broadcast content could be "I shared a push notification with you".

[0264] Thus, in the example shown in Figure 25, the text information that can be used for voice broadcast after processing the tweet 2501 shown in Figure 25(a) can be as shown in Figure 25(b). If the tweet 2501 is information in a private conversation, the processed information can be as shown in information 2502 in Figure 25(b), which is "I shared a push with you: Highlights of the 2024 Speaker and Headphone Exhibition"; if the tweet 2501 is information in a group conversation, i.e., a group chat, the processed information can be as shown in information 2503 in Figure 25(b), which is "I shared a push with you: Highlights of the 2024 Speaker and Headphone Exhibition".

[0265] By processing non-text information through electronic devices, transitional phrases can be added to non-text information, reducing the abruptness of directly using the contact's voice to read the message, supplementing background information, and enhancing the user's understanding of the information; considering the context in which the contact sent the message (such as private chat or group chat), the transitional phrases are made more in line with the current context, avoiding inconsistencies between the context and the transitional phrases.

[0266] In the real-time information merging process, the electronic device can decide whether to merge multiple messages and the merging method based on the receiving time interval between adjacent messages, the number of messages, and whether there is any message being broadcast.

[0267] Figure 26 illustrates a real-time information merging process provided in an embodiment of this application. According to the process shown in Figure 26, when an electronic device receives a new message, it can determine whether there is already any message in the broadcasting process that has not yet ended. If there is no previously received message being broadcast, the electronic device can broadcast the newly received message according to the normal broadcasting process. This normal broadcasting process can include the processes described in the preceding embodiments, such as the processes for determining the target contact, retrieving voice features and managing permissions, and generating broadcast control strategies as shown in Figure 1.

[0268] If an electronic device receives new information while simultaneously broadcasting previously received information, it can determine if the new information shares any similarities with the broadcast information. For example, it might determine if the information comes from the same contact person or if the content is identical or similar. If the new information does not share any similarities with the broadcast information, the electronic device can broadcast the new information using the normal broadcast process. If the new information shares similarities with the currently broadcast information, the electronic device summarizes the unbroadcast information and uses an information simplification and judgment process to determine the information to be broadcast. This unbroadcast information includes both currently received information and previously received but not yet broadcast information.

[0269] Figures 27 and 28 show two examples of real-time message merging. In the group chat scenario shown in Figure 27, multiple identical messages sent by multiple contacts can be merged into a single message through merging processing, forming a corresponding voice broadcast message.

[0270] Specifically, in the example shown in Figure 27(a), contact Tom sent a red envelope 2701 in the work communication group. Subsequently, several other contacts sent messages. For example, contact Mike sent a message 2702 saying "Thank you for the red envelope, Tom," while contacts Lucy and Bella sent emojis, i.e., messages 2703 and 2704 in Figure 27, to express their gratitude for Tom's red envelope 2701. The electronic device can determine that the messages sent by contacts Mike, Lucy, and Bella all express gratitude for Tom's red envelope 2701. Therefore, the aforementioned contacts Mike, Lucy, and Bella can be considered as related contacts. After processing the messages 2702, 2703, and 2704, the electronic device can merge the messages sent by the multiple related contacts into message 2705 shown in Figure 27(b), i.e., "Everyone is thanking Tom for the red envelope."

[0271] In the chat scenario shown in Figure 28, as shown in Figure 28(a), the same contact Lucy sends multiple messages in a short period of time, namely messages 2801-2804. By summarizing, generalizing, or merging these multiple messages, a complete voice broadcast message can be formed. For example, message 2805, as shown in Figure 28(b), is obtained by merging and summarizing the multiple messages 2801-2804 in Figure 28(a).

[0272] By merging real-time information through electronic devices, when multiple messages are sent by the same or related contacts within a short period of time, the time consumed by playing notification sounds between messages can be avoided through the aggregation and merging operation, thus improving the efficiency of information acquisition. By aggregating the information from the same or related contacts together, users can better understand the information and reduce the impact of omissions on information comprehension.

[0273] In one possible implementation of this application, for non-real-time information, the electronic device can extract information demand features according to user instructions through steps such as information retrieval and summary summarization, generate search-based filtered information, implicitly complete missing features, broadcast information summaries accordingly, and trigger active questioning.

[0274] Figure 29 illustrates an information retrieval and summary processing flow provided in this embodiment of the application. This flow shows the relevant steps for processing non-real-time information. As shown in Figure 29, the above flow can be triggered based on a user's instruction to obtain information. In one example, the user's instruction can be a voice command. For instance, a user can actively send an instruction to an electronic device via voice to request relevant information. After receiving the user's instruction, the electronic device can extract the demand features from the user's request. These demand features can represent the specific content of the information requested by the user. The electronic device can determine whether the demand features in the request are comprehensive to execute the subsequent information retrieval and summary processing flow. If comprehensive or complete demand features can be extracted from the user's request, the electronic device can search and filter one or more pieces of information that meet the aforementioned demand features within the corresponding range based on the extracted features. If the demand features extracted from the user's request are not comprehensive or complete, and information cannot be directly retrieved and filtered based on the extracted demand features, the electronic device can supplement the missing features based on other relevant information. For example, electronic devices can complete missing features based on recently received information to obtain comprehensive or complete requirements features, thereby retrieving and filtering one or more pieces of information that meet the above requirements features within the corresponding range.

[0275] As shown in Figure 29, after filtering out one or more pieces of information, the electronic device can determine whether the filtered information includes non-text information. For non-text information, the electronic device can convert it into text information. The process of converting non-text information into text information by the electronic device can be found in the relevant descriptions in the foregoing embodiments of this application, and will not be repeated here.

[0276] The electronic device can summarize and extract relevant summaries from the selected and converted text information. Based on this, it can also determine whether any important information related to the summarized information exists in the received non-real-time information. If so, the electronic device can proactively ask the user questions to confirm whether the relevant important information should be broadcast to the user as well. After broadcasting the information, the electronic device can also receive further questions from the user regarding the broadcasted information, thus enabling interaction between the electronic device and the user.

[0277] Figure 30 illustrates an example of information retrieval and summary aggregation provided in this application. Figure 30 shows an example of retrieving messages within a specific range based on semantic understanding and summarizing them.

[0278] Specifically, Figure 30 shows a group chat scenario. In the work communication group shown in Figure 30(a), contact Tom sent multiple messages 3001-3003. Messages 3001 and 3002 are two tweets forwarded by contact Tom, and message 3003 is a text message sent by contact Tom. After the electronic device receives these multiple messages, the user can actively ask the electronic device about the content of the messages sent by contact Tom. For example, as shown in Figure 30(b), the user can ask the electronic device via voice, "What did Tom say in the work communication group?" In response to the user's request, the electronic device can extract relevant demand features, including contact Tom, the work communication group, etc. Based on the extracted demand features, the electronic device can perform a search within the relevant scope, that is, search for the information sent by contact Tom in the work communication group. The result obtained by the electronic device is the multiple messages 3001-3003 sent by contact Tom as shown in Figure 30(a). At this time, the electronic device can process the above multiple messages to form the target information to be broadcast, as shown in message 3005 in Figure 30(b). The electronic device uses a target voice that is the same as or similar to the real voice of the contact Tom to broadcast the message 3005.

[0279] In one possible implementation of this application, when the requirement features extracted by the electronic device from the user's request are not comprehensive or complete, the electronic device can supplement the incomplete or incomplete requirements.

[0280] As an example of an embodiment of this application, a user can proactively ask an electronic device, "What did Jack say this morning?" After processing the request, the electronic device can determine that the user wants to obtain information sent by contact Jack this morning. This request is not comprehensive; by supplementing the request features, the electronic device can determine the scope of information retrieved from contact Jack this morning, including information in group chats and private chats. Therefore, in the above example, the information retrieved by the electronic device according to the supplemented request features can include the information sent by contact Jack in the Saturday team-building group, "The destination for Saturday's team-building is Dameisha," and the information sent by contact Jack individually to the user, "Are you free to have dinner together tonight?" For these two pieces of information, the final target information to be broadcast by the electronic device can be the text, "I told everyone in the Saturday team-building group that the destination is Dameisha. I also sent you a separate message asking if you're free to have dinner together tonight?" This text information can be broadcast using a voice that is the same as or similar to contact Jack's real voice.

[0281] By using information retrieval and summary processing, electronic devices can use semantic understanding to obtain the range of information that users are interested in for non-real-time information. They can then summarize the information within that range, avoiding the time-consuming process of reading each item and the tedious process of manual searching and browsing, thus improving the efficiency of users in obtaining information.

[0282] By applying this application and reducing unnecessary information sources and contact name announcements, broadcast time can be reduced while avoiding ambiguity. Furthermore, by summarizing multiple messages, adding transitional phrases, replacing meaningless information, and streamlining lengthy announcements, the efficiency and accuracy of information delivery are improved, thus optimizing and enhancing the user's listening experience. For non-real-time information, by retrieving and summarizing the information needed by the user, the time-consuming process of reading each message individually and the tedious manual browsing are avoided, improving the efficiency of information retrieval for users. Using a voice that is the same as or similar to the contact person's when broadcasting information can bridge the psychological distance between the user and the information, helping the user understand the information more accurately and increasing user satisfaction.

[0283] As a specific application example of this application, this method can also be applied to accessible reading scenarios. For example, when an electronic device receives a message from a contact of a relevant user, the electronic device can execute the method provided in this application to convert the message into target information and broadcast the target information using a voice synthesized or cloned that is identical or similar to that of the contact person. This facilitates timely information acquisition for the relevant user, achieving the effect of recognizing the person by their voice. In this way, the relevant user can intuitively determine the source of the information and quickly identify the contact person who sent the message based on the voice data used in the broadcast. In one example, the relevant user can include a blind user. In another example, when converting the received information into target information, the electronic device can process the information according to the aforementioned broadcast control strategy to improve the efficiency of information broadcasting.

[0284] As another specific application example of this application, this method can also be applied to scenarios using smart homes.

[0285] In one example of a smart home scenario, the user of the electronic device can be a child or a minor. This method can be applied to the control process of the electronic device in child mode. For example, a child user can use the electronic device for learning, playing games, or other entertainment activities. Information about a child user's use of the electronic device in child mode can be sent to their parents or other guardians. Parents or other guardians can send messages to the child user as needed. For example, after a child user has used the electronic device for a certain period of time, parents or other guardians can receive relevant prompts. Thus, parents or other guardians can send a message to the child user to remind them to stop using the electronic device. At this time, the electronic device receiving the message can execute the method provided in this application embodiment, using the same or similar voice as the parent or guardian who sent the message, to verbally announce the message and remind the child user to stop using the electronic device.

[0286] In another possible implementation, relevant prompts can be built into the electronic device. In child mode, when a child user meets the corresponding prompt conditions using the electronic device, the device can automatically extract the relevant prompts and broadcast them in the same or similar voice as the parent or one of the child user's guardians.

[0287] In another example of a smart home scenario, the electronic device can be any smart home device with voice broadcast functionality. When any smart home device receives a message sent by a user's contact, it can convert the message into voice information by executing the method provided in this application embodiment, and broadcast the voice information using the same or similar voice as the contact. For example, while a user is watching a video on a smart TV at home, messages sent to the user by family members can be broadcast through the smart TV, and the same or similar voice as the family member can be used during the broadcast. For instance, if a wife is watching TV at home, a message sent by her husband to her mobile phone can be broadcast through the smart TV in a smart home scenario, such as the voice message "I'll be home in 10 minutes, please make me a pot of tea." Alternatively, after a user has cooked rice using a smart rice cooker, the user can send a command to the smart rice cooker via voice or button operation, prompting the smart rice cooker to remind other family members to come to the dining room for dinner. The smart rice cooker can send the aforementioned "go to the restaurant for dinner" message to each family member. Upon receiving this message, each family member's electronic device can convert it into a voice message using the method described in this application, and broadcast it in the same or similar voice as the user cooking, reminding each family member to go to the restaurant for dinner. In the above example, the types of electronic devices used by each family member can be different. For example, family member A, who is outdoors, can use a mobile phone to execute this method and broadcast the voice message "go to the restaurant for dinner"; family member B, who is watching TV in the living room, can use a smart TV to broadcast the voice message "go to the restaurant for dinner"; and family member C, who is indoors in a room, can use a smartwatch, smart bracelet, or a smart speaker in the room to execute this method and broadcast the message "go to the restaurant for dinner" to family member C via voice.

[0288] In the example above where a family member is reminded to come to the dining room, the entire process can also be implemented using a smart speaker. For instance, after preparing the meal, the user can directly say to the smart speaker, "Call everyone to come eat." The smart speaker can process this voice and send the processed information to smart speakers in other rooms of the house. These smart speakers can then use the same or a similar voice as the user to remind everyone to come to the dining room. For environments where information cannot be transmitted via other smart speakers, such as family member A who is outdoors in the aforementioned example, the processed information can be sent to other smart electronic devices associated with family member A, such as mobile phones, smartwatches, etc. These associated electronic devices will then execute this method and use the same or a similar voice as the user to remind family member A to come to the dining room.

[0289] In another example of a smart home scenario, an electronic device could be a smart alarm clock. This smart alarm clock can remind a user at a set time, or based on instructions from another user, using a voice that is the same as or similar to that of a particular contact. For example, a student could use a smart alarm clock for a wake-up call, where the alarm clock could use their father's or mother's voice to remind them "Get up now!" at 6:00 AM. Alternatively, a smart alarm clock could also provide reminders in other scenarios requiring timed processing. For instance, in a scenario where a student is taking a homework test, a smart alarm clock could use a teacher's voice to remind the student "You can start answering the questions" or "The test is over."

[0290] This application embodiment can divide an electronic device into functional modules based on the above method example. For example, each function can be divided into a separate functional module, or one or more functions can be integrated into a single functional module. The integrated module can be implemented in hardware or as a software functional module. It should be noted that the module division in this application embodiment is illustrative and only represents one logical functional division; other division methods may be used in actual implementation. The following description uses the example of dividing each function into a separate functional module.

[0291] Referring to FIG31, a structural block diagram of an information processing device provided in this application embodiment is shown, corresponding to the above embodiments. This device can be applied to the electronic devices in the foregoing embodiments. Specifically, the device may include the following modules: a contact identification module 3101, a voice message retrieval module 3102, a target voice generation module 3103, a target information generation module 3104, and a voice broadcasting module 3105, wherein:

[0292] The contact identification module 3101 is used to identify the target contact who sent the first information in response to the received first information.

[0293] The voice message retrieval module 3102 is used to retrieve voice messages of the target contact.

[0294] The target voice generation module 3103 is used to generate a target voice similar to the voice of the target contact based on the voice message, wherein the voice message includes the voice message in the first page;

[0295] Target information generation module 3104 is used to generate target information corresponding to the first information;

[0296] The voice broadcasting module 3105 is used to broadcast the target information using the target voice.

[0297] The aforementioned device may be an electronic device in the foregoing embodiments, or it may be a unit or component in the aforementioned electronic device that can perform the corresponding function.

[0298] It should be noted that all relevant content of each step involved in the above method embodiments can be referenced from the functional description of the corresponding functional module, and will not be repeated here.

[0299] This application also provides an electronic device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it can implement the information processing methods described in the foregoing embodiments.

[0300] This application also provides a computer-readable storage medium storing computer instructions. When the computer instructions are executed on an electronic device, the electronic device performs the aforementioned related method steps to implement the information processing methods in the foregoing embodiments.

[0301] This application also provides a computer program product that, when run on a computer, causes the computer to perform the aforementioned related steps to implement the information processing methods in the foregoing embodiments.

[0302] This application also provides a chip, which can be a processor, or the chip includes a processor. The processor can be a general-purpose processor or a dedicated processor; wherein, the processor is used to support the foldable screen device in performing the above-mentioned related steps to implement the information processing methods in the foregoing embodiments.

[0303] Optionally, the chip further includes a transceiver for receiving control from the processor to support the electronic device in performing the aforementioned steps to implement the information processing methods in the foregoing embodiments.

[0304] Optionally, the chip may also include a storage medium.

[0305] The chip can be implemented using one or more field programmable gate arrays (FPGAs), programmable logic devices (PLDs), controllers, state machines, gate logic, discrete hardware components, any other suitable circuits, or any combination of circuits capable of performing the various functions described throughout this application.

[0306] Finally, it should be noted that the above description is only a specific implementation of this application, but the protection scope of this application is not limited thereto. Any changes or substitutions within the technical scope disclosed in this application should be covered within the protection scope of this application.

Claims

1. An information processing method, characterized in that, include: In response to the received first information, determine the target contact to whom the first information was sent; Retrieve voice messages of the target contact and generate a target voice similar to the voice of the target contact based on the voice messages, wherein the voice messages include voice messages in the first page; Generate target information corresponding to the first information; The target information is broadcast using the target sound.

2. The method according to claim 1, characterized in that, The retrieval of the target contact's voice messages includes: Display a first page associated with the target contact, the first page including the target contact's conversation page; Retrieve voice messages sent by the target contact on the first page.

3. The method according to claim 2, characterized in that, The first page also includes the target contact's history conversation page, and displaying the first page associated with the target contact includes: In response to the first action, the target contact's history conversation page is displayed.

4. The method according to claim 1, characterized in that, The voice messages also include voice messages on the second page, and the voice messages for retrieving the target contact include: In response to the second operation, a second page associated with the target contact is displayed, the second page including a history session retrieval page; Retrieve voice messages sent by the target contact on the second page.

5. The method according to any one of claims 1 to 4, characterized in that, The step of generating a target voice similar to the target contact's voice based on the voice message includes: Extract the voice features of the target contact from the voice messages of the target contact; Based on the sound characteristics and the received first information, sound data is synthesized to obtain a target sound similar to the voice of the target contact person.

6. The method according to claim 5, characterized in that, Extracting the voice features of the target contact from the voice messages of the target contact includes: Click to play the voice message from the target contact that was retrieved; During the playback of the target contact's voice message, audio data is captured; Extract the voice features of the target contact from the audio data.

7. The method according to any one of claims 1 to 6, characterized in that, The voice messages also include voice messages pre-recorded by the target contact or recorded and stored during a call with the target contact. The retrieval of the target contact's voice messages further includes: Based on the contact information of the target contact, retrieve the voice messages of the target contact from the database storing the voice messages of the target contact.

8. The method according to any one of claims 1 to 7, characterized in that, After retrieving the voice messages of the target contact, the process also includes: If the retrieved voice message contains the voices of multiple contacts, then factors for voice filtering are determined; wherein, the factors include at least one of the following: spectral information, sound intensity information, or duration information of the contact's voice; Based on the aforementioned factors, the voice belonging to the target contact is extracted from the voice message.

9. The method according to any one of claims 1 to 8, characterized in that, The generation of target information corresponding to the first information includes: Determine the source of the information; The target information is generated based on the information source and the content of the first information received.

10. The method according to claim 9, characterized in that, The determination of the information source includes: Determine the differences between the first piece of information and the adjacent previous piece of information; the differences include the receiving time interval, the platform to which it belongs, the session type, the group to which it belongs, and the contact to which it belongs; The specific content of the information source to be broadcast is determined based on the differences mentioned above.

11. The method according to claim 10, characterized in that, The specific content of determining the source of the information to be broadcast based on the difference includes: If the reception time interval between the first message and the adjacent previous message is greater than a preset interval, then it is determined that the content of the information source to be broadcast includes the complete information source; If the reception time interval between the first message and the adjacent previous message is less than or equal to the preset interval, then the changes in the platform, session type, group, and contact person between the first message and the adjacent previous message are determined sequentially; the specific content of the information source to be broadcast is determined based on the changes.

12. The method according to claim 11, characterized in that, The specific content of determining the source of the information to be broadcast based on the transformation includes: If any of the following changes between the first message and the preceding message: platform, session type, group, or contact, then the content of the message to be broadcast will include the changed content.

13. The method according to any one of claims 10 to 12, characterized in that, The specific content of determining the source of the information to be broadcast based on the difference also includes: Determine the level of familiarity between the current user and the target contact; The specific details of the name of the target contact person included in the information source to be broadcast are determined based on the level of familiarity.

14. The method according to claim 13, characterized in that, Determining the familiarity between the user and the target contact includes: Obtain the interaction behavior characteristics between the current user and the target contact; Based on the characteristics of the interactive behavior, the degree of familiarity between the user and the target contact is determined.

15. The method according to claim 13 or 14, characterized in that, The specific details of determining the name of the target contact person included in the information source to be broadcast based on the level of familiarity include: If the target contact is determined to be a familiar contact of the current user based on the level of familiarity, then the name of the target contact can be omitted from the source of the information to be broadcast. If, based on the level of familiarity, the target contact is determined to be an unfamiliar contact of the current user, then the name of the target contact is simplified, and the name of the target contact included in the information source to be broadcast is determined to be the simplified name of the target contact.

16. The method according to any one of claims 9 to 15, characterized in that, The step of generating the target information based on the information source and the content of the received first information includes: Estimate the duration of the first information to be read aloud; If the duration exceeds a preset value, the first information is simplified. The target information is generated based on the information source and the content of the simplified first information.

17. The method according to any one of claims 9 to 16, characterized in that, The first information also includes non-text information, and the step of generating the target information based on the information source and the content of the received first information further includes: Determine the information type of the non-text information; The non-text information is converted into text based on the information type. The target information is generated based on the information source and the first information after text conversion.

18. The method according to claim 17, characterized in that, The non-text information includes link information; the text conversion of the non-text information according to the information type includes: The link content corresponding to the link information is determined, and the link content is summarized to obtain a summary text in text form.

19. The method according to claim 17 or 18, characterized in that, After converting the non-text information to text according to the information type, the method further includes: Determine the session type that received the non-text information; Add transitional phrases to the first information after text conversion based on the session type.

20. The method according to any one of claims 1 to 19, characterized in that, The first information includes multiple messages sent within a preset time period, and the generation of target information corresponding to the first information further includes: Merge multiple messages sent within a preset time period; Generate the target information corresponding to the merged multiple pieces of information.

21. The method according to claim 20, characterized in that, The merging of multiple messages sent within a preset time period includes: Determine the conversation type of multiple messages sent by the target contact within a preset time period; Multiple messages sent by the target contact within a preset time period and belonging to the same session type are merged.

22. The method according to claim 20, characterized in that, The first information includes multiple group conversation messages sent by multiple associated contacts within a preset time period. The merging of these multiple messages sent within the preset time period includes: The content of multiple messages sent by the associated contacts within a preset time period is determined respectively; Multiple messages with similar content sent by multiple associated contacts within a preset time period are merged.

23. The method according to any one of claims 1 to 22, characterized in that, The first information also includes non-real-time information, and the method further includes: In response to a user instruction, retrieve one or more non-real-time messages corresponding to the user instruction; Generate target information corresponding to one or more of the aforementioned non-real-time information, and broadcast the target information via voice.

24. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the information processing method as described in any one of claims 1 to 23.

25. A computer program product, characterized in that, When the computer program product is run on a computer, the computer performs the information processing method as described in any one of claims 1 to 23.

Citation Information

Patent Citations

  • Method for reading text short message

    CN101175272A

  • Information processing method and device

    CN104900226A

  • Character anthropomorphic broadcasting method and system thereof

    CN111261139A

  • Method and device for broadcasting message through voice and medium

    CN113012679A