Interaction method for call interface, and related apparatus
By supporting switching and interaction between AI assistant and human mode in the call interface, and using a large language model to generate summaries, the problem of AI assistants struggling to cope with complex call scenarios has been solved, improving user experience and work efficiency.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-05-28
- Publication Date
- 2026-04-02
AI Technical Summary
Existing AI assistants struggle to effectively handle complex and ever-changing call scenarios, resulting in poor user experience and low work efficiency.
This paper provides a call interface interaction method that supports switching between AI assistant and human mode. It generates call voice signals by detecting operations and prompts, displays information cards and summary cards, allows users to interact and adjust the state of the AI assistant in real time, and generates key information summaries using a large language model.
It improves the user experience and work efficiency in complex call scenarios, ensuring that the AI assistant can better handle complex scenarios, generate key information summaries, reduce erroneous responses, and improve information recording efficiency.
Smart Images

Figure CN2024095790_02042026_PF_FP_ABST
Abstract
Description
A call interface interaction method and related device TECHNICAL FIELD
[0001] The present application relates to the technical field of terminal devices, and in particular to a call interface interaction method and related device. BACKGROUND
[0002] With the development of society and the revolution of manufacturing technology, the world is accelerating into a digital and intelligent society of everything sensing, everything connecting and everything intelligent, and the digitization of electronic devices has gradually become mainstream. In the context of the digital age, people have higher requirements for efficient work and life, such as the multi-task scenario of having an important phone call while participating in an online meeting. Users can use an AI assistant to help answer the phone and record important information during the call to improve efficiency.
[0003] Currently, using an AI assistant to answer a call has become a mainstream scenario. However, due to the complexity of the call scenario and the diversity of the call type, the AI assistant has difficulty in handling complex and variable call scenarios. Therefore, a call mode that supports user interaction with the AI assistant is essential.
[0004] SUMMARY
[0005] The present application provides a call interface interaction method that can help users answer calls and generate key information summary cards, thereby improving user experience and work efficiency.
[0006] The first aspect of the present application provides a call interface interaction method. The method is applied to an electronic device with an AI assistant application, and the electronic device is not limited to a smart terminal electronic device such as a mobile phone, computer, tablet, watch, etc. The method includes: detecting that an AI call mode is enabled, the AI call mode being a mode in which an AI assistant answers a call with a caller; and detecting a first operation, the first operation being used to indicate that the AI call mode is switched to a manual call mode, the manual call mode being a mode in which a user answers a call with the caller.
[0007] The method uses an AI assistant to answer a call, helps users receive some key call information, and realizes interaction during a call through interaction between the AI call mode and the manual mode, thereby reducing the incorrect response of the AI assistant to the call content and solving the problem that the AI assistant cannot handle complex call scenarios. This can improve user experience and work efficiency.
[0008] In one possible implementation, the method further includes: detecting input of a prompt word, generating a call voice signal based on the prompt word, and sending the call voice signal to the caller.
[0009] In the scheme, the user can set the reply content and manner of the AI assistant based on the prompt word popped up by the interface. The AI assistant can adjust the state and content in a timely manner according to the actual scene requirements, thereby providing better use experience.
[0010] In a possible implementation, the method further includes: the call interface displays an information card and a summary card, the information card is used to indicate key content text in the call information, and the summary card is used to indicate summary content text in the call information.
[0011] In the scheme, the call interface displays the information card and the summary card to record the content in the call process, thereby improving the use experience and work efficiency of the user.
[0012] In a possible implementation, the method further includes: the call interface displays an AI assistant icon, and the AI assistant icon is used to switch the interaction mode with the user.
[0013] In the scheme, a scene of interface interaction is provided, and the user switches the call mode by operating the AI assistant icon.
[0014] In a possible implementation, the prompt word specifically includes a fixed recommended word, and the fixed recommended word includes a tone recommended word and a control recommended word; the tone recommended word is used to indicate the tone of the output voice text; and the control recommended word is used to output the voice text to control the duration of the call.
[0015] In the scheme, the prompt word includes a tone recommended word and a control recommended word, wherein the tone recommended word can be selected by selecting a tone type, including a cute tone, a lively tone, a serious tone, a gentle tone, an angry tone, and the like, to help the AI assistant adapt to the current call context by adjusting the tone and better cope with complex scenes to express the user's emotions; and the control recommended word is mainly used to control and remind the call duration by selecting the call language.
[0016] In a possible implementation, the prompt word specifically includes a real-time recommended word, and the real-time recommended word is text data generated according to real-time call content and is used to reply to the caller.
[0017] In the scheme, a real-time recommendation word is proposed, which is mainly based on the reply to the caller object Some key information prompt words generated by the call content of the caller object, such as the caller object asks which restaurant, A restaurant or B restaurant? The call interface can appear the options of "A restaurant" and "B restaurant" two restaurants, or the caller object proposes to reserve the meeting room time, and the call interface gives options based on the user's choice for selection. This monkey way can well provide user participation in AI call interaction, which can well help AI assistant better express user ideas and intentions, and thus can effectively improve user experience and work efficiency.
[0018] In a possible implementation, the prompt word specifically includes user input text, which is manually input by the user and is text data for expressing user intentions and ideas.
[0019] In the scheme, a way of supporting user manual input of text is proposed to support user real-time expression of ideas and interaction with the caller object, which can well help AI assistant better cope with complex scenarios and better express user intentions, and in some scenarios can be more tactful and humble, and more helpful to interpersonal relationship maintenance.
[0020] In a possible implementation, the prompt word specifically includes user input voice, and the method includes: detecting a second operation, the second operation being used to indicate receiving the user input voice; converting the user input voice into AI assistant voice; and sending the AI assistant voice to the caller object.
[0021] In the scheme, an AI assistant answers the call, and the human helps the AI assistant better reply to the caller object by inputting voice.
[0022] In a possible implementation, the method further includes: detecting a third operation, and adjusting the call speed of the AI assistant.
[0023] In the scheme, a method for adjusting the call speed of the AI assistant is proposed, which can improve the user experience of using the AI assistant, so that the AI assistant can better simulate human behavior and language rhythm.
[0024] In a possible implementation, the call interface displays an information card and an abstract card, the information card is used to indicate key content text in the call information, and the abstract card is used to indicate summary content text in the call information.
[0025] In a possible implementation, the call interface display information card and the summary card specifically include: detecting an audio signal, identifying information of a preset type in the audio signal to obtain key content text and summary text; generating an information card based on the key content text; and generating a summary card based on the summary text.
[0026] In this solution, the user can identify key information, preset information, and the like from the audio information of the incoming call object, and based on these information, important summary information can be extracted and a summary card can be generated. In addition, information in the call content can be summarized to generate a summary text for recording and reminding the user.
[0027] In a possible implementation, the preset type in the audio signal includes, but is not limited to, a telephone number, an address, and a schedule.
[0028] In a possible implementation, the summary text is a text content that summarizes the call content in a preset time length.
[0029] In this solution, the AI assistant can generate a content summary text based on a large language model according to the audio call content in a preset time length.
[0030] The second aspect of the present application provides an electronic device of an AI assistant, the device comprising: a detection module configured to detect an AI call mode, the AI call mode being a mode in which the AI assistant calls an incoming call object; and the detection module is further configured to detect a first operation, the first operation being used to indicate that the AI call mode is switched to a manual call mode, the manual call mode being a mode in which the user calls the incoming call object.
[0031] In a possible implementation, the device further comprises: an input module configured to input a prompt word, generate a call voice signal based on the prompt word, and send the call voice signal to the incoming call object.
[0032] In a possible implementation, the device further comprises: a display module configured to display an information card and a summary card, the information card being used to indicate key content text in call information, and the summary card being used to indicate summary content text in call information.
[0033] In a possible implementation, the display module is further configured to display an AI assistant icon, the AI assistant icon being used to switch the interaction mode with the user.
[0034] In a possible implementation, the prompt words specifically include fixed recommendation words, and the fixed recommendation words include tone class recommendation words and control class recommendation words; the tone class recommendation words are used to indicate the tone of the output voice text; and the control class recommendation words are used to output the voice text to control the duration of the call progress.
[0035] In a possible implementation, the prompt words specifically include real-time recommendation words, and the real-time recommendation words are text data generated according to real-time call content, and are used to reply to the caller.
[0036] In a possible implementation, the prompt words specifically include user input text, and the user input text is text data manually input by the user, and is used to express the user's intention and idea.
[0037] In a possible implementation, the prompt words specifically include user input voice, and the method comprises: detecting a second operation, the second operation being used to indicate that the user input voice is received; converting the user input voice into AI assistant voice; and sending the AI assistant voice to the caller.
[0038] In a possible implementation, the detection module is further used to detect a third operation, the third operation being used to indicate that the call speed of the AI assistant is adjusted.
[0039] In a possible implementation, the display module is used to display an information card and an abstract card, the information card being used to indicate key content text in the call information, and the abstract card being used to indicate summary content text in the call information.
[0040] In a possible implementation, the device further comprises a text generation module, and the text generation module is used to generate text data from the detected audio signal, specifically comprising: the detection module detects the audio signal, identifies information of a preset type in the audio signal, to obtain key content text and summary text; the text generation module generates the information card based on the key content text; and the text generation module generates the abstract card based on the summary text.
[0041] In a possible implementation, the preset type in the audio signal includes, but is not limited to, a telephone number, an address, and a schedule.
[0042] In a possible implementation, the summary text is text content summarizing the call content within a preset duration.
[0043] The third aspect of the present application provides an electronic device of an AI assistant, which can include a processor, the processor and a memory are coupled, and the memory stores program instructions. When the program instructions stored in the memory are executed by the processor, the method of the first aspect or any implementation manner of the first aspect is implemented. For the processor to execute the steps in each possible implementation manner of the first aspect, specific details can be referred to the first aspect, which will not be repeated here.
[0044] The fourth aspect of the present application provides a computer readable storage medium, and the computer readable storage medium stores a computer program. When the computer program is run on a computer, the computer is caused to execute the method of any implementation manner of the first aspect.
[0045] The fifth aspect of the present application provides a computer program product, and when the computer program product is run on a computer, the computer is caused to execute the method of any implementation manner of the first aspect.
[0046] The beneficial effects of the second aspect to the fifth aspect described above can be referred to the introduction of the first aspect, which will not be repeated here. BRIEF DESCRIPTION OF DRAWINGS
[0047] FIG. 1 is a schematic diagram of an architecture of an application scenario provided by an embodiment of the present application;
[0048] FIG. 2 is a schematic diagram of a call receiving interface provided by an embodiment of the present application, wherein (a) of FIG. 2 and (b) of FIG. 2 are schematic diagrams of a call application interface of artificial receiving, and (c) of FIG. 2 is a schematic diagram of a call application interface of AI receiving;
[0049] FIG. 3 is a schematic diagram of starting an AI call mode provided by an embodiment of the present application;
[0050] FIG. 4 is a schematic diagram of generating and typing a call voice signal based on a prompt word provided by an embodiment of the present application, wherein (a) of FIG. 4 and (b) of FIG. 4 respectively show two different presentation modes;
[0051] FIG. 5 is a schematic diagram of a call interface after calling out a virtual keyboard provided by an embodiment of the present application;
[0052] FIG. 6 is a schematic diagram of a call interface after calling out a virtual keyboard provided by an embodiment of the present application;
[0053] FIG. 7 is a schematic diagram of generating a summary card provided by an embodiment of the present application;
[0054] FIG. 8 is a schematic diagram of entering a call card interface from a receiving interface provided by an embodiment of the present application;
[0055] FIG. 9 is a schematic diagram of an electronic device 100 carrying an AI assistant provided by an embodiment of the present application. DETAILED DESCRIPTION
[0056] In order to make the purposes, technical solutions and advantages of the present application clearer, the embodiments of the present application are described below with reference to the drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Those skilled in the art can know that the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems as new application scenarios appear.
[0057] The terms "first", "second", and the like in the specification and claims of the present application and the above-described drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the descriptions used in this way can be interchanged under appropriate circumstances, so that the embodiments can be implemented in an order other than that illustrated or described in the present application. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device that includes a series of steps or modules does not have to be limited to only those steps or modules clearly listed, but can include other steps or modules that are not clearly listed or inherent to these processes, methods, products or devices. The naming or numbering of the steps appearing in the present application does not mean that the steps in the method flow must be executed in the time / logical order indicated by the naming or numbering, and the flow steps that have been named or numbered can change the execution order according to the technical purpose to be achieved, as long as the same or similar technical effects can be achieved.
[0058] The division of units appearing in the present application is a logical division, and in actual application, there can be another division mode, for example, multiple units can be combined or integrated in another system, or some features can be ignored or not executed, in addition, the coupling or direct coupling or communication connection between the units shown or discussed can be through some interface, the indirect coupling or communication connection between the units can be electrical or other similar forms, which are not limited in the present application. And the units or sub-units described as separate components can or can not be physically separated, can or can not be physical units, or can be distributed to multiple circuit units, and some or all of the units can be selected according to actual needs to achieve the purpose of the present application scheme.
[0059] For the sake of understanding, some technical terms related to the embodiments of the present application are introduced first.
[0060] (1) AI assistant
[0061] The AI assistant is an artificial intelligence assistant with functions such as AI question and answer assistant, AI drawing assistant, AI analysis assistant, and intelligent voice translation robot, which provides an efficient and convenient artificial intelligence system platform for users.
[0062] (2) Message card
[0063] The categories of the message card can be divided in various ways. For example, for WeChat, which belongs to instant messaging information software, unread messages can be divided into text messages, picture messages, video messages, voice messages, and voice call / video call messages according to content. They can be divided into personal messages and group messages according to the attributes of the sending object. These different types can all be used as categories of the message card.
[0064] (3) Large language model
[0065] Large language models are one of the applications of deep learning, especially in the field of natural language processing (NLP). Large language models need to be trained on a large amount of text data to learn various patterns and structures of language. The goal of these models is to understand and generate human language in order to have effective conversations and answer various questions.
[0066] Currently, with the progress of technology and the development of society, telephone conferences in daily work should become overwhelming, often missing important calls because of the numerous things. In order to make people's life more convenient and work more efficient, using AI assistants to answer calls is an important means. AI assistants can have a conversation with the caller according to the preset tone, and can also generate a call summary information for the key information in the conversation content, so that even if the user is not present, the user can still clearly understand the content of the call.
[0067] This method can effectively avoid missing important calls by using AI assistants to answer the phone; by supporting AI mode and manual mode switching, it can effectively avoid the problem that AI assistants are difficult to handle complex scenarios and complex call content; this method also supports user guidance of AI call tone and call content in AI call mode, so that the AI assistant can better handle the call content and can reply to the questions in the call according to the user's wishes; in addition, the summary information generated by the large language model based on the call audio data can be used to obtain the key information of the telephone conference and the summary information for recording the call content within the preset call duration, which can effectively improve work efficiency and user experience.
[0068] Therefore, the embodiments of the present application provide a call interface interaction method, which allows the use of AI assistants to answer important incoming calls by setting up important phone calls. Within a certain period of time after the unanswered incoming call rings, the AI assistant automatically answers the call and generates a key information summary card based on the call content. During the AI assistant answering process, the user can intervene at any time to interact with the AI assistant, guide the AI assistant to better handle the call, and also can turn off the AI assistant function at any time during the call to continue the call using the manual answering method.
[0069] Please refer to FIG. 1, which is an architecture diagram of an application scenario provided by an embodiment of the present application. As shown in FIG. 1, the application scenario includes a user 001 and an electronic device 002. The electronic device 002 can be a smart terminal electronic device such as a mobile phone, a tablet, a notebook computer, a watch, and the like. Specifically, the user 001 is usually the owner of the electronic device 002, and obtains network information and communicates with others through the electronic device 002. The electronic device 002 is equipped with an operating system and installed with various software such as a call application, a short message, an instant messaging software, an email application, and the like. The electronic device obtains satellite / network call signals through a communication module such as an antenna, and enables the call application to display a call interface.
[0070] In a possible implementation, the user 001 sets the AI assistant in the electronic device 002 to automatically answer the phone after the ring for a period of time (the period of time can be preset, and it is assumed that the period of time is set to 40 seconds). Exemplarily, it is assumed that the user does not answer the phone after the electronic device rings for 40 seconds, and at this time, the AI assistant automatically answers the phone.
[0071] In a possible implementation, the user 001 receives a call from “Zhang San”, at which time the electronic device 002 interface displays a call answering button and a call hanging up button, and displays an identification of “swiping up to answer the call by the AI assistant” at the bottom. Swiping up the AI assistant button can enable the AI assistant to answer the call.
[0072] The above scheme is described by taking the AI assistant answering the phone as an example. In other embodiments of the present application, the AI assistant can also help to answer a voice call and the like, and can communicate with other users in a voice communication manner.
[0073] Please refer to FIG. 2, which is a schematic diagram of a call answering interface provided by an embodiment of the present application.
[0074] S101: AI assistant icon is displayed on a call interface
[0075] The electronic device detects a call signal, displays a call interface, and the call interface displays an AI assistant icon. The electronic device is installed with a call application, which can be a satellite call application provided by the system or a network call application. The call application receives the call signal through a communication module, for example, receives a satellite signal through an antenna. After receiving the satellite signal, the call application is started and displayed on the front end to display a call interface.
[0076] In the scheme, the incoming call interface includes a call receiving interface and a call receiving interface. The call receiving interface is the call application interface before the user performs the preset operation. The call receiving interface is the interface of the call being connected after the user performs the preset operation. The preset operation is the operation defined by the system to connect the call. For example, the call is connected by dragging the call receiving button to move.
[0077] The call receiving interface and the call receiving interface both display the AI assistant icon. In the scheme, the preset operation can also be an interactive operation with the AI assistant, such as dragging the AI assistant icon to move, interacting with the AI assistant through voice input, etc.
[0078] Exemplarily, FIG. 2 shows the call receiving interface and the call receiving interface respectively. Wherein (a) in FIG. 2 and (b) in FIG. 2 are artificial call application interface diagrams, and (c) in FIG. 2 is an AI call application interface diagram. (a) in FIG. 2 is a call receiving interface, and (b) in FIG. 2 and (c) in FIG. 2 are call receiving interfaces. In the call receiving interface of (a) in FIG. 2, in addition to the call receiving button and the hang-up button, and in the call receiving interface of (b) in FIG. 2, in addition to the hang-up button, the AI assistant icon is displayed at the bottom of the screen. Before the user interacts with the AI assistant icon, only part of the AI assistant icon is displayed on the screen, and part of the AI assistant icon is outside the screen, that is, the AI assistant is not completely displayed on the call receiving interface. The user can drag the AI assistant icon to move to completely display the AI assistant icon on the screen. In (c) in FIG. 2, in the AI call mode, the AI assistant icon is completely displayed in the preset position on the screen, and the AI assistant application synchronously receives the audio signal, understands the audio, and outputs the AI voice according to the user's input and the semantic understanding of the AI assistant to the audio converted to text, and sends the AI voice to the incoming call object.
[0079] S102: Enable AI call mode
[0080] The AI call mode is enabled when the AI assistant icon is detected to move from the bottom of the screen to the middle of the screen, and the call audio is synchronously transmitted to the AI assistant application. In the call receiving interface, the user can interact with any one of the call receiving, hang-up, and AI assistant icons to process the call. The interaction of the call receiving and hang-up buttons is consistent with the prior art. The user can drag the AI assistant icon to move, and drag the AI assistant icon from the bottom of the screen to the preset area in the middle of the screen, or the distance of dragging is greater than a preset value, so as to enable the AI call mode (the AI call mode is opposite to the artificial call mode).
[0081] Please refer to FIG. 3, which is a schematic diagram of starting an AI call mode according to an embodiment of the present application. As an example, FIG. 3 shows a process of enabling the AI call mode. After the user selects the AI assistant icon, the user drags the AI assistant to move in the screen. When the AI assistant moves to a preset area of the screen (such as the area of the dashed box in the middle figure), if the user releases the gesture at this time, it is considered that the AI call answering condition is met, and the AI call is enabled. After the gesture is released, the AI assistant icon moves from the gesture release position to a preset position of the answering interface.
[0082] In a possible implementation, in the AI call mode, the electronic device sends the received audio data to the call application and synchronously sends the received audio data to the AI assistant application. After the AI assistant application calls a language model based on the received audio data to understand the audio data, the AI assistant application outputs reply text corresponding to the audio data, converts the reply text into a voice signal, and sends the voice signal to the call application, which then sends the voice signal to the caller. At the same time, in the AI call mode, the call application synchronously sends the voice signal to the audio output module for output when sending the voice signal to the caller, for example, directly outputting the voice signal through a loudspeaker or sending the voice signal to a Bluetooth earphone through a Bluetooth module. In this way, the user can obtain all conversations between the AI assistant and the caller in real time as a third party.
[0083] In a possible implementation, the input received by the electronic device is a call voice signal generated based on a prompt word and input by the user. Specifically, the input of the prompt word is detected, the call voice signal is generated based on the prompt word, and the call voice signal is sent to the caller. In the AI call answering mode, the user can input the prompt word, the AI assistant generates reply text based on the input prompt word and the understanding of the received audio information, converts the reply text into a call voice, and sends the call voice to the call application, which then sends the call voice to the caller / audio output module.
[0084] In a possible implementation, the input of the prompt word can be a recommended word selected by the user in the call interface or an input sentence input by the user in the call interface by calling a virtual keyboard. Specifically, after the AI call answering mode is enabled, the AI assistant can display recommended words in the call interface, and the user can select the recommended words to affect the call sentences output by the AI. In addition, the call interface in the AI call mode can also display a virtual keyboard icon. The user clicks the virtual keyboard icon to enable the virtual keyboard to input directly, and the input sentence is used as the prompt word.
[0085] In this scheme, the recommended words can be generated in the following ways:
[0086] The recommended words can include fixed recommended words and real-time recommended words. The fixed recommended words are fixed words applicable to all AI calls. The real-time recommended words are generated by the AI assistant based on the received voice signal, and are directly related to the audio sentence of the incoming call object. In other words, the real-time recommended words are generated by real-time recognition according to the current call content, and are generated for the semantics of the current incoming call object.
[0087] 1) Fixed recommended words
[0088] The fixed recommended words include tone-based recommended words and process control-based recommended words. The tone-based recommended words are used to affect the wording of the output voice text, for example, selecting a formal tone and a lively tone, and the final output text will be different. The text corresponding to the severe tone is more written, and the text corresponding to the lively tone is more colloquial.
[0089] The process control-based recommended words are used to control the progress length of the call, including quick ending, maintaining the call, etc. For example, selecting quick ending to let the AI assistant add a sentence indicating the expected end of the call while completing the current output or executing the current output, so as to convey the intention of ending the conversation to the incoming call object (instead of directly hanging up the phone).
[0090] 2) Real-time recommended words
[0091] The real-time recommended words can be generated by the AI assistant based on the received audio signal, or can be generated by the AI assistant based on the expected reply content output for the received audio signal. The expected reply content is the text generated by the AI assistant based on the received audio signal and not yet sent to the call application.
[0092] Specifically, the real-time recommended words can be keywords extracted from the text obtained by voice-to-text conversion of the received audio signal. For example, the text obtained by voice conversion is "Where to eat in the evening, A restaurant or B restaurant?", and the keywords "A restaurant" and "B restaurant" are extracted as recommended words. In addition, the derived words of the keywords can also be obtained based on the keywords and the context according to the preset rules. For example, the foregoing voice is a question sentence, and two choices are given - A restaurant and B restaurant. In addition to selecting an answer from A restaurant and B restaurant, the other party can also be allowed to decide, and therefore, the derived word "decided by the other party" can be obtained.
[0093] Alternatively, the real-time recommended words can also be keywords extracted from the expected reply content, or derived words obtained from the keywords extracted from the expected reply content based on the preset rules. For example, the text obtained by voice conversion is "Eat in the evening?", and the AI can output the reply words "agree" and "refuse" in addition to "see later", and therefore, the derived word "see later" is obtained.
[0094] 3) Presentation of recommended words
[0095] The recommended words can be presented by recommended word category, or by the dimension of individual recommended words. Each recommended word corresponds to a bubble and is presented in a preset area of the call interface. The recommended word bubble can be presented statically or in a dynamic effect, for example, the recommended word bubble can move within the preset area, for example, a dynamic effect of presenting a moving bubble from inside the screen to outside the screen, and an individual recommended word bubble can be presented for a period of time and then disappear for a period of time.
[0096] All recommended words can be displayed on the call interface, or only recommended words of a preset type can be displayed on the call interface, and recommended words of a non-preset type are folded and only displayed on the call interface after the user performs a preset operation.
[0097] Real-time recommended words have a direct impact on the output result of the AI call, mood recommended words only have a modifying effect on the output result, and process control recommended words are only called when the user has a demand, therefore, all real-time recommended words are displayed on the call interface in the form of recommended word bubbles, and mood recommended words and process control recommended words are only displayed in the form of two type controls, and clicking the type control displays the recommended words of the type on the call interface.
[0098] Please refer to FIG. 4, which is a schematic diagram of generating and typing a call speech signal based on a prompt word according to an embodiment of the present application. Exemplarily, (a) in FIG. 4 and (b) in FIG. 4 respectively show two different presentation modes. In (a) in FIG. 4, two rows of recommended words are displayed in a static display mode, one row of real-time recommended words, such as “politely refuse”, “accept the invitation of the other party”, “give a conclusion later” and the like, which are related to the current call content and need to be recommended in real time according to the current call content; and one row of fixed recommended words, such as “light mood”, “vibrant mood”, “formal mood”, which are guiding and suggestive selection of the reply mood of the AI assistant. All types of recommended words are folded and directly displayed in the form of recommended word bubbles (limited by the screen area, some bubbles are displayed outside the screen), and more recommended word bubbles can be displayed by swiping left and right. In (b) in FIG. 4, the real-time recommended words are displayed around the icon of the AI assistant and dynamically change position over time, and are configured with a dynamic effect (for example, the bubble moves from inside the screen to outside the screen and gradually appears, with the transparency continuously decreasing), and the fixed recommended words are only displayed in a fixed position in the form of classification controls—“mood selection” and “call control”, and clicking the corresponding classification control displays the recommended words of the type on the call interface in the form of recommended word bubbles, and the result after display is referred to the fixed recommended words in (a) in FIG. 4.
[0099] In an implementation, if the recommended word displayed does not meet the user's needs, the user can also click the virtual keyboard icon in the lower left corner to call up the virtual keyboard and manually input a sentence as a recommended word.
[0100] Referring to FIG. 5, FIG. 5 is a schematic diagram of a post-call interface provided by an embodiment of the present application to call up a virtual keyboard. As an example, as shown in FIG. 5, after the user inputs a custom text, the user clicks send, the virtual keyboard transmits the user input text to the AI assistant application, and the AI assistant application receives the input text and generates a reply text using the input text as a recommended word. For example, the input text is directly used as the reply text.
[0101] After the AI assistant application receives the recommended word, the AI assistant application inputs the recommended word and the text converted from the received audio into a built-in / cloud language model to obtain a reply text, and converts the reply text into audio and outputs the audio to the call application.
[0102] When the user selects a tonal recommended word, different voices can be selected to synthesize the audio when the reply text is converted into audio according to the tonal recommended word. The correspondence between the tonal recommended word and the voice can be predefined.
[0103] For example, different tonal recommended words can correspond to different voice models, and the corresponding voice model is selected according to the tonal recommended word; or different tonal recommended words correspond to different voice labels, and the voice model is selected based on the label. It can be understood that, in order to unify the output audio in sound, different voice models can be different tonal voice models obtained by fine-tuning based on the same reference voice.
[0104] When the call application sends the received audio, the audio output module is also called to play the received audio, so that the user can receive the voice conversation between the AI assistant and the caller in real time.
[0105] In the AI call answering interface, in addition to the recommended word interaction, the following interactions can also be included:
[0106] ① Long press AI assistant to send voice: this can be achieved in two ways. One way is to long press the AI assistant to trigger the AI assistant to receive a voice signal and send the voice signal to the call application, and then send the voice signal to the opposite end by the call application. Another way is to long press the AI assistant to trigger the AI call mode to be interrupted, and temporarily switch to a manual answering mode, so that the call application directly receives and sends the user's voice. When it is detected that the long press gesture is released, the AI call mode is switched back.
[0107] ② Dragging the AI assistant to move, adjusting the speech speed of the AI assistant by controlling the moving direction and example. For example, dragging the AI assistant to move left / right to adjust the speed of the speech. Of course, it can also be up and down. The speech speed can be adjusted by directly dragging the AI assistant to move, or by activating the speech speed adjustment function through a preset gesture (such as a long press time exceeding a preset value), and then dragging the AI assistant to move to adjust the speech speed.
[0108] In a possible implementation, the electronic device can implement switching the AI call mode to the manual answering mode. Specifically, detecting that the AI assistant icon is moved from the middle of the screen to the bottom of the screen, and switching from the AI call mode to the manual answering mode. The user can switch the AI call mode to the manual answering mode by moving the AI assistant icon to the bottom of the screen during the AI call. Specifically, if the electronic device detects the movement of the AI assistant icon again, the electronic device obtains the position of the AI assistant icon in real time according to the operation gesture, and if the icon of the AI assistant at least partially falls in a preset area (such as outside the screen display area), the electronic device ends the AI call mode, and the call application no longer sends the audio data to the AI assistant application, but switches to the manual answering mode and only outputs through the audio output module.
[0109] Please refer to FIG. 6, which is a schematic diagram of a call interface after calling out a virtual keyboard provided by an embodiment of the present application. Exemplarily, FIG. 6 shows an interactive process of switching from the AI call mode to the manual answering mode. After selecting the AI assistant icon, the AI assistant icon is dragged to move downward, and the AI assistant icon is moved to at least partially outside the screen display area, and the AI call mode is switched to the manual answering mode.
[0110] In a possible implementation, the AI assistant application can also generate a summary card and an information card during the call process, wherein the summary card is used to record the key information in the call process, and the information card is used to summarize the call information in the call process.
[0111] In a possible implementation, the electronic device can generate a summary card. Specifically, displaying the answering interface, the AI assistant application obtains the audio signal of the incoming call object in real time, identifies the preset type of information in the audio signal to obtain the key content text, and generates a summary text based on the audio signal, the call application generates an information card based on the key content text, generates a summary card based on the summary text, and displays the summary card on the answering interface.
[0112] After the call application is connected (whether in AI answering mode or manual answering mode), the call application enters the answering interface. After entering the answering interface, the AI assistant application obtains the audio signal received by the call application in real time and analyzes the audio signal to identify the preset type of information (key content text) and obtain summary information (summary text) within a preset time. For example, after performing speech-to-text operation on the audio signal, the obtained text is input into the large language model to obtain the key content text and the summary text.
[0113] In this scheme, the preset type of information is: telephone number, address, schedule, etc. When the audio contains the preset type of information, the AI assistant application extracts the preset type of information separately to obtain the key content text. Specifically, the large language model can be trained to have the ability to identify the telephone number format of each country and the ability to identify address descriptions, so that when the text is input into the large language model, the text content conforming to the telephone number format / address description is extracted separately.
[0114] In an implementation manner, the AI assistant also generally has the ability to summarize and summarize the speaking content of the caller to obtain the summary text. The summary text can be generated in time units, for example, a summary text is generated every preset time (such as 1 minute). The audio pause time of the caller in the audio can also be used to perform audio sentence segmentation on the pause time exceeding the preset time (such as 5s), and generate a summary text according to the audio sentence. Of course, the AI assistant can also generate a summary text according to the sentence after self-sentence breaking. Alternatively, the AI assistant can judge the topic according to the audio content and generate a summary text according to the topic.
[0115] In an implementation manner, the AI assistant application sends the extracted key content text and summary text to the call application. After receiving the key content text and summary text, the call application generates different cards according to different content types and presents them on the call interface. Specifically, an information card corresponding to the key content text is generated, and a summary card corresponding to the summary text is generated, so that the information card and the summary card are presented on the answering interface.
[0116] In an implementation manner, due to the limitation of screen area, the content presented on the answering interface is limited and cannot display all summary cards and information cards. Since the information card usually has less content and is important information, and the summary card is relatively not important information, the information card can be displayed completely, and the summary card can be displayed in layers. The layered summary card only displays the latest summary card completely.
[0117] Please refer to FIG. 7, which is a schematic diagram of generating a summary card according to an embodiment of the present application. As an example, FIG. 7 shows a receiving interface, and an information card A01 "Place: AA District BB Road B Restaurant" is displayed above the head portrait of the caller, and a summary card B01 "Department organizes dinner at 19:00 tonight, place: AA District BB Road B Restaurant" is displayed above the AI assistant icon in the middle of the screen.
[0118] In a possible implementation, the electronic device can further generate the summary card and the information card. Specifically, a preset operation is detected, and the call card interface is entered, and all information cards and summary cards of the call process are displayed in the call card interface. If the user expects to view all summary cards, the information cards and summary cards of the entire call process can be displayed by a preset operation. For example, swiping to the right enters the call card interface, and all information cards and summary cards are displayed in the call card interface.
[0119] Please refer to FIG. 8, which is a schematic diagram of entering the call card interface from the receiving interface according to an embodiment of the present application. As an example, FIG. 8 shows a schematic diagram of entering the call card interface from the receiving interface, and swiping to the right by a preset distance on the left side (or any position) of the screen of the receiving interface enters the call card interface, and the call card interface displays all information cards and summary cards, in which the key information corresponds to the information card, and the call summary corresponds to the summary card.
[0120] After the information card and the summary card are generated, the call is immediately hung up, and the information card and the summary card also exist until the call application is exited. After the call application is exited, the information card / summary card is destroyed, and of course, a folder can be allocated for each call, and the information card and the summary card are stored in the folder.
[0121] The information card and the summary card can also support interaction, for example, long pressing to add the information card / summary card to the notes / notes / contacts. Specifically, after long pressing the card, a secondary menu appears at the position of the card, and the content of the information card / the content of the summary card is newly created as a note / contact according to the selected secondary menu option.
[0122] The above mainly introduces the scheme provided by the embodiments of the present application from the method perspective. It can be understood that the electronic device includes the hardware structure and / or software module corresponding to the execution of each function in order to realize the above functions. Those skilled in the art should easily realize that, in combination with the modules and algorithm steps of each example described in the embodiments disclosed herein, the present application can be realized in the form of hardware or the combination of hardware and computer software. Whether a certain function is executed in the form of hardware or computer software driven hardware depends on the specific application and design constraints of the technical solution. The skilled person can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0123] The embodiments of the present application can divide the function modules of the electronic device according to the above method examples. For example, each function module can be divided according to each function, or two or more functions can be integrated in one processing module. The above integrated module can be realized in the form of hardware or software function module. It should be noted that the division of modules in the embodiments of the present application is illustrative, and is only a logical function division. Actual implementation can have another division manner.
[0124] The electronic device in the embodiments of the present application is described in detail as follows:
[0125] Please refer to FIG. 9, which is a schematic diagram of an electronic device 100 carrying an AI assistant provided by the embodiments of the present application, including an input module 101, a detection module 102, a text generation module 103 and a display module 104.
[0126] The detection module 101 is configured to detect an AI call mode, which is a mode of applying the AI assistant to call the incoming call object. The detection module is also configured to detect a first operation, which is used to indicate switching the AI call mode to a manual call mode, which is a mode of calling the incoming call object by the user.
[0127] The input module 102 is configured to input a prompt word, generate a call voice signal based on the prompt word, and send the call voice signal to the incoming call object.
[0128] The text generation module 103 is configured to generate a prompt word. The text generation module 103 generates text information according to the received audio signal, identifies audio data of a preset type, and generates an abstract text and a summary text based on a large prediction model. The text generation module generates an information card based on the abstract text, and generates an abstract card based on the summary text.
[0129] The display module 104 is configured to display an information card and an abstract card. The information card is used to indicate key content text in the call information, and the abstract card is used to indicate summarized content text in the call information. The display module is also configured to display an AI assistant icon, which is used to switch the interaction mode with the user.
[0130] In a possible implementation, the detection module 102 detects that the AI call mode is enabled, and the AI assistant performs a voice call with the caller. During the call, the input module 101 allows the user to input text data, voice data, and the content of a selection control according to actual needs and call content. The detection module 102 detects the input data in real time, and adjusts the call content and call state of the AI assistant according to the input data, so that the user can interact with the AI assistant in real time during the call between the AI and the caller. In addition, the display interface displays the AI assistant icon and displays a prompt word in real time. The prompt word can be dynamically generated according to audio data, can be input in real time by the user, or can be of a fixed type. The prompt word can control the call process and call time by controlling the call reply content and tone of the AI assistant.
[0131] In a possible implementation, the electronic device supports switching the call mode through the AI assistant icon, for example, switching from the AI call mode to the manual call mode, or switching from the manual call mode to the AI call mode.
[0132] The electronic device described above is generally a device with an AI assistant, which supports the AI assistant to answer a call and allows the user to interact with the AI assistant in the AI call mode, ensures that the AI assistant performs the call according to the user's intention, and effectively ensures the quality of the call. In addition, the electronic device can also generate corresponding abstract text and summary text according to the audio data of the call, so as to facilitate the user to view and record the meeting content.
[0133] In the above embodiments, all or part of the embodiments can be implemented by software, hardware, firmware, or any combination thereof. When implemented by software, all or part of the embodiments can be implemented in the form of a computer program product.
[0134] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on the computer, the processes or functions according to the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium, for example, the computer instructions can be transmitted from one website site, computer, training device or data center to another website site, computer, training device or data center through wired (such as coaxial cable, optical fiber, digital subscriber line) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that the computer can store or be integrated into a data storage device such as a training device, a data center, etc. which includes one or more available media sets. The available media can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid state disk (SSD)), etc.
Claims
1. An interaction method for a call interface, characterized in that, Applied to an electronic device carrying an AI assistant application, comprising: detecting that an AI call mode is enabled, the AI call mode being a mode in which the AI assistant is applied to call a caller; detecting a first operation, the first operation being used to indicate switching the AI call mode to a manual call mode, the manual call mode being a mode in which a user calls the caller.
2. The method of claim 1, wherein, The method further comprises: detecting input of a prompt word, generating a call voice signal based on the prompt word, and sending the call voice signal to the caller.
3. The method of claim 1, wherein, The method further comprises: The call interface displays an information card and an abstract card, the information card being used to indicate key content text in call information, and the abstract card being used to indicate summary content text in the call information.
4. The method according to any one of claims 1 to 3, characterized in that, The method further comprises: The call interface displays an AI assistant icon, the AI assistant icon being used to interact with the user to switch modes.
5. The method of claim 2, wherein, The prompt word specifically includes a fixed recommendation word, and the fixed recommendation word includes a tone recommendation word and a control recommendation word; The tone recommendation word is used to indicate the tone of the output voice text; The control recommendation word is used to output voice text to control the duration of the call progress.
6. The method of claim 2, wherein, The prompt word specifically includes a real-time recommendation word, which is text data generated according to real-time call content and is used to reply to the caller.
7. The method of claim 2, wherein, The prompt word specifically includes user input text, which is manually input by the user and is text data used to express the user's intentions and ideas.
8. The method of claim 2, wherein, The prompt word specifically includes user input voice, and the method comprises: detecting a second operation, the second operation being used to indicate receiving the user input voice; convert the user input voice into AI assistant voice; and send the AI assistant voice to the caller.
9. The method of claim 1, wherein, The method further comprises: detecting a third operation, and adjusting the call speed of the AI assistant.
10. The method of claim 3, wherein, The call interface displays an information card and an abstract card, the information card being used to indicate key content text in call information, and the abstract card being used to indicate summary content text in the call information.
11. The method of claim 10, wherein, The call interface displays an information card and an abstract card specifically includes: detecting an audio signal, identifying information of a preset type in the audio signal to obtain key content text and summary text; generating an information card based on the key content text; generating an abstract card based on the summary text.
12. The method of claim 10, wherein, The preset type in the audio signal includes but is not limited to a phone number, an address, and a schedule.
13. The method of claim 10, wherein, The summary text is text content summarizing call content within a preset duration.
14. An electronic device carrying an Al assistant, the electronic device comprising: The device comprises: a detection module for detecting that an AI call mode is enabled, the AI call mode being a mode in which the AI assistant is applied to call a caller; The detection module is also used to detect a first operation, the first operation being used to indicate switching the AI call mode to a manual call mode, the manual call mode being a mode in which a user calls the caller.
15. The electronic device of claim 14, wherein, The device further comprises: The input module is configured to input a prompt word, generate a call voice signal based on the prompt word, and send the call voice signal to the incoming call object.
16. The electronic device of claim 14, wherein, The device further comprises: The display module is configured to display an information card and an abstract card, the information card being configured to indicate key content text in the call information, and the abstract card being configured to indicate summary content text in the call information.
17. The electronic device of any of claims 14-16, wherein, The display module is further configured to display an AI assistant icon, the AI assistant icon being configured to switch the interaction mode with the user.
18. The electronic device of claim 15, wherein, The prompt word specifically includes a fixed recommendation word, and the fixed recommendation word includes a tone recommendation word and a control recommendation word. The tone recommendation word is configured to indicate a tone of output voice text. The control recommendation word is configured to output voice text to control a duration of the call.
19. The electronic device of claim 15, wherein, The prompt word specifically includes a real-time recommendation word, and the real-time recommendation word is text data generated based on real-time call content and configured to reply to the incoming call object.
20. The electronic device of claim 15, wherein, The prompt word specifically includes user input text, and the user input text is text data manually input by the user and configured to express the user's intention and idea.
21. The electronic device of claim 15, wherein, The prompt word specifically includes user input voice, and the method comprises: detecting a second operation configured to indicate that the user input voice is received; converting the user input voice into AI assistant voice; and sending the AI assistant voice to the incoming call object.
22. The electronic device of claim 14, wherein, The detection module is further configured to detect a third operation and adjust a call speed of the AI assistant.
23. The electronic device of claim 14, wherein, The display module is configured to display an information card and an abstract card, the information card being configured to indicate key content text in the call information, and the abstract card being configured to indicate summary content text in the call information.
24. The electronic device of claim 23, wherein, The device further comprises a text generation module configured to generate text information from the detected audio information, and specifically comprises: detecting the audio signal, identifying information of a preset type in the audio signal to obtain key content text and summary text; the text generation module generates an information card based on the key content text; the text generation module generates an abstract card based on the summary text.
25. The electronic device of claim 23, wherein, The preset type in the audio signal includes but is not limited to a phone number, an address, and a schedule.
26. The electronic device of claim 23, wherein, The summary text is text content summarizing call content within a preset duration.
27. An electronic device carrying an Al assistant, the electronic device comprising: The device comprises a memory and a processor, the memory stores a code, and the processor is configured to execute the code, when the code is executed, the device executes the method of any one of claims 1 to 13.
28. A computer storage medium, comprising, The computer storage medium stores instructions, and the instructions cause the computer to implement the method of any one of claims 1 to 13 when executed by the computer.
29. A computer program product, characterised in that, The computer program product stores instructions, and the instructions cause the computer to implement the method of any one of claims 1 to 13 when executed by the computer.