Call processing method and device
By using real-time speech-to-text and semantic analysis, key information in telephone transactions is extracted and displayed in a differentiated manner, solving the problem of users forgetting key details during telephone transactions and achieving efficient and convenient business processing.
Patent Information
- Application Number
- CN202511104847.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-07
- Publication Date
- 2025-11-11
AI Technical Summary
During telephone service transactions, users have to listen to lengthy audio messages from customer service representatives, which can lead to forgetting key details, confusion, and inefficiency.
Through real-time speech-to-text conversion and semantic analysis, the call content is segmented into multiple semantic segments, key business information is extracted, displayed in chronological order, and key information is highlighted using differentiated display styles.
Users can quickly identify and respond to key business information, reducing errors and improving the efficiency and accuracy of telephone business processing.
Smart Images

Figure CN120935299A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of electronic equipment technology, specifically relating to a call processing method and apparatus. Background Technology
[0002] In today's era of rapid technological advancement, despite the widespread application of internet technology and the emergence of various online service platforms, telephone services remain an indispensable key method in the customer service systems of many industries due to their immediacy, directness, and universal applicability. Especially in sectors closely related to people's daily lives, such as banking, telecommunications, and insurance, telephone channels handle a large number of core service functions, including business consultation, processing, and inquiries.
[0003] In related technologies, in the telephone service handling mode, after making or receiving a service call, users need to listen to a series of voice messages broadcast by the customer service representative on the other end. These messages may include service introductions, handling procedures, and option explanations.
[0004] However, customer service voice broadcasts often contain a large amount of information and are delivered at a fast pace. After listening to a long string of content, users can easily forget key business details, leading to confusion when making subsequent choices. This requires repeatedly asking customer service to repeat the previous content, which not only wastes a lot of time and energy for both users and customer service, but also makes the telephone business handling process lengthy and cumbersome, greatly reducing the efficiency of business handling. Summary of the Invention
[0005] The purpose of this application is to provide a call processing method and apparatus that can solve the technical problem of low business processing efficiency in related technologies.
[0006] In a first aspect, embodiments of this application provide a call processing method, the method comprising:
[0007] During the call, receive the user's initial input;
[0008] In response to the first input, the voice content of both parties in the call is converted into text data;
[0009] Semantic analysis is performed on the text data to segment it into multiple semantic segments, and key business information is extracted from each semantic segment; wherein, the key business information includes at least one of the following: business operation instructions, numerical information, and contact information;
[0010] The semantic paragraphs are displayed in chronological order; the display style of the key business information is different from the display style of the remaining information in the semantic paragraphs.
[0011] Secondly, embodiments of this application provide a call processing apparatus, the apparatus comprising:
[0012] The receiving module is used to receive the user's initial input during a call;
[0013] The processing module is configured to, in response to the first input, convert the voice content of both parties in the call into text data; perform semantic analysis on the text data, segment the text data into multiple semantic segments, and extract key business information from each semantic segment; wherein, the key business information includes at least one of the following: business operation instructions, numerical information, and contact information;
[0014] The display module is used to display the semantic paragraphs in chronological order; wherein the display style of the key business information is different from the display style of the remaining information in the semantic paragraphs.
[0015] Thirdly, embodiments of this application provide an electronic device including a processor and a memory, wherein the memory stores a program or instructions executable on the processor, and the program or instructions, when executed by the processor, implement the steps of the call processing method as described in the first aspect.
[0016] Fourthly, embodiments of this application provide a readable storage medium storing a program or instructions that, when executed by a processor, implement the steps of the call processing method as described in the first aspect.
[0017] Fifthly, embodiments of this application provide a chip, the chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the call processing method as described in the first aspect.
[0018] In a sixth aspect, embodiments of this application provide a computer program product stored in a storage medium, which is executed by at least one processor to implement the call processing method as described in the first aspect.
[0019] In this embodiment of the application, during a call, a first input from the user is received; in response to the first input, the voice content of both parties in the call is converted into text data; semantic analysis is performed on the text data, the text data is divided into multiple semantic segments, and key business information in each semantic segment is extracted; wherein, the key business information includes at least one of the following: business operation instructions, numerical information, and contact information; the semantic segments are displayed in chronological order; wherein, the display style of the key business information is different from the display style of the remaining information in the semantic segment.
[0020] As can be seen, in this embodiment of the application, by using real-time speech-to-text and semantic analysis, lengthy call content is transformed into structured semantic segments, and key business information in the semantic segments is highlighted in a differentiated manner, enabling users to intuitively identify and respond quickly, reducing misoperations, and helping users quickly locate key business information such as operation instructions, amounts, and contact information. This avoids the inefficiency caused by repeatedly listening to voice messages and the problem of easily missing key information, thereby improving the efficiency of business processing. Attached Figure Description
[0021] Figure 1 This is a flowchart of a call processing method provided in an embodiment of this application;
[0022] Figure 2 This is one of the example diagrams of a call processing method provided in the embodiments of this application;
[0023] Figure 3 This is a second example diagram of a call processing method provided in an embodiment of this application;
[0024] Figure 4 This is the third example diagram of a call processing method provided in the embodiments of this application;
[0025] Figure 5 This is the fourth example diagram of a call processing method provided in the embodiments of this application;
[0026] Figure 6 This is the fifth example diagram of a call processing method provided in the embodiments of this application;
[0027] Figure 7 This is a structural block diagram of a call processing device provided in an embodiment of this application;
[0028] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application;
[0029] Figure 9 This is a schematic diagram of the hardware structure of an electronic device that implements the various embodiments of this application. Detailed Implementation
[0030] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.
[0031] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0032] Telephone services offer significant advantages such as immediacy, directness, and universal applicability. They are particularly valuable in sectors closely intertwined with daily life, such as banking, telecommunications, and insurance, where telephone channels serve numerous core service functions. For example, in banking, users can use the phone to inquire about account balances, transfer funds, apply for and activate credit cards, and inquire about and process loan applications. In telecommunications, users can change service plans, top up phone credit, report faults, and apply for broadband services. In insurance, users can consult about insurance products, purchase policies, and file claims. The convenience and directness of telephone services enable customers in these industries to quickly obtain professional assistance and solutions when encountering problems, effectively protecting customer rights and improving customer satisfaction.
[0033] In related technologies, when conducting business via telephone, after a user makes or receives a business call, a service process primarily based on voice interaction begins. During this process, the user needs to concentrate on listening to a series of voice messages broadcast by the customer service representative. These voice messages are diverse and cover business introductions, such as detailed explanations of a financial product, processing procedures, and option descriptions.
[0034] However, customer service voice prompts often contain a large amount of information and are delivered at a fast pace. After listening to a long string of messages, users can easily forget key business details, leading to confusion when making subsequent choices. This necessitates repeatedly asking customer service to repeat the information, significantly reducing the efficiency of business processing. Furthermore, users cannot selectively listen to key content. If they only want to understand a specific step or option in the business process, they must spend time listening to a lot of irrelevant introductory information, wasting their time and negatively impacting the user experience. Clearly, this technology suffers from the technical problem of low efficiency in telephone business processing.
[0035] To address the aforementioned technical problems, embodiments of this application provide a call processing method and apparatus to offer a more convenient, efficient, and intelligent solution for handling telephone services, thereby improving the efficiency of telephone service processing.
[0036] The call processing method provided in the embodiments of this application will now be described in detail with reference to the accompanying drawings.
[0037] It should be noted that the call processing method provided in this application is applicable to electronic devices. In practical applications, such electronic devices include, but are not limited to, mobile terminals such as smartphones, tablets, smartwatches, and personal digital assistants. This application does not limit these devices.
[0038] Figure 1 This is a flowchart of a call processing method provided in an embodiment of this application, such as... Figure 1 As shown, the method may include the following steps: step 101, step 102, step 103 and step 104.
[0039] In step 101, during the call, the user's first input is received.
[0040] In this embodiment of the application, during the process of handling business by telephone, the user can trigger the electronic device to start real-time voice parsing and text conversion of the voice content of both parties in the call through the first input.
[0041] In this embodiment of the application, the first input includes, but is not limited to: click operation, long press operation, swipe operation, air gesture or voice command.
[0042] In some embodiments, step 101 may specifically include the following step: step 1011.
[0043] In step 1011, when the call interface is displayed, the first input to the call parsing control in the call interface is received.
[0044] In this embodiment of the application, after a user makes or receives a business call, in addition to displaying standard functional controls such as mute control, recording control, hands-free control, virtual keyboard control, and hang-up control on the call interface, a new functional control called "call parsing control" can also be added; wherein, the call parsing control is used to trigger real-time parsing and display of the call voice content.
[0045] For example, in the process of conducting business by telephone, such as Figure 2 As shown, the screen 21 of the electronic device 20 displays a call interface 22. In addition to standard function controls such as mute control 221, recording control 222, hands-free control 223, virtual keyboard control 224, and hang-up control 225, the call interface 22 also includes a call parsing control 225.
[0046] When a user clicks the call parsing control 225, the electronic device 20 receives the operation, treats it as the first input, and immediately activates the background voice analysis module, which then performs real-time parsing of the call audio content. Furthermore, after the call parsing control 225 is clicked, its color can dynamically change, for example, from gray to blue, to indicate to the user that the function has been enabled.
[0047] In some embodiments, step 101 may specifically include the following step: step 1012.
[0048] In step 1012, when exiting the call interface, the first input to the call parsing control in the floating interactive window is received.
[0049] It should be noted that exiting the call interface and exiting the call are two completely different concepts. The former only involves changes in the interface interaction layer, that is, switching from the call interface to other interfaces, while the latter involves the termination of the call connection layer, that is, ending the call.
[0050] In this embodiment of the application, after the user exits the call interface, the electronic device can generate a floating interactive window at the edge of the screen (e.g., the top), and display a call parsing control in the floating interactive window. The call parsing control is used to trigger the real-time parsing and display of the call voice content.
[0051] In this embodiment of the application, when a user actively or passively exits the call interface, such as returning to the home screen or switching applications, but still needs to view or operate the call parsing content, the real-time parsing and display function can be triggered through the call parsing control in the floating interactive window.
[0052] In this embodiment, the floating interactive window can be a semi-transparent floating window. For compact devices, the default size of the floating interactive window can be 80px × 80px, which typically occupies about 12% of the screen width; for standard devices, the size of the floating interactive window can be 100px × 100px, which typically occupies about 15% of the screen width; for large-screen devices, the size of the floating interactive window can be 120px × 120px, which typically occupies about 18% of the screen width.
[0053] In practical applications, when the call interface is displayed, the floating interactive window defaults to a minimized icon, such as a circular icon. This design avoids obscuring the main content of the call interface.
[0054] In this embodiment, the floating interactive window can be represented by an "atomic island".
[0055] In this embodiment, after the user makes the first input to the call parsing control in the floating interactive window, the floating interactive window can simultaneously display the status "Intelligent voice analysis in progress..." to avoid user misoperation.
[0056] For example, during a telephone transaction, if a user exits the call but remains on the call, such as... Figure 3 As shown, a floating interactive window 32 is displayed on the screen 31 of the electronic device 30. The floating interactive window 32 contains a hang-up control 321, a call duration indicator 322, and a call parsing control 323.
[0057] When a user clicks the call parsing control 323, the electronic device 30 receives the operation, treats it as the first input, and immediately activates the background voice analysis module, which then performs real-time parsing of the call audio content. Furthermore, after the call parsing control 323 is clicked, its color can dynamically change, for example, from gray to blue, to indicate to the user that the function has been enabled.
[0058] In this embodiment, the user can return to the call interface by clicking on a specific control in the floating interactive window, or by using a specific gesture or voice command, thus realizing the interaction from the floating interactive window to the call interface.
[0059] As can be seen, in this embodiment of the application, during the process of handling business by telephone, a call parsing control can be provided as an explicit control for different interactive scenarios such as displaying the call interface or exiting the call interface. This ensures that the intelligent voice processing function can be quickly discovered and invoked. By operating the call parsing control, users can trigger the real-time parsing and display of the call voice content, which can reduce the user's learning cost.
[0060] In step 102, in response to the first input, the voice content of both parties in the call is converted into text data.
[0061] In this embodiment, both the user's voice and the customer service voice can be collected simultaneously. Sound source separation technology is used to distinguish the content from both sides, avoiding audio mixing interference. An end-to-end speech recognition model, such as the Conformer or Transformer model, is employed to support high concurrency and low latency real-time transcription. Furthermore, colloquial filler words such as "uh" and "ah" can be filtered, and misrecognized words like "four" and "ten" can be corrected through contextual verification. The transcribed text data is stored in segments according to timestamps, with each segment associated with the start and end times of the corresponding audio segment.
[0062] In step 103, semantic analysis is performed on the text data to divide it into multiple semantic segments, and key business information is extracted from each semantic segment.
[0063] Considering that text data is usually based on raw text generated from speech-to-text, and that raw text generated from speech-to-text contains redundant information such as interjections and repetitive content, direct use would affect information acquisition efficiency. Therefore, in this embodiment of the application, natural language processing technology is used to perform deep semantic understanding on the text data, dividing it into semantic segments with complete business intent. Each semantic segment focuses on a single business theme such as "account inquiry" or "transfer operation," thereby achieving a precise and structured presentation of the call content.
[0064] In this embodiment of the application, an intent recognition model such as the BERT model or a business knowledge graph can be used to divide semantic segments. The division rules of the intent recognition model may include: the speaker switches and the silence interval is greater than or equal to 500ms, and the semantic topic changes, such as switching from "query balance" to "transfer".
[0065] Since semantic paragraphs often suffer from information overload, this application embodiment extracts key business information from semantic paragraphs to significantly improve information acquisition efficiency.
[0066] In this embodiment, natural language processing technology, business rules, and contextual understanding can be combined to accurately extract key business information from semantic paragraphs through hierarchical parsing, multimodal fusion, and dynamic optimization.
[0067] In this embodiment of the application, key business information may include at least one of the following: business operation instructions, digital information, and contact information.
[0068] In this embodiment of the application, the business operation instruction refers to the action instruction initiated by the user or the system that requires an immediate response or subsequent execution, such as applying for a return, activating a membership, or changing a password.
[0069] For example, when a user says "I want to return an item," the electronic device needs to extract the "return" instruction and associate it with digital information such as the order number and purchase time, as well as contact information such as the shipping address, to complete the return process. This process, by quickly recognizing and executing instructions, can significantly improve user satisfaction.
[0070] In this application embodiment, digital information refers to numerical data and associated units related to business; wherein, numerical data includes but is not limited to: amount, quantity, time, percentage, etc., and associated units include but are not limited to: yuan, day, % etc.
[0071] For example, if a user says "I want to buy 3 items, each costing 100 yuan", the electronic device needs to extract the numerical information "quantity = 3, unit price = 100 yuan", calculate the total price, and associate it with the operation instruction "place order" to complete the transaction.
[0072] In this embodiment of the application, contact information refers to identification information used for communication between business stakeholders such as users, customer service, and suppliers. Contact information includes, but is not limited to, telephone numbers, email addresses, addresses, and user IDs.
[0073] For example, if a user says "My address is No. 123, XX Road, Chaoyang District, Beijing", the electronic device needs to extract the contact information "address" and associate it with the operation command "modify address" to complete the update of the user profile.
[0074] As can be seen, in this embodiment, the three types of information—business operation instructions, digital information, and contact information—cover the core decision-making points, data support elements, and communication and collaboration links in the business scenario, respectively, and together constitute the key elements of the business closed loop, thereby ensuring that it can cover a variety of different telephone business handling scenarios.
[0075] In step 104, semantic paragraphs are displayed in chronological order; the display style of key business information is different from the display style of the remaining information in the semantic paragraphs.
[0076] Considering that chronological order is the main thread of event evolution and implies cause-and-effect relationships (e.g., after a user places an order, the system generates an order, providing a basis for subsequent decisions), this embodiment of the application displays semantic paragraphs in chronological order. This can improve the efficiency and accuracy of information processing, especially in scenarios requiring high accuracy and low latency, such as intelligent customer service and financial risk control, where its value is even more significant.
[0077] In this embodiment, users can choose to view semantic segments independently, without having to passively listen to all the voice, thus optimizing the call experience.
[0078] In this embodiment, differentiated display styles such as text highlighting, region highlighting, icon marking, font enhancement, and animated prompts can be used to display key business information in semantic paragraphs.
[0079] In this embodiment of the application, when displaying semantic paragraphs, key business information and ordinary text are distinguished by style in order to guide users to focus on core information and reduce information overload.
[0080] In some embodiments, step 104 may specifically include the following steps: step 1041.
[0081] In step 1041, when the call interface is displayed, semantic paragraphs are dynamically displayed in a preset area of the call interface.
[0082] In this embodiment of the application, considering that during a call, if the call interface is displayed, the user's visual focus is usually on the call interface, therefore, when the call interface is displayed, semantic segments are dynamically displayed in a preset area of the call interface to realize real-time synchronous mapping of voice content and text information, thereby improving the efficiency of users in obtaining business information.
[0083] For example, in Figure 2 Based on the example shown, such as Figure 4 As shown, semantic segments are displayed in chronological order within the preset area 23 of the call interface 22, namely semantic segment 231, semantic segment 232, semantic segment 233, ... The key business information in each semantic segment is displayed in a style that is different from other text information in the semantic segment. For example, the key business information "Press 1 for credit card business" in semantic segment 231 is displayed in a bold, italic and large font style, and the key business information "Press 2 for manual service" in semantic segment 232 is displayed in a bold, italic and large font style.
[0084] In some embodiments, step 104 may specifically include the following step: step 1042.
[0085] In step 1042, when exiting the call interface, semantic paragraphs are displayed in layers in a floating interactive window, and the history can be viewed by swiping.
[0086] In this embodiment, considering that during a call, if the user exits the call interface, their visual focus is usually on the interface outside the call interface or the desktop, semantic segments are displayed in layers in a floating interactive window when the call interface is exited, so as to realize real-time synchronous mapping of voice content and text information within a limited display space; in addition, it can also support swiping to view history, improving the efficiency of users obtaining business information.
[0087] For example, in Figure 3 Based on the example shown, such as Figure 5As shown, semantic paragraphs are displayed in chronological order within the floating interactive window 32, namely semantic paragraph 324, semantic paragraph 325, semantic paragraph 326, ..., where semantic paragraph 324 is earlier than semantic paragraph 325, and semantic paragraph 325 is earlier than semantic paragraph 326. Semantic paragraph 324 is located at the top layer of the display interface, semantic paragraph 325 is located in the middle layer, and semantic paragraph 326 is located at the bottom layer. Key business information in each semantic paragraph is displayed in a style that distinguishes it from other text information in that semantic paragraph. For example, the key business information "Press 1 for credit card services" in semantic paragraph 324 is displayed in bold, italic, and large font size, and the key business information "Press 2 for customer service" in semantic paragraph 325 is also displayed in bold, italic, and large font size. Users can view semantic paragraph 325 by swiping right on semantic paragraph 324, and view semantic paragraph 326 by swiping right on semantic paragraph 325.
[0088] As can be seen, in this embodiment of the application, a differentiated display strategy of semantic paragraphs can be adopted for different interaction scenarios such as displaying the call interface and exiting the call interface, so as to match the user's attention focus, interaction efficiency and information retention needs, thereby improving the efficiency of users to obtain business information.
[0089] In this embodiment of the application, by converting the voice content of both parties in a call to text in real time, accurately performing semantic segmentation and extracting key business information, and displaying key business information in the semantic segments in a differentiated manner, users can get rid of the dilemma of "forgetting after listening" in traditional telephone services and achieve efficient and autonomous business processing.
[0090] As can be seen from the above embodiments, in this embodiment, during a call, the user's first input is received; in response to the first input, the voice content of both parties in the call is converted into text data; semantic analysis is performed on the text data, the text data is divided into multiple semantic segments, and key business information in each semantic segment is extracted; wherein, the key business information includes at least one of the following: business operation instructions, numerical information, and contact information; the semantic segments are displayed in chronological order; wherein, the display style of the key business information is different from the display style of the remaining information in the semantic segment.
[0091] As can be seen, in this embodiment of the application, by using real-time speech-to-text and semantic analysis, lengthy call content is transformed into structured semantic segments, and key business information in the semantic segments is highlighted in a differentiated manner, enabling users to intuitively identify and respond quickly, reducing misoperations, and helping users quickly locate key business information such as operation instructions, amounts, and contact information. This avoids the inefficiency caused by repeatedly listening to voice messages and the problem of easily missing key information, thereby improving the efficiency of telephone business processing.
[0092] In some embodiments provided in this application, the provided call processing method can also be used in... Figure 1 Based on the embodiment shown, the following step is added to step 104 above: step 105.
[0093] In step 105, when displaying semantic paragraphs in chronological order, dynamic visual tags are added to the key business information in the displayed semantic paragraphs.
[0094] In this embodiment, dynamic visual markers are used to improve the convenience for users to obtain or manipulate information.
[0095] In this embodiment of the application, considering that business operation instructions usually need to be completed by inputting numbers through a virtual keyboard, dynamic visual tags will be used to associate the virtual keyboard when the key business information includes business operation instructions.
[0096] For example, the dynamic visual marker is a "blue highlighted border + virtual keyboard icon". The "blue highlighted border" in the dynamic visual marker prominently reminds the user. After the user clicks the "virtual keyboard icon" in the dynamic visual marker, the virtual keyboard input interface is automatically displayed to facilitate the user's input of numbers. For example, when customer service announces "Please press 1 to confirm the transfer", the electronic device automatically generates a clickable "Press 1" dynamic visual marker. The user clicks this dynamic visual marker to directly jump to the virtual keyboard input interface.
[0097] In this embodiment of the application, dynamic visual markers are added to the business operation instructions in the semantic paragraphs to assist users in handling telephone business, thereby reducing business operation steps and the error rate, and thus improving the efficiency of telephone business handling.
[0098] In this embodiment of the application, considering that in telephone business scenarios, the accurate input of digital information such as amount, verification code, account, etc. directly affects business security, when the key business information includes digital information, dynamic visual markers are used to verify the legality of the input. If there is an error, submission is prohibited and correction suggestions are displayed next to the dynamic visual markers.
[0099] In this embodiment of the application, input validity verification includes, but is not limited to: format verification such as length and character type, business verification such as balance and limit, and risk control verification such as remote operation warning.
[0100] In this embodiment of the application, active protection of digital input is achieved by adding dynamic visual markers to the digital information in semantic segments, thereby improving the security of telephone business processing.
[0101] In this embodiment of the application, considering that users have a need to efficiently obtain contact information such as phone numbers, addresses, and email addresses in telephone service scenarios, dynamic visual tags are used to trigger the copying of contact information when key business information includes contact information.
[0102] In this embodiment, dynamic visual markers include, but are not limited to: floating copy icons, dynamic color coding such as blue for phone numbers, purple for addresses, and orange for email addresses, and pressure-sensitive feedback such as pressure-level response supported by 3D Touch devices.
[0103] In this embodiment of the application, dynamic visual markers are added to the contact information in semantic paragraphs to help users quickly copy contact information, thereby improving the efficiency of telephone service processing.
[0104] Considering that the rapid identification and response of key business information such as operation instructions, numbers, and contact information directly impacts user experience during telephone service transactions, traditional solutions rely solely on voice broadcasting or simple text transcription, requiring users to manually filter information, which is inefficient and prone to errors. Therefore, this embodiment utilizes dynamic visual tagging technology to achieve intelligent enhanced display and interactive optimization of key business information, thereby improving the efficiency of telephone service transactions.
[0105] In some embodiments provided in this application, the provided call processing method, Figure 1 Based on the embodiment shown, after step 104 above, the following step can be added: step 106.
[0106] In step 106, a second input from the user for at least one semantic segment is received, and in response to the second input, the original speech segment corresponding to the semantic segment selected by the second input is played.
[0107] In this embodiment of the application, the second input includes, but is not limited to, click operation, long press operation, swipe operation, etc.
[0108] In this embodiment, on the one hand, the corresponding voice segment can be played back by simply manipulating the text paragraph, realizing the backtracking from text to voice, which solves the problem that information is easily forgotten and difficult to backtrack in traditional calls; on the other hand, it supports multimodal interaction of text and voice, which can adapt to the usage habits of different users and improve the user's call experience.
[0109] In some embodiments provided in this application, the provided call processing method, Figure 1 Based on the embodiment shown, after step 104 above, the following step can be added: step 107.
[0110] In step 107, if the voice pause of the other end of the call exceeds a preset duration threshold or a preset keyword is detected when exiting the call interface, the display area of the floating interactive window is expanded, and a prompt message is displayed in the expanded floating interactive window; wherein, the prompt message includes the duration of the voice pause, keywords, or operation suggestions.
[0111] In this embodiment of the application, when the user exits the call interface during a call, the electronic device will continuously monitor the status of the other end of the call. If it detects that the pause in the other end's voice exceeds a preset duration threshold, or if it identifies preset keywords, it will automatically expand the display area of the floating interactive window, that is, increase the size of the floating interactive window, and display relevant prompt information in the expanded floating interactive window to help the user better understand the call situation and perform corresponding operations.
[0112] In this embodiment of the application, the preset duration threshold can be flexibly set according to the actual application scenario and user needs, such as 3 seconds, 5 seconds, etc.
[0113] In this embodiment of the application, the preset keywords include, but are not limited to: business processing terms, status query terms, and terms waiting for response.
[0114] For example, terms related to business processing include: "Please enter," and "Please confirm." Terms related to status inquiry include: "Query results," and "Current balance." Terms related to waiting for a response include: "Continue?" and "Please reply."
[0115] In this embodiment, the floating interactive window can expand in a horizontal, vertical, or both directions simultaneously. The expansion size of the floating interactive window can be reasonably set according to the screen size and the requirements of the displayed content.
[0116] In this embodiment, the electronic device displays corresponding prompts within the expanded floating interactive window. These prompts may include the following: 1) If the trigger condition is a voice pause exceeding a preset duration threshold, the actual voice pause duration of the other end of the call is displayed in the prompt, precisely in seconds, allowing the user to clearly understand the duration of the pause. 2) If the trigger condition is the recognition of a preset keyword, the recognized keyword is highlighted in the prompt, facilitating the user's quick capture of important information during the call. 3) Based on different trigger conditions and call scenarios, corresponding operation suggestions are provided to the user. For example, when the keyword "meeting" is recognized, the operation suggestion could be "Remind you to prepare meeting-related materials"; when the voice pause is long, the operation suggestion could be "Suggest asking the other end whether to continue the call," etc. The operation suggestions are presented in concise and clear language for easy understanding and execution by the user.
[0117] For example, in Figure 3 Based on the example shown, when a user exits the call interface and performs other operations, if the backend analysis detects that the other end's audio is paused or recognizes the keyword "Are you still there?", such as... Figure 6 As shown, the display area of the expanded floating interactive window 32 is extended, and relevant prompt information 327 is displayed in the expanded floating interactive window 32 to remind the user to handle the call content.
[0118] As can be seen, in this embodiment, when the user exits the call interface, by monitoring the status of the other end in real time and displaying prompts promptly, the user can quickly understand the important details of the call without having to re-enter the call interface. This avoids the cumbersome operation and visual interference caused by frequent interface switching, improving the efficiency and convenience of information acquisition. Furthermore, it enables the user to react quickly and proactively communicate with the other end, preventing communication interruptions or misunderstandings caused by untimely information transmission, thereby improving call fluency and efficiency. In addition, in specific scenarios such as business and customer service, preset keywords and corresponding operation suggestions can help users quickly grasp the key points of the call, operate according to business processes, reduce unnecessary communication steps, improve the speed and accuracy of business processing, and thus enhance the overall operational efficiency of the business system.
[0119] In some embodiments provided in this application, the information manipulation functions during call processing can be further optimized, enabling users to more conveniently process key information in call-related text, thereby improving call processing efficiency and user experience. Accordingly, the provided call processing method, in Figure 1 Based on the embodiment shown, after step 104 above, the following step can be added: step 108.
[0120] In step 108, a third input from the user for at least one semantic paragraph is received, and in response to the third input, a copy operation is performed on the contact information or numerical information in the semantic paragraph selected by the third input.
[0121] In this embodiment of the application, the third input includes, but is not limited to: click operation, long press operation, swipe operation, etc.
[0122] In this embodiment, after copying the contact information or numerical information in the semantic paragraph selected by the third input, this information can be copied to the clipboard. The clipboard is an area for temporarily storing copied or cut information, which the user can then paste into the desired location in subsequent operations. Furthermore, after the copying operation is complete, a corresponding feedback prompt can be provided to inform the user that the copying operation has been successfully completed. This feedback prompt can be a pop-up message displaying "Contact information / numerical information copied," or an audio prompt to inform the user.
[0123] As can be seen, in this embodiment, considering that users often need to extract key contact or numerical information from a large amount of text during call processing, through the above step 108, the user only needs to select the semantic paragraph with a simple input operation, and the electronic device can automatically identify and copy the key information, saving the user's time in manually searching and copying information and improving the efficiency of information acquisition. Furthermore, compared to the traditional information copying method, which requires the user to accurately select the required information before copying, a process that is relatively cumbersome, in this embodiment, the user only needs to select the semantic paragraph as a whole, and the electronic device will automatically identify and copy the key information, reducing the user's operation steps and making information processing more convenient and efficient.
[0124] The call processing method provided in this application can be executed by a call processing device. This application uses the example of a call processing device executing the call processing method to illustrate the call processing device provided in this application.
[0125] Figure 7 This is a structural block diagram of a call processing device provided in an embodiment of this application, such as... Figure 7 As shown, the call processing device 700 may include: a receiving module 701, a processing module 702, and a display module 703;
[0126] The receiving module 701 is used to receive the user's first input during a call;
[0127] The processing module 702 is configured to, in response to the first input, convert the voice content of both parties in the call into text data; perform semantic analysis on the text data, segment the text data into multiple semantic segments, and extract key business information from each semantic segment; wherein, the key business information includes at least one of the following: business operation instructions, numerical information, and contact information;
[0128] The display module 703 is used to display the semantic paragraphs in chronological order; wherein the display style of the key business information is different from the display style of the remaining information in the semantic paragraphs.
[0129] As can be seen from the above embodiments, in this embodiment, during a call, the user's first input is received; in response to the first input, the voice content of both parties in the call is converted into text data; semantic analysis is performed on the text data, the text data is divided into multiple semantic segments, and key business information in each semantic segment is extracted; wherein, the key business information includes at least one of the following: business operation instructions, numerical information, and contact information; the semantic segments are displayed in chronological order; wherein, the display style of the key business information is different from the display style of the remaining information in the semantic segment.
[0130] As can be seen, in this embodiment of the application, by using real-time speech-to-text and semantic analysis, lengthy call content is transformed into structured semantic segments, and key business information in the semantic segments is highlighted in a differentiated manner, enabling users to intuitively identify and respond quickly, reducing misoperations, and helping users quickly locate key business information such as operation instructions, amounts, and contact information. This avoids the inefficiency caused by repeatedly listening to voice messages and the problem of easily missing key information, thereby improving the efficiency of telephone business processing.
[0131] Optionally, as an embodiment, the call processing device 700 may further include:
[0132] The processing module is further configured to add dynamic visual markers to key business information in the semantic paragraph; wherein, when the key business information includes business operation instructions, the dynamic visual markers are used to associate with a virtual keyboard; when the key business information includes numerical information, the dynamic visual markers are used to verify the legality of the input, prohibiting submission when there is an error and displaying correction suggestions next to the dynamic visual markers; when the key business information includes contact information, the dynamic visual markers are used to trigger the copying of the contact information.
[0133] Optionally, as an embodiment, the receiving module 701 is specifically used to receive a first input to the call parsing control in the call interface when the display module displays the call interface; or, when exiting the call interface, to receive a first input to the call parsing control in the floating interactive window; wherein the call parsing control is used to trigger real-time parsing and display of the call voice content.
[0134] Optionally, as an embodiment, the display module 703 is specifically used to dynamically display the semantic paragraph in a preset area of the call interface when the call interface is displayed; or, when exiting the call interface, to display the semantic paragraph in layers in a floating interactive window and support swiping to view the history.
[0135] Optionally, as an embodiment, the receiving module 701 is further configured to receive a second input from the user for at least one semantic paragraph;
[0136] The processing module 702 is further configured to, in response to the second input, play the original speech segment corresponding to the semantic segment selected by the second input.
[0137] Optionally, as an embodiment, the display module 703 is further configured to, when exiting the call interface, if the processing module detects that the voice pause of the other end of the call exceeds a preset duration threshold, or identifies a preset keyword, expand the display area of the floating interactive window and display prompt information in the expanded floating interactive window; wherein, the prompt information includes the duration of the voice pause, keywords, or operation suggestions.
[0138] The call processing device in this application embodiment can be an electronic device or a component within an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or a device other than a terminal. For example, the electronic device can be a mobile phone, tablet computer, laptop computer, PDA, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc. It can also be a server, network attached storage (NAS), personal computer (PC), television (TV), ATM, or self-service machine, etc. This application embodiment does not specifically limit the scope of the device.
[0139] The call processing device in this application embodiment can be a device with an operating system. This operating system can be Android, iOS, or any other possible operating system; this application embodiment does not impose specific limitations.
[0140] The call processing apparatus provided in this application embodiment can implement the various processes implemented in the above method embodiments, and will not be described again here to avoid repetition.
[0141] Optionally, such as Figure 8 As shown, this application embodiment also provides an electronic device 800, including a processor 801 and a memory 802. The memory 802 stores a program or instructions that can run on the processor 801. When the program or instructions are executed by the processor 801, they implement the various steps of the above-described call processing method embodiment and can achieve the same technical effect. To avoid repetition, they will not be described again here.
[0142] It should be noted that the electronic devices in the embodiments of this application include the aforementioned mobile electronic devices and non-mobile electronic devices.
[0143] Figure 9 This is a schematic diagram of the hardware structure of an electronic device that implements the various embodiments of this application.
[0144] The electronic device 900 includes, but is not limited to, components such as: radio frequency unit 901, network module 902, audio output unit 903, input unit 904, sensor 905, display unit 906, user input unit 907, interface unit 908, memory 909, and processor 910.
[0145] Those skilled in the art will understand that the electronic device 900 may also include a power supply (such as a battery) for supplying power to various components. The power supply may be logically connected to the processor 910 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. Figure 9 The electronic device structure shown does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.
[0146] The user input unit 907 is used to receive the user's first input during the call;
[0147] The processor 910 is configured to, in response to the first input, convert the voice content of both parties in the call into text data; perform semantic analysis on the text data, segment the text data into multiple semantic segments, and extract key business information from each semantic segment; wherein the key business information includes at least one of the following: business operation instructions, numerical information, and contact information;
[0148] Display unit 906 is used to display the semantic paragraphs in chronological order; wherein the display style of the key business information is different from the display style of the remaining information in the semantic paragraphs.
[0149] As can be seen, in this embodiment of the application, by using real-time speech-to-text and semantic analysis, lengthy call content is transformed into structured semantic segments, and key business information in the semantic segments is highlighted in a differentiated manner, enabling users to intuitively identify and respond quickly, reducing misoperations, and helping users quickly locate key business information such as operation instructions, amounts, and contact information. This avoids the inefficiency caused by repeatedly listening to voice messages and the problem of easily missing key information, thereby improving the efficiency of telephone business processing.
[0150] Optionally, as an embodiment, the processor 910 is further configured to add dynamic visual markers to key business information in the semantic paragraph; wherein, when the key business information includes business operation instructions, the dynamic visual markers are used to associate with a virtual keyboard; when the key business information includes numerical information, the dynamic visual markers are used to verify the legality of the input, prohibiting submission when there is an error and displaying correction suggestions next to the dynamic visual markers; when the key business information includes contact information, the dynamic visual markers are used to trigger the copying of the contact information.
[0151] Optionally, as an embodiment, the user input unit 907 is specifically used to receive a first input to the call parsing control in the call interface when the call interface is displayed; or, when exiting the call interface, to receive a first input to the call parsing control in the floating interactive window; wherein the call parsing control is used to trigger real-time parsing and display of the call voice content.
[0152] Optionally, as an embodiment, the display unit 906 is specifically used to dynamically display the semantic paragraph in a preset area of the call interface when the call interface is displayed; or, when exiting the call interface, to display the semantic paragraph in layers in a floating interactive window and support swiping to view the history.
[0153] Optionally, as an embodiment, the user input unit 907 is also configured to receive a second input from the user for at least one semantic paragraph;
[0154] The processor 910 is also configured to, in response to the second input, play the original speech segment corresponding to the semantic segment selected by the second input.
[0155] Optionally, as an embodiment, the display unit 906 is further configured to, when exiting the call interface, if the processor 910 detects that the voice pause of the other end of the call exceeds a preset duration threshold, or identifies a preset keyword, expand the display area of the floating interactive window and display prompt information in the expanded floating interactive window; wherein, the prompt information includes the duration of the voice pause, keywords, or operation suggestions.
[0156] It should be understood that, in this embodiment, the input unit 904 may include a graphics processing unit (GPU) 9041 and a microphone 9042. The GPU 9041 processes image data of still images or videos obtained by an image capture device (such as a camera) in video capture mode or image capture mode. The display unit 906 may include a display panel 9061, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, or the like. The user input unit 907 includes at least one of a touch panel 9071 and an input device 9072. The touch panel 9071 is also called a touch screen. The touch panel 9071 may include a touch detection device and a touch controller. The input device 9072 may include, but is not limited to, a physical keyboard, function keys (such as volume control buttons, power buttons, etc.), a trackball, a mouse, and a joystick, which will not be described in detail here.
[0157] The memory 909 can be used to store software programs and various data. The memory 909 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, the memory 909 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous linked dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM). The memory 909 in the embodiments of this application includes, but is not limited to, these and any other suitable types of memory.
[0158] Processor 910 may include one or more processing units; optionally, processor 910 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into processor 910.
[0159] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described call processing method embodiments and achieve the same technical effect. To avoid repetition, they will not be described again here.
[0160] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.
[0161] This application also provides a chip, which includes a processor and a communication interface. The communication interface and the processor are coupled. The processor is used to run programs or instructions to implement the various processes of the above-described call processing method embodiments and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0162] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.
[0163] This application also provides a computer program product, which is stored in a storage medium and executed by at least one processor to implement the various processes of the call processing method embodiments described above, and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0164] It should be noted that, in this document, the terms "comprising," "including," or any variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in the examples.
[0165] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (e.g., ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (e.g., a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of this application.
[0166] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.
Claims
1. A call processing method, characterized in that, The method includes: During the call, receive the user's initial input; In response to the first input, the voice content of both parties in the call is converted into text data; Semantic analysis is performed on the text data to segment it into multiple semantic segments, and key business information is extracted from each semantic segment; wherein, the key business information includes at least one of the following: business operation instructions, numerical information, and contact information; The semantic paragraphs are displayed in chronological order; the display style of the key business information is different from the display style of the remaining information in the semantic paragraphs.
2. The method according to claim 1, characterized in that, The method further includes: Add dynamic visual markers to key business information in the semantic paragraphs; Wherein, when the key business information includes business operation instructions, the dynamic visual marker is used to associate with the virtual keyboard; When the key business information includes numerical information, the dynamic visual marker is used to verify the legality of the input. If there is an error, submission is prohibited and a correction suggestion is displayed next to the dynamic visual marker. When the key business information includes contact information, the dynamic visual tag is used to trigger the copying of the contact information.
3. The method according to claim 1, characterized in that, The first input received from the user includes: When the call interface is displayed, receive the first input to the call parsing control in the call interface; or... When exiting the call interface, receive the first input to the call parsing control in the floating interactive window; The call parsing control is used to trigger the real-time parsing and display of the call voice content.
4. The method according to claim 1, characterized in that, The method of displaying the semantic paragraphs in chronological order includes: When the call interface is displayed, the semantic paragraph is dynamically displayed in a preset area of the call interface; or, When exiting the call interface, the semantic paragraphs are displayed in layers in a floating interactive window, and the history can be viewed by swiping.
5. The method according to claim 1, characterized in that, After the step of displaying the semantic paragraphs in chronological order, the method further includes: Receive a second input from the user for at least one semantic paragraph; In response to the second input, the original speech segment corresponding to the semantic segment selected by the second input is played.
6. The method according to claim 1, characterized in that, The method further includes: If, upon exiting the call interface, a pause in the other end's voice is detected to exceed a preset duration threshold, or a preset keyword is identified, the display area of the floating interactive window is expanded, and a prompt message is displayed within the expanded floating interactive window. The prompts include the duration of the voice pause, keywords, or operation suggestions.
7. A call processing device, characterized in that, The device includes: The receiving module is used to receive the user's initial input during a call; The processing module is configured to, in response to the first input, convert the voice content of both parties in the call into text data; perform semantic analysis on the text data, segment the text data into multiple semantic segments, and extract key business information from each semantic segment; wherein, the key business information includes at least one of the following: business operation instructions, numerical information, and contact information; The display module is used to display the semantic paragraphs in chronological order; wherein the display style of the key business information is different from the display style of the remaining information in the semantic paragraphs.
8. The apparatus according to claim 7, characterized in that, The device further includes: The processing module is further configured to add dynamic visual markers to key business information in the semantic paragraph; wherein, when the key business information includes business operation instructions, the dynamic visual markers are used to associate with a virtual keyboard; when the key business information includes numerical information, the dynamic visual markers are used to verify the legality of the input, prohibiting submission when there is an error and displaying correction suggestions next to the dynamic visual markers; when the key business information includes contact information, the dynamic visual markers are used to trigger the copying of the contact information.
9. The apparatus according to claim 7, characterized in that, The receiving module is specifically used to receive a first input to the call parsing control in the call interface when the display module displays the call interface; or, when exiting the call interface, to receive a first input to the call parsing control in the floating interactive window; wherein the call parsing control is used to trigger real-time parsing and display of the call voice content.
10. The apparatus according to claim 7, characterized in that, The display module is specifically used to dynamically display the semantic paragraph in a preset area of the call interface when the call interface is displayed; or, when exiting the call interface, to display the semantic paragraph in layers in a floating interactive window and support swiping to view the history.
11. The apparatus according to claim 7, characterized in that, The receiving module is also configured to receive a second input from the user for at least one semantic paragraph; The processing module is also configured to, in response to the second input, play the original speech segment corresponding to the semantic segment selected by the second input.
12. The apparatus according to claim 7, characterized in that, The display module is further configured to, when exiting the call interface, if the processing module detects that the voice pause of the other end of the call exceeds a preset duration threshold, or identifies a preset keyword, expand the display area of the floating interactive window and display prompt information in the expanded floating interactive window; wherein, the prompt information includes the duration of the voice pause, keywords, or operation suggestions.