Session group processing method, session management device and session management system
By collecting and processing session data on wearable devices, constructing target session groups, and adding intelligent assistants, the problem of information management is solved, and intelligent management of session content and improvement of user experience are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN ARCENSION TECHNOLOGY CO LTD
- Filing Date
- 2026-01-13
- Publication Date
- 2026-04-14
AI Technical Summary
In existing technologies, how to manage recorded information in a unified manner has become an urgent problem to be solved, especially how to effectively manage the recording and organization of important information such as meeting highlights and sudden inspirations on mobile terminals and smart devices.
This paper provides a method for processing conversation groups. It collects conversation data through wearable devices, performs text conversion and speaker classification, constructs target conversation groups, and displays detailed information on terminal devices. It also supports adding smart assistants and binding contacts to achieve intelligent management of conversation groups.
It enhances the intelligence of conversation group management and improves user experience. Users can add intelligent assistants at any time for analysis and organization, making it easier to understand and communicate conversation content.
Smart Images

Figure CN121864512A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of session management technology, and in particular to methods, devices and systems for processing session groups. Background Technology
[0002] With the development of mobile terminals and smart devices, users often need to record important information in a timely manner in daily life, work communication and unexpected scenarios, such as meeting highlights, sudden inspirations, verbal promises or temporary communication content.
[0003] How to manage the recorded information in a unified manner has become an urgent problem to be solved. Summary of the Invention
[0004] This application provides a method for processing conversation groups, a conversation management device, and a system. Users can add intelligent assistants to different conversation groups at any time while viewing conversations. The intelligent assistants can be used to assist users in completing corresponding analysis and organization tasks, thereby improving the intelligence of conversation group management and enhancing the user experience.
[0005] To address the aforementioned technical problems, this application provides a method for processing a conversation group, comprising: displaying detailed information of a first target conversation group on a display interface; the first target conversation group is obtained by converting conversation data to be converted collected by a wearable device; the conversation data to be converted includes audio data during the conversation; in response to a speaker addition operation, adding a first smart assistant as a speaker to the first target conversation group; and the first smart assistant interacting with the other speakers in the first target conversation group.
[0006] To solve the above-mentioned technical problems, another technical solution adopted in this application is to provide a session management device, which includes a memory and a processor coupled to each other, wherein the memory stores program instructions and the processor executes the program instructions to implement the processing method provided by the above-mentioned technical solution.
[0007] To solve the above-mentioned technical problems, another technical solution adopted in this application is to provide a session management system, which includes a wearable device and a session management device as provided in the above technical solution.
[0008] To solve the above-mentioned technical problems, another technical solution adopted in this application is to provide a computer-readable storage medium storing program instructions that can be executed by a processor, the program instructions being used to implement the processing method provided by the above-mentioned technical solution.
[0009] The method, device, and system for processing conversation groups provided in this application display detailed information of a first target conversation group on a display interface. The first target conversation group is obtained by converting conversation data to be converted collected by a wearable device. The conversation data to be converted includes audio data during the conversation. In response to a speaker addition operation, a first intelligent assistant is added as a speaker to the first target conversation group. The first intelligent assistant is used to interact with other speakers in the first target conversation group. In this way, users can add the intelligent assistant to different conversation groups at any time while viewing the conversation, and use the intelligent assistant to assist users in completing corresponding analysis and organization tasks, thereby improving the intelligence of conversation group management and enhancing the user experience. Attached Figure Description
[0010] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Wherein: Figure 1 This is a flowchart illustrating an embodiment of the session management method provided in this application; Figure 2 This is a flowchart illustrating another embodiment of the session management method provided in this application; Figure 3 This is a flowchart illustrating another embodiment of the session management method provided in this application; Figure 4 This is a flowchart illustrating another embodiment of the session management method provided in this application; Figure 5 This is a flowchart illustrating another embodiment of the session management method provided in this application; Figure 6 This is a flowchart illustrating another embodiment of the session management method provided in this application; Figure 7 This is a schematic diagram of the emotion / stress display interface provided in this application; Figure 8 This is a flowchart illustrating an embodiment of the session group processing method provided in this application; Figure 9 This is a flowchart illustrating an embodiment of the interaction method for a conversation group provided in this application; Figure 10 This is a flowchart illustrating an embodiment of the session group processing method provided in this application; Figure 11 This is a flowchart illustrating another embodiment of the session group processing method provided in this application; Figure 12This is a flowchart illustrating another embodiment of the session group processing method provided in this application; Figure 13 This is a flowchart illustrating another embodiment of the session group processing method provided in this application; Figure 14 This is a flowchart illustrating another embodiment of the session management method provided in this application; Figure 15 This is a flowchart illustrating another embodiment of the session management method provided in this application; Figure 16 This is a flowchart illustrating another embodiment of the session management method provided in this application; Figure 17 This is a flowchart illustrating another embodiment of the session management method provided in this application; Figure 18 This is a schematic diagram of the session group 1 display interface provided in this application; Figure 19 This is a schematic diagram of another display interface of the conversation group provided in this application; Figure 20 This is a schematic diagram of another display interface of the conversation group provided in this application; Figure 21 This is a schematic diagram of another display interface of the conversation group provided in this application; Figure 22 This is a schematic diagram of another display interface of the conversation group provided in this application; Figure 23 This is a schematic diagram of another display interface of the conversation group provided in this application; Figure 24 This is a schematic diagram of another display interface of the conversation group provided in this application; Figure 25 This is a schematic diagram of the structure of an embodiment of the session management device provided in this application; Figure 26 This is a schematic diagram of the structure of an embodiment of the session management system provided in this application; Figure 27 This is a schematic diagram of an embodiment of the computer-readable storage medium provided in this application. Detailed Implementation
[0011] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. It is understood that the specific embodiments described herein are only for explaining this application and not for limiting it. Furthermore, it should be noted that, for ease of description, only the parts related to this application are shown in the accompanying drawings, not all structures. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.
[0012] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0013] See Figure 1 , Figure 1 This is a flowchart illustrating an embodiment of the session management method provided in this application. The session management method includes: Step 11: Obtain the session data to be converted collected by the wearable device.
[0014] The session data to be converted includes audio data during the session and the timestamp corresponding to when the audio data was generated.
[0015] In some embodiments, the wearable device may be a smart bracelet, smartwatch, smart ring, smart button, or other device that can be conveniently worn on the user's body. The wearable device has recording and storage functions. It can record audio data during a conversation and store it in its internal storage module. The wearable device can choose an appropriate time to send the recorded audio data to a cloud server and / or a terminal device. In some embodiments, the audio data may include at least one of the following: human voice, background noise, music, machine sounds, wind sounds, and silent clips.
[0016] Wearable devices record a timestamp corresponding to the moment the audio data was generated during the recording process. This timestamp can be recorded in a year-month-day format.
[0017] In this embodiment, the executing entity can be a cloud server or a terminal device. The terminal device can be a smartphone, tablet computer, computer, or other device with an operating system and logical processing capabilities.
[0018] In some embodiments, the conversation can be a conversation between several participants in a meeting setting. In some embodiments, the meeting setting can be an offline meeting setting, where the audio data of the entire conversation is generated by several participants in an offline meeting. For example, during an offline meeting, participants can use wearable devices to record the audio data of the conversation.
[0019] In some embodiments, the session process can be a multi-party dialogue session in a non-conference scenario.
[0020] In some embodiments, the entire session can be conducted in an offline session scenario or an online session scenario.
[0021] Step 12: Perform text conversion and speaker classification on the audio data to obtain the text information and audio segment of each speaker under the corresponding timestamp, and construct the target conversation group based on the text information and audio segment.
[0022] After receiving / acquiring the session data to be converted collected by the wearable device, the audio data in the session data to be converted is converted into text and the speaker is classified.
[0023] In some embodiments, since the audio data consists of conversations between different people, it is necessary to categorize the different voices in the audio data according to the speaker to obtain the audio data corresponding to each speaker. Then, text conversion is performed on the audio data corresponding to each speaker to obtain the corresponding text information. Furthermore, since the audio data consists of conversations between different people, the audio data corresponding to each speaker is a series of audio segments ordered by timestamps.
[0024] In some embodiments, the audio data can first be converted to text to obtain corresponding text information and audio segments. Then, different voices in the audio data can be categorized by speaker to obtain audio data corresponding to each speaker. Finally, based on the correspondence between text information and audio data, several pieces of text information are assigned to the corresponding speakers.
[0025] In some embodiments, the audio data can be converted to text to obtain text information and audio segments at each timestamp. Speaker identification can also be performed on the audio data to identify several speakers corresponding to the audio data. The text information and audio segments are then bound to their corresponding speaker identifiers to construct a target session group based on the text information and audio segments.
[0026] In some embodiments, after obtaining text information and audio segments, a target session group is constructed based on the text information and audio segments. For example, in the form of dialogues with different speaker identifiers, the text information, the timestamp corresponding to each text information, and the audio segment icon are presented in chronological order by timestamp in the target session group, along with the speaker identifier. The speaker identifier is used to distinguish different speakers.
[0027] Step 13: In response to the selection of a target session group, display the detailed information of the target session group.
[0028] The detailed information includes: text messages of different speakers displayed in dialogue format in time-stamp order, as well as the timestamp and audio segment icon corresponding to each text message.
[0029] In this embodiment, wearable devices are used to collect conversation data to be converted. This conversation data includes audio data during the conversation and the timestamp corresponding to the time the audio data was generated. The audio data is converted into text and categorized by speaker to obtain the text information and audio segment of each speaker under the corresponding timestamp. A target conversation group is constructed based on the text information and audio segment, so that users can quickly view the constructed conversation group on their terminal devices. In response to the selection of a target conversation group, detailed information of the target conversation group is displayed. The detailed information includes: text information of different speakers displayed in dialogue form according to timestamp order, as well as the timestamp and audio segment icon corresponding to each text information. This allows users to review previous conversation content on their terminal devices and communicate within the conversation group through conversations. The timestamp corresponding to each text information is displayed in the target conversation group, presenting a complete record of the actual conversation process, such as the date and time, thereby realizing online management of conversation data.
[0030] See Figure 2 , Figure 2 This is a flowchart illustrating another embodiment of the session management method provided in this application. The session management method includes: Step 21: Obtain the session data to be converted collected by the wearable device.
[0031] The session data to be converted includes audio data during the session and the timestamp corresponding to when the audio data was generated.
[0032] Step 22: Perform text conversion and speaker classification on the audio data to obtain the text information and audio segment of each speaker under the corresponding timestamp, and construct the target conversation group based on the text information and audio segment.
[0033] In some embodiments, steps 21 to 22 have the same or similar technical solutions as the other embodiments of this application, and will not be described in detail here.
[0034] Step 23: In response to the selection of a target session group, display the detailed information of the target session group.
[0035] The detailed information includes: text information of different speakers displayed in dialogue format according to timestamp order, as well as the timestamp corresponding to each text information, audio segment icons, and virtual icons corresponding to different speakers.
[0036] In some embodiments, the virtual identifier is only used to distinguish different speakers and cannot represent specific speakers. For example, in an actual conversation, speaker A and speaker B are having a dialogue. Based on this, when constructing a target conversation group, different virtual identifiers can be set for speaker A and speaker B. However, this virtual identifier cannot explicitly refer to the specific information of speaker A and speaker B. From the perspective of the target conversation group, it can only know that the two speakers have conducted a conversation based on the text information displayed in the target conversation group, but it does not know the actual identities of speaker A and speaker B, such as their names, positions, and other identity information.
[0037] Based on this, this embodiment provides a contact binding function, which allows users to bind contacts to these virtual identifiers when viewing target conversation groups. This makes it easier for users to clearly know the identity of the speakers in the target conversation group and to better understand the content and purpose of the conversation.
[0038] Step 24: In response to the selection operation of the target virtual identifier, display the contact list; the contact list is used to display several existing contact identifiers.
[0039] In some embodiments, the user can select a target virtual identifier, such as by touch or by using a mouse. After the target virtual identifier is selected, a contact list is displayed on the screen. The contact list displays several existing contact identifiers. The contacts in this contact list are obtained by the user adding and saving contacts in advance. In some embodiments, this contact list can be obtained by accessing the device's address book.
[0040] In some embodiments, the contact list can be obtained by retrieving the address book in the conversation group.
[0041] Step 25: In response to the selection of the first target contact identifier, replace the target virtual identifier with the first target contact identifier and establish a contact binding relationship.
[0042] In some embodiments, after the contact list is displayed, existing contact identifiers displayed in the contact list are available for selection by the user. Based on this, in response to the selection of a first target contact identifier, the target virtual identifier is replaced with the first target contact identifier, and a contact binding relationship is established.
[0043] For example, a target conversation group may contain virtual identifiers A and B. The user knows that virtual identifier B corresponds to themselves, and can then bind a contact to virtual identifier A. For instance, if the user selects virtual identifier A, a contact list will be displayed on the corresponding interface. The user can then select a first target contact identifier (e.g., "Xiao Mouhong") from the contact list. This first target contact identifier can then replace the target virtual identifier, establishing a contact binding relationship. The first target contact identifier will then be displayed in the target conversation group.
[0044] In other embodiments, in response to the absence of a first target contact identifier in the contact list, a first target contact identifier is added. After adding the first target contact identifier, the first target contact identifier replaces the target virtual identifier, and a contact binding relationship is established. For example, the target session group has virtual identifiers A and B. The user knows that virtual identifier B corresponds to themselves, so they can bind virtual identifier A to a contact. For example, if the user selects virtual identifier A, a contact list is displayed on the corresponding interface. If the user finds that there is no first target contact identifier corresponding to virtual identifier A in the contact list, they can add the first target contact identifier to the contact list. After adding the first target contact identifier, the user can select the first target contact identifier, replace the target virtual identifier, and establish a contact binding relationship. At this time, the first target contact identifier is displayed in the target session group.
[0045] In this embodiment, in addition to creating a target conversation group for interaction, a contact binding function is also provided, so that users can bind the speaker's identifier in the target conversation group to an existing contact. This makes it easier for users to understand the content in the target conversation group by referring to the existing contact, and also facilitates quick communication with the existing contact in the future.
[0046] See Figure 3 , Figure 3 This is a flowchart illustrating another embodiment of the session management method provided in this application. The session management method includes: Step 31: Obtain the session data to be converted collected by the wearable device.
[0047] The session data to be converted includes audio data during the session and the timestamp corresponding to when the audio data was generated.
[0048] Step 32: Perform text conversion and speaker classification on the audio data to obtain the text information and audio segment of each speaker under the corresponding timestamp, and construct the target conversation group based on the text information and audio segment.
[0049] Step 33: In response to the selection of a target session group, display the detailed information of the target session group.
[0050] The detailed information includes: text messages of different speakers displayed in dialogue format in time-stamp order, as well as the timestamp and audio segment icon corresponding to each text message.
[0051] In some embodiments, steps 31 to 33 have the same or similar technical solutions as the other embodiments of this application, and will not be described in detail here.
[0052] Step 34: In response to the selection operation of the added control, display the contact list; the contact list is used to display several existing contact identifiers.
[0053] In some embodiments, since the target session group currently only contains speakers involved in previous sessions, this embodiment provides an add function. Users can choose to invite new speakers to the target session group so that the new speakers can view the information involved in the target session group.
[0054] In some embodiments, an add control can be provided in the display interface of the target conversation group. After the user selects the add control, a contact list is displayed. The user can view the contact list and select the contact they wish to add.
[0055] Step 35: In response to the selection of the second target contact ID, add the second target contact ID to the target session group.
[0056] After the second target contact ID is added to the target conversation group, the user corresponding to the second target contact ID can enter the target conversation group through the terminal device to view it and interact with it in the target conversation group.
[0057] It is understandable that any user who can view the target session group can interact within the target session group.
[0058] In this embodiment, in addition to creating a target conversation group for interaction, a speaker addition function is also provided, so that users can add new contacts to the target conversation group. The new contacts can view the content of the conversation group and communicate in the target conversation group.
[0059] See Figure 4 , Figure 4 This is a flowchart illustrating another embodiment of the session management method provided in this application. The session management method includes: Step 41: Obtain the session data to be converted collected by the wearable device.
[0060] The session data to be converted includes audio data during the session and the timestamp corresponding to when the audio data was generated.
[0061] Step 42: Perform text conversion and speaker classification on the audio data to obtain the text information and audio segment of each speaker under the corresponding timestamp, and construct the target conversation group based on the text information and audio segment.
[0062] Step 43: In response to the selection of a target session group, display the detailed information of the target session group.
[0063] The detailed information includes: text messages of different speakers displayed in dialogue format in time-stamp order, as well as the timestamp and audio segment icon corresponding to each text message.
[0064] In some embodiments, steps 41 to 43 have the same or similar technical solutions as the other embodiments of this application, and will not be described in detail here.
[0065] Step 44: In response to the selection of the added control, display the contact list; the contact list is used to display several existing contact identifiers.
[0066] In some embodiments, because the target session group currently only contains speakers involved in previous sessions, and the speaker identifiers used when constructing the target session group may be virtual identifiers. For example, in an actual session, speaker A and speaker B converse, but the virtual identifier cannot explicitly identify the specific information of speaker A and speaker B. From the perspective of the target session group, it can only know that the two speakers have engaged in a conversation based on the text information displayed in the target session group, without knowing the actual identities of speaker A and speaker B, such as their names, positions, or other identity information. Furthermore, the real speakers corresponding to these speaker identifiers cannot see this target session group.
[0067] Therefore, this embodiment provides an add function. Users can choose to invite speakers to the target conversation group so that speakers can view the information involved in the target conversation group.
[0068] In some embodiments, an add control can be provided in the display interface of the target conversation group. After the user selects the add control, a contact list is displayed. The user can view the contact list and select the contact to be added.
[0069] Step 45: In response to the selection of a third target contact ID, add the third target contact ID to the target session group.
[0070] Among them, the contact corresponding to the third target contact identifier is the actual speaker corresponding to one of the speaker identifiers in the target session group.
[0071] After the third target contact identifier is added to the target conversation group, the user corresponding to the third target contact identifier can enter the target conversation group through the terminal device to view it and interact with it in the target conversation group.
[0072] It is understandable that any user who can view the target session group can interact within the target session group.
[0073] In this embodiment, in addition to constructing a target conversation group for interaction, a speaker addition function is also provided, so that users can add the real contact (third target contact identifier) corresponding to the speaker identifier in the target conversation group. Since the contact corresponding to the third target contact identifier is the real speaker corresponding to one of the speaker identifiers in the target conversation group, it can explain the previous conversation content in more detail during the interaction process and interact with the other speakers more accurately.
[0074] See Figure 5 , Figure 5 This is a flowchart illustrating another embodiment of the session management method provided in this application. The session management method includes: Step 51: Obtain the session data to be converted collected by the wearable device.
[0075] The session data to be converted includes audio data during the session and the timestamp corresponding to when the audio data was generated.
[0076] Step 52: Perform text conversion and speaker classification on the audio data to obtain the text information and audio segment of each speaker under the corresponding timestamp, and construct the target conversation group based on the text information and audio segment.
[0077] Step 53: In response to the selection of a target session group, display the detailed information of the target session group.
[0078] The detailed information includes: text messages of different speakers displayed in dialogue format in time-stamp order, as well as the timestamp and audio segment icon corresponding to each text message.
[0079] In some embodiments, steps 51 to 53 have the same or similar technical solutions as the other embodiments of this application, and will not be described in detail here.
[0080] Step 54: In response to the selection operation of the first icon, play the audio segment corresponding to the first icon.
[0081] In some embodiments, once the first icon is selected, it is used to play its corresponding audio segment individually.
[0082] In some embodiments, after the first icon is selected, a playback window can be displayed, within which the corresponding audio segment is played. This playback window can display information such as the duration of the audio segment, the speaker's mood, and stress level. Furthermore, the playback window can include corresponding audio adjustment controls, such as playback speed control and language conversion control. For example, if the user's current language differs from the language of the audio, they can use the language conversion control to select a language they understand, and then the audio will be played in the selected language.
[0083] In other embodiments, since the text information is obtained by converting audio data, there may be some conversion errors. Therefore, a first icon is provided to allow users to directly play the corresponding audio segment and understand the original information of the text.
[0084] Step 55: In response to the selection operation of the second icon, start playing all subsequent audio segments sequentially from the audio segment corresponding to the second icon.
[0085] In some embodiments, once the second icon is selected, all subsequent audio segments are played sequentially, starting from the audio segment corresponding to the second icon.
[0086] In some embodiments, selecting the second icon displays a playback window where its corresponding audio segments are played sequentially. This playback window can display the duration of each audio segment, the speaker's mood, and stress level. Furthermore, the playback window can include corresponding audio adjustment controls, such as playback speed control, language switching control, skip control, pause control, and close control. For example, if the user's language differs from the audio's language, they can use the language switching control to select a language they understand, and the audio will be played in that language. Similarly, if the user selects the skip control, the currently playing audio segment can be skipped sequentially, and the next segment can be played.
[0087] In other embodiments, since the text information is obtained by converting audio data, there may be some conversion errors. Therefore, a first icon is provided to allow users to directly play the corresponding audio segment and understand the original information of the text.
[0088] Based on this, users can play back audio from previous conversations at any point in the target conversation group with a single click, improving the user experience.
[0089] See Figure 6 , Figure 6 This is a flowchart illustrating another embodiment of the session management method provided in this application. The session management method includes: Step 61: Obtain the session data to be converted collected by the wearable device.
[0090] The session data to be converted includes audio data during the session and the timestamp corresponding to when the audio data was generated.
[0091] Step 62: Perform text conversion and speaker classification on the audio data to obtain the text information and audio segment of each speaker under the corresponding timestamp, and construct the target conversation group based on the text information and audio segment.
[0092] Step 63: In response to the selection of a target session group, display the detailed information of the target session group.
[0093] The detailed information includes: text messages of different speakers displayed in dialogue format in time-stamp order, as well as the timestamp and audio segment icon corresponding to each text message.
[0094] In some embodiments, steps 61 to 63 have the same or similar technical solutions as the other embodiments of this application, and will not be described in detail here.
[0095] Step 64: In response to the selection operation of the forwarding control, display the list to be forwarded; the list to be forwarded is used to display several targets to be forwarded, including existing conversation groups and / or existing contacts.
[0096] In this embodiment, a session group forwarding function is also provided. Users can choose to forward a target session group to other session groups and / or a single contact.
[0097] In some embodiments, a forwarding control can be set in the display interface of the target session group. After the user selects the forwarding control, a list of messages to be forwarded is displayed. The user can view the list of messages to be forwarded and select the target message to be forwarded from the list.
[0098] Step 65: In response to at least one target being selected, forward the information of the target session group to at least one target.
[0099] In some embodiments, a user can select at least one target to be forwarded from a list of targets to be forwarded. After determining the forwarding, the terminal device forwards the information of the target session group to at least one target to be forwarded.
[0100] If the target to be forwarded is a conversation group, then speakers in that conversation group can see the content of the target conversation group.
[0101] If the target to be forwarded is a contact, then that contact can view the content of the target conversation group.
[0102] In this embodiment, in addition to creating a target session group for interaction, a forwarding function is also provided, so that users can forward the target session group to other targets, allowing other targets to view the content of the target session group and improving the user experience.
[0103] In some embodiments, an emotion / stress control is provided on the display interface of the target conversation group. After displaying the detailed information of the target conversation group, the following process may occur: in response to the selection operation of the emotion / stress control, display the emotion information and / or stress information corresponding to at least one speaker identifier during the conversation; wherein, the emotion information and / or stress information is obtained by performing emotion recognition and / or stress recognition on the audio data.
[0104] Furthermore, the way to display the emotional / stress information corresponding to at least one speaker's identifier during the conversation can be to display the emotional distribution and / or stress distribution corresponding to at least one speaker's identifier at different time periods during the conversation on the same screen.
[0105] like Figure 7As shown, the interface contains two corresponding information tracks: the upper track, 920, is the emotion index track. Based on the parsed type identifier data, the system generates different colored blocks on the timeline. For example, orange blocks correspond to the "Thrilled" state, indicating that the user's speech rate significantly increases and heart rate variability decreases during this time period; blue blocks correspond to the "Nervous" state. This color coding allows users to clearly see the emotional trend throughout the meeting. The lower track, 922, is the feature parameter track. It plots physiological data collected by the wearable device (such as the STR index, a stress index that combines heart rate intensity and speech intensity) as a dynamic bar chart. The bar charts corresponding to most time slices are displayed in the first color (e.g., blue bars), indicating that the acoustic or heart rate characteristics at that moment are within the normal fluctuation range. Orange bars visually reflect that at that specific moment in the recording, the rate of change of the processed acoustic feature parameters and / or the fluctuation amplitude of the heart rate feature parameters simultaneously exceed the preset first and second thresholds. This visually differentiated design allows users to quickly locate the moment with the highest "stress" or "excitement" during the entire meeting simply by visual scanning, without needing to play the recording. The bar chart below will highlight this moment (e.g., an orange bar), intuitively reflecting the trigger moment when the fluctuation of the characteristic parameter exceeds the threshold. A red cursor indicates the current playback progress. When the user clicks on an orange bar or an orange block above it, the application responds by instantly jumping the red cursor to the starting time position corresponding to that orange bar, and the player will automatically jump to the starting time point of that block, enabling a quick review of the high-energy moments.
[0106] See Figure 8 , Figure 8 This is a flowchart illustrating an embodiment of the session group processing method provided in this application. The processing method includes: Step 81: Obtain the session data to be converted collected by the wearable device.
[0107] The session data to be converted includes audio data from the session.
[0108] Step 82: Perform text conversion and speaker classification on the audio data to obtain the text information corresponding to each speaker, and mark the text information according to its importance.
[0109] In some embodiments, different colors can be used to mark keywords, phrases, sentences, paragraphs, etc., in the text information. For example, different colors can be used to set keywords, phrases, sentences, and paragraphs in the text information. For example, the color can be the font color or the text highlighting color. In some embodiments, the importance can vary within the same piece of text; for example, a sentence may only be marked with a keyword, keyword phrase, or a portion of the sentence.
[0110] In some embodiments, the system can acquire a collected audio data stream (which may be transmitted in real time or not) and store the audio data stream in a buffer queue; receive a marking instruction triggered by the user through an input interface; in response to the marking instruction, determine the target audio segment from the audio data stream according to a preset segment extraction rule; generate a key mark associated with the target audio segment and establish an index relationship between the key mark and the target audio segment; wherein, the segment extraction rule includes extraction rules based on the time dimension and / or extraction rules based on the content dimension.
[0111] It's understandable that a circular buffer technique can be used to ensure that the triggered operation unit (e.g., pressing a button) covers the content prior to the press. Once recording begins, the device continuously retains the most recent N time units of audio data, such as milliseconds, seconds, or minutes, in a temporary buffer. When the user triggers the operation unit by pressing a button, the processor does not start recording from the current moment but instead "retrieves" historical data from the buffer. The operation unit can be a physical button, a virtual button on a touchscreen, a double-tap gesture (such as double-tapping a ring / glasses temple), or even a specific voice command (such as saying "Mark it"), etc.
[0112] In some embodiments, the segment extraction rule is a time-dimensional extraction rule, and determining the target audio segment includes: determining the first time point at which the marking instruction is received; based on the first time point, backtracking backward for a preset first duration as the starting point of the target audio segment; based on the first time point, extending backward for a preset second duration as the ending point of the target audio segment; and extracting the audio data between the starting point and the ending point as the target audio segment.
[0113] It's understandable that a user might hear an important sentence in a conversation, but by the time they realize it and press a key, the sentence has already been partially or completely spoken. Time-based mechanical truncation logic could include: the user presses a key at time point T0. The system is configured with preset rules: before T... pre Seconds (e.g., 3 seconds, etc.) and afterT post The system locates the address from T0 to 3s in the cache and begins copying data until recording reaches T0+3s. The final generated "highlight segment" is 6 seconds long, with the highlighted point in the center. This solution offers fast response, low computational cost, and network independence, making it suitable for low-power wearable devices. It's understood that the time intervals described above can also be in other units or orders of magnitude, such as 3000 milliseconds, 5 seconds, 10 seconds, etc., and can be set as needed.
[0114] In some embodiments, the segment extraction rule is a content-dimensional extraction rule. Determining the target audio segment includes: performing speech recognition and semantic analysis on the audio data stream; analyzing the semantic integrity before and after the time point of receiving the marking instruction; identifying the semantic start boundary and semantic end boundary based on the semantic integrity; and using the audio data corresponding to the semantic start boundary to the semantic end boundary as the target audio segment. Identifying the semantic start boundary and semantic end boundary based on semantic integrity includes: detecting syntactic features in the audio text, including conjunctions, transition words, summarizing words, or preset key information guiding words; and / or, detecting acoustic features in the audio data, including changes in the speaker's speech rate, the duration of intonation pauses, or changes in volume; when a feature indicating the start of a topic or paragraph is detected, it is determined as the semantic start boundary; when a feature indicating a topic transition or paragraph end is detected, it is determined as the semantic end boundary.
[0115] This is understandable. For example, a speaker is explaining a core point, and the user, pressing a button at any point during the explanation, wants to save the entire logical paragraph, not just a fragmented sentence. The content-based intelligent segmentation processing logic includes: Text conversion: The device converts the audio stream into text in real-time or near real-time. Trigger point location: The user presses a button at T0. The system analyzes the text at T0. Semantic boundary expansion: Forward search: The algorithm analyzes the text forward, looking for paragraph start signals. For example: Detecting conjunctions / introductory words: "Firstly," "The important point is," "I want to emphasize that." Detecting long pauses (such as silence exceeding 2 seconds) is considered the end of the previous paragraph. Backward search: The algorithm records backward until a topic end signal is detected. For example: Detecting summarizing words: "In conclusion." Detecting topic switching words: "Next, let's look at," "On the other hand." Dynamic segmentation: The system, based on the timestamps of the identified semantic start and end sentences, trims out a semantically complete audio segment. The above solution makes the extracted content more accurate, providing an excellent reading / listening experience and reducing the likelihood of incomplete sentences.
[0116] Considering the complex usage scenarios of the devices (such as noisy environments, dialect communication, and network fluctuations), a single "content-based" semantic analysis may fail due to low speech recognition accuracy. Therefore, this embodiment proposes an adaptive hybrid processing logic that prioritizes semantic segmentation and automatically downgrades to time segmentation when conditions are not met. That is, the target audio segment is determined according to preset segment extraction rules, including: obtaining the signal quality parameters of the audio data stream or the confidence parameters of speech recognition; determining whether the signal quality parameters meet preset quality thresholds, and / or whether the confidence parameters meet preset confidence thresholds; if they meet, the content-based extraction rules are used to determine the target audio segment; if they do not meet, the time-based extraction rules are used to determine the target audio segment, or if the content-based extraction rules fail to identify a valid semantic boundary, the system automatically switches to the time-based extraction rules to determine the target audio segment.
[0117] It can be understood that the above may include: Step S1: The instruction receiving and data buffering processor writes the audio data collected by the microphone into a circular buffer in real time. When a user's marked instruction (such as pressing a button) is detected, the current timestamp is recorded as the trigger time point T0. Step S2: The environment and quality prediction processor first performs signal-to-noise ratio (SNR) detection on the audio segments before and after T0 for a preset duration (such as the first 3-5 seconds). Judgment logic: If the SNR is lower than the preset noise threshold (for example, SNR <10dB, indicating that the environment is extremely noisy), the system determines that it is not suitable for semantic analysis and directly jumps to step S5 (time truncation mode). If the SNR is higher than the threshold, it proceeds to step S3. Step S3: Semantic recognition and confidence assessment performs automatic speech recognition (ASR) on the audio data before and after T0 and converts it into a text stream. The recognition confidence returned by the ASR engine is obtained. Judgment logic: If the average confidence is lower than the preset score (for example, 60%), it means that although the environment is not noisy, there may be dialects, unclear pronunciation, or multiple people speaking overlapping, resulting in unreliable text. At this point, proceed directly to step S5 (time truncation mode). If the confidence level is satisfactory, proceed to step S4 (semantic truncation mode). Step S4: Content-based semantic boundary search (preferred path) forward search: Traverse backward from the text position corresponding to T0 to find the nearest semantic start feature (such as a period, long pause, or the guiding word "first point"). Record the timestamp of this position as T. start Backward search: Traverse backwards from T0, searching for the nearest semantic end feature (such as paragraph closing phrases or topic transition words). Record the timestamp of this position as T. end Boundary rationality verification: Calculate the interception duration T. d = T end -T start If T dIf the result exceeds a reasonable range (e.g., semantic analysis shows the segment is less than 3 seconds or more than 90 seconds), then the semantic analysis is deemed abnormal, the result is discarded, and the process proceeds to step S5. If T d If the result is within a reasonable range (e.g., 10 seconds to 3 minutes), then this is taken as the final result, and step S6 is executed. Step S5: Time-based fallback interception (degradation path) This is a safety net mechanism to ensure that "if the user presses it, it will definitely be remembered." It directly uses the trigger time T0 as the baseline and traces back a fixed duration T. pre (e.g., 45 seconds or 30 seconds), continuing for a fixed duration T. post (e.g., 15 seconds or 30 seconds). Taking a time interval symmetrical before and after as an example, T can be determined. start = T0-30s, T end = T0 + 30s. Of course, the above method can also determine T using asymmetric interception. start Step S6: Tag Generation and Indexing Based on the Determined T start and T end An index tag is written into the audio metadata. In some embodiments, if step S4 is executed, the application's UI displays a "smart semantic tag" icon; if step S5 is executed, the application's UI displays a "timed capture" icon, prompting the user that the segment may require manual fine-tuning.
[0118] In some embodiments, marking text information according to its importance can be done through the following process: Step 91: Determine the importance of the text information corresponding to the physiological data based on the physiological data.
[0119] In some embodiments, each speaker may wear a wearable device during the session, which can collect the speaker's physiological data, such as heart rate, respiratory rate, body temperature, electrocardiogram data, and blood pressure.
[0120] For example, different levels can be assigned to different physiological data, with each level corresponding to a degree of importance. Once the levels of the physiological data are determined, the degree of importance of the corresponding textual information can be determined.
[0121] It is understandable that during communication, a speaker's emotions may change depending on the importance of the content, leading to corresponding changes in physiological data. Based on this, these transient changes in physiological data can be used to distinguish the importance of audio data delivered at the same time. Since text information is derived from audio data, the importance of the text information corresponding to the physiological data can be determined.
[0122] Step 92: Obtain the corresponding tagging method according to the importance, and use the tagging method to tag the text information.
[0123] In some embodiments, different colors can be assigned to different levels of importance. For example, there can be three levels of importance: first importance, second importance, and third importance. The first level of importance is higher than the second level, and the second level is higher than the third level. The first level of importance can be marked in red, the second level in yellow, and the third level in green. If something is not important, it is not marked.
[0124] In some embodiments, marking text information according to its importance can also be done through the following process: Step 101: Perform emotion recognition on the audio data corresponding to the text information to obtain the emotion information corresponding to the text information.
[0125] During a conversation, the speaker's emotions may change depending on the importance of the content. Based on this, emotion recognition can be performed on the audio data to obtain corresponding emotional information.
[0126] For example, the tone, speed, and volume of audio data can be converted into emotional fluctuations to generate emotion tags. Emotion tags can be relaxed, excited, tense, angry, negative, etc. In this way, different text information can correspond to different emotion tags. These emotion tags can serve as emotional information. In some embodiments, the conversation can be categorized by emotion within a preset period (e.g., one hour) for easy review and analysis.
[0127] Step 102: Determine the importance of the text information based on emotional information.
[0128] In some embodiments, different levels of importance can be assigned to different emotional information.
[0129] Step 103: Obtain the corresponding tagging method according to the importance, and use the tagging method to tag the text information.
[0130] In some embodiments, different colors can be assigned to different levels of importance. For example, there can be three levels of importance: first importance, second importance, and third importance. The first level of importance is higher than the second level, and the second level is higher than the third level. The first level of importance can be marked in red, the second level in yellow, and the third level in green. If something is not important, it is not marked.
[0131] In some embodiments, the session data to be converted further includes importance-based operation instructions. These instructions are generated by the user operating the wearable device. Each speaker can wear the wearable device during the session and can operate it independently to generate importance-based operation instructions. Based on this, marking text information according to its importance can also follow the following process: Step 111: Obtain the importance operation instructions corresponding to the text information.
[0132] Since both text information and importance-based operation instructions have timestamps, the text information and importance-based operation instructions can be aligned based on the timestamps, thereby identifying which text information corresponds to importance-based operation instructions.
[0133] The importance of operation commands can be distinguished by double-click, single-click, triple-click, long press and shake, etc. Different operations performed by users on wearable devices can correspond to different levels of importance.
[0134] Step 112: Determine the importance of the text information based on the importance operation instructions.
[0135] In some embodiments, different levels of importance can be assigned to operation instructions of different importance levels. For example, a wearable device can provide three operation modes with different levels of importance, and based on this, three operation instructions with different levels of importance can be formed. Each operation instruction is assigned a corresponding level of importance. For example, the three levels of importance operation instructions include a first level of importance operation instruction, a second level of importance operation instruction, and a third level of importance operation instruction. The first level of importance corresponds to the first level of importance. The second level of importance corresponds to the second level of importance. The third level of importance corresponds to the third level of importance.
[0136] It's understandable that the specific importance level of the operation instructions can be set according to the functionality of the wearable device. For example, if a wearable device can provide two levels of importance operation, then two corresponding importance levels can be set.
[0137] Wearable devices can offer four levels of operation, each with its own level of importance. Alternatively, wearable devices can offer six levels of operation, each with its own level of importance.
[0138] Step 113: Obtain the corresponding tagging method according to the importance, and use the tagging method to tag the text information.
[0139] In some embodiments, different colors can be assigned to different levels of importance. For example, there can be three levels of importance: first importance, second importance, and third importance. The first level of importance is higher than the second level, and the second level is higher than the third level. The first level of importance can be marked in red, the second level in yellow, and the third level in green. If something is not important, it is not marked.
[0140] Step 83: Construct the target session group based on the tagged text information.
[0141] After the text information is tagged, the text information in the constructed target session group will display the corresponding tags so that users can know the importance of different texts.
[0142] Step 84: In response to the selection of a target conversation group, display the detailed information of the target conversation group; wherein the detailed information includes: text information of different speakers displayed in the form of a conversation and the corresponding tag for each text information.
[0143] In some embodiments, step 84 has the same or similar technical solutions as the other embodiments of this application, and will not be described in detail here.
[0144] In this embodiment, the device acquires conversation data to be converted, which includes audio data from the conversation. The audio data is then converted to text and categorized by speaker to obtain text information for each speaker, and the text information is marked according to its importance. A target conversation group is constructed based on the marked text information. In response to the selection of a target conversation group, detailed information about the target conversation group is displayed. This detailed information includes text information of different speakers displayed in a dialogue format, along with the corresponding markers for each text information. This allows users to review previous conversation content on their terminal device and quickly understand the important content considered important by different speakers throughout the conversation through the corresponding markers, facilitating subsequent targeted processing.
[0145] In some embodiments, a tag display control is provided on the display interface of the target session group; after displaying the detailed information of the target session group, in response to the selection operation of the tag display control, the current display interface is switched to the tag display interface; the tag display interface displays the text information marked in the target session group.
[0146] In some embodiments, some text information in the target session group may be marked, while some text is not marked. Therefore, the marked text information is scattered in different locations, making it difficult to view centrally. Based on this, this embodiment proposes setting up a separate marking display interface to centrally display the marked text information. For example, a marking display control can be set on the display interface of the target session group. After the user clicks the marking display control, the current display interface switches to the marking display interface; the marking display interface displays all the marked text information in the target session group. If the current display interface cannot display all of it, the remaining marked text information can be viewed by swiping.
[0147] In some embodiments, a record display control is provided on the display interface of the target session group; after displaying the detailed information of the target session group, in response to the selection operation of the record display control, the current display interface is switched to the record display interface; the record display interface displays the record information corresponding to the target session group; wherein, the record information includes at least the interaction information between the user and the smart assistant in the target session group.
[0148] In some embodiments, users can interact with a smart assistant within a target conversation group, such as asking the smart assistant questions or instructing it to perform tasks. The smart assistant can then respond accordingly, providing content for the user to view. In this embodiment, a record display interface is provided to facilitate the review and viewing of these interactions later. The record display interface displays the record information corresponding to the target conversation group. For example, a record display control is provided on the target conversation group's display interface. When the user clicks this control, the current display interface switches to the record display interface; the record display interface displays the record information corresponding to the target conversation group; wherein the record information includes at least the interaction information between the user and the smart assistant within the target conversation group.
[0149] In some embodiments, users can choose whether to record their interactions within a target session group and display the recordings on the recording display interface. If the user chooses to record the interactions, these records can be displayed on the recording display interface. If the user chooses not to record the interactions, these records will not be recorded.
[0150] In some embodiments, since the target session group corresponds to a smart assistant, the smart assistant can be used as an auxiliary tool. For example, it can receive pending tasks sent by wearable devices; use the smart assistant to process the pending tasks; and record the processing results of the pending tasks, thereby improving the intelligence of the relevant applications for session management on the terminal device.
[0151] In some embodiments, users can operate wearable devices to ask questions, thereby generating pending tasks that are sent to a terminal device or a cloud server. The terminal device can then use a smart assistant to process these tasks and record the results. These results can be stored in the user's personal space for later viewing. For example, the wearable device can remind the user to check the results during non-working hours.
[0152] In some embodiments, after recording the processing result of the pending matter, a prompt message is sent to the wearable device in response to the current time being within a preset time period; the prompt message is used to prompt the user of the wearable device to view the processing result of the pending matter.
[0153] In one application scenario, the pending task received from the wearable device is to check tomorrow's weather. The user uses a smart assistant to check tomorrow's weather and records the specific weather forecast. Responding to the user's departure time, a notification is sent to the wearable device, prompting the user to check the weather, thereby enhancing the intelligence of the related applications managing this session on the terminal device.
[0154] In one application scenario, a wearable device sends a pending task as a memo, which is then recorded using a smart assistant. In response to the current time being the time specified in the memo, a notification is sent to the wearable device, prompting the user to review the memo. This enhances the intelligence of the related applications managing this session on the terminal device.
[0155] See Figure 9 , Figure 9 This is a flowchart illustrating an embodiment of the interaction method for a conversation group provided in this application. The interaction method includes: Step 121: Display the detailed information of the target session group on the display interface; wherein, the display interface is equipped with a smart assistant control.
[0156] The target session group is obtained by converting the session data to be converted collected by the wearable device; the session data to be converted includes audio data during the session.
[0157] In some embodiments, the method for constructing the target session group can be found in other embodiments of this application, and will not be repeated here.
[0158] In some embodiments, this application can introduce a smart assistant into a target session group to assist the user. Based on this, a smart assistant control is set on the display interface, and the user can select the smart assistant control to activate its functionality.
[0159] Step 122: In response to the selection operation of the smart assistant control, display a smart assistant interaction interface on the same screen as the display interface.
[0160] In some embodiments, when a user needs assistance from a smart assistant, they can select a smart assistant control. In response to the selection of the smart assistant control, a smart assistant interaction interface is displayed simultaneously on the screen. That is, the current display interface can simultaneously show the target conversation group display interface and the smart assistant interaction interface. At this time, the user can both view relevant information in the target conversation group and operate related controls within it, and also interact with the smart assistant in the smart assistant interaction interface. For example, when viewing content in the target conversation group, the user can copy that content into the smart assistant interaction interface to submit a task to the smart assistant based on that content, so that the smart assistant can process it and provide a result.
[0161] By displaying the target conversation group display interface and the smart assistant interaction interface on the same screen, the number of user switching between interfaces can be reduced. When a problem that needs to be handled is seen in the target conversation group display interface, the user can immediately ask the smart assistant a question in the smart assistant interaction interface so that the smart assistant can answer and handle it, thereby improving the user experience.
[0162] Step 123: In response to the user's target operation on the smart assistant's interactive interface, use the smart assistant to process the target task corresponding to the target operation.
[0163] In some embodiments, the intelligent assistant's interface displays several recommendation word controls. For example, after displaying an intelligent assistant interface, the interface can show recommendation word controls for the user to use. The user can select a corresponding recommendation word control so that the intelligent assistant can directly process the task based on the recommendation word corresponding to that control. That is, in response to the user's selection of a target recommendation word control, the intelligent assistant processes the target task corresponding to the target recommendation word control.
[0164] In some embodiments, the recommendation word control includes at least one of the following: a conversation summary control, a plan summary control, a decision summary control, a key marker extraction control, and an information extraction control.
[0165] In some embodiments, in response to a user's selection of a conversation summary control, a smart assistant processes the conversation summary task corresponding to the control. After processing the task, the smart assistant displays the corresponding summary content on its interactive interface. In some embodiments, this summary content can be recorded for later viewing on a record display interface.
[0166] In some embodiments, in response to a user's selection of a plan summary control, a smart assistant processes the plan summary task corresponding to the control. After processing the plan summary task, the smart assistant displays the corresponding summary content on its interactive interface. In some embodiments, this summary content can be recorded for later viewing on a recording display interface.
[0167] In some embodiments, in response to a user's selection of a decision summary control, a smart assistant processes the decision summary task corresponding to the control. After processing the decision summary task, the smart assistant displays the corresponding summary content on its interactive interface. In some embodiments, this summary content can be recorded for later viewing on a recording display interface.
[0168] In some embodiments, in response to a user's selection of a key marker extraction control, a smart assistant processes the marker extraction task corresponding to the key marker extraction control. After processing the marker extraction task, the smart assistant displays the extracted marker content on the smart assistant's interactive interface. In some embodiments, this marker content can be recorded for viewing on a recording display interface.
[0169] In some embodiments, in response to a user's selection of an information extraction control, a smart assistant processes the information extraction task corresponding to the control. After processing the task, the smart assistant displays the extracted information on its interface. In some embodiments, this extracted information may be recorded for later viewing on a record display interface.
[0170] In some embodiments, in response to text content entered by the user on the intelligent assistant's interactive interface, the intelligent assistant is used to process the target task corresponding to the text content.
[0171] In some embodiments, in response to text content entered by the user on the smart assistant's interactive interface, the smart assistant processes the target task corresponding to the text content based on the content currently displayed in the interface of the target conversation group.
[0172] In some embodiments, the detailed information of the target session group includes: text information of different speakers displayed in a dialogue format in timestamp order, and tag information corresponding to each text message. Step 143 can also be the following process: Step 1231: Receive the text content entered by the user on the intelligent assistant's interactive interface; the text content is used to instruct the intelligent assistant to organize the tagging information of a single speaker and summarize the tagging information.
[0173] In some embodiments, users can also use voice input in the smart assistant interface to interact with the smart assistant via voice.
[0174] Step 1232: Use the smart assistant to display the tagging information of individual speakers in the interface of the target conversation group, and display a summary of the tagging information in the smart assistant's interactive interface.
[0175] For example, the target conversation group includes the conversation content of speaker A and speaker B. The text corresponding to this conversation content has relevant tagging information. For instance, the user enters text content on the intelligent assistant's interface; this text content instructs the intelligent assistant to organize and summarize speaker A's tagging information. Based on this, the intelligent assistant displays speaker A's tagging information in the target conversation group's interface and displays the summary information of the tagging information in the intelligent assistant's interactive interface. In this way, speaker A's tagging information and the summary information can be displayed simultaneously on the screen, allowing users to view them synchronously for easier understanding and comparison.
[0176] Step 124: Display the task results of the target task on the intelligent assistant's interactive interface.
[0177] In some embodiments, after the task result of the target task is displayed on the smart assistant's interactive interface, the interaction information between the user and the smart assistant in the target conversation group is recorded; the interaction information can be displayed on the record display interface.
[0178] In some embodiments, in response to an adjustment operation on the boundary of the smart assistant's interactive interface, the size of the smart assistant's interactive interface and the interface size of the target conversation group are adjusted synchronously. Based on this, users can adjust the boundary of the smart assistant's interactive interface according to their own viewing habits, and synchronously adjust the size of the smart assistant's interactive interface and the interface size of the target conversation group to achieve a user-favorable interface size, facilitating user operation on the interface.
[0179] In this embodiment, detailed information of the target conversation group is displayed on the display interface. The display interface includes a smart assistant control. The target conversation group is obtained by converting conversation data collected by a wearable device. The conversation data includes audio data from the conversation. In response to a selection operation on the smart assistant control, a smart assistant interaction interface is simultaneously displayed on the display interface. In response to a user's target operation on the smart assistant interaction interface, the smart assistant handles the target task corresponding to the target operation. The task result of the target task is displayed on the smart assistant interaction interface. This allows users to review previous conversation content on their terminal device and invoke the smart assistant during the conversation to assist in completing user tasks. This helps users quickly organize relevant information in the target conversation group without needing to use other methods, thus improving the intelligence of conversation group management and user experience.
[0180] See Figure 10, Figure 10 This is a flowchart illustrating an embodiment of the session group processing method provided in this application. The processing method includes: Step 131: Display the details of the first target session group on the display interface.
[0181] The first target session group is obtained by converting the session data to be converted collected by the wearable device; the session data to be converted includes audio data during the session.
[0182] In some embodiments, the method for constructing the target session group can be found in other embodiments of this application, and will not be repeated here.
[0183] Step 132: In response to the speaker addition operation, add the first smart assistant as a speaker to the first target conversation group; the first smart assistant is used to interact with the other speakers in the first target conversation group.
[0184] In some embodiments, the smart assistant can exist not only as a smart assistant control as mentioned in other embodiments, but also as a speaker in a conversation group. For example, a user selects an add control on the display interface of the first target conversation group to display a contact list; the contact list can display a smart assistant identifier. The user can select this smart assistant identifier (first smart assistant) to add the first smart assistant as a speaker to the first target conversation group. At this time, the first smart assistant can exist as a speaker in the first target conversation group. The first smart assistant can interact with other speakers in the target conversation group. For example, the user can interact with the first smart assistant, and the first smart assistant can reply to the corresponding interaction content in a conversational manner.
[0185] The display interface shows detailed information about the first target conversation group; the first target conversation group is obtained by converting the conversation data to be converted collected by the wearable device; the conversation data to be converted includes audio data during the conversation; in response to the speaker addition operation, the first smart assistant is added as a speaker to the first target conversation group; the first smart assistant is used to interact with the other speakers in the first target conversation group, so that users can add the smart assistant to different conversation groups at any time during the conversation, and use the smart assistant to assist users in completing the corresponding analysis and organization work, thereby improving the intelligence of conversation group management and enhancing the user experience.
[0186] In some embodiments, since different conversation groups may involve different industry sectors, and considering the different sectors involved, the intelligence level of the smart assistant will also vary. If a general-purpose smart assistant is used, some problems may arise, such as the large amount of data it consumes and its inaccurate understanding of the technologies in different industry sectors. Based on this, this application proposes to set up targeted smart assistants according to industry sectors, allowing the domain-specific smart assistant to focus on solving tasks and problems within that domain.
[0187] In some embodiments, step 132 described above may be the following process: Step 1321: In response to the speaker addition operation, obtain the first industry sector corresponding to the first target session group.
[0188] In some embodiments, in response to a speaker addition operation, industry sector identification can be performed on the text information involved in the first target session group to obtain the first industry sector corresponding to the first target session group.
[0189] In some embodiments, in response to a speaker addition operation, a list of existing industry sectors can be displayed, allowing the user to select the first industry sector corresponding to the first target session group, thereby obtaining the first industry sector corresponding to the first target session group.
[0190] In some embodiments, in response to a speaker addition operation, an industry field input field can also be provided, allowing the user to enter the first industry field of the first target session group, thereby obtaining the first industry field corresponding to the first target session group.
[0191] Step 1322: Obtain the first intelligent assistant corresponding to the first industry sector.
[0192] In some embodiments, after obtaining the first industry sector corresponding to the first target session group, the corresponding first smart assistant can be obtained based on the first industry sector. For example, the first industry sector can be sent to a cloud server so that the cloud server can determine the corresponding first smart assistant and deploy the first smart assistant to the first target session group.
[0193] In some embodiments, the smart assistant can be deployed locally. After obtaining the first industry sector corresponding to the first target session group, the first smart assistant corresponding to the first industry sector can be obtained from the locally deployed smart assistant.
[0194] Step 1323: Add the first intelligent assistant as a speaker to the first target conversation group.
[0195] For example, if the first industry sector is the patent industry sector, then the first intelligent assistant in the patent industry sector can be added to the first target conversation group.
[0196] For example, if the first industry sector is the automotive industry sector, then the first intelligent assistant in the automotive industry sector can be added to the first target conversation group.
[0197] For example, if the first industry sector is the apparel industry sector, then the first smart assistant in the apparel industry sector can be added to the first target conversation group.
[0198] In some embodiments, step 1323 has the same or similar technical solutions as the other embodiments of this application, and will not be described in detail here.
[0199] In some embodiments, after the first intelligent assistant is added as a speaker to the first target session group, when the first intelligent assistant is unable to resolve the first task raised in the first target session group, the second industry field corresponding to the first task is obtained; the first intelligent assistant is updated so that the first intelligent assistant has the ability to resolve the first task related to the second industry field.
[0200] For example, during user interaction with the first intelligent assistant, if the first intelligent assistant cannot resolve the first task raised in the first target conversation group, it indicates that the user's question exceeds the first industry domain of the first intelligent assistant. In this case, the first intelligent assistant can be updated to enable it to resolve the first task related to a second industry domain. After updating, the first intelligent assistant can then interact with the user to resolve tasks related to the second industry domain.
[0201] Similarly, if the updated first intelligent assistant cannot solve the new task proposed in the first target conversation group, the new industry field corresponding to the new task is taken; the first intelligent assistant is updated so that it has the ability to solve new tasks related to the new industry field.
[0202] In some embodiments, after the first intelligent assistant is added as a speaker to the first target conversation group, see [link / reference]. Figure 11 Alternatively, the process can be as follows: Step 141: When the first intelligent assistant cannot solve the first task proposed in the first target conversation group, obtain the second industry field corresponding to the first task.
[0203] Step 142: Obtain the second intelligent assistant corresponding to the second industry sector.
[0204] Since this application can provide intelligent assistants for several industry sectors, a second intelligent assistant corresponding to the second industry sector can be obtained.
[0205] In some embodiments, after obtaining the second industry domain corresponding to the first task, a corresponding second intelligent assistant can be obtained based on the second industry domain. For example, the second industry domain is sent to a cloud server so that the cloud server determines the corresponding second intelligent assistant and deploys the second intelligent assistant to the first target session group.
[0206] In some embodiments, the smart assistant can be deployed locally. After obtaining the second industry sector corresponding to the first task, a second smart assistant corresponding to the second industry sector can be obtained from the locally deployed smart assistant.
[0207] Step 143: Add the second smart assistant as a speaker to the first target conversation group.
[0208] In some embodiments, step 143 has the same or similar technical solutions as other embodiments of this application, and will not be described in detail here. The second intelligent assistant is used to interact with other speakers in the first target conversation group to solve the interaction tasks corresponding to the second industry field.
[0209] In some embodiments, after adding the second intelligent assistant as a speaker to the first target conversation group, see [link to documentation]. Figure 12 Alternatively, the process can be as follows: Step 151: In response to the second task proposed in the first target session group, obtain the industry sector corresponding to the second task.
[0210] After adding the second smart assistant as a speaker to the first target conversation group, the user can continue to interact with these smart assistants in the first target conversation group.
[0211] Once the second task provided by the user is detected, the industry sector corresponding to the second task can be obtained.
[0212] Step 152: In response to the industry sector belonging to the first industry sector, use the first intelligent assistant to handle the second task.
[0213] In some embodiments, in response to the industry field belonging to the first industry field, the first intelligent assistant can be invoked to handle the second task and the corresponding processing result can be given in the first target session group.
[0214] Step 153: In response to the industry field belonging to the second industry field, use the second intelligent assistant to handle the second task.
[0215] In some embodiments, in response to the industry field belonging to a second industry field, a second intelligent assistant can be invoked to handle a second task and provide the corresponding processing result in the first target session group.
[0216] By using the above methods, intelligent assistants from different industry sectors can be used to handle their corresponding tasks, thereby avoiding the problem of tasks being unsolvable due to the mismatch between the intelligent assistant and the task, and improving the task processing efficiency.
[0217] In some embodiments, in response to an industry sector that does not belong to the second or third industry sector, a smart assistant belonging to that industry sector can be added to the first target session group. Alternatively, the first or second smart assistant can be updated to enable it to handle that industry sector.
[0218] In the above embodiments, the smart assistant is treated as a contact in an address book. Users can add the smart assistant to different conversation groups at any time during the viewing process, and use the smart assistant to assist users in completing corresponding analysis and organization tasks.
[0219] See Figure 13 , Figure 13 This is a flowchart illustrating another embodiment of the session group processing method provided in this application. The processing method includes: Step 161: In response to the existence of several session groups, initiate a session merging operation using a third-party smart assistant.
[0220] In some embodiments, if multiple sessions (meetings) are conducted successively based on the same project, several session groups may exist according to the technical solution of this application. Since these session groups all concern the same project's content, scattering them across different session groups may be inconvenient for organization, management, and viewing. Therefore, this application provides a session merging operation, which can merge related session groups to combine scattered groups into one. The merged group can then operate according to any of the above embodiments, thereby integrating the entire project-related content.
[0221] In some embodiments, conversation groups can be merged to form new conversation groups through manual selection. These conversation groups can be built based on online conversations or offline conversations.
[0222] In some embodiments, since this application provides a smart assistant, the smart assistant can be used to initiate a session merging operation, thereby improving the efficiency of merging session groups.
[0223] Step 162: Use a third-party intelligent assistant to merge at least two related conversation groups to form a new conversation group.
[0224] In some embodiments, in response to the existence of several conversation groups, a third intelligent assistant is used to identify several conversation groups and determine the relevant second target conversation groups; a conversation group selection interface is provided; the conversation group selection interface provides all the second target conversation groups for the user to select; in response to the user's selection operation, the third intelligent assistant is used to merge at least two relevant second target conversation groups selected by the user to form a new conversation group.
[0225] In some embodiments, although the third-party intelligent assistant can identify relevant second target conversation groups, there may be some accuracy issues, leading to the merging of some less relevant conversation groups. Therefore, a conversation group selection interface is provided, allowing users to select the conversation groups to be merged, thereby manually correcting potential merging anomalies caused by the third-party intelligent assistant. Furthermore, users can select which conversation groups to merge according to their needs, without having to merge all relevant conversation groups.
[0226] In some embodiments, step 162 may be the following process: using a third intelligent assistant to merge at least two related conversation groups in chronological order of timestamps to form a new conversation group.
[0227] Since the text and other information in different conversation groups all have timestamps, they can be merged according to the order of the timestamps to form new conversation groups. This makes it easier for users to view the content or for smart assistants to process it and arrive at reasonable conclusions based on the chronological order.
[0228] In this embodiment, a smart assistant is used to merge conversation groups, thereby simplifying the manual merging process.
[0229] See Figure 14 , Figure 14 This is a flowchart illustrating another embodiment of the session management method provided in this application. The session management method includes: Step 171: In response to the establishment of communication between at least two wearable devices, determine the master device and the slave device from the at least two wearable devices.
[0230] In conversational scenarios, if it's necessary to collect physiological data and related operations from multiple speakers, each speaker needs to wear a wearable device. How to effectively manage these wearable devices in conversational scenarios involving multiple devices is a problem that needs to be solved.
[0231] Based on this, in a conversational scenario, communication can be established between at least two wearable devices to determine the master and slave devices among them. For example, the wearable device with the most remaining battery power can be designated as the master device, and the others as slave devices. Since the master device may require more operations, such as collecting audio data and transmitting it to the terminal device and / or cloud server, its power consumption is higher, resulting in greater power requirements. Alternatively, the wearable device with the most up-to-date functional modules can be designated as the master device, and the others as slave devices. Because the master device may require more operations, such as collecting audio data, it needs to utilize the best functional modules to improve the quality of the collected audio data.
[0232] In some embodiments, prior to step 171, refer to Figure 15 The method also includes: Step 181: Obtain the session information of the offline session in advance. The session information includes the session start time, the preset session duration, and the speaker.
[0233] In some embodiments, the session information for offline sessions can be set in advance via a memo. The session information includes the session start time, the preset session duration, and the speaker for the session.
[0234] In some embodiments, the session information for offline sessions can be pre-set in the application of the session group. The session information includes the session start time, the preset session duration, and the session speaker.
[0235] It's understandable that each speaker in the conversation has a wearable device, and they will bring the wearable device to the conversation. The speaker can then establish a connection with their wearable device, allowing for easy identification of the device and subsequent charging reminders.
[0236] During a conversation, the wearable device needs to collect audio, gather the speaker's physiological data, and respond to the speaker's actions. All of these require the wearable device to provide power throughout the entire conversation. Therefore, this embodiment proposes pre-emptive power management for the wearable device.
[0237] Step 182: In response to the time interval from the current time to the start time of the session meeting the preset time interval, obtain the current battery level of the wearable device of each speaker in the session.
[0238] In some embodiments, the minimum remaining battery power of the wearable device can be determined based on a preset session duration. This minimum remaining battery power enables the wearable device to be used as the primary device within the preset session duration. The timing for executing step 182 can be determined by the session start time.
[0239] This time interval ensures that the wearable device, after being charged within this interval, has enough power to support the entire session.
[0240] Step 183: In response to the fact that the current battery level of the target speaker's wearable device does not meet the session requirements, send a charging reminder to the target speaker's wearable device.
[0241] The session requirements include: the wearable device's battery power must be sufficient to support its use as the primary device for the preset session duration.
[0242] In some embodiments, after the target speaker sees a charging prompt on their corresponding wearable device, the wearable device needs to be charged. When the session start time arrives, the wearable device can be used to participate in the session, either as a master device or a slave device.
[0243] Step 172: Collect session data using the master device and collect relevant user information using the slave device.
[0244] The session data includes audio data during the session, and related information includes user physiological data and / or user operation information on the wearable device.
[0245] In some embodiments, a redundant wearable device can be configured as the master device, and the wearable device worn by the user can be configured as the slave device. For example, in a meeting scenario, a wearable device can be configured as the master device in the meeting room, and the wearable device worn by the user can be configured as the slave device. After all users have entered the meeting room, the wearable device in the meeting room and the wearable device worn by the user communicate to determine the master device and the slave device.
[0246] In some embodiments, both the master device and the slave device are wearable devices worn by the user. Based on this, the master device collects session data and relevant information of the first user, while the slave device collects relevant information of the other users. Here, the slave device may only collect relevant information of the other users and is not responsible for collecting audio data during the session. The master device, however, is responsible for collecting audio data during the session and collecting relevant information of the first user. The master device is worn by the first user.
[0247] Step 173: Send the session data and relevant user information to the terminal device and / or cloud server so that the terminal device and / or cloud server can construct a session group based on the session data and relevant user information.
[0248] In some embodiments, the slave device can interact with the master device and send the user-related information it collects to the master device, which then sends this data to the terminal device and / or cloud server.
[0249] In some embodiments, the slave device can interact with the terminal device and / or cloud server to send the user-related information it collects to the terminal device and / or cloud server.
[0250] After a session group is created, users can view and interact with it in the corresponding application on their terminal device.
[0251] In this embodiment, in response to at least two wearable devices establishing communication, a master device and a slave device are determined from the at least two wearable devices. The master device collects session data, and the slave device collects relevant user information. The session data includes audio data during the session, and the relevant information includes user physiological data and / or user operation information on the wearable devices. The session data and relevant user information are sent to a terminal device and / or a cloud server, so that the terminal device and / or the cloud server can construct a session group based on the session data and relevant user information. By assigning master and slave devices to the wearable devices, the use of wearable devices during the session is managed rationally, avoiding the problem of repeated collection of session data in multiple wearable device scenarios and saving power consumption of the wearable devices.
[0252] In some embodiments, considering the storage capacity limitations of wearable devices, if the session duration is long, the wearable device's storage capacity may be insufficient to store the session data collected during that session. Therefore, the following approach can be used to address this: In response to the remaining available storage capacity of the master and slave devices being less than a preset storage capacity, the existing session data in the master device and the existing relevant information in the slave device are sent to the terminal device and / or the cloud server. After transmission is complete, the existing session data in the master device and the existing relevant information in the slave device are cleared. This frees up more remaining available storage capacity in the master and slave devices for continuous collection of session data during the session.
[0253] In some embodiments, the preset storage capacity can be determined based on the total storage capacity of the wearable device. For example, 10% of the total storage capacity of the wearable device can be used as the preset storage capacity. For example, 15% of the total storage capacity of the wearable device can be used as the preset storage capacity. For example, 20% of the total storage capacity of the wearable device can be used as the preset storage capacity.
[0254] In some embodiments, considering the storage capacity limitations of wearable devices, if the session duration is long, the wearable device's storage capacity may be insufficient to store the session data collected during that session. Therefore, the following approach can be used to address this: In response to the master device and slave device having collected session data and related information for a preset duration, the collected session data in the master device and the collected related information in the slave device are sent to the terminal device and / or cloud server. After transmission is complete, the collected session data in the master device and the collected related information in the slave device are cleared.
[0255] In some embodiments, the preset duration can be determined based on the preset session duration. For example, the preset duration can be obtained according to a ratio. If the ratio is 10% and the preset session duration is 2 hours, then the preset duration can be 12 minutes. If the ratio is 20% and the preset session duration is 2 hours, then the preset duration can be 24 minutes.
[0256] In some embodiments, the preset duration can be a fixed duration, such as 10 minutes, 20 minutes, 30 minutes, etc.
[0257] In some embodiments, in response to the master device and slave device receiving a user operation instruction, session data and relevant user information are sent to the terminal device and / or cloud server; the user operation instruction is used to indicate the end of the session. After transmission is complete, the session data in the master device and the relevant information in the slave device are cleared.
[0258] In some embodiments, in response to the master device and slave device receiving a user operation instruction, session data and relevant user information are sent to the terminal device and / or cloud server; the user operation instruction is used to instruct the master device and slave device to send data. After transmission is complete, the session data in the master device and the relevant information in the slave device are cleared. In this embodiment, the user can actively initiate the data transmission operation between the master device and slave device. For example, if the topic of the current session ends during the session, the user can operate on the wearable device to cause the master device and slave device to send the data of the current topic to the terminal device and / or cloud server so as to continue collecting data for the next topic.
[0259] See Figure 16 , Figure 16 This is a flowchart illustrating another embodiment of the session management method provided in this application. The session management method includes: Step 191: During the session, obtain the remaining battery power of each wearable device.
[0260] In some embodiments, the scheme of this embodiment can be implemented according to a preset cycle.
[0261] In some embodiments, the preset period can be the preset duration described above. After the session data collected in the master device and the relevant information collected in the slave device are sent to the terminal device and / or cloud server, step 191 is executed.
[0262] In some embodiments, step 191 may be performed after the collected session data from the master device and the collected relevant information from the slave device are sent to the terminal device and / or cloud server.
[0263] In some embodiments, the remaining power of the main device can be monitored in real time during the session, and step 191 is executed when the remaining power of the main device is less than the preset power.
[0264] Step 192: Re-determine the master and slave devices based on the remaining power.
[0265] Among them, the main equipment that was re-determined had the most remaining power.
[0266] After the master and slave devices are re-determined, data acquisition and command response operations during the session are performed according to the latest master and slave devices.
[0267] In some embodiments, due to power consumption issues, the power consumption of the master device may be significantly greater than that of the slave device. Therefore, the remaining power of the master device may not be sufficient to support data collection until the end of the session. The technical solution of this application can be adopted to re-determine the master and slave devices among these wearable devices, thereby maximizing the remaining power of the re-determined master device to support data collection until the end of the session.
[0268] See Figure 17 , Figure 17 This is a flowchart illustrating another embodiment of the session management method provided in this application. The session management method includes: Step 201: In response to the establishment of communication between the terminal device and at least two wearable devices, the terminal device is identified as the master device and the at least two wearable devices are identified as slave devices.
[0269] In conversational scenarios, if it's necessary to collect physiological data and related operations from multiple speakers, each speaker needs to wear a wearable device. How to effectively manage these wearable devices in conversational scenarios involving multiple devices is a problem that needs to be solved.
[0270] Based on this, in a session scenario, the terminal device establishes communication with at least two wearable devices, thereby identifying the terminal device as the master device and the at least two wearable devices as slave devices.
[0271] In a meeting setting, a terminal device is set up in the meeting room as the master device, and wearable devices worn by users are used as slave devices. After all users enter the meeting room, the terminal device in the meeting room and the wearable devices worn by the users communicate, thereby designating the terminal device as the master device and at least two wearable devices as slave devices.
[0272] Step 202: Collect session data using the master device and collect relevant user information using the slave device.
[0273] The session data includes audio data during the session, and related information includes user physiological data and / or user operation information on the wearable device.
[0274] Step 203: Send the session data and relevant user information to the terminal device and / or cloud server so that the terminal device and / or cloud server can construct a session group based on the session data and relevant user information.
[0275] In some embodiments, the slave device can interact with the master device and send the user-related information it collects to the master device, which then sends this data to the terminal device and / or cloud server.
[0276] In some embodiments, the slave device can interact with the terminal device and / or cloud server to send the user-related information it collects to the terminal device and / or cloud server.
[0277] In this embodiment, in response to the establishment of communication between a terminal device and at least two wearable devices, the terminal device is designated as the master device, and the at least two wearable devices are designated as slave devices. The master device collects session data, and the slave devices collect relevant user information. The session data includes audio data during the session, and the relevant information includes user physiological data and / or user operation information on the wearable devices. The session data and relevant user information are sent to the terminal device and / or the cloud server, enabling the terminal device and / or the cloud server to construct a session group based on the session data and user information. By assigning master and slave roles to the terminal device and wearable devices, the use of wearable devices during the session is managed rationally, avoiding the problem of repeated collection of session data in multiple wearable device scenarios and saving power consumption of the wearable devices.
[0278] In some embodiments, after a terminal device malfunctions, the master-slave relationship is determined from the wearable device.
[0279] In some embodiments, a corresponding session group can be constructed for each session scenario according to the technical solutions provided in this application. Users can view these session groups through an application on their terminal device. When a session group is selected, users can enter that session group to view the corresponding session content and apply the technical solutions of any of the above embodiments within that session group to perform the corresponding operations.
[0280] In one application scenario, combined Figures 18-24 Please explain the technical solutions involved in the application: After several session groups are created, users can view them on their terminal devices, such as... Figure 18 As shown, you can view the basic information of 6 conversation groups, such as the conversation time, conversation topic, number of participants, and number of recorded (note) items. For example, the conversation group with the topic "Operation Accompaniment - West Ring Road" was created on October 28, with 2 participants and 2 recorded (note) items.
[0281] Users can select the conversation topics they wish to view on this interface. For example, selecting the conversation group "Operational Accompaniment - West Ring Road" will switch to... Figure 19 The interface shown contains various controls and content. For example, a represents an information display control, b represents a tag display control, c represents a record display control, d represents a speaker's identifier, e represents a timestamp, f represents the first icon, g represents the second icon, h represents a smart assistant control, i represents a mood tag control, and j represents an AI insight control. Figure 19 The interface shown primarily displays the text information of the dialogue between two speakers at different times. Different colored text indicates the importance of that text segment. The emotion marker control `i` represents the current emotion of the dialogue and can be associated with markers. Users can select the emotion marker control `i` to display the emotion. The AI insight control `j` is used to proactively gain insights throughout the conversation. For example, it can automatically summarize and automatically identify key phrases. Users can also customize conditions for the AI insight control; once a condition is met, it will be marked. For example, setting keywords such as "ces" or "stocks" will enable automatic marking. In subsequent processing (marking display interface or record display interface), the content of these AI insight controls `j` can be quickly filtered out separately.
[0282] like Figure 19As shown, there are five dialogues with two speakers in the target conversation group. Each dialogue area displays a corresponding timestamp (e), a first icon (f), a second icon (g), text information, and a speaker identifier (d). When the user selects the first icon (f) in the second dialogue, the audio corresponding to the text information in that dialogue will play. When the user selects the second icon (g) in the second dialogue, the audio corresponding to the text information in each subsequent dialogue will play sequentially, starting from the second dialogue, until the audio corresponding to the text information in every five dialogues has been played. The intelligent assistant control (h) can still be present on this interface to assist the user.
[0283] When the user selects Figure 19 When displaying the marker control b in the middle, from Figure 19 The displayed interface switches to Figure 20 The interface shown. Figure 20 The interface shown primarily presents Figure 19 The dialogue content of the tagged speakers is displayed. The smart assistant control (h) can still be present on this interface to assist the user.
[0284] When the user selects Figure 19 When displaying records in control c, from Figure 19 The displayed interface switches to Figure 21 The interface displayed. Or when the user selects... Figure 20 When displaying records in control c, from Figure 20 The displayed interface switches to Figure 21 The interface shown. Figure 21 The displayed interface primarily presents the user's actions and recorded content during the viewing process. For example... Figure 21 The document records information related to the "automatically generated summary" and improvements to the security briefing. The intelligent assistant control (h) can still be present on this interface to assist users.
[0285] When the user selects Figure 19 When using the smart assistant control h in the middle, from Figure 19 The displayed interface switches to Figure 22 The interface shown. Figure 22 The interface shown primarily presents the intelligent assistant's interactive interface and the chat group information interface. Users can interact with the intelligent assistant in the interactive interface to help it answer and process corresponding tasks. These tasks can be recorded (as notes) for later use. Figure 21 The interface shown is displayed separately. The intelligent assistant's interface can also provide recommended word controls, such as summaries, action items, and key decisions. Users can select these recommended word controls so that the intelligent assistant can perform the corresponding tasks.
[0286] When the user selects Figure 22 When displaying the marker control b in the middle, from Figure 22 The displayed interface switches to Figure 23 The interface shown.
[0287] When the user selects Figure 22 When displaying records in control c, from Figure 22 The displayed interface switches to Figure 24 The interface shown.
[0288] When the user selects Figure 20 When using the smart assistant control h in the middle, from Figure 20 The displayed interface switches to Figure 23 The interface shown.
[0289] When the user selects Figure 21 When using the smart assistant control h in the middle, from Figure 21 The displayed interface switches to Figure 24 The interface shown.
[0290] When the user selects Figure 19 When the speaker identifier d is specified, a contact can be bound to the speaker identifier d.
[0291] In other embodiments, Figure 19 You can also set up the addition of controls, which will appear when the user selects... Figure 19 When adding controls, you can add contacts within the session group. Bind a contact to the speaker's identifier d.
[0292] See Figure 25 , Figure 25 This is a schematic diagram of an embodiment of the session management device provided in this application. The session management device 300 includes a memory 301 and a processor 302 coupled to each other. The memory 301 stores program instructions, and the processor 302 is used to execute the program instructions to implement the method provided in any of the above embodiments.
[0293] See Figure 26 , Figure 26 This is a schematic diagram of an embodiment of the session management system provided in this application. The session management system 400 includes: a wearable device 500 and a session management device 300.
[0294] See Figure 27 , Figure 27 This is a schematic diagram of an embodiment of the computer-readable storage medium provided in this application. The computer-readable storage medium 600 stores program instructions 601 that can be executed by a processor. The program instructions 601 are used to implement the methods provided in any of the above embodiments.
[0295] In the several embodiments provided in this application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For instance, the division of circuits or units is merely a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed.
[0296] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment, depending on actual needs.
[0297] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0298] The above description is merely an embodiment of this application and does not limit the patent scope of this application. Any equivalent structural or procedural changes made based on the description and drawings of this application, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.
Claims
1. A method for processing a session group, characterized in that, The processing method includes: The display interface shows the detailed information of the first target session group; the first target session group is obtained by converting the session data to be converted collected by the wearable device; the session data to be converted includes audio data during the session; In response to the speaker addition operation, the first intelligent assistant is added as a speaker to the first target session group; the first intelligent assistant is used to interact with the other speakers in the first target session group.
2. The processing method according to claim 1, characterized in that, The step of adding the smart assistant as a speaker to the first target conversation group in response to the speaker addition operation includes: In response to the speaker addition operation, obtain the first industry sector corresponding to the first target session group; Obtain the first intelligent assistant corresponding to the first industry sector; Add the first smart assistant as a speaker to the first target conversation group.
3. The processing method according to claim 1 or 2, characterized in that, After adding the first intelligent assistant as a speaker to the first target conversation group, the method further includes: When the first intelligent assistant cannot resolve the first task raised in the first target conversation group, the second industry sector corresponding to the first task is obtained. The first intelligent assistant is updated to enable it to solve a first task related to the second industry sector.
4. The processing method according to claim 1 or 2, characterized in that, After adding the first intelligent assistant as a speaker to the first target conversation group, the method further includes: When the first intelligent assistant cannot resolve the first task raised in the first target conversation group, the second industry sector corresponding to the first task is obtained. Obtain the second intelligent assistant corresponding to the second industry sector; Add the second smart assistant as a speaker to the first target conversation group.
5. The processing method according to claim 1 or 2, characterized in that, After adding the second smart assistant as a speaker to the first target conversation group, the method further includes: In response to the second task proposed in the first target session group, obtain the industry sector corresponding to the second task; In response to the fact that the industry sector belongs to the first industry sector, the second task is processed using the first intelligent assistant; In response to the industry sector belonging to the second industry sector, the second intelligent assistant is used to handle the second task.
6. The processing method according to claim 1 or 2, characterized in that, The method further includes: In response to the existence of several conversation groups, a conversation merging operation is initiated using a third-party intelligent assistant; Use a third-party intelligent assistant to merge at least two related conversation groups to form a new conversation group.
7. The processing method according to claim 6, characterized in that, The process of initiating a session merging operation using a third-party intelligent assistant includes: The third intelligent assistant is used to identify several conversation groups and determine the relevant second target conversation groups; A session group selection interface is provided; the session group selection interface provides all the second target session groups for the user to select. The method of merging at least two related conversation groups to form a new conversation group using a third-party intelligent assistant includes: In response to the user's selection action, the third intelligent assistant is used to merge at least two related second target conversation groups selected by the user to form a new conversation group.
8. The processing method according to claim 6, characterized in that, The method of merging at least two related conversation groups to form a new conversation group using a third-party intelligent assistant includes: The third intelligent assistant is used to merge at least two related conversation groups in chronological order according to their timestamps to form a new conversation group.
9. A session management device, characterized in that, The session management device includes a memory and a processor coupled to each other, the memory storing program instructions, and the processor executing the program instructions to implement the processing method according to any one of claims 1 to 8.
10. A session management system, characterized in that, The session management system includes: a wearable device and the session management apparatus as described in claim 9.