System and method for determining semantic points in person-to-person session

By identifying and analyzing human-person conversations in multiple conversation rounds, determining natural language attributes and deriving instantaneous states and conversation nuances, the problem of difficult to understand and extract semantic points of human-person conversation in the prior art is solved, and efficient information search and navigation is achieved.

CN120019381APending Publication Date: 2025-05-16SAMSUNG ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380071890.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-10-10
Filing Date
2023-10-10
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The prior art is difficult to effectively understand and extract semantic points in human-to-person conversations, especially in chat rooms or groups where multiple users participate, resulting in difficulty in searching and navigation of information.

Method used

By identifying human-to-person conversations across multiple conversation rounds, determining natural language attributes, deriving instantaneous states and session nuances, and dynamically storing relevant information to determine semantic relations and generate semantic points.

Benefits of technology

It realizes effective identification and extraction of semantic points in human-to-person conversations, simplifies information search and navigation, and improves user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120019381A_ABST
    Figure CN120019381A_ABST
Patent Text Reader

Abstract

A system and method for determining semantic points in a person-to-person session is provided. The method includes identifying a person-to-person session including a plurality of conversation rounds, and determining natural language (NL) attributes from each conversation round. Further, the method includes deriving an instantaneous state based on the one or more NL attributes. Further, the method includes deriving one or more session nuisance associated with the person-to-person session based on the one or more NL attributes. Further, the method includes dynamically storing information associated with the person-to-person session based on the one or more NL attributes, the transient state, and the one or more session nuance associated with each session round, and determining one or more semantic relationships and associated conversation timelines within the person-to-person session based on the dynamically stored information. In addition, the method includes generating semantic points corresponding to the one or more semantic relationships and associated conversation timelines within the determined person-to-person session.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to natural language processing. More specifically, the present disclosure relates to systems and methods for determining semantic points in human-to-human conversations. Background Art

[0002] With the recent advancement of mobile communications, online messaging platforms have gained popularity as an easy mode of communication. Such messaging platforms enable users to send messages including text or graphics between users. In addition, some messaging platforms allow for the implementation of chat rooms and / or groups where multiple users can simultaneously participate in discussing one or more common topics.

[0003] However, when multiple users send messages in such chat rooms or groups, it is sometimes difficult to extract or search for relevant information. In addition, traditional search techniques can only perform time-consuming keyword-based searches, and are cumbersome because users must analyze all search results and draw expected conclusions.

[0004] Some dialogue summary tools are also available, which simply combine one or more dialogues of one or more users to form a single dialogue segment. However, such tools cannot find the dialogue responsible for the expected conclusion. Usually, dialogue understanding is classified into three categories, i.e., human-robot conversation (HBC), human-human conversation (HHC) and robot-robot conversation (BBC). Among the above three dialogue categories, HBC is highly goal-oriented, structured and predictable in nature, while BBC is rarely used. However, HHC is a very unstructured conversation form with many uncertainties. Therefore, conventional systems cannot effectively understand HHC.

[0005] Furthermore, as described above, techniques for attempting to understand HHC are very time consuming. In particular, such techniques involve processing of each message within a conversation session. Furthermore, it is difficult to extract and / or navigate to specific information within such a conversation session because such a session includes a message chain that includes information related to multiple topics. Moreover, such techniques generate too many notifications and undesired suggestions that may cause user anxiety and are highly undesirable.

[0006] Therefore, there is a need for a system that can process human-to-human conversations and identify semantic points in the conversations to effectively identify the intended conclusion of the conversation.

[0007] The above information is presented as background information only to assist with an understanding of the present disclosure. No determination has been made, and no assertion is made, as to whether any of the above might be applicable as prior art with respect to the present disclosure. Summary of the invention

[0008] Technical Solution

[0009] Aspects of the present disclosure are to at least address the above-mentioned problems and / or disadvantages, and to at least provide the advantages described below. Therefore, one aspect of the present disclosure is to provide a selection of concepts introduced in a simplified format, which are further described in the detailed description of the present disclosure. The present disclosure is neither intended to identify the key or essential inventive concepts of the present disclosure, nor to determine the scope of the present disclosure.

[0010] Additional aspects will be set forth in part in the description which follows and, in part, will be obvious from the description, or may be learned by practice of the presented embodiments.

[0011] According to one aspect of the present disclosure, a method for determining semantic points in a person-to-person conversation is provided. The method includes identifying a person-to-person conversation including a plurality of conversation turns on an electronic device. In addition, the method includes determining one or more natural language (NL) attributes for each of the plurality of conversation turns. In addition, the method includes deriving a transient state based on the one or more NL attributes for each conversation turn. In addition, the method includes deriving one or more conversation nuances associated with the person-to-person conversation based on the one or more NL attributes for each conversation turn. In addition, the method includes: after each conversation turn, dynamically storing information associated with the person-to-person conversation at one or more memories based on the one or more NL attributes, the transient state, and the one or more conversation nuances associated with each conversation turn. In addition, the method includes determining one or more semantic relationships and an associated conversation timeline within the person-to-person conversation based on the dynamically stored information. In addition, the method includes generating semantic points corresponding to the one or more semantic relationships and the associated conversation timeline within the determined person-to-person conversation.

[0012] According to another aspect of the present disclosure, a system for determining semantic points in a human-to-human conversation is provided. The system includes an identification module configured to identify a human-to-human conversation including a plurality of conversation turns on an electronic device. In addition, the system includes an NL attribute generator module configured to determine one or more natural language (NL) attributes for each of the plurality of conversation turns. In addition, the system includes a transient state estimator module configured to derive a transient state based on the one or more NL attributes for each conversation turn. In addition, the system includes a conversation nuance (CN) classifier module, which is configured to derive one or more conversation nuances associated with the human-to-human conversation based on the one or more NL attributes and one or more conversation turns for each conversation turn. In addition, the system includes a turn memory update module, which is configured to dynamically store information associated with the human-to-human conversation at one or more memories after each conversation turn based on the one or more NL attributes, the transient state, and the one or more conversation nuances associated with each conversation turn. In addition, the system includes a hierarchical semantic point module, which is configured to determine one or more semantic relationships and associated conversation timelines within a person-to-person conversation based on the dynamically stored information. The hierarchical semantic point module is also configured to generate semantic points corresponding to the one or more semantic relationships and associated conversation timelines within the determined person-to-person conversation.

[0013] Other aspects, advantages, and salient features of the present disclosure will become apparent to those skilled in the art from the following detailed description which, in conjunction with the accompanying drawings, discloses various embodiments of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] The above and other aspects, features and advantages of certain embodiments of the present disclosure will become more apparent through the following description in conjunction with the accompanying drawings, in which:

[0015] Figure 1 An environment of a system for determining semantic points in a conversation between people according to an embodiment of the present disclosure is shown;

[0016] Figure 2 A schematic block diagram of a system for determining semantic points in a conversation between people according to an embodiment of the present disclosure is shown;

[0017] Figure 3A A schematic block diagram showing modules and a session memory of a system for determining semantic points in a conversation between people according to an embodiment of the present disclosure;

[0018] Figure 3B A diagram showing a human-to-human conversation for determining semantic points according to an embodiment of the present disclosure;

[0019] Figure 3C A diagram showing a human-to-human conversation for determining semantic points according to an embodiment of the present disclosure;

[0020] Figure 4 An embodiment of a natural language (NL) attribute generator module according to an embodiment of the present disclosure is shown;

[0021] Figure 5 An embodiment of a transient state estimator module according to an embodiment of the present disclosure is shown;

[0022] Fig. 6A illustrates an embodiment of a conversation nuance (CN) classifier module according to various embodiments of the present disclosure;

[0023] Figure 6B illustrates an embodiment of a conversation nuance (CN) classifier module according to various embodiments of the present disclosure;

[0024] Fig. 7A An embodiment of a round memory update module according to various embodiments of the present disclosure is shown;

[0025] Figure 7B An embodiment of a round memory update module according to various embodiments of the present disclosure is shown;

[0026] Figure 7C An embodiment of a round memory update module according to various embodiments of the present disclosure is shown;

[0027] Fig.7D An embodiment of a round memory update module according to various embodiments of the present disclosure is shown;

[0028] Fig. 8A An embodiment of a hierarchical semantic point module according to various embodiments of the present disclosure is shown;

[0029] Figure 8B An embodiment of a hierarchical semantic point module according to various embodiments of the present disclosure is shown;

[0030] Figure 8C An embodiment of a hierarchical semantic point module according to various embodiments of the present disclosure is shown;

[0031] Fig.9A Various usage scenarios of a system for determining semantic points in a conversation between people according to various embodiments of the present disclosure are shown;

[0032] Fig. 9B Various usage scenarios of the system for determining semantic points in a conversation between people according to various embodiments of the present disclosure are shown;

[0033] Fig. 9C Various usage scenarios of the system for determining semantic points in a conversation between people according to various embodiments of the present disclosure are shown;

[0034] Fig.9D Various usage scenarios of the system for determining semantic points in a conversation between people according to various embodiments of the present disclosure are shown;

[0035] Fig.9E Various usage scenarios of the system for determining semantic points in a conversation between people according to various embodiments of the present disclosure are shown;

[0036] Fig.9F Various usage scenarios of the system for determining semantic points in a conversation between people according to various embodiments of the present disclosure are shown;

[0037] Figure 9G Various usage scenarios of the system for determining semantic points in a conversation between people according to various embodiments of the present disclosure are shown;

[0038] Figure 9H Various usage scenarios of the system for determining semantic points in a conversation between people according to various embodiments of the present disclosure are shown; and

[0039] Fig.10 The following is a processing flow for determining semantic points in a conversation between people according to an embodiment of the present disclosure.

[0040] Throughout the drawings, like reference numerals will be understood to refer to like parts, components and structures. DETAILED DESCRIPTION

[0041] The following description with reference to the accompanying drawings is provided to assist in a comprehensive understanding of the various embodiments of the present disclosure as defined by the claims and their equivalents. It includes various specific details to assist in understanding, but these details should be considered as merely exemplary. Therefore, it will be appreciated by those of ordinary skill in the art that various changes and modifications may be made to the various embodiments described herein without departing from the scope and spirit of the present disclosure. In addition, descriptions of well-known functions and configurations may be omitted for clarity and conciseness.

[0042] The terms and words used in the following description and claims are not limited to the bibliographical meanings, but are merely used by the inventor to enable a clear and consistent understanding of the present disclosure. Therefore, it will be apparent to those skilled in the art that the following description of various embodiments of the present disclosure is provided for illustrative purposes only and not for the purpose of limiting the present disclosure as defined by the appended claims and their equivalents.

[0043] It will be understood that singular forms include plural referents unless the context clearly dictates otherwise. Thus, for example, reference to "a component surface" includes reference to one or more of such surfaces.

[0044] References throughout this disclosure to "on one hand," "on the other hand," or similar language mean that a particular feature, structure, or characteristic described in conjunction with an embodiment is included in at least one embodiment of the present disclosure. Thus, appearances of the phrases "in an embodiment," "in another embodiment," and similar language throughout this disclosure may, but do not necessarily, all refer to the same embodiment.

[0045] The terms "comprises", "includes" or any other variation thereof are intended to cover a non-exclusive inclusion, such that a process or method that includes a series of steps includes not only those steps but may also include other steps not expressly listed or inherent to such process or method. Similarly, one or more devices or subsystems or elements or structures or components followed by "comprises ... one" do not exclude the presence of other devices or other subsystems or other elements or other structures or other components or additional devices or additional subsystems or additional elements or additional structures or additional components without more constraints.

[0046] The terms "multi-party session", "human-to-human session" and "session" are used interchangeably throughout the specification. The terms "user device", "device" and "electronic device" and their inherent variations are used interchangeably throughout the specification.

[0047] The present disclosure relates to a method and system for determining semantic points in a human-to-human conversation (HHC) based on conversation turns, natural language (NL) attributes, transient states, and one or more conversation nuances in the conversation.

[0048] Parameters such as conversation turns, NL attributes, transient states, and conversational nuances play an important role in identifying semantic points in HHC and determine the expected outline of the conversation.

[0049] In some embodiments, the methods and systems of the present disclosure may enable advanced features in group conversation applications, such as semantic search and display of important information, related conversation recommendations, tracking conversation nuances, and suggesting changes to conversation summaries.

[0050] Figure 1 An environment of a system for determining semantic points in a conversation between people according to an embodiment of the present disclosure is shown.

[0051] Reference Figure 1, which shows an environment 100 including a system 104 of a multi-party conversation 102 with multiple participants (i.e., user A, user B, user C, and user D). In one embodiment, the multi-party conversation 102 may be implemented on a user device associated with each participant via any suitable social media messaging platform. The user devices of the participants may include, but are not limited to, smartphones, tablets, laptops, personal computers, smart watches, smart TVs, IoT devices, and any other electronic devices configured to facilitate communication between users via a messaging platform. The multi-party conversation 102 refers to a person-to-person conversation in which the participants may be discussing one or more topics (e.g., meeting plans).

[0052] The multi-party conversation 102 may be received at the system 104 as input from the respective user devices of the users AD. In another embodiment, the system 104 may be an independent entity located at a remote location and connected to the user devices of the participants of the multi-party conversation 102 via any suitable network. For example, the system 104 may be located at a physical server ( Figure 1 In another embodiment, the system 104 may be implemented in a corresponding user device of one or more participants / users AD.

[0053] The system 104 may be configured to receive a multi-party conversation 102 as input, and process the multi-party conversation 102 to determine one or more semantic points in the multi-party conversation 102. The multi-party conversation 102 may include multiple conversation turns, such as user A's conversation "So where do we meet tonight?" can be considered as one conversation turn in the illustrated embodiment. The system 104 may also be configured to identify each conversation turn from the multi-party conversation 102. Thereafter, the system 104 may be configured to determine one or more natural language (NL) attributes for each of the multiple conversation turns of the multi-party conversation 102. NL attributes may be defined as building blocks (e.g., natural sentences) of a conversation turn. In an embodiment, the NL attributes may include grammatical components such as, but not limited to, verbs or nouns. In another embodiment, the NL attributes may include, but are not limited to, intent, conversation acts, named entities, and relationships between one or more NL attributes. Intent may indicate the purpose of the conversation. In the context of a conversational dialogue, a conversation act may be defined as an utterance that serves a function in the dialogue. The types of conversation acts may include questions, statements, or request actions. Named entities may refer to various nouns mentioned in a conversation turn. In another embodiment, any information extracted directly or indirectly from the natural language text can be attributed to NL attributes. Figure 1 In the illustrated embodiment, the NL attributes corresponding to the first conversation turn from user A (i.e., "So where do we meet tonight?") may include an intent such as "arrange a meeting" and a conversation act such as "request," as well as a named entity such as "user A."

[0054] Next, the system 104 may be configured to derive a transient state for each conversation turn and one or more conversation nuances associated with the multi-party conversation 102 based on the one or more NL attributes and the one or more conversation turns. In an embodiment, the transient state may refer to a state associated with each of the determined NL attributes. Based on the context of the multi-party conversation, the transient state may indicate a target memory to which the NL attribute must be converted. In some other embodiments, the transient state may also indicate the lifespan of the NL attribute. For example, the transient state includes, but is not limited to, confirmed, temporary, and ignored. Conversation nuances may be defined as categories or labels that reflect the level of uncertainty in a person's conversation. Examples of conversation nuances and / or labels may include, but are not limited to, requesting information, suggestions, optional suggestions, agreeing, rejecting, and concluding. For example, in Figure 1 In an illustrative embodiment of , for the first conversation turn, the instantaneous state of the action "arrange a meeting" can be "confirmed" and the conversational nuance can be "request information". Similarly, for the second conversation turn "Burger restaurant on 4th Street?" conducted by B, the instantaneous state of the conversation turn can be "provisional" and the conversational nuance can be "suggestion". In a similar manner, the system 104 can identify the instantaneous state and conversational nuance of each conversation turn in the multi-party conversation 102. Therefore, in an embodiment, deriving the instantaneous state for each conversation turn may include assigning one of a provisional label, a confirmation label, and an ignored label to each of the one or more NL attributes based on one or more conversation turns.

[0055] The system 104 may be configured to perform each of the above steps after each conversation turn in the multi-party conversation 102, and dynamically store information associated with the multi-party conversation 102 based on one or more NL attributes, transient states, and one or more conversation nuances associated with each conversation turn. Thereafter, the system 104 may be configured to determine one or more semantic relationships and associated conversation timelines within the multi-party conversation 102 based on the dynamically stored information. The system 104 may also be configured to generate semantic points corresponding to the determined one or more semantic relationships and associated conversation timelines.

[0056] Based on the information related to the semantic points, semantic relations, and the associated conversation timeline, the system 104 can implement features such as semantic search, aggregation, response suggestions, alert generation for the user in an effective and efficient manner.

[0057] For example, if user B searches for “plans for tonight,” the system 104 is configured to generate an output 106 showing important and relevant conversations 106a from user B within the multi-party conversation 102, relevant semantic points 106b of multiple conversation turns, and an overall conclusion 106c.

[0058] Furthermore, the illustrated embodiments are essentially and the system 104 may be implemented to minimize chat in a chat room and / or messaging platform, conversation suggestions based on high-level semantic points, navigation and extraction of specific information based on semantic point-based searches, or navigation to areas of interest in recorded video.

[0059] Further, the system 104 may be configured to enable a faster and easier chat consumption experience by enabling (one or more) users to easily track specific topics in the multi-party conversation 102, enabling users to easily find and navigate to specific pieces of information discussed in the multi-party conversation 102, and providing an easy-to-read compact view of the overall conversation with key information marked, while also providing the overall flow and timeline of the conversation.

[0060] In other embodiments, the system 104 may also be configured to provide reliable and intelligent artificial intelligence (AI) assistance to users by correctly understanding intent parameters from the multi-party conversation 102, providing appropriately timed proactive help in intent completion, and providing relevant “suggested replies” to reduce the required cognitive and user workload.

[0061] The system 104 may be configured to perform at least Figure 2 , Figure 3A , Figure 3B and Figure 3C One or more operations are described in detail to achieve the above technical advantages.

[0062] Figure 2 A schematic block diagram of a system for determining semantic points in a conversation between people according to an embodiment of the present disclosure is shown.

[0063] Reference Figure 2 , the system 201 can be used with Figure 1 . In another embodiment, the system 201 may be included in an electronic / user device associated with a user involved in a person-to-person conversation. In another embodiment, the system 201 may be configured to operate as a standalone device or system based on a server / cloud architecture that is communicatively coupled to an electronic device. Examples of electronic devices may include, but are not limited to, mobile phones, smart watches, laptop computers, desktop computers, personal computers (PCs), notebooks, tablet computers, mobile phones, and / or any other smart device configured to support person-to-person conversations via a messaging platform as discussed throughout this disclosure.

[0064] System 201 may be configured to receive and process human-to-human conversations to determine corresponding semantic points. System 201 may include a processor / controller 202 , an input / output (I / O) interface 204 , one or more modules 206 , a transceiver 208 , and a memory 210 .

[0065] In an embodiment, the processor / controller 202 may be operably coupled to each of the I / O interface 204, the module 206, the transceiver 208, and the memory 210. In one embodiment, the processor / controller 202 may include at least one data processor for executing processes in a virtual storage area network. The processor / controller 202 may include a dedicated processing unit, such as an integrated system (bus) controller, a memory management control unit, a floating point unit, a graphics processing unit, a digital signal processing unit, etc. In one embodiment, the processor / controller 202 may include a central processing unit (CPU), a graphics processing unit (GPU), or both. The processor / controller 202 may be one or more general-purpose processors, digital signal processors, application-specific integrated circuits, field programmable gate arrays, servers, networks, digital circuits, analog circuits, combinations thereof, or other devices now known or later developed for analyzing and processing data. The processor / controller 202 may execute a software program (such as manually generated (i.e., programmed) code) to perform the desired operation.

[0066] The processor / controller 202 may be configured to communicate with one or more input / output (I / O) devices via the I / O interface 204. The I / O interface 204 may employ communication code division multiple access (CDMA), high speed packet access (HSPA+), global system for mobile communications (GSM), long term evolution (LTE), world interoperability for microwave access (WiMax), etc.

[0067] Using I / O interface 204, system 201 can communicate with one or more I / O devices, specifically, user devices associated with a person-to-person session. For example, an input device can be an antenna, a microphone, a touch screen, a touch pad, a storage device, a transceiver, a video device / source, etc. An output device can be a printer, a fax machine, a video display (e.g., a cathode ray tube (CRT), a liquid crystal display (LCD), a light emitting diode (LED), a plasma, a plasma display panel (PDP), an organic light emitting diode display (OLED), etc.), an audio speaker, etc. In an embodiment, system 201 can use I / O interface 204 to communicate with an electronic device associated with a user.

[0068] The processor / controller 202 may be configured to communicate with a communication network via a network interface. In another embodiment, the network interface may be an I / O interface 204. The network interface may be connected to a communication network to enable the connection of the system 201 to an external environment and / or device / system. The network interface may use connection protocols, including but not limited to direct connection, Ethernet (e.g., twisted pair 10 / 100 / 1000BaseT), transmission control protocol / Internet protocol (TCP / IP), token ring, Institute of Electrical and Electronics Engineers (IEEE) 802.11a / b / g / n / x, etc. The communication network may include but is not limited to direct interconnection, local area network (LAN), wide area network (WAN), wireless network (e.g., using wireless application protocol), the Internet, etc. Using the network interface and the communication network, the voice assistant device can communicate with other devices. The network interface may use connection protocols, including but not limited to direct connection, Ethernet (e.g., twisted pair 10 / 100 / 1000BaseT), transmission control protocol / Internet protocol (TCP / IP), token ring, IEEE802.11a / b / g / n / x, etc.

[0069] In an embodiment, the processor / controller 202 may receive a person-to-person conversation from at least one user 218. In some embodiments where the system 201 is implemented as a standalone entity at a server / cloud architecture, the person-to-person conversation may be received from a user device associated with the user 218. Figure 2 Only one user 218 is depicted, but it is apparent that the system 201 may be configured to Figure 1 The plurality of user devices participating in the group conversation discussed receive the person-to-person conversation. The processor / controller 202 may execute a set of instructions on the received person-to-person conversation information to identify corresponding semantic points in the conversation. The processor / controller 202 may implement various technologies, such as but not limited to natural language processing (NLP), data extraction, artificial intelligence (AI), etc., to achieve the desired goal.

[0070] In some embodiments, the memory 210 is communicatively coupled to at least one processor / controller 202. The memory 210 may be configured to store data, instructions executable by at least one processor / controller 202. In one embodiment, the memory 210 may communicate via a bus within the system 201. The memory 210 may include, but is not limited to, non-transitory computer-readable storage media, such as various types of volatile and non-volatile storage media, wherein various types of volatile and non-volatile storage media include, but are not limited to, random access memory, read-only memory, programmable read-only memory, electrically programmable read-only memory, electrically erasable read-only memory, flash memory, tape or disk, optical media, etc. In one example, the memory 210 may include a cache or random access memory for the processor / controller 202. In an optional example, the memory 210 is separate from the processor / controller 202, such as a cache memory of a processor, a system memory, or other memory. The memory 210 may be an external storage device or a database for storing data. The memory 210 may be operable to store instructions executable by the processor / controller 202. The functions, actions, or tasks shown or described in the figures may be performed by a programmed processor / controller 202 to execute instructions stored in the memory 210. The functions, actions, or tasks are independent of the specific type of instruction set, storage medium, processor, or processing strategy, and may be performed by software, hardware, integrated circuits, firmware, microcode, etc., operating alone or in combination. Likewise, processing strategies may include multi-processing, multi-tasking, parallel processing, etc.

[0071] In some embodiments, module 206 may be included in memory 210. Memory 210 may also include database 212 to store data. One or more modules 206 may include a set of instructions that can be executed to cause system 201 to perform any one or more of the methods / processes disclosed herein. One or more modules 206 may be configured to use the data stored in database 212 to perform the steps of the present disclosure to determine the semantic points in the person-to-person conversation as discussed herein. In an embodiment, each of one or more modules 206 may be a hardware unit that may be external to memory 210. In addition, memory 210 may include an operating system 214 for performing one or more tasks of system 201 as performed by a general operating system in the communication domain. Memory 210 may also include a session memory 216 configured to store person-to-person conversations and associated parameters at each conversation turn. The associated parameters of the person-to-person conversation may include, but are not limited to, user preferences, intents, actions, transient states, conversation nuances, etc. associated with each conversation turn in the person-to-person conversation. The transceiver 208 may be configured to receive signals from and / or send signals to electronic devices associated with the user. In one embodiment, the database 212 may be configured to store information required by the one or more modules 206 and the processor / controller 202 to perform one or more functions for determining semantic points in a human-to-human conversation.

[0072] In an embodiment, I / O interface 204 may enable input to and output from system 201 using suitable devices such as, but not limited to, a display, keyboard, mouse, touch screen, microphone, speakers, and the like.

[0073] In addition, the present disclosure contemplates a computer-readable medium including instructions or receiving and executing instructions in response to a propagation signal. In addition, instructions can be sent or received via a network via a communication port or interface or using a bus (not shown). A communication port or interface can be a part of a processor / controller 202, or can be a separate component. A communication port can be created with software, or can be a physical connection in hardware. A communication port can be configured to be connected to any other component or combination thereof in a network, an external medium, a display, or a system. The connection to the network can be a physical connection (such as a wired Ethernet connection), or can be established wirelessly. Similarly, additional connections to other components of the system 201 can be physical or can be established wirelessly. The network can optionally be directly connected to the bus. For the sake of brevity, the architecture and standard operation of operating system 214, memory 210, database 212, processor / controller 202, transceiver 208, and I / O interface 204 are not discussed in detail.

[0074] Figure 3AA schematic block diagram showing modules of a system for determining semantic points in a conversation between people and a conversation memory according to an embodiment of the present disclosure is shown.

[0075] Figure 3B and Figure 3C A diagram showing a conversation between people for determining semantic points according to another embodiment of the present disclosure is shown.

[0076] Reference Figure 3A , Figure 3B and Figure 3C , showing the module 206 and session storage 216 involved in achieving the desired objectives of the present disclosure. Figure 3A The illustrated embodiment of FIG. 1 depicts a sequential flow of processing between modules 206 for determining semantic points in a conversation between people. Figure 3B and Figure 3C illustrate Figure 3A Module 206 may include, but is not limited to, a natural language (NL) attribute generator module 302, a transient state estimator module 304, a conversation nuance (CN) classifier module 306 (also referred to as “CN module 306”), a turn memory update module 308, and a hierarchical semantic point module 310. Module 206 may be implemented by suitable hardware and / or software applications.

[0077] In an embodiment, system 201 may be configured to receive a person-to-person conversation (HHC) and context information 301 as initial input. In another embodiment, HHC and context information 301 may be appropriately converted to text format before being input to system 201. For example, if a person-to-person conversation is performed in a media format (such as an audio or video format), the conversation may first be converted to a text format. The person-to-person conversation may be performed via any suitable platform (such as, but not limited to, a messaging device and / or a software mobile application). Context information 301 may include external information related to HHC. Context information may be received from a user device and / or an external source (such as, but not limited to, a social networking service, a web search engine, or a website). For example, context information includes information from user device applications (such as navigation applications, web browsers, messaging applications, sensors, and device contacts). Similarly, context information from an external source may include user profile information from a social networking service or a user's browser history that may be attributed to context. In an embodiment, HHC and context information 301 may be used as input through NL attribute generator module 302. The NL attribute generator module 302 may be configured to process the received HHC and context information 301 to identify each conversation turn in the received conversation and generate one or more NL attributes corresponding to each conversation turn. In some optional embodiments, the recognition module 300 may be configured to recognize a person-to-person conversation including multiple conversation turns on the electronic device. The NL attribute generator module 302 may be configured to implement any suitable technology, such as but not limited to natural language processing (NLP), natural language understanding (NLU) and natural language generation (NLG), artificial intelligence (AI), etc., to identify conversation turns and associated NL attributes from the HHC and context information 301.

[0078] In an embodiment, the NL attribute generator module 302 may be configured to perform topic determination to determine the previous topic and the current topic from the context information 301. Topics may include, but are not limited to, flight booking, meeting, etc. In addition, the NL attributes determined by the NL attribute generator module 302 may include the intention of the conversation, the behavior of the conversation, and the slot information of the conversation. The intention may refer to the reason why the conversation occurred. For example, when we talk to a friend to meet for dinner, the intention is to have dinner with him / her. Other examples of the intention of the conversation may include, but are not limited to, arranging a meeting, movie plans, travel plans, etc. The behavior of the conversation may classify the conversation turn in a given conversation to indicate whether the conversation turn is a request, response, proposal, confirmation, rejection, etc. A slot may refer to a named entity identified in a conversation turn. For example, Burger King is a slot in the conversation turn "Let's meet at Burger King". Other examples of slot information of a conversation may include a place, a point of interest, a date, a time, etc. In an embodiment, a conversation turn may be associated with a single behavior and may include one or more slot information.

[0079] Thereafter, the determined one or more NL attributes corresponding to each conversation turn may be passed through the transient state estimator module 304 and the CN classifier module 306. The transient state estimator module 304 may be configured to derive the transient state corresponding to each conversation turn based on the corresponding NL attributes. The transient state derived by the transient state estimator module 304 may include state information, such as but not limited to confirmed, temporary, and ignored. For example, the transient state estimator module 304 determines the transient state as "the meeting is confirmed" and "the location is temporary." The transient state estimator module 304 may correlate one or more NL attributes to derive the corresponding transient state. For example, for the slot information of the proposed behavior and location, the transient state estimator module 304 determines the corresponding transient state as the location is temporary.

[0080] The CN classifier module 306 may be configured to derive one or more conversation nuances associated with the conversation based on the corresponding NL attributes. The conversation nuances may include information such as, but not limited to, start, request information, suggestion, optional suggestion, agreement, agreement of optional suggestion, rejection, conclusion, etc. For example, for the behavior of confirmation, and the slot information indicating the point of interaction (POI) as a hamburger restaurant, the conversation nuance may be determined to be agreement. Therefore, the CN classifier module 306 may be configured to establish a relationship between one or more NL attributes to derive the corresponding conversation nuances. In addition, the conversation nuances may be classified into three broader categories: positive, negative, and neutral, each category indicating the intentions of various users regarding each conversation turn.

[0081] The derived instantaneous state and session nuances corresponding to the conversation turn may be passed through the turn memory update module 308. The turn memory update module 308 may be configured to dynamically update the user preference memory 312, the cache memory 314, and the final destination memory 316 based on the received instantaneous state and session nuances. The user preference memory 312, the cache memory 314, and the final destination memory 316 are part of the session memory 216. In an embodiment, the user preference memory 312 may correspond to the user, and the cache memory 314 and the final destination memory 316 may correspond to the subject of the conversation. In an embodiment, the turn memory update module 308 may be configured to transfer information between the cache memory 314 and the final destination memory 316 based on the session nuances associated with the conversation turn. Specifically, the three categories of CN may be responsible for converting information from the cache memory 314 to the final destination memory 316 and vice versa. In an embodiment, a neutral value of CN may indicate that the final destination memory 316 has not changed, a positive value of CN may indicate that the information should be moved from the cache memory 314 to the final destination memory 316, and a negative value of CN may indicate that the information should be moved from the final destination memory 316 to the cache memory 314. In another embodiment, the turn memory update module 308 may transfer information between the cache memory 314 and the final destination memory 316 based on the rules defined above. In addition, the turn memory update module 308 may be configured to update the user preference memory 312 with information corresponding to each of the conversation turns. In an embodiment, the turn memory update module 308 may be configured to update the cache memory 314 with information stored in the user preference memory 312 based on the instantaneous state associated with each conversation turn. In an embodiment, the turn memory update module 308 may also be configured to update the user preference memory 312, the cache memory 314, and the final destination memory 316 with the NL attribute. Thus, the turn memory update module 308 may be configured to dynamically store information associated with the person-to-person conversation after each conversation turn based on one or more NL attributes and transient states associated with each conversation turn in the conversation turn, and nuances associated with one or more conversation turns of the conversation. Furthermore, the turn memory update module 308 may be configured to create an update timeline 318 based on the dynamically stored information. The update timeline 318 may include update points associated with each conversation turn. Furthermore, the update point may include information associated with one or more NL attributes, transient states, conversation nuance labels, NL representations of the update points, and one or more changes. In an embodiment, the turn memory update module 308 may be configured to generate a conversation timeline based on one or more update points.

[0082] The update timeline 318 may be fed to the hierarchical semantic point module 310. In other embodiments, the information updated by the turn memory update module 308 may be passed through the hierarchical semantic point module 310. The hierarchical semantic point module 310 may be configured to determine one or more semantic relationships and associated conversation timelines within a person-to-person conversation based on dynamically stored information. In another embodiment, a transient state associated with each conversation turn may be used to identify a semantic relationship. In addition, the semantic relationship may indicate one or more semantic points and a corresponding conversation turn. In addition, the hierarchical semantic point module 310 may be configured to generate semantic points corresponding to the determined one or more semantic relationships and associated conversation timelines within the person-to-person conversation. In an embodiment, the semantic relationship may be used to generate a hierarchical semantic point (HSP). For example, in the process of generating an HSP, two or more semantic points (SPs) from the same or different levels of semantic points may be compared to establish common characteristics. The determined common characteristics may be referred to as semantic relationships. In an embodiment, the hierarchical semantic point module 310 may be configured to generate one or more hierarchical semantic points across multiple levels of a conversation based on the update points in the update timeline 318. For example, the update information from each conversation turn can be labeled as a level 0 semantic point (SP). Level 0 SPs can be combined to generate level 1 SPs. In addition, level 1 and level 0 SPs can be combined to generate level 2 SPs. In addition, level 0-2 SPs can be combined in pairs to generate level 3 SPs, etc. In addition, the hierarchical semantic point module 310 can be configured to combine the generated one or more hierarchical semantic points across multiple levels to generate more high-level semantic points. In addition, the hierarchical semantic point module 310 can be configured to associate one or more NL attributes based on the combined one or more semantic points to represent the high-level semantic point.

[0083] In some embodiments, the hierarchical semantic point module 310 may also be configured to determine the scope of the NL representation and the associated one or more conversation turns for the user of the HHC based on the generated semantic point. In an embodiment, the NL representation may refer to a result, that is, a summary associated with one or more conversation turns generated based on the generated semantic point. In addition, the "scope" may be associated with the number of conversation turns that can be used to determine the NL representation. In addition, the hierarchical semantic point module 310 may be configured to determine at least one conversation turn that directly contributes to the semantic point from one or more conversation turns of the person-to-person conversation based on the semantic point and one or more update points associated with the conversation turn. In addition, the hierarchical semantic point module 310 may be configured to determine a compressed version of one or more conversation turns based on the semantic point and at least one conversation turn. The compressed version is displayed on a user interface for the user of the person-to-person conversation.

[0084] Module 206 may be implemented by any suitable hardware and / or instruction set. Figure 3AThe sequential flow shown in is self-explanatory, and embodiments may include additions / omissions of steps as required. In some embodiments, one or more operations performed by module 206 may be performed by processor / controller 202 based on requirements.

[0085] Figure 3B and Figure 3C A diagram of a person-to-person conversation for determining semantic points is shown. Specifically, Figure 3B and Figure 3C An example of a human-to-human conversation including ten conversation turns from four users (i.e., user A, user B, user C, and user D) is shown. In addition, the various outputs corresponding to each conversation turn generated by each of the NL attribute generator module 302, the transient state estimator module 304, and the CN classifier module 306 are clearly disclosed in the table shown. Therefore, Figure 3B and Figure 3C The operation of the system 201 on a human-to-human conversation is clearly shown.

[0086] Figure 4 An embodiment of a natural language (NL) attribute generator module according to an embodiment of the present disclosure is shown.

[0087] Reference Figure 4 , the NL attribute generator module 402 can be used with Figure 3A The NL attribute generator module 302 shown corresponds. As shown, the NL attribute generator module 402 can receive multiple dialogue turns as part of a human-to-human conversation. An example of such a dialogue turn can be "Hamburger Coffee Shop on 4th Street". The NL attribute generator module 402 can process the received dialogue turns using techniques such as, but not limited to, classification based on machine learning (ML) and deep learning (DL), relationship extraction, dependency parsing, etc. to determine NL attributes. In addition, NL attributes may include intent, named entities, topics, and dialogue behaviors, etc. For example, in the illustrated embodiment for the dialogue "Hamburger Coffee Shop on 4th Street", the NL attribute generator module 402 may determine the intent as "arrange a meeting", the named entity as "POI: Hamburger Coffee Shop, Location: 4th Street", the topic as "Meeting, Intent: Arrange a Meeting", and the dialogue behavior as "Request". In general, HHC can start with an intent, and many times there may be several such sub-conversations with different intents. Therefore, the NL attribute generator module 402 can be configured to classify a set of dialogues within a session with a specific intent and identify corresponding NL attributes. In some embodiments, the output of the NL attribute generator module 402 can be provided as feedback input in the next determination of NL attributes.

[0088] Figure 5 An embodiment of a transient state estimator module according to an embodiment of the present disclosure is shown.

[0089] Reference Figure 5 , the transient state estimator module 502 can be used with Figure 3A The instantaneous state estimator module 502 corresponds to the instantaneous state estimator module 304 shown. The instantaneous state estimator module 502 may be configured to receive a plurality of conversation turns as part of the HHC and the NL attributes determined by the NL attribute generator modules 402 and 302. The instantaneous state estimator module 502 may also be configured to receive context information indicating the previous state of the conversation turn in the HHC. The instantaneous state estimator module 502 may be configured to process the received information using techniques such as, but not limited to, reinforcement learning, ML and DL classification, rule-based tables and various models to determine the instantaneous state of the conversation turn as confirmed and / or determined (fixed), temporary or ignored. In another embodiment, the confirmed and / or determined instantaneous state may indicate that the generated NL attribute is agreed to by the participants of the HHC, the temporary instantaneous state may indicate that the generated NL attribute is not agreed to by the participants, and the ignored instantaneous state may indicate that the generated NL attribute is not linked to the current conversation topic. In addition, the instantaneous state determined by the instantaneous state estimator module 304 may be used to update the user preference memory 312 and the cache memory 314.

[0090] Fig. 6A and Figure 6B An embodiment of a conversation nuance (CN) classifier module according to various embodiments of the present disclosure is shown.

[0091] Reference Fig. 6A and Figure 6B , the CN classifier module 602 can be used with Figure 3A The CN classifier module 602 corresponds to the CN classifier module 306 shown. The CN classifier module 602 may receive multiple conversation turns as part of the HHC and the NL attributes determined by the NL attribute generator modules 402 and 302. The CN classifier module 602 may also be configured to receive context information indicating a previous state or a previously identified conversation nuance level as input. Thereafter, the CN classifier module 602 may be configured to process the received information using techniques such as rule-based tables, ML and / or DL ​​classification, and various other models to determine a corresponding conversation nuance label for each conversation turn. In an embodiment, the conversation nuance label may indicate that the generated NL attribute was not agreed to by the participants of the conversation. Specifically, the conversation nuance may provide a high-level understanding of the user's intent regarding one or more NL attributes associated with the conversation turn of the HHC. In addition, the conversation nuance determined by the CN classifier module 602 may be used to update the cache memory 314.

[0092] In addition, conversation nuances may be broadly categorized into three categories, namely, positive CN, negative CN, and neutral CN. Positive CN may include conversation nuances such as agreement, disagreement, and agreement, positive emotions and feelings, etc. Negative CN may include conversation nuances such as disagreement, agreement, and disagreement, negative emotions and feelings, etc. In addition, neutral CN may include conversation nuances such as suggestions, questions, and requests. In an embodiment, positive CN may be responsible for sending information from cache memory 314 to final destination memory 316. Negative CN may be responsible for sending information from final destination memory 316 back to cache memory 314. Neutral CN may not change any information already stored in either cache memory 314 or final destination memory 316. However, neutral CN may be responsible for creating a new instance in cache memory 314. In some embodiments, conversation nuances may be combined with other emotional attributes of a person.

[0093] Additionally, the recognition of conversational nuances helps generate accurate understanding, summaries, and conversation suggestions. Figure 6B A table is shown with session nuances associated with different HHC sessions. Figure 6B Also shown is a component "score" which may also be generated by the CN classifier module 602 to efficiently generate accurate understanding, summaries, and conversation suggestions. The score associated with conversational nuance may be based on one or more NL attributes associated with a conversation turn.

[0094] Fig. 7A , Figure 7B , Figure 7C and Fig.7D An embodiment of a round memory update module according to various embodiments of the present disclosure is shown.

[0095] Reference Fig. 7A , the round memory update module 702 can be used as follows Figure 3AThe round memory update module 702 corresponds to the round memory update module 308 shown. The round memory update module 702 may use the session nuances, transient state, and NL attributes as input to update one or more memories. The round memory update module 702 may be configured to update the memory based on the received information using techniques such as, but not limited to, ML and / or DL ​​classification and / or various other models. In an embodiment, the round memory update module 702 may be configured to make a memory update decision based on the value of the transient state and the session nuance class label. The round memory update module 702 may be configured to manage the movement of NL attributes and other related information between different memories (i.e., the user preference memory 312, the cache memory 314, and the final destination memory 316). For example, the round memory update module 702 may be configured to use the transient state to update the NL attribute value in the user preference memory 312 and / or the cache memory 314. In addition, the round memory update module 702 may be configured to use the session nuance label information to manage the movement of information between the cache memory 314 and the final destination memory 316.

[0096] Reference Figure 7B , showing the information flow between different memories that can be managed by the round memory update module 702. Initially, the multi-party conversation and the associated NL attribute values ​​can be fed into the architectural pipeline shown as "model output (MO)". Thereafter, the round memory update module 702 can update the user preference memory (UPM) using the determined transient state. For example, if the transient state is "ignore", the MO value is unused. If the transient state is "temporary", the MO value is only updated to the UPM. If the transient state is "confirmed", the MO value is updated to both the UPM and the cache memory (CM). Thereafter, the round memory update module 702 can be configured to use the determined conversation nuances classified in three categories, namely positive CN, negative CN and neutral CN. For neutral type CN, the round memory update module 702 can leave the information in the CM without further processing. For positive type CN, the round memory update module 702 can be configured to move the round MO from the CM to the final destination memory (FGM). For negative type CN, the round memory update module 702 may move the round MO from the FGM back to the CM. In an embodiment, NL attribute values ​​may be represented as MOs, user-specific information may be stored in the UPM, unconfirmed information may be stored in the CM, and all agreed information may be stored in the FGM. In addition, any updates to the UPM, CM, and FGM are stored in an update history (UH) timeline (also referred to as an "update timeline") that can be used for semantic point determination.

[0097] Reference Figure 7C and Fig.7D, showing an example of a person-to-person conversation and the flow of information and / or NL attributes to different memories through the round storage update module 702. In addition, Figure 7C and Fig.7D Also shown is an update timeline generated based on the delivery flow of information in different memories. Specifically, Figure 7C and Fig.7D An HHC session between four participants and their corresponding conversation turns are shown.In an embodiment, a user preference store may be associated with each participant.

[0098] Fig. 8A , Figure 8B and Figure 8C An embodiment of a Hierarchical Semantic Point (HSP) module according to various embodiments of the present disclosure is shown.

[0099] Reference Fig. 8A , Figure 8B and Figure 8C , the HSP module 802 (interchangeably referred to as "semantic point module 802" and / or "HSP module 802") can be used with Figure 3A corresponds to the hierarchical semantic point module 310 shown in .

[0100] The HSP module 802 may receive update timeline information as input. The HSP module 802 may be configured to process the input information using techniques such as, but not limited to, ML and / or DL ​​classification, similarity detection, reasoning, and / or various other models to generate semantic points. The semantic points may be based on NL attribute similarity, dialog turn range, and conversation nuances.

[0101] In another embodiment, any updates to the final destination memory 316 may be stored in the update timeline 318. The HSP module 802 may be configured to perform a similarity check between each update in the update history timeline point stored in the update timeline 318. The HSP module 802 may also determine the similarity between NL attributes from the same semantic point and / or NL attributes from semantic points across different levels. In addition, the HSP module 802 may associate a conversation turn range with a semantic point for each determined similarity. In addition, in an embodiment, the NL attribute similarity results may be passed to the NL attribute generator module 302 to generate a description of the semantic point. The NL attribute generator module 302 may generate various variations of the semantic point description. The HSP module 802 may utilize an inference model to generate more advanced semantic points. The generated semantic points may be used for applications such as, but not limited to, compressed display of conversation turns, advanced understanding of search sessions, and the like. The HSP module 802 may also be configured to analyze various topics within the same HHC session or multiple HHC sessions. Figure 8B and Figure 8CThe determination and interaction of various key semantic points in the HHC session by the HSP module 802 are shown. In the illustrated embodiment, the HSP module 802 may determine three levels of semantic points, namely, level 0 SP, level 1 SP, and level 2 SP. However, such semantic points may be determined based on different points on the update timeline 318. In an embodiment, each conversation turn and associated information may be indicated as a level 0 SP. For example, "User A suggests arranging a meeting", "User B suggests a hamburger coffee shop on 4th Street.", etc., are considered to be level 0 SPs. Various information provided at the level 0 SP may be used to define a level 1 SP. For example, "Users A, B, and C agree to meet at a hamburger restaurant at 7 p.m." is considered to be a level 1 SP. In addition, information at level 0 SP and level 1 SP may be used to define information at level 2 SP. For example, "Users A, B, and C are interested in hamburger restaurants" is defined as a level 2 SP. In addition, the above illustrated embodiment with three SP levels is natural, and the HSP module 802 may determine any number of levels of semantic points based on requirements.

[0102] In addition, in an embodiment, in order to determine the HSP, two or more semantic points (SPs) from the same or different levels of semantic points may be compared to establish common characteristics. The determined common characteristics may be referred to as semantic relationships. For example, a level 1 SP "User B changed his mind and went to a hamburger restaurant" is determined from two level 0 SPs "User B suggested a hamburger coffee shop on 4th Street" and "User B agreed to the hamburger restaurant". In the two level 0 SPs, "User B" is a common actor, and the slot information of "POI" is changed from "hamburger coffee shop" to "hamburger restaurant". Therefore, both the combined actor and POI information give clues that user B agrees to the change of location, which corresponds to the semantic relationship between the semantic points under consideration.

[0103] Fig.9A , Fig. 9B , Fig. 9C , Fig.9D , Fig.9E , Fig.9F , Figure 9G and Figure 9H Various usage scenarios of a system for determining semantic points in a conversation between people according to various embodiments of the present disclosure are illustrated.

[0104] Fig.9A A scenario illustrating convenient topic tracking and improved glance-ability through semantic point search according to an embodiment of the present disclosure.

[0105] Reference Fig.9A, among the users (i.e., user A, user B, user C, and user D), a meeting discussion occurred in the user's friend chat group not long ago. Now, it is shown that user B wants to remember the details of the meeting plan (e.g., where, what time, who will join, etc.). However, the user is busy at work, and he does not have time to read all the messages and find information from them. In this scenario, on the day of the meeting, user B can simply search for "plans tonight", and the system 104, 201 can only show the conversation turns related to the determined semantic points. Specifically, in response to the user search, the system 104, 201 can only display the conversation turns that directly contribute to the ultimate purpose of searching for information on the user screen. In addition, the system 104, 201 can also highlight the semantic points, which can provide improved readability of the content. In addition, in response to the user search, the system 104, 201 can also provide a summary that can be generated based on the semantic point tracking. The generated summary can further enhance the reading of the expected conclusion.

[0106] Fig. 9B A scenario of using high-level semantic points to implement a compact view of a conversation according to an embodiment of the present disclosure is shown.

[0107] Reference Fig. 9B , showing a scenario where a user opens a chat group after a time gap and wants to see what progress has been made in the plan within a short time span. Therefore, to achieve this, the user simply pinches inward on the chat using the touch screen display, and the system 104, 201 can reduce the chat and display the semantic points and other important information. The action can be reversed by simply receiving a pinch out command from the user. In some embodiments, the user can further reduce the chat by performing the pinch inward action again, and the system 104, 201 can only display the semantic points. The system 104, 201 can use advanced semantic point generation to display the results to the user. Therefore, the system 104, 201 can provide the user with overall control of the conversation to quickly and easily read relevant information.

[0108] Fig. 9C A scenario of dialogue suggestion generation using high-level semantic points according to an embodiment of the present disclosure is shown.

[0109] Reference Fig. 9C, shows a scenario where users A, B, C, and D are discussing in a group chat about meeting at a restaurant and user D is currently in an office meeting and cannot check his phone. Therefore, in this scenario, the system 104, 201 can display the expected conclusion of the group chat to the user through the user's smart watch connected to the user's phone. The system 104, 201 can also display suggested replies based on the conclusion of the group chat. The system 104, 201 can use higher-level semantic points to display the expected conclusion and suggested replies. In further embodiments, the system 104, 201 can also learn and improve the suggested replies based on the user's response to the reply suggested by the system 104, 201 in the first instance.

[0110] Fig.9D A scenario of navigating and extracting specific information using a semantic point-based search according to an embodiment of the present disclosure is illustrated.

[0111] Reference Fig.9D , showing a scenario where a large number of chats have occurred since the last time the user opened the chat group. Now, the user wants to know whether his best friend "Claire" will attend the party tonight. But the user does not have the patience to read more than 40 messages. In this case, the system 104, 201 can use the semantic point related information to identify the specific conversation turn corresponding to the search. In the illustrated embodiment, the user search has been implemented by an AI assistant installed in the user's electronic device.

[0112] Fig.9E A scene of navigating to an area of ​​interest within a video using semantic points according to an embodiment of the present disclosure is shown.

[0113] Reference Fig. 9C , showing a scenario where a user wants to refer to a specific portion of a multi-party conversation video, where a specific member disagrees with another member on a specific topic. However, the user does not remember when it happened (in the video). In response to such a query by the user, the system 104, 201 can utilize the subtitle text associated with the video and determine the semantic points to provide the user with the desired result. Specifically, the system 104, 201 can utilize the semantic points associated with the HHC conversation in the video to highlight the specific portion of the video where a specific member disagrees with another member on a specific topic as desired by the user.

[0114] Fig.9F A scenario of reducing notification frequency by using a trigger based on a semantic point according to an embodiment of the present disclosure is shown.

[0115] Reference Fig.9F, showing a scenario where a user is discussing a trip with his "Forever Friends" chat group. But since the group has very frequent messages (and therefore notifications), the user wants to mute the group. At the same time, the user may also want to know if the location of the trip has been confirmed so that he can book his ticket as soon as possible. So, he can't mute the group either. In such a scenario, the user can simply request the system to notify when the travel location is confirmed on the group. The user can use an AI assistant installed in his user device to provide the request. The system 104, 201 can use semantic points to check "location confirmation". When the semantic points are determined, the system 104, 201 can trigger the desired notification. Therefore, the system 104, 201 can reduce the notification frequency and improve the notification relevance.

[0116] Figure 9G A scenario is shown in which relevant active prompts are implemented from an AI assistant through triggering based on semantic points according to an embodiment of the present disclosure.

[0117] Reference Figure 9G , shows a scenario in which users A, B, C, and D are discussing setting up a review meeting with stakeholders for a project they are working on in a work meeting. During the meeting, they may have discussed multiple points, including the agenda for the meeting, who needs to be invited, the date and time, etc. In such a scenario, the system 104, 201 can track the intent and related parameters associated with the users, including who has confirmed, who has not confirmed, etc. In addition, the system 104, 201 may be able to actively help the user create events related to the meeting with minimal effort using the tracked intent and related parameters. In some embodiments, the system 104, 201 can help an AI assistant installed in a user device to create such events.

[0118] Figure 9H A scenario of reliable AI assistance for intent completion using semantic point understanding according to an embodiment of the present disclosure is shown.

[0119] Reference Figure 9H , showing a scenario where users are discussing a meeting with their friends. 3 of the 4 participants agree on the location. At this point, the user may try to book a taxi via the AI ​​assistant. In this scenario, the system 104, 201 may call and help the AI ​​assistant display a message indicating one of the participants that has not yet confirmed, and whether the user still wants to book a taxi. Therefore, the system 104, 201 can make the AI ​​assistant more reliable.

[0120] Fig.10 The processing flow of a method for determining semantic points in a conversation between people according to an embodiment of the present disclosure is shown.

[0121] Reference Fig.10The steps of method 1000 may be performed by system 104, 201, and system 104, 201 may be integrated into a user's electronic device or provided separately.

[0122] At operation 1002 , method 1000 includes identifying, on an electronic device, a human-to-human conversation including a plurality of conversation turns.

[0123] At operation 1004, method 1000 includes determining, for each of the plurality of dialogue turns, one or more natural language (NL) attributes. The one or more natural language attributes include at least one of the following: an intent of one or more dialogue turns from a human-to-human conversation, a dialogue act, a named entity, and a relationship between the one or more natural language attributes.

[0124] At operation 1006, method 1000 includes: for each conversation turn, deriving a transient state based on the one or more NL attributes. In addition, deriving the transient state for each conversation turn may include: assigning one of a temporary label, a confirmed label, and an ignored label to each of the one or more NL attributes based on the one or more conversation turns of the human-to-human conversation.

[0125] At operation 1008, method 1000 includes: for each conversation turn, deriving one or more conversation nuances associated with the human-to-human conversation based on the one or more NL attributes. In addition, deriving the one or more conversation nuances for each conversation turn includes: generating one of agree, disagree, change of mind, optional proposal, and rejection for each conversation turn to model uncertainty in the human-to-human conversation.

[0126] At operation 1010, method 1000 includes dynamically storing information associated with a person-to-person conversation at one or more memories after each conversation turn based on one or more NL attributes, a momentary state, and one or more conversation nuances associated with each conversation turn. In another embodiment, the method includes dynamically updating stored information associated with one or more NL attributes of the person-to-person conversation at one or more memories after each conversation turn based on the momentary state and one or more conversation nuances associated with each conversation turn. The one or more memories include a user preference memory, a cache memory, and a final destination memory.

[0127] In an embodiment, dynamically updating the stored information includes transferring information between a cache memory and a final destination memory based on the conversation nuances. In another embodiment, dynamically updating the stored information includes updating information at a user preference memory based on a transient state associated with each conversation turn. In addition, the method includes dynamically updating the stored information associated with the person-to-person conversation at one or more memories after each conversation turn based on one or more NL attributes, the transient state, and one or more conversation nuance tags, and creating an update timeline based on the dynamic update of the stored information. The update timeline includes update points associated with each conversation turn. In another embodiment, each of the update points includes information associated with one or more NL attributes, the transient state, the conversation nuance tags, the NL representation of the update point, and one or more changes.

[0128] At operation 1012 , method 1000 includes determining one or more semantic relationships and an associated conversation timeline within a human-to-human conversation based on the dynamically stored information.

[0129] At operation 1014, method 1000 includes generating semantic points corresponding to the one or more semantic relationships within the determined human-to-human conversation and the associated conversation timeline. In an embodiment, the step of generating the semantic points may include: generating one or more hierarchical semantic points across multiple levels based on the update point in the update timeline, combining the one or more hierarchical semantic points across multiple levels to generate more high-level semantic points, and associating one or more NL changes based on the combined one or more high-level semantic points to represent the semantic points.

[0130] Although shown and described in a particular order Fig.10 The steps discussed above may occur in a variation of the order, but according to various embodiments, these steps may occur in a variation of the order.

[0131] The present disclosure provides various technical advances based on the key features discussed above. In addition, the present disclosure can achieve a faster and easier chat consumption experience by enabling users to easily track specific topics in a multi-party conversation. In addition, the present disclosure enables users to easily find and navigate to a specific piece of information discussed in a multi-party conversation, and provides an easily readable compact view of the entire conversation with key information highlighted.

[0132] The present disclosure may also enable reliable and intelligent artificial intelligence (AI) assistance to users by correctly understanding expected parameters from multi-party conversations.

[0133] While the present disclosure has been shown and described with reference to various embodiments thereof, it will be understood by those skilled in the art that various changes in form and details may be made therein without departing from the spirit and scope of the present disclosure as defined by the appended claims and their equivalents.

Claims

1. A method for determining a semantic point in a conversation between people, the method comprising: Recognizing 1002 a person-to-person conversation including a plurality of conversation turns on an electronic device; For each dialogue turn of the plurality of dialogue turns, determining 1004 one or more natural language NL attributes; For each dialogue turn, deriving 1006 a transient state based on the one or more NL attributes; For each conversation turn, deriving 1008 one or more conversation nuances associated with the human-to-human conversation based on the one or more NL attributes; dynamically storing 1010 information associated with the human-to-human conversation at one or more memories after each conversation turn based on the one or more NL attributes, the transient state, and the one or more conversation nuances associated with each conversation turn; determining 1012 one or more semantic relationships and associated conversation timelines within a person-to-person conversation based on the dynamically stored information; and One or more semantic points corresponding to the one or more semantic relationships and the associated conversation timeline within the determined human-to-human conversation are generated 1014 .

2. The method according to claim 1, wherein: The one or more NL attributes include at least one of: intents from the plurality of dialog turns of a human-to-human conversation, dialog acts, named entities, and relationships between the one or more NL attributes.

3. The method according to any one of claims 1 and 2, wherein: The instantaneous state derived for each dialogue turn includes: assigning one of a temporary label, a confirmed label, and an ignored label to each of the one or more NL attributes based on the plurality of dialogue turns of the human-to-human conversation.

4. The method according to claim 1, wherein: Deriving the one or more conversational nuances for each conversational turn includes generating one or more labels for each conversational turn to model a level of uncertainty in a human-to-human conversation.

5. The method according to one of claims 1 to 4, further comprising: dynamically updating, at the one or more memories, stored information associated with the one or more NL attributes of the human-to-human conversation after each conversation turn based on a momentary state associated with each conversation turn and the one or more conversation nuances, The one or more memories include a user preference memory 312 , a cache memory 314 , and a final destination memory 316 .

6. The method according to claim 5, wherein: Dynamically updating the stored information includes transferring information between the cache memory 314 and the final destination memory 316 based on the one or more session nuances.

7. The method according to one of claims 5 and 6, wherein: Dynamically updating the stored information includes updating the information at the user preference store 312 based on the instantaneous state associated with each conversation turn.

8. The method according to one of claims 5 to 7, further comprising: dynamically updating, after each conversation turn, stored information associated with the human-to-human conversation at the one or more memories based on the one or more NL attributes, the transient state, and the one or more conversation nuance tags; as well as Create an update timeline based on dynamic updates of stored information, The update timeline includes an update point associated with each dialogue turn.

9. The method according to claim 8, further comprising: generating the one or more semantic points based on update points in the update timeline; as well as The one or more semantic points are combined to generate one or more hierarchical semantic points.

10. The method according to one of claims 8 and 9, wherein: Each of the update points includes information associated with the one or more NL attributes, a transient state, a session nuance tag, a NL representation of the update point, and one or more changes.

11. The method according to one of claims 1 to 10, further comprising: A NL representation of a user of a human-to-human conversation and a scope of associated multiple conversation turns are determined based on the generated one or more semantic points.

12. The method according to one of claims 1 to 11, further comprising: Based on the one or more semantic points and the one or more update points associated with the conversation turns, determining at least one conversation turn that directly contributes to the one or more semantic points from the one or more conversation turns of the human-to-human conversation; as well as determining a compressed version of the one or more conversation turns based on the one or more semantic points and the at least one conversation turn, Wherein, the compressed version of the one or more conversation turns is displayed on a user interface of a user for a human-to-human conversation.

13. The method according to one of claims 1 to 12, further comprising: After each conversation turn, a conversation timeline is generated based on one or more update points associated with dynamically storing information associated with the human-to-human conversation.

14. A system for determining semantic points in a conversation between people, the system comprising: The recognition module 300 is configured to recognize a human-to-human conversation including a plurality of conversation turns on the electronic device; A natural language NL attribute generator module 302, configured to determine one or more NL attributes for each of the plurality of dialogue turns; a transient state estimator module 304 configured to derive a transient state based on the one or more NL attributes for each dialogue turn; A conversation nuance CN classifier module 306 , configured to derive, for each conversation turn, one or more conversation nuances associated with the human-to-human conversation based on the one or more NL attributes; a turn memory update module 308 configured to dynamically store information associated with the human-to-human conversation at one or more memories after each conversation turn based on the one or more NL attributes, the transient state, and the one or more conversation nuances associated with each conversation turn; as well as The HSP module 310 is configured to: Determining one or more semantic relationships and associated conversation timelines within a person-to-person conversation based on the dynamically stored information; Semantic points corresponding to the one or more semantic relationships and the associated conversation timeline within the determined human-to-human conversation are generated.

15. The system according to claim 14, in, The round memory update module 308 is configured to: dynamically updating, at the one or more memories, stored information associated with the one or more NL attributes of the human-to-human conversation after each conversation turn based on a momentary state associated with each conversation turn and the one or more conversation nuances; The one or more memories include a user preference memory 312 , a cache memory 314 , and a final destination memory 316 .