Dialogue processing method and device, computer device and computer readable storage medium
By acquiring and analyzing multiple dialogue contents, extracting keywords and associating them with emotional intent, the problem of low accuracy in human-computer dialogue is solved, and the naturalness and fluency of the dialogue are improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- VOICEAI TECH CO LTD
- Filing Date
- 2023-08-03
- Publication Date
- 2026-05-12
AI Technical Summary
In existing technologies, the accuracy of human-computer dialogue is low, resulting in poor fluency and naturalness of the dialogue, making it difficult to effectively respond to users' chat content.
By acquiring at least two conversations with a user, analyzing their text content, extracting conversation keywords, and associating these keywords with the conversation content to determine the target text and emotional intent, a conversation can then be initiated.
It improves the accuracy of recognizing user chat content and emotions, and increases the naturalness and fluency of human-computer dialogue.
Smart Images

Figure CN117149965B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, specifically to a dialogue processing method, apparatus, computer device, and computer-readable storage medium. Background Technology
[0002] With the continuous development of artificial intelligence, humans and machines can now engage in fluent conversations. In related technologies, this dialogue is achieved by recognizing user chat content and determining the appropriate responses. However, current methods for recognizing user chat content suffer from low accuracy. Consequently, responses are often ineffective in addressing user concerns, resulting in poor naturalness in human-machine dialogue. Therefore, the low accuracy of chat content recognition and the lack of naturalness in human-machine dialogue lead to issues such as low fluency and poor dialogue quality in human-machine interactions. Summary of the Invention
[0003] This application provides a dialogue processing method, apparatus, computer device, and computer-readable storage medium, which can improve the fluency and effectiveness of human-computer dialogue.
[0004] In a first aspect, embodiments of this application disclose a dialogue processing method, the method comprising:
[0005] Obtain at least two conversations from the user;
[0006] Analyze the content of at least two dialogues to obtain the corresponding text content;
[0007] Extract dialogue keywords from the text content;
[0008] Based on the dialogue keywords, associate the at least two dialogue contents to determine the target text and the user's emotional intent;
[0009] Engage with the user based on the target text and the emotional intent.
[0010] Secondly, embodiments of this application disclose a dialogue processing apparatus, the dialogue processing apparatus comprising:
[0011] The acquisition unit is used to acquire at least two dialogue entries from the user.
[0012] The recognition unit is used to analyze the content of the at least two dialogues to obtain the corresponding text content;
[0013] The extraction unit is used to extract dialogue keywords from the text content;
[0014] The first determining unit is used to determine the target text and the user's emotional intent by associating the at least two dialogue contents with the dialogue keywords.
[0015] A dialogue unit is used to engage in dialogue with the user based on the target text and the emotional intent.
[0016] Thirdly, embodiments of this application disclose a computer device, which includes a processor and a memory. The memory stores a computer program, and the processor calls the computer program to implement the above-mentioned dialogue processing method.
[0017] Fourthly, embodiments of this application disclose a readable storage medium storing program code, which is invoked by a processor to implement the aforementioned dialogue processing method.
[0018] Fifthly, this application discloses a computer program product comprising computer program code, which, when executed by a processor, causes the aforementioned communication method to be performed.
[0019] In this embodiment, at least two dialogues with a user are acquired; the corresponding text content is obtained by analyzing the two dialogues; dialogue keywords are extracted from the text content; the dialogue keywords are associated with the two dialogues to determine the target text and the user's emotional intent; and a dialogue is conducted with the user based on the target text and the emotional intent. In this way, the technical solution of this application improves the accuracy of recognizing user chat content and emotions by associating dialogue keywords in at least two dialogues, and increases the naturalness of human-computer dialogue by conducting dialogue with the user based on the target text and emotional intent. This improves the fluency and effectiveness of human-computer dialogue. Attached Figure Description
[0020] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0021] Figure 1 This is a schematic diagram of the architecture of an interactive system disclosed in an embodiment of this application;
[0022] Figure 2 This is a flowchart illustrating a dialogue processing method disclosed in an embodiment of this application;
[0023] Figure 3 This is a schematic diagram of a dialogue content scenario disclosed in an embodiment of this application;
[0024] Figure 4 This is a schematic diagram of another dialogue content scenario disclosed in an embodiment of this application;
[0025] Figure 5 This is a flowchart illustrating another dialogue processing method disclosed in an embodiment of this application;
[0026] Figure 6 This is a schematic diagram of the structure of a dialogue processing device disclosed in an embodiment of this application;
[0027] Figure 7 This is a schematic diagram of the structure of a computer device disclosed in an embodiment of this application;
[0028] Figure 8 This is a schematic diagram of the structure of a computer-readable storage medium disclosed in an embodiment of this application. Detailed Implementation
[0029] The embodiments of this application are described in detail below. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain this application, and should not be construed as limiting this application.
[0030] To enable those skilled in the art to better understand the solutions of this application, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.
[0031] Currently, when implementing human-computer dialogue using natural language models (e.g., ChatGPT), the process typically involves identifying the user's emotional intent within a single segment or sentence of dialogue input, and then matching the corresponding response based on that intent. However, in human-computer dialogue, a single conversation can include multiple segments or sentences. Identifying the user's emotional intent from a single segment or sentence often fails to accurately convey the user's emotions and needs. Consequently, responses matched based on the user's emotional intent cannot provide an effective and natural response to the dialogue. Therefore, the low accuracy of chat content recognition and poor naturalness in human-computer dialogue result in low fluency and poor dialogue quality between humans and machines.
[0032] To address the aforementioned issues, this application's embodiments involve: acquiring at least two dialogue messages from a user; analyzing these two messages to obtain corresponding text content; extracting dialogue keywords from the text content; associating these keywords with the two dialogue messages to determine the target text and the user's emotional intent; and engaging in dialogue with the user based on the target text and emotional intent. In this way, the technical solution of this application improves the accuracy of recognizing user chat content and emotions by associating dialogue keywords from at least two dialogue messages, and by engaging in dialogue with the user based on the target text and emotional intent, increases the naturalness of human-computer dialogue. This, in turn, enhances the fluency and effectiveness of human-computer dialogue.
[0033] To better understand the embodiments of this application, the system architecture diagram will be introduced below.
[0034] Please see Figure 1 , Figure 1 This is a schematic diagram of the structure of a dialogue system 100 disclosed in an embodiment of this application. Figure 1 As shown, the dialogue system 100 may include a terminal device 110 and a server 120.
[0035] Terminal device 110 has a client corresponding to server 120 installed. Terminal device 110 may include a mobile device or a PC (personal computer) device. Terminal device 110 may include an input device and an output device. The input device may include a keyboard and a mouse. The input device can be used to receive user dialogue content. The output device may include a display screen or a monitor. The output device can be used to display the dialogue content input by the user. Terminal device 110 can receive user dialogue content input through the input device via the client; terminal device 110 can also return corresponding reply content to the output device via the client to realize dialogue with the user.
[0036] Server 120 can be used to execute the method of this application, specifically to analyze dialogue content of different dialogue types using different analysis modules to obtain text content corresponding to the dialogue content. It can also be used to associate dialogue keywords with human-centered content to determine the user's emotional intent. Furthermore, it can enable dialogue between humans and machines. Server 120 may also include a storage module for storing natural language processing models. This storage module can also store dialogue content of image, video, audio, and emoticon types, as well as the corresponding text content for each type of dialogue content. Server 120 can be connected to terminal device 110.
[0037] In one implementation, a user can input dialogue content on the client through the input device of terminal device 110. Server 120 analyzes the dialogue content using different analysis modules for different types of dialogue content to determine the target text and emotional intent. Based on the target text and emotional intent, it determines the corresponding response content. This response content is then returned to the client of terminal device 110 and displayed through the output device of terminal device 110.
[0038] It should be noted that, Figure 1 The system architecture shown is merely an example and does not constitute a limitation on the technical solutions disclosed in the embodiments of this application. As interactive system architectures evolve and new application scenarios emerge, the technical solutions disclosed in the embodiments of this application are also applicable to similar technical problems.
[0039] Please see Figure 2 , Figure 2 This is a schematic flowchart of a dialogue processing method disclosed in an embodiment of this application. The dialogue processing method may include the following steps.
[0040] 201. Obtain at least two conversations with the user.
[0041] This application can be applied to the process of users engaging in human-computer dialogue with natural language models (e.g., ChatGPT). The terminal device can acquire the user's dialogue content through an input device and can also send the user's dialogue content to a server. After the server acquires the user's dialogue content, it can obtain at least two dialogue messages sent by the user.
[0042] The user's dialogue content can be either the initial chat content sent by the system (e.g., a greeting triggered by specific conditions, a topic related to trending information, etc.) or the user's first reply. This first reply is a response from the user to the initial chat content sent by the system, and it is related to the initial chat content. The user's dialogue content can also be a second reply generated by the system based on the user's first reply (the system's reply content), and the user responds to that second reply (the third reply content). This third reply is a response from the user to the system's second reply, and it is related to the second reply content.
[0043] It should be noted that the human-computer dialogue is achieved based on the initial chat content (chat content actively sent by the system), the first reply content (user's first reply content), the second reply content (system's reply content), and the third reply content (user's second reply content) mentioned above. The first and third reply contents can each include at least two dialogues (at least two conversations), and correspondingly, the second reply content can also include at least two dialogues (at least two conversations). The above human-computer dialogue process is illustrative of the method described in this application, and the number of replies in the human-computer dialogue is not specifically limited here. The dialogue can end when the system no longer receives dialogue content from the user.
[0044] In some other implementations, the user's dialogue content can also be the initial chat content initiated by the user to the system. The initial chat content initiated by the user can be the dialogue content expressing the user's opinion or the dialogue content carrying a question. There is no specific limitation on the initial chat content initiated by the user.
[0045] It should be noted that a user's dialogue content (either actively initiated or a response to content sent by the system) can include multiple dialogues (multiple conversations). These multiple dialogues can be of the same type but express different emotional intentions from the user. These multiple dialogues representing different emotional intentions can reveal complex emotional intentions from the user.
[0046] In related technologies, when identifying user dialogue content, a one-to-one analysis of each dialogue segment from multiple dialogues is typically performed to determine the user's emotional intent corresponding to a specific dialogue. However, the emotional intent of the user corresponding to a specific dialogue is not entirely the same as the complex emotional intent of the user across multiple dialogue segments. Therefore, summarizing the user's complex emotional intent based on dialogue fragments leads to low accuracy in identifying the user's emotional intent.
[0047] To address the aforementioned technical issues, this embodiment acquires multiple dialogues from the user, correlates the corresponding text content within these dialogues to determine the target text the user actually needs to express within the multiple dialogues, and then determines the user's emotional intent based on this correlated target text. Determining the user's emotional intent by correlating the target text obtained from at least two of the user's multiple dialogues improves the accuracy of identifying the user's emotional intent. Accordingly, engaging in dialogue with the user based on this emotional intent enhances the fluency and effectiveness of the human-computer interaction.
[0048] For example, at least two dialogue messages sent by the user to the system can be:
[0049] [User 1]: I really love this architectural style!
[0050] [User 1]: But I don't know where this place is, sob sob~
[0051] [User 1]: Do you know where this is?
[0052] In the aforementioned multiple dialogues between users, each dialogue represents a different emotional intent. This application's embodiments correlate these multiple dialogues to determine the actual text content and emotional intent the user intends to express. This improves the accuracy of identifying the user's emotional intent, thereby enhancing the fluency of human-computer dialogue.
[0053] Furthermore, the at least two dialogues from the aforementioned user can be of different dialogue types. Dialogue types can be text, images, audio, video, or emojis. When identifying multiple dialogues of different types, it's common to overlook dialogues representing various emotional intentions of the user. Alternatively, multiple dialogue types can be individually identified to determine the user's emotional intention corresponding to a single dialogue. However, the above technical solutions, when determining a user's emotional intention based on multiple dialogues, suffer from low accuracy due to the one-sided nature of the dialogue content.
[0054] To address the aforementioned technical issues, this embodiment can also associate multiple dialogues based on the content of multiple dialogues of a corresponding dialogue type, and determine the user's emotional intent based on the associated target text, thereby improving the accuracy of identifying the user's emotional intent. Accordingly, dialogue is conducted with the user based on this emotional intent, improving the fluency and effectiveness of human-computer dialogue.
[0055] For example, at least two dialogue messages of at least one type sent by the user to the system may also be:
[0056] [User 2]: I really love this architectural style!
[0057] [User 2]: See attached Figure 3 , attached Figure 3 Example dialogue content for image type
[0058] [User 2]: But I don't know where this place is, sob sob~
[0059] [User 2]: See attached Figure 4 , attached Figure 4 Example dialogue content for emoticons
[0060] [User 2]: Do you know where this is?
[0061] Therefore, after identifying the dialogue type of at least two dialogue messages of at least one type sent by the user to the system, the corresponding dialogue type can be determined to be:
[0062] [User 2]: [text]
[0063] [User 2]: [picture]
[0064] [User 2]: [text]
[0065] [User 2]: [emoticon]
[0066] [User 2]: [text]
[0067] It should be noted that each dialogue type can be associated with a corresponding analysis module. Once the dialogue type of the dialogue content is determined, the analysis model corresponding to that dialogue type is used to analyze the dialogue content and obtain the corresponding text content.
[0068] 202. Analyze at least two dialogues to obtain the corresponding text content.
[0069] The user's dialogue content includes various different types. After obtaining at least two dialogue contents from the user, the server can identify the dialogue type of the dialogue content and determine the corresponding text content based on the analysis module corresponding to the dialogue type.
[0070] The user's dialogue content can be text-based, or it can be image-based and / or audio-based and / or video-based and / or emoji-based. After obtaining at least two dialogue messages containing at least one type from the user, the user's at least two dialogue messages of at least one type can be converted into corresponding text content.
[0071] It should be noted that for each type of dialogue content, a corresponding analysis module is associated. This analysis module can perform analysis methods for dialogue types different from the dialogue content itself. When dialogue content of various types appears in the user's dialogue content, the dialogue types of at least two dialogue pieces of at least one type can be identified. The analysis module corresponding to the dialogue type of the dialogue content is used to identify each dialogue piece input by the user, obtaining the text content corresponding to each dialogue piece.
[0072] In one implementation, when a user's dialogue content contains an image-type dialogue, an image analysis module is used to perform image analysis on the image-type dialogue content. The module identifies image information within the dialogue content and determines at least one target object within the image information. Based on the at least one target object in the image information and the target object's state information, the descriptive text corresponding to the image-type dialogue content is determined, thus obtaining the text content corresponding to the image-type dialogue content.
[0073] For example, when the user's conversation content includes... Figure 3 The dialogue content is shown as an image. Using a preset image analysis module to identify the image information in this dialogue content, it can be determined that the image information contains multiple target objects. For example: buildings, waterside, trees, people, etc. In addition, identifying the image information in the dialogue content can also determine the state information of multiple target objects. For example: antique buildings, murky water, lush trees, bustling crowds, etc. Therefore, based on at least one target object identified from the image information and the state information corresponding to at least one target object, the descriptive text of the image information can be determined, and the corresponding text content could be: This is a Chinese garden, built by the water, with many antique buildings, such as pavilions and archways, and many tourists.
[0074] In another implementation, when a user's dialogue content contains a video-type dialogue, a video analysis module is used to perform video analysis on the video-type dialogue content. The video information in the dialogue content is parsed, and at least one target object and at least one corresponding action information of the target object are identified. Based on the at least one target object and the action information of the target object, a descriptive text for the video-type dialogue content is determined, resulting in the text content corresponding to the video-type dialogue content.
[0075] In another implementation, when the user's dialogue content contains audio-type dialogue, an audio analysis module is used to perform audio analysis on the audio-type dialogue content. The audio information in the dialogue content is parsed and converted into corresponding text to obtain the text content corresponding to the video-type dialogue content. Specifically, the method for converting audio information into corresponding text information can be a speech recognition method based on an acoustic model.
[0076] In another implementation, when a user's dialogue content contains emoticons, it can be determined whether the emoticon content is a dynamic or static emoticon. If the dialogue content is a dynamic emoticon, a video analysis module is used to perform video analysis on the dynamic emoticon content to determine the corresponding text content; if the dialogue content is a static emoticon, an image analysis module is used to perform image analysis on the static emoticon content to determine the corresponding text content.
[0077] For example, when the user's conversation content includes... Figure 4 The example shows dialogue content with static emoticons. Using a preset image analysis module to identify the emoticon information within the dialogue content, the corresponding text content can be determined. Furthermore, by extracting the meaning expressed in the text content corresponding to the emoticon information, the user's emotion conveyed by this static emoticon type of dialogue content can be determined. For example, a kitten crying indicates sadness.
[0078] At least two dialogue entries of at least one type input by the user can be converted into corresponding text content.
[0079] For example, at least two dialogue messages of at least one type sent by a user to the system are analyzed using the analysis module corresponding to the dialogue type, and the corresponding text content can be:
[0080] [User 2]: I really love this architectural style!
[0081] [User 2]: This is a Chinese garden built by the water, with many antique buildings such as pavilions and archways, and it attracts a lot of tourists.
[0082] [User 2]: But I don't know where this place is, sob sob~
[0083] [User 2]: The kitten is crying, which means it is sad.
[0084] [User 2]: Do you know where this is?
[0085] 203. Extract dialogue keywords from the text content.
[0086] Based on the text content corresponding to different types of user dialogue, the server can extract dialogue keywords from the text content. It can also determine the user's emotional intent based on the dialogue keywords in the text content. It is understood that the text content can include multiple segments of text content corresponding to at least two dialogue pieces, and these multiple segments of text content can include different dialogue keywords. These different dialogue keywords can express the same or different emotional intents of the user. Therefore, multiple dialogue keywords can be extracted from the text content corresponding to at least two dialogue pieces of at least one type, and multiple emotional intents corresponding to multiple segments of text content can be determined based on each dialogue keyword.
[0087] Specifically, extracting dialogue keywords from multiple text segments can be done by determining the corresponding dialogue keywords based on the syntactic structure of the text content. For example, if the text content has a subject-verb-object syntactic structure, the object in this syntactic structure can be extracted as the dialogue keyword for that segment. This dialogue keyword with the object as the dialogue keyword can represent the referential information in the user's multiple text segments. Alternatively, the predicate in this syntactic structure can be extracted as the dialogue keyword for that segment. This dialogue keyword with the predicate as the dialogue keyword can represent the user's emotions.
[0088] For example, the dialogue keywords corresponding to the text content obtained from at least two dialogue messages of at least one type sent by the user to the system can be:
[0089] [User 2]: Architectural Style
[0090] [User 2]: Chinese Garden
[0091] [User 2]: Location, where
[0092] [User 2]: Sad
[0093] [User 2]: Where
[0094] 204. Based on the keywords in the dialogue, associate at least two dialogue contents to determine the target text and the user's emotional intent.
[0095] The server can also effectively associate text content based on dialogue keywords to integrate the text and obtain the corresponding target text. Based on the dialogue keywords, it can associate text content corresponding to at least two dialogue items of at least one type from the user, obtaining the associated target text. Specifically, the method for associating text content corresponding to at least two dialogue items to obtain the target text can be to determine whether there is referential information in the context of the text content based on the dialogue keywords. If obvious referential information exists, key information in the text content is extracted based on this referential information, and the key information in the text content is associated to obtain the target text.
[0096] The server can also determine the corresponding user's emotional intent based on the target text. This can be achieved through semantic understanding of the target text. Furthermore, it can be done by extracting user emotions from dialogue keywords and determining whether those emotions can be interpreted as the user's emotional intent based on the tone of the target text.
[0097] For example, in the above text content, the referential information in the context could be "architectural style" corresponding to "Chinese garden," and "place" corresponding to "Liujiang Ancient Town." The target text associated with the above dialogue keywords could be: I really like the architectural style of Chinese gardens, but I don't know where Liujiang Ancient Town is. The user's emotional intent determined based on this target text could be: sadness.
[0098] 205. Engage in dialogue with users based on the target text and emotional intent.
[0099] After determining the target text and the user's emotional intent in the dialogue, the server can generate corresponding response content based on this information. This response content can be multiple dialogue messages of various types. The server can then send this response content to a terminal device with the user's client installed. The user's client on the terminal device displays the response content through its output device, thus enabling a dialogue with the user.
[0100] exist Figure 2 In the described method embodiments, at least two dialogue contents from a user are acquired; the at least two dialogue contents are analyzed to obtain corresponding text content; dialogue keywords are extracted from the text content; the dialogue keywords are associated with at least two dialogue contents to determine the target text and the user's emotional intent; and a dialogue is conducted with the user based on the target text and emotional intent. In this way, the technical solution of this application improves the accuracy of recognizing user chat content and emotions by associating dialogue keywords in at least two dialogue contents, and increases the naturalness of human-computer dialogue by conducting dialogue with the user based on the target text and emotional intent. This improves the fluency and effectiveness of human-computer dialogue.
[0101] Please see Figure 5 , Figure 5 This is a schematic flowchart of another dialogue processing method disclosed in an embodiment of this application. The dialogue processing method may include the following steps.
[0102] 501. Obtain at least two conversations with the user.
[0103] For details on the implementation of step 501, please refer to the above. Figure 2 The description of step 201 in the corresponding embodiments will not be repeated here.
[0104] 502. Determine the dialogue type corresponding to at least two dialogue contents, associate the dialogue type with the corresponding analysis module, and the analysis module executes the corresponding analysis method.
[0105] At least two conversations between users can be of the same type or different types. Each conversation has a corresponding conversation type. It should be noted that conversation types can include text, images, videos, emoticons, and audio.
[0106] Therefore, after acquiring at least two dialogue pieces from the user, the dialogue type of each dialogue piece can be identified using a form recognition module. For example, when a dialogue piece is received, it can be represented by x, and the identified dialogue type can be represented by y, where S(x) = y, and y can be one of the following values: text, picture, video, audio, or emoticon. By determining the value of y, the dialogue type of the dialogue piece is determined.
[0107] 503. Use the analysis module corresponding to the dialogue type to analyze at least two dialogue contents and obtain the corresponding text content.
[0108] Text-based dialogue is the most common form of conversation and serves as the foundation for dialogue. Text-based dialogue can be semantically understood based on model recognition, allowing for the conversion of different dialogue types into corresponding text-based dialogue. A pre-defined model then performs semantic understanding based on this text-based dialogue, enabling human-computer dialogue.
[0109] Image or video dialogue content can be used to supplement text-based dialogue content. Image or video analysis modules can be used to identify image or video information within the dialogue content, determining the corresponding descriptive text. Emoticons can express user emotions, enhancing the emotional impact of text-based dialogue content. The corresponding user emotion can be determined by extracting the emotional meaning from emoticons. Audio dialogue content is the voice representation of text-based dialogue content, and can be directly converted into corresponding text content.
[0110] 504. Integrate the text content from the text and the dialogue content to obtain the integrated text content.
[0111] It should be noted that after analyzing at least two dialogues, the analysis can include both text-based dialogue content and the resulting text content. When there is overlap between the text content derived from dialogues of other types besides text-based dialogue content, the text-based dialogue content and the resulting text content need to be integrated to obtain the integrated text content.
[0112] 505. Query whether there is referential content in the context of the integrated text content, where the referential content is at least two keywords of the same category but different scopes.
[0113] Image or video dialogue can supplement text-based dialogue. By identifying the descriptive text corresponding to the image or video dialogue, and examining the modifiers, objects, and complements in the complete syntactic structure of the descriptive text, we can determine if there is referential content in the context of the integrated text. Emoticons can express user emotions and can enhance the emotional impact of text-based dialogue. We can extract emotional meanings from emoticons to identify dialogue keywords expressing user emotions. Audio dialogue is the auditory representation of text-based dialogue; it can be directly converted into corresponding text content, and dialogue keywords based on the syntactic structure of the text content can be extracted.
[0114] It should be noted that the referential content consists of at least two keywords within the same category but with different scopes. The referential content can include at least two keywords. The two keywords in the referential content must have a corresponding relationship and can be keywords within the same category but with different scopes. That is, the referential content can be two hierarchical concepts, and the two keywords in the referential content can strengthen the importance of the keywords in the dialogue. For example, location name 1 and the concrete location name 2 have a referential relationship. Within the context of the text content, referential content with a referential relationship exists in a corresponding manner and cannot exist independently.
[0115] 506. If they exist, extract the dialogue keywords from the integrated text content.
[0116] If present, extract the referential content from the integrated text content to obtain at least two corresponding dialogue keywords. When corresponding referential content exists in the text content, that corresponding referential content can be extracted as a dialogue keyword. Associate at least two text contents containing at least two dialogue keywords. Integrate at least two text contents based on at least two dialogue keywords to obtain the target text.
[0117] 507. Based on the keywords in the dialogue, associate at least two dialogue contents to determine the target text and the user's emotional intent.
[0118] 508. Engage in dialogue with the user based on the target text and the emotional intent.
[0119] The specific implementation process of steps 507 and 508 can be found above. Figure 2 The descriptions of steps 204 and 205 in the corresponding embodiments will not be repeated here.
[0120] exist Figure 5 In the described method embodiments, at least two dialogue contents from a user are acquired; the at least two dialogue contents are analyzed to obtain corresponding text content; dialogue keywords are extracted from the text content; the dialogue keywords are associated with at least two dialogue contents to determine the target text and the user's emotional intent; and a dialogue is conducted with the user based on the target text and emotional intent. In this way, the technical solution of this application improves the accuracy of recognizing user chat content and emotions by associating dialogue keywords in at least two dialogue contents, and increases the naturalness of human-computer dialogue by conducting dialogue with the user based on the target text and emotional intent. This improves the fluency and effectiveness of human-computer dialogue.
[0121] It should be understood that the same or corresponding information in the different embodiments described above can be referenced in relation to each other.
[0122] Please see Figure 6 , Figure 6 This is a schematic diagram of the structure of a dialogue processing device 600 disclosed in an embodiment of this application. Figure 6 As shown, the dialogue processing device 600 may include:
[0123] Acquisition unit 601 is used to acquire at least two dialogue contents from the user;
[0124] The recognition unit 602 is used to analyze the at least two dialogue contents to obtain the corresponding text content;
[0125] Extraction unit 603 is used to extract dialogue keywords from the text content;
[0126] The first determining unit 604 is used to determine the target text and the user's emotional intent by associating the at least two dialogue contents with the dialogue keywords.
[0127] Dialogue unit 605 is used to engage in dialogue with the user based on the target text and the emotional intent.
[0128] In some embodiments, the dialogue processing device 600 may further include:
[0129] The second determining unit 606 is used to determine the dialogue type corresponding to the at least two dialogue contents, the dialogue type is associated with the corresponding analysis module, and the analysis module executes the corresponding analysis method.
[0130] The recognition unit 602 is specifically used to analyze the content of the at least two dialogues using the analysis module corresponding to the dialogue type to obtain the corresponding text content.
[0131] In some embodiments, the identification unit 602 is specifically used for:
[0132] When the dialogue type is an image type, the image analysis module is used to identify the image information in the dialogue content in order to determine at least one target object in the image information;
[0133] The text content corresponding to the dialogue content is determined based on the at least one target object.
[0134] In some embodiments, the identification unit 602 is specifically used for:
[0135] When the dialogue type is video, the video analysis module is used to parse the video information in the dialogue content to determine at least one target object and at least one action information of the target object in the video information.
[0136] The text content corresponding to the dialogue content is determined based on the target object of the at least one and the action information of the at least one corresponding to the target object.
[0137] In some embodiments, the identification unit 602 is specifically used for:
[0138] When the dialogue type is audio, the audio analysis module is used to parse the audio information in the dialogue content to obtain the text content corresponding to the audio information.
[0139] In some embodiments, the dialogue processing device 600 may further include:
[0140] Integration unit 607 is used to integrate the text content and the text content in the dialogue content to obtain integrated text content;
[0141] The extraction unit 603 is specifically used for:
[0142] Query whether there is referential content in the context of the integrated text content, wherein the referential content is at least two keywords of the same category but different scopes;
[0143] If it exists, the referential content is extracted from the integrated text content to obtain at least two corresponding dialogue keywords.
[0144] In some embodiments, the first determining unit 605 is specifically used for:
[0145] Associate the at least two text contents containing the at least two dialogue keywords;
[0146] The target text is obtained by integrating the at least two text contents based on the at least two dialogue keywords;
[0147] Determine the user's emotional intent based on the target text.
[0148] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the above-described device and module can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0149] In several embodiments disclosed in this application, the coupling between modules can be electrical, mechanical, or other forms of coupling.
[0150] Furthermore, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module. The integrated modules described above can be implemented in hardware or as software functional modules.
[0151] like Figure 7 As shown in the figure, this application also discloses a structural schematic diagram of a computer device 700. The computer device 700 includes a processor 710 and a memory 720. The memory 720 stores computer program instructions. When the computer program instructions are invoked by the processor 710, they can actually execute the various method steps disclosed in the above embodiments. Those skilled in the art will understand that the structure of the computer device shown in the figures does not constitute a limitation on the computer device, and may include more or fewer components than shown, or combine certain components, or have different component arrangements. Wherein:
[0152] The processor 710 may include one or more processing cores. The processor 710 connects to various parts of the entire battery management system using various interfaces and lines. It executes instructions, programs, code sets, or instruction sets stored in the memory 720, calls data stored in the memory 720, performs various functions and processes data within the battery management system, and performs various functions and processes data within the computer device, thereby providing overall monitoring of the computer device. Optionally, the processor 710 may be implemented using at least one of the following hardware forms: Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), and Programmable Logic Array (PLA). The processor 710 may integrate one or more of the following: Central Processing Unit (CPU), Graphics Processing Unit (GPU), and modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the displayed content; and the modem handles wireless communication. It is understood that the modem may also not be integrated into the processor 710 and may be implemented separately using a communication chip.
[0153] The memory 720 may include random access memory (RAM) or read-only memory (ROM). The memory 720 may be used to store instructions, programs, code, code sets, or instruction sets. The memory 720 may include a program storage area and a data storage area. The program storage area may store instructions for implementing an operating system, instructions for implementing at least one function (such as touch functionality, sound playback functionality, image playback functionality, etc.), instructions for implementing the various method embodiments described below, etc. The data storage area may also store data created during the use of the computer device (such as phonebooks, audio and video data, chat log data, etc.). Accordingly, the memory 720 may also include a memory controller to expose processor 710's access to the memory 720.
[0154] Although not shown, the computer device 700 may also include a display unit, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 710 in the computer device loads the executable files corresponding to the processes of one or more applications into the memory 720 according to the following instructions, and the processor 710 runs the applications stored in the memory 720, thereby implementing the various method steps disclosed in the foregoing embodiments.
[0155] like Figure 8 As shown in the embodiments of this application, a computer-readable storage medium 800 is also disclosed, which stores computer program instructions 810, which can be called by a processor to execute the methods described in the above embodiments.
[0156] Computer-readable storage media can be electronic storage devices such as flash memory, EEPROM (Electrically Erasable Programmable Read-Only Memory), EPROM, hard disk, or ROM. Optionally, computer-readable storage media include non-volatile computer-readable storage medium. Computer-readable storage medium 800 has storage space for program code that performs any of the method steps described above. This program code can be read from or written to one or more computer program products. The program code can be compressed, for example, in a suitable form.
[0157] According to one aspect of this application, a computer program product or computer program is disclosed, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods disclosed in the various optional implementations of the above embodiments.
[0158] The above are merely preferred embodiments of this application and are not intended to limit this application in any way. Although this application has disclosed preferred embodiments as above, it is not intended to limit this application. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the technical solution of this application. Any simple modifications, equivalent changes and alterations made to the above embodiments based on the technical essence of this application without departing from the scope of the technical solution of this application shall still fall within the scope of the technical solution of this application.
Claims
1. A dialogue processing method, characterized in that, The method includes: Obtain at least two conversations from the user; Analyze the at least two dialogue contents to obtain corresponding text content, the text content including different dialogue keywords, the different dialogue keywords expressing the user's different emotional intentions; Extract dialogue keywords from the text content: Query whether there is referential content in the context of the text content. The referential content is at least two keywords of the same category but different scopes. If they exist, extract the referential content from the text content to obtain at least two corresponding dialogue keywords. Based on the dialogue keywords, associate the at least two dialogue contents to determine the target text and the user's emotional intent: associate the at least two text contents containing the at least two dialogue keywords, integrate the at least two text contents based on the at least two dialogue keywords to obtain the target text, and determine the user's emotional intent based on the target text; Engage with the user based on the target text and the emotional intent.
2. The method as described in claim 1, characterized in that, The method further includes: Determine the dialogue type corresponding to the at least two dialogue contents, associate the dialogue type with the corresponding analysis module, and the analysis module executes the corresponding analysis method; The analysis of the at least two dialogue contents yields the corresponding text content, including: The analysis module corresponding to the dialogue type is used to analyze the content of the at least two dialogues to obtain the corresponding text content.
3. The method as described in claim 2, characterized in that, When the dialogue type is image type, the step of using the analysis module corresponding to the dialogue type to analyze the at least two dialogue contents to obtain the corresponding text content includes: The image analysis module is used to identify image information in the dialogue content in order to determine at least one target object in the image information; The text content corresponding to the dialogue content is determined based on the at least one target object.
4. The method as described in claim 2, characterized in that, When the dialogue type is video, the step of using the analysis module corresponding to the dialogue type to analyze the at least two dialogue contents to obtain the corresponding text content includes: The video analysis module is used to parse the video information in the dialogue content to determine at least one target object in the video information and the action information of at least one target object. The text content corresponding to the dialogue content is determined based on the target object of the at least one and the action information of the at least one corresponding to the target object.
5. The method as described in claim 2, characterized in that, When the dialogue type is audio, the step of using the analysis module corresponding to the dialogue type to analyze the at least two dialogue contents to obtain the corresponding text content includes: The audio analysis module is used to parse the audio information in the dialogue content to obtain the text content corresponding to the audio information.
6. The method as described in claim 1, characterized in that, The method further includes: By integrating the text content from the text and the text content from the dialogue, the integrated text content is obtained. The system queries whether referential content exists in the context of the text content. The referential content is at least two keywords within the same category but different scopes. If it exists, the referential content is extracted from the text content to obtain at least two corresponding dialogue keywords, including: Query whether there is referential content in the context of the integrated text content, wherein the referential content is at least two keywords of the same category but different scopes; If it exists, the referential content is extracted from the integrated text content to obtain at least two corresponding dialogue keywords.
7. The method as described in claim 6, characterized in that, The step of determining the target text and the user's emotional intent by associating the at least two dialogue contents with the dialogue keywords includes: Associate the at least two text contents containing the at least two dialogue keywords; The target text is obtained by integrating the at least two text contents based on the at least two dialogue keywords; The user's emotional intent is determined based on the target text.
8. A dialogue processing device, characterized in that, The dialogue processing device includes: The acquisition unit is used to acquire at least two dialogue entries from the user. The identification unit is used to analyze the at least two dialogue contents to obtain corresponding text content, wherein the text content includes different dialogue keywords, and the different dialogue keywords express different emotional intentions of the user. The extraction unit is used to extract dialogue keywords from the text content: query whether there is referential content in the context of the text content, wherein the referential content is at least two keywords of the same category but different scopes; if they exist, extract the referential content from the text content to obtain at least two corresponding dialogue keywords. The first determining unit is configured to determine the target text and the user's emotional intent by associating the at least two dialogue contents with the dialogue keywords: associating the at least two text contents containing the at least two dialogue keywords, integrating the at least two text contents according to the at least two dialogue keywords to obtain the target text, and determining the user's emotional intent according to the target text. A dialogue unit is used to engage in dialogue with the user based on the target text and the emotional intent.
9. A computer device, characterized in that, It includes a memory and a processor, the memory storing a computer program, and the processor calling the computer program to implement the method as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program or computer instructions that, when executed by a processor, implement the method as described in any one of claims 1-7.