Conversational information acquisition method and apparatus, storage medium, and electronic device
Patent Information
- Application Number
- CN202211676142.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-26
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2042-12-26
AI Technical Summary
[0005]本申请实施例提供了一种对话信息的获取方法和装置、存储介质及电子装置,以至少解决相关技术中,无法结合不同模态信息确定用户的对话信息等问题
[0016]在本申请实施例中,对接收到的第一对话信息进行语义解析,确定所述第一对话信息对应的语义解析结果,其中,所述语义解析结果用于指示所述第一对话信息中是否存在具备目标词性的第一关键词和/或表征第一多媒体信息的第二关键词;根据所述语义解析结果确定与所述第一对话信息对应的目标范围内的历史对话信息;在所述历史对话信息中存在第一多媒体信息的情况下,对所述第一多媒体信息进行多媒体信息解析,得到多媒体信息解析结果;根据所述多媒体信息解析结果和所述第一对话信息确定所述第一对话信息对应的第二对话信息;采用上述技术方案,解决了无法结合不同模态信息确定用户的对话信息等问题,本发明实施例结合上下文信息和多模态多轮状态,帮忙用户进行跨模态的指代消解,实现多模态多轮对话。
Smart Images

Figure CN116795957B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of smart home technology, and more specifically, to a method and apparatus for acquiring dialogue information, a storage medium, and an electronic device. Background Technology
[0002] Current smart home dialogue systems are primarily built on a single-modal text architecture and cannot answer cross-modal questions such as those involving video or images. For example, if a user sends a photo of an air conditioner model 328 and asks the system, "What modes does this air conditioner have?", the current system cannot answer such multimodal questions because it cannot simultaneously combine information from different modalities.
[0003] In related technologies, text-based referential resolution methods are the primary approach. These methods utilize information such as the first keyword in the user's question, combined with the context text of the dialogue, to construct a series of deep learning models for resolving referential issues. For example, translation models directly translate the first keyword into its corresponding referential content in the preceding text during prediction, or they syntactically parse the preceding text and then construct rules to replace the first keyword of the current user's question with the corresponding noun or subject from the preceding text, thus resolving the referential issue. However, regardless of whether the implementation is based on deep learning models or rule-based methods such as syntactic parsing, it is based on a text-centric monomodal approach. Such systems cannot parse the image content information appearing in the preceding text, and therefore cannot achieve referential resolution.
[0004] There is still no effective solution to the problem that related technologies cannot combine different modal information to determine the user's dialogue information. Summary of the Invention
[0005] This application provides a method and apparatus for acquiring dialogue information, a storage medium, and an electronic device to at least solve the problem in related technologies that it is impossible to determine a user's dialogue information by combining different modal information.
[0006] According to one embodiment of this application, a method for obtaining dialogue information is provided, comprising: performing semantic parsing on received first dialogue information to determine a semantic parsing result corresponding to the first dialogue information, wherein the semantic parsing result is used to indicate whether there is a first keyword with a target part of speech and / or a second keyword representing first multimedia information in the first dialogue information; determining historical dialogue information within a target range corresponding to the first dialogue information based on the semantic parsing result; if the first multimedia information exists in the historical dialogue information, performing multimedia information parsing on the first multimedia information to obtain a multimedia information parsing result; and determining second dialogue information corresponding to the first dialogue information based on the multimedia information parsing result and the first dialogue information.
[0007] In an exemplary embodiment, determining historical dialogue information within a target range corresponding to the first dialogue information based on the semantic parsing result includes: determining historical dialogue information within a first range corresponding to the first dialogue information when the first keyword and / or the second keyword exist in the first dialogue information, wherein the target range includes the first range; and determining historical dialogue information within a second range corresponding to the first dialogue information when the first keyword exists in the first dialogue information but the second keyword does not exist, wherein the target range includes the second range, and the first range is larger than the second range.
[0008] In an exemplary embodiment, performing multimedia information parsing on the first multimedia information to obtain a multimedia information parsing result includes: determining whether there are keywords in the first dialogue information used to indicate a first object, wherein the first object includes at least one of the following: text, object, user; and if there are keywords in the first dialogue information used to indicate the first object, performing multimedia information parsing on the first multimedia information according to the parsing method corresponding to the first object to obtain a multimedia information parsing result.
[0009] In an exemplary embodiment, the multimedia information of the first multimedia information is parsed according to the parsing method corresponding to the first object to obtain a multimedia information parsing result, which includes at least one of the following: when the first object is text, the first multimedia information is parsed using a text recognition method to determine the text information in the first multimedia information, wherein the multimedia information parsing result includes: the text information; when the first object is an object, the first multimedia information is parsed using an object recognition method to determine the object information in the first multimedia information, wherein the multimedia information parsing result includes: the object information; when the first object is a user, the first multimedia information is parsed using a human body recognition method to determine the user information in the first multimedia information, wherein the multimedia information parsing result includes: the user information.
[0010] In an exemplary embodiment, performing multimedia information parsing on the first multimedia information to obtain a multimedia information parsing result includes: performing multimedia information parsing on the first multimedia information using a parsing method to obtain a second object in the first multimedia information and object information of the second object, wherein the parsing method includes at least one of the following: text recognition method, object recognition method, human body recognition method, and the second object includes at least one of the following: text, object, user, and the multimedia information parsing result includes the object information.
[0011] In an exemplary embodiment, determining the second dialogue information corresponding to the first dialogue information based on the multimedia information parsing result and the first dialogue information includes: determining the noun corresponding to the first keyword and / or the second keyword in the first dialogue information based on the multimedia information parsing result; replacing the first keyword and / or the second keyword in the first dialogue information with the noun, and obtaining the replaced first dialogue information as the second dialogue information.
[0012] In an exemplary embodiment, before performing multimedia information parsing on the first multimedia information to obtain the multimedia information parsing result, the method further includes: determining whether the first multimedia information exists in the historical dialogue information; if the first multimedia information does not exist in the historical dialogue information, obtaining third dialogue information input by a second object, wherein the input time of the third dialogue information is later than the input time of the first dialogue information; if the first multimedia information does not exist in the third dialogue information, sending a prompt message to the second object to indicate that the first dialogue information is incomplete.
[0013] According to another embodiment of this application, a device for acquiring dialogue information is also provided, comprising: a first parsing module, configured to perform semantic parsing on received first dialogue information and determine a semantic parsing result corresponding to the first dialogue information, wherein the semantic parsing result is used to indicate whether there is a first keyword with a target part of speech and / or a second keyword representing first multimedia information in the first dialogue information; a first determining module, configured to determine historical dialogue information within a target range corresponding to the first dialogue information based on the semantic parsing result; a second parsing module, configured to perform multimedia information parsing on the first multimedia information when the first multimedia information exists in the historical dialogue information, and obtain a multimedia information parsing result; and a second determining module, configured to determine second dialogue information corresponding to the first dialogue information based on the multimedia information parsing result and the first dialogue information.
[0014] According to another aspect of the embodiments of this application, a computer-readable storage medium is also provided, wherein a computer program is stored in the computer program, and the computer program is configured to execute the above-described method for obtaining dialogue information when it is run.
[0015] According to another aspect of the embodiments of this application, an electronic device is also provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the above-described method for obtaining dialogue information through the computer program.
[0016] In this embodiment, semantic parsing is performed on the received first dialogue information to determine the semantic parsing result corresponding to the first dialogue information. The semantic parsing result indicates whether the first dialogue information contains a first keyword with a target part of speech and / or a second keyword representing first multimedia information. Based on the semantic parsing result, historical dialogue information within the target range corresponding to the first dialogue information is determined. If the historical dialogue information contains first multimedia information, multimedia information parsing is performed on the first multimedia information to obtain a multimedia information parsing result. Based on the multimedia information parsing result and the first dialogue information, second dialogue information corresponding to the first dialogue information is determined. This technical solution solves the problem of being unable to determine user dialogue information by combining different modal information. This embodiment of the invention combines contextual information and multimodal multi-turn states to help users perform cross-modal referencing resolution and realize multimodal multi-turn dialogue. Attached Figure Description
[0017] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0018] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 This is a schematic diagram of the hardware environment for a method of obtaining dialogue information according to an embodiment of this application;
[0020] Figure 2 This is a flowchart of a method for obtaining dialogue information according to an embodiment of this application;
[0021] Figure 3 This is a schematic diagram of a method for obtaining dialogue information according to an embodiment of this application;
[0022] Figure 4 This is a structural block diagram (a) of a dialogue information acquisition device according to an embodiment of this application;
[0023] Figure 5 This is a structural block diagram (II) of a dialogue information acquisition device according to an embodiment of this application. Detailed Implementation
[0024] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.
[0025] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0026] According to one aspect of the embodiments of this application, a method for acquiring dialogue information is provided. This method for acquiring dialogue information is widely applicable to whole-house intelligent digital control application scenarios such as smart homes, smart home ecosystems, and intelligencehouse ecosystems. Optionally, in this embodiment, the above-mentioned method for acquiring dialogue information can be applied to, for example... Figure 1 The hardware environment shown consists of terminal device 102 and server 104. For example... Figure 1 As shown, server 104 is connected to terminal device 102 via a network and can be used to provide services (such as application services) to the terminal or clients installed on the terminal. A database can be set up on the server or independently of the server to provide data storage services for server 104. Cloud computing and / or edge computing services can be configured on the server or independently of the server to provide data processing services for server 104.
[0027] The aforementioned network may include, but is not limited to, at least one of the following: wired network, wireless network. The aforementioned wired network may include, but is not limited to, at least one of the following: wide area network, metropolitan area network, local area network. The aforementioned wireless network may include, but is not limited to, at least one of the following: Wi-Fi (Wireless Fidelity), Bluetooth. The terminal device 102 may not be limited to PC, mobile phone, tablet computer, smart air conditioner, smart range hood, smart refrigerator, smart oven, smart stove, smart washing machine, smart water heater, smart washing equipment, smart dishwasher, smart projector, smart TV, smart clothes rack, smart curtains, smart audio-visual equipment, smart socket, smart speaker, smart speaker box, smart fresh air equipment, smart kitchen and bathroom equipment, smart bathroom equipment, smart robot vacuum cleaner, smart window cleaning robot, smart mopping robot, smart air purifier, smart steam oven, smart microwave oven, smart water heater, smart air purifier, smart water dispenser, smart door lock, etc.
[0028] This embodiment provides a method for obtaining dialogue information, applied to a computer terminal. Figure 2 This is a flowchart of a method for obtaining dialogue information according to an embodiment of this application, which includes the following steps:
[0029] Step S202: Perform semantic parsing on the received first dialogue information to determine the semantic parsing result corresponding to the first dialogue information, wherein the semantic parsing result is used to indicate whether there is a first keyword with target part of speech and / or a second keyword representing the first multimedia information in the first dialogue information.
[0030] For example, the first dialogue information could be "Who is the little girl in this picture?" or "What are the functions of that air conditioner?"
[0031] It should be noted that when the first dialogue information is "Who is the little girl in this picture?", the first dialogue information contains both the first keyword and the second keyword; when the first dialogue information is "Who is this little girl?", the first dialogue information contains the first keyword but not the second keyword.
[0032] It should be noted that the first keyword can be understood as a demonstrative pronoun, which includes but is not limited to: this, that, this one, that one; the first keyword can be understood as a personal pronoun, which includes but is not limited to: you, I, he, it; the second keyword can be understood as a multimodal information keyword, which includes but is not limited to: image, photo, video, audio.
[0033] Step S204: Determine the historical dialogue information within the target range corresponding to the first dialogue information based on the semantic parsing result;
[0034] It should be noted that, in this embodiment of the invention, the historical dialogue record can be the most recent historical dialogue record with the first object, or the historical dialogue record of the most recent five rounds with the first object, or all historical dialogue records with the first object. This embodiment of the invention does not limit this.
[0035] Step S206: If the first multimedia information exists in the historical dialogue information, perform multimedia information parsing on the first multimedia information to obtain the multimedia information parsing result;
[0036] Step S208: Determine the second dialogue information corresponding to the first dialogue information based on the multimedia information parsing result and the first dialogue information.
[0037] Through the above steps, semantic parsing is performed on the received first dialogue information to determine the semantic parsing result corresponding to the first dialogue information. The semantic parsing result indicates whether there is a first keyword with a target part-of-speech attribute and / or a second keyword representing first multimedia information in the first dialogue information. Based on the semantic parsing result, historical dialogue information within the target range corresponding to the first dialogue information is determined. If first multimedia information exists in the historical dialogue information, multimedia information parsing is performed on the first multimedia information to obtain a multimedia information parsing result. Based on the multimedia information parsing result and the first dialogue information, second dialogue information corresponding to the first dialogue information is determined. This solves the problem in related technologies where it is impossible to combine different modal information to determine the user's dialogue information. This embodiment of the invention combines contextual information and multimodal multi-turn states to help users perform cross-modal referencing resolution and realize multimodal multi-turn dialogue.
[0038] In an exemplary embodiment, determining historical dialogue information within a target range corresponding to the first dialogue information based on the semantic parsing result includes: determining historical dialogue information within a first range corresponding to the first dialogue information when the first keyword and / or the second keyword exist in the first dialogue information, wherein the target range includes the first range; and determining historical dialogue information within a second range corresponding to the first dialogue information when the first keyword exists in the first dialogue information but the second keyword does not exist, wherein the target range includes the second range, and the first range is larger than the second range.
[0039] It should be noted that if the first keyword exists in the first dialogue information and the second keyword also exists, or if the first keyword does not exist in the first dialogue information but the second keyword does exist, it indicates that there is a high probability that the first multimedia information exists in the historical dialogue records with the first object. Therefore, a large range of historical dialogue records are obtained. If the first keyword exists in the first dialogue information but the second keyword does not exist, it indicates that there is a low probability that the first multimedia information exists in the historical dialogue records with the first object. Therefore, a small range of historical dialogue records are obtained.
[0040] It should be noted that the historical dialogue information within the second scope of the embodiments of the present invention can be the most recent historical dialogue record with the first object, and the historical dialogue information within the first scope can be the most recent five historical dialogue records with the first object, or it can be all historical dialogue records with the first object. The embodiments of the present invention do not limit this.
[0041] In an exemplary embodiment, performing multimedia information parsing on the first multimedia information to obtain a multimedia information parsing result includes: determining whether there are keywords in the first dialogue information used to indicate a first object, wherein the first object includes at least one of the following: text, object, user; and if there are keywords in the first dialogue information used to indicate the first object, performing multimedia information parsing on the first multimedia information according to the parsing method corresponding to the first object to obtain a multimedia information parsing result.
[0042] Specifically, the multimedia information is parsed according to the parsing method corresponding to the first object to obtain a multimedia information parsing result, which includes at least one of the following: when the first object is text, the multimedia information is parsed using a text recognition method to determine the text information in the multimedia information, wherein the multimedia information parsing result includes: the text information; when the first object is an object, the multimedia information is parsed using an object recognition method to determine the object information in the multimedia information, wherein the multimedia information parsing result includes: the object information; when the first object is a user, the multimedia information is parsed using a human body recognition method to determine the user information in the multimedia information, wherein the multimedia information parsing result includes: the user information.
[0043] It's important to note that image understanding encompasses a wide range of tasks, including object detection, image character recognition (OCR), face recognition, and human detection. Therefore, it's necessary to combine the part-of-speech tagging of the first keyword in the initial dialogue information to help narrow down the scope of image understanding. For example, if a user sends a photo of an air conditioner model 328 and asks the system, "What modes does this air conditioner have?", the word "this" modifies "air conditioner." Since an air conditioner is a household appliance, the image should undergo object recognition. Object detection identifies it as an air conditioner and identifies the model number as 328. This improves the efficiency of image recognition.
[0044] For example, if the first dialogue information is "What is the text in this picture and who said it?", it means that there is text in the first multimedia information. Therefore, the image is parsed by text recognition to obtain the text in the image.
[0045] In an exemplary embodiment, performing multimedia information parsing on the first multimedia information to obtain a multimedia information parsing result includes: performing multimedia information parsing on the first multimedia information using a parsing method to obtain a second object in the first multimedia information and object information of the second object, wherein the parsing method includes at least one of the following: text recognition method, object recognition method, human body recognition method, and the second object includes at least one of the following: text, object, user, and the multimedia information parsing result includes the object information.
[0046] In this embodiment of the invention, when a user sends a multimedia message such as an image or a video, the received multimedia message is parsed using detection methods such as object detection, image text recognition (OCR), face recognition, and human body detection. The parsing results are then stored in a database, and the detection results can be retrieved directly from the database if needed in the future.
[0047] In an exemplary embodiment, determining the second dialogue information corresponding to the first dialogue information based on the multimedia information parsing result and the first dialogue information includes: determining the noun corresponding to the first keyword and / or the second keyword in the first dialogue information based on the multimedia information parsing result; replacing the first keyword and / or the second keyword in the first dialogue information with the noun, and obtaining the replaced first dialogue information as the second dialogue information.
[0048] For example, a user sends a photo of an air conditioner model 328 and asks the system, "What modes does this air conditioner have?" The word "this" modifies "air conditioner." Since an air conditioner is a home appliance, the image should undergo object recognition. Object detection should identify that this is a home appliance, an air conditioner, and that the model number is 328. Therefore, the question "What modes does this air conditioner have?" should be replaced with "What modes does the model 328 air conditioner have?"
[0049] In an exemplary embodiment, before performing multimedia information parsing on the first multimedia information to obtain the multimedia information parsing result, the method further includes: determining whether the first multimedia information exists in the historical dialogue information; if the first multimedia information does not exist in the historical dialogue information, obtaining third dialogue information input by a second object, wherein the input time of the third dialogue information is later than the input time of the first dialogue information; if the first multimedia information does not exist in the third dialogue information, sending a prompt message to the second object to indicate that the first dialogue information is incomplete.
[0050] In other words, a user may send a dialogue message first, and then send a multimedia message. Therefore, if the first multimedia message is not present in the historical dialogue record, a third dialogue message that is later than the input time of the first dialogue message can be entered. Then, the noun corresponding to the first keyword can be determined based on the second multimedia message in the third dialogue message. If there is no multimedia message in either the third dialogue message or the historical dialogue message, a prompt message indicating that the first dialogue message is incomplete is sent to the second object, thereby enabling the second object to supplement the first dialogue message.
[0051] To better understand the process of obtaining the above-mentioned dialogue information, the implementation flow of obtaining the above-mentioned dialogue information will be described below in conjunction with optional embodiments, but this is not intended to limit the technical solution of the embodiments of this application.
[0052] This embodiment provides a method for obtaining dialogue information. Figure 3 This is a schematic diagram of a method for obtaining dialogue information according to an embodiment of this application, such as... Figure 3 As shown, the specific steps are as follows: Step S301: Obtain the text dialogue input by the user;
[0053] Step S302: Part-of-speech and syntactic parsing of the text dialogue;
[0054] Step S303: If the first keyword is present in the text dialogue, trigger the multimodal reference resolution function;
[0055] It should be noted that the first keyword includes: the indicative first keyword and the personal first keyword. The indicative first keyword includes, but is not limited to: this, that, this one, that one; the personal first keyword includes, but is not limited to: you, I, he, it; the second keyword includes, but is not limited to: picture, photo, etc. Of course, users may not say the keywords such as picture or photo, but directly send a photo first and then initiate a conversation, which will trigger the multimodal reference resolution function.
[0056] Step S304: If the multimodal reference resolution function is triggered, obtain the historical dialogue records;
[0057] It should be noted that the most recent conversation might contain an image, or it might be several rounds later. If there are no multimodal triggers and the most recent conversation was not an image, the user will be prompted that the conversation is incomplete. If multimodal triggers are present, five rounds of conversation history will be retrieved; if the user's current conversation does not contain a multimodal trigger, the most recent conversation history will be retrieved.
[0058] Step S305: If images exist in the historical dialogue records, perform image content understanding;
[0059] Because image understanding encompasses a wide range of functions, including object detection, image text recognition (OCR), face recognition, and human detection, it's necessary to combine this with the part-of-speech tagging of the primary keyword to be resolved in the current text dialogue to help narrow down the scope of image understanding. For example, if a user sends a photo of an air conditioner model 328 and asks the system, "What modes does this air conditioner have?", the word "this" modifies the air conditioner. Since an air conditioner is a type of home appliance, the image should undergo object recognition. Image understanding should identify that this is a home appliance, an air conditioner, and specifically, the model number 328. Alternatively, when a user sends an image, object detection, image text recognition (OCR), face recognition, and human detection can be used to analyze the image and then store it in the dialogue history.
[0060] Step S306: Replace the first keyword in the corresponding position of the text dialogue with the noun corresponding to the image based on the parsing results.
[0061] It should be noted that, considering syntactic structure and sentence fluency, word deduplication is also necessary. For example, the word "this" is deduplicated from "air conditioner" and the image-recognized air conditioner 328, and appropriate post-processing is performed. Finally, multimodal referencing resolution is completed.
[0062] This invention combines part-of-speech tagging, syntactic parsing, and triggering multimodal reference resolution based on a second keyword. Part-of-speech tagging and syntactic parsing narrow down the image content recognition range. Syntactic parsing is then used to perform post-processing on the final resolved sentence. This invention, by combining contextual information and multimodal multi-turn states, helps users perform cross-modal reference resolution, enabling multimodal multi-turn dialogue. Simultaneously, by incorporating grammatical knowledge within the context of vision, the scope of multimodal tasks such as image understanding is narrowed, improving image recognition efficiency. Since image recognition is slow and resource-intensive, combining grammatical information significantly improves efficiency. Triggering multimodal words or limiting the most recent turn to images reduces meaningless multimodal reference resolution triggers, improving system efficiency. Finally, post-processing based on grammatical rules makes the reference resolution results more coherent. This invention is less costly and easier to implement compared to pure deep learning methods.
[0063] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of this application.
[0064] Figure 4 This is a structural block diagram (a) of a dialogue information acquisition device according to an embodiment of this application; as shown... Figure 4 As shown, it includes:
[0065] The first parsing module 42 is used to perform semantic parsing on the received first dialogue information and determine the semantic parsing result corresponding to the first dialogue information. The semantic parsing result is used to indicate whether there is a first keyword with target part of speech and / or a second keyword representing the first multimedia information in the first dialogue information.
[0066] The first determining module 44 is used to determine historical dialogue information within the target range corresponding to the first dialogue information based on the semantic parsing result;
[0067] The second parsing module 46 is used to perform multimedia information parsing on the first multimedia information when the first multimedia information exists in the historical dialogue information, and to obtain the multimedia information parsing result.
[0068] The second determining module 48 is used to determine the second dialogue information corresponding to the first dialogue information based on the multimedia information parsing result and the first dialogue information.
[0069] The above-described device performs semantic parsing on the received first dialogue information to determine the semantic parsing result corresponding to the first dialogue information. The semantic parsing result indicates whether the first dialogue information contains a first keyword with a target part-of-speech attribute and / or a second keyword representing first multimedia information. Based on the semantic parsing result, historical dialogue information within the target range corresponding to the first dialogue information is determined. If the historical dialogue information contains first multimedia information, multimedia information parsing is performed on the first multimedia information to obtain a multimedia information parsing result. Based on the multimedia information parsing result and the first dialogue information, second dialogue information corresponding to the first dialogue information is determined. This solves the problem in related technologies where it is impossible to combine different modal information to determine the user's dialogue information. This embodiment of the invention combines contextual information and multimodal multi-turn states to help users perform cross-modal referencing resolution and realize multimodal multi-turn dialogue.
[0070] In an exemplary embodiment, a first determining module is configured to determine historical dialogue information within a first range corresponding to the first dialogue information when the first keyword and / or the second keyword are present in the first dialogue information, wherein the target range includes the first range; and to determine historical dialogue information within a second range corresponding to the first dialogue information when the first keyword is present in the first dialogue information but the second keyword is not present, wherein the target range includes the second range and the first range is greater than the second range.
[0071] In an exemplary embodiment, the second parsing module is configured to determine whether there are keywords in the first dialogue information that indicate a first object, wherein the first object includes at least one of the following: text, object, user; if there are keywords in the first dialogue information that indicate the first object, the first multimedia information is parsed according to the parsing method corresponding to the first object to obtain a multimedia information parsing result.
[0072] In one exemplary embodiment, the second parsing module is further configured to perform at least one of the following: when the first object is text, parsing the first multimedia information by means of text recognition to determine text information in the first multimedia information, wherein the multimedia information parsing result includes: the text information; when the first object is an object, parsing the first multimedia information by means of object recognition to determine object information in the first multimedia information, wherein the multimedia information parsing result includes: the object information; when the first object is a user, parsing the first multimedia information by means of human body recognition to determine user information in the first multimedia information, wherein the multimedia information parsing result includes: the user information.
[0073] In an exemplary embodiment, the second parsing module is configured to perform multimedia information parsing on the first multimedia information through a parsing method to obtain a second object in the first multimedia information and object information of the second object, wherein the parsing method includes at least one of the following: text recognition method, object recognition method, human body recognition method, and the second object includes at least one of the following: text, object, user, and the multimedia information parsing result includes the object information.
[0074] In an exemplary embodiment, the second determining module is configured to determine the noun corresponding to the first keyword and / or the second keyword in the first dialogue information based on the multimedia information parsing result; replace the first keyword and / or the second keyword in the first dialogue information with the noun, and obtain the replaced first dialogue information as the second dialogue information.
[0075] In one exemplary embodiment, Figure 5 This is a structural block diagram (II) of a dialogue information acquisition device according to an embodiment of this application; as shown Figure 5 As shown, the method includes: a sending module 52, wherein, before performing multimedia information parsing on the first multimedia information to obtain the multimedia information parsing result, the method further includes: a first determining module, used to determine whether the first multimedia information exists in the historical dialogue information; a first parsing module, used to obtain third dialogue information input by a second object when the first multimedia information does not exist in the historical dialogue information, wherein the input time of the third dialogue information is later than the input time of the first dialogue information; and a sending module 52, used to send a prompt message indicating that the first dialogue information is incomplete to the second object when the first multimedia information does not exist in the third dialogue information.
[0076] Embodiments of this application also provide a storage medium including a stored program, wherein the program executes any of the methods described above when it is run.
[0077] Optionally, in this embodiment, the storage medium may be configured to store program code for performing the following steps:
[0078] S1, perform semantic parsing on the received first dialogue information to determine the semantic parsing result corresponding to the first dialogue information, wherein the semantic parsing result is used to indicate whether there is a first keyword with target part of speech and / or a second keyword representing the first multimedia information in the first dialogue information;
[0079] S2, determine the historical dialogue information within the target range corresponding to the first dialogue information based on the semantic parsing result;
[0080] S3, if the first multimedia information exists in the historical dialogue information, perform multimedia information parsing on the first multimedia information to obtain the multimedia information parsing result;
[0081] S4, determine the second dialogue information corresponding to the first dialogue information based on the multimedia information parsing result and the first dialogue information.
[0082] Embodiments of this application also provide an electronic device including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform the steps in any of the above method embodiments.
[0083] Optionally, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor and the input / output device is connected to the processor.
[0084] Optionally, in this embodiment, the processor can be configured to perform the following steps via a computer program:
[0085] S1, perform semantic parsing on the received first dialogue information to determine the semantic parsing result corresponding to the first dialogue information, wherein the semantic parsing result is used to indicate whether there is a first keyword with target part of speech and / or a second keyword representing the first multimedia information in the first dialogue information;
[0086] S2, determine the historical dialogue information within the target range corresponding to the first dialogue information based on the semantic parsing result;
[0087] S3, if the first multimedia information exists in the historical dialogue information, perform multimedia information parsing on the first multimedia information to obtain the multimedia information parsing result;
[0088] S4, determine the second dialogue information corresponding to the first dialogue information based on the multimedia information parsing result and the first dialogue information.
[0089] Optionally, in this embodiment, the storage medium may include, but is not limited to, various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.
[0090] Optionally, specific examples in this embodiment can refer to the examples described in the above embodiments and optional implementations, and will not be repeated here.
[0091] Obviously, those skilled in the art should understand that the modules or steps of this application described above can be implemented using general-purpose computing devices. They can be centralized on a single computing device or distributed across a network of multiple computing devices. Optionally, they can be implemented using computer-executable program code, thereby storing them in a storage device for execution by a computing device. In some cases, the steps shown or described can be performed in a different order than those presented here, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, this application is not limited to any particular combination of hardware and software.
[0092] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. A method for acquiring dialogue information, characterized in that, include: Semantic parsing is performed on the received first dialogue information to determine the semantic parsing result corresponding to the first dialogue information, wherein the semantic parsing result is used to indicate whether there is a first keyword with target part of speech and / or a second keyword representing the first multimedia information in the first dialogue information; Based on the semantic parsing results, determine the historical dialogue information within the target range corresponding to the first dialogue information; If the first multimedia information exists in the historical dialogue information, the first multimedia information is parsed to obtain the multimedia information parsing result; Based on the multimedia information parsing result and the first dialogue information, determine the second dialogue information corresponding to the first dialogue information; The determination of historical dialogue information within the target range corresponding to the first dialogue information based on the semantic parsing result includes: If the first keyword and / or the second keyword exist in the first dialogue information, determine the historical dialogue information within a first range corresponding to the first dialogue information, wherein the target range includes: the first range, and the first keyword includes: an indicative first keyword and a personal first keyword; If the first keyword exists in the first dialogue information but the second keyword does not exist, determine the historical dialogue information within a second range corresponding to the first dialogue information, wherein the target range includes the second range and the first range is larger than the second range.
2. The method for obtaining dialogue information according to claim 1, characterized in that, The first multimedia information is parsed to obtain the multimedia information parsing result, including: Determine whether there are keywords in the first dialogue information used to indicate a first object, wherein the first object includes at least one of the following: text, object, user; If the first dialogue information contains keywords that indicate the first object, the first multimedia information is parsed according to the parsing method corresponding to the first object to obtain the multimedia information parsing result.
3. The method for obtaining dialogue information according to claim 2, characterized in that, The multimedia information is parsed according to the parsing method corresponding to the first object to obtain the multimedia information parsing result, which includes at least one of the following: When the first object is text, the first multimedia information is parsed by text recognition to determine the text information in the first multimedia information, wherein the multimedia information parsing result includes: the text information; When the first object is an object, the first multimedia information is parsed using an object recognition method to determine the object information in the first multimedia information, wherein the multimedia information parsing result includes: the object information; When the first object is a user, the first multimedia information is parsed using human body recognition to determine the user information in the first multimedia information, wherein the multimedia information parsing result includes the user information.
4. The method for acquiring dialogue information according to any one of claims 1-3, characterized in that, The first multimedia information is parsed to obtain the multimedia information parsing result, including: The first multimedia information is parsed using a parsing method to determine the second object in the first multimedia information and the object information of the second object. The parsing method includes at least one of the following: text recognition method, object recognition method, and human body recognition method. The second object includes at least one of the following: text, object, and user. The multimedia information parsing result includes the object information.
5. The method for acquiring dialogue information according to any one of claims 1-3, characterized in that, Determining the second dialogue information corresponding to the first dialogue information based on the multimedia information parsing result and the first dialogue information includes: Based on the multimedia information parsing results, determine the nouns corresponding to the first keyword and / or the second keyword in the first dialogue information; Replace the first keyword and / or the second keyword in the first dialogue information with the noun, and obtain the replaced first dialogue information as the second dialogue information.
6. The method for acquiring dialogue information according to any one of claims 1-3, characterized in that, Before performing multimedia information parsing on the first multimedia information to obtain the multimedia information parsing result, the method further includes: Determine whether the first multimedia information exists in the historical dialogue information; If the first multimedia information is not present in the historical dialogue information, the third dialogue information input by the second object is obtained, wherein the input time of the third dialogue information is later than the input time of the first dialogue information. If the first multimedia information is not present in the third dialogue information, a prompt message indicating that the first dialogue information is incomplete is sent to the second object.
7. A device for acquiring dialogue information, characterized in that, include: The first parsing module is used to perform semantic parsing on the received first dialogue information and determine the semantic parsing result corresponding to the first dialogue information. The semantic parsing result is used to indicate whether there is a first keyword with target part of speech and / or a second keyword representing the first multimedia information in the first dialogue information. The first determining module is used to determine historical dialogue information within the target range corresponding to the first dialogue information based on the semantic parsing result; The second parsing module is used to perform multimedia information parsing on the first multimedia information when the first multimedia information exists in the historical dialogue information, and to obtain the multimedia information parsing result. The second determining module is used to determine the second dialogue information corresponding to the first dialogue information based on the multimedia information parsing result and the first dialogue information. The first determining module is further configured to: determine historical dialogue information within a first range corresponding to the first dialogue information when the first keyword and / or the second keyword exist in the first dialogue information, wherein the target range includes the first range and the first keyword includes an indicative first keyword and a personal first keyword; and determine historical dialogue information within a second range corresponding to the first dialogue information when the first keyword exists in the first dialogue information but the second keyword does not exist, wherein the target range includes the second range and the first range is larger than the second range.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein the program, when executed, performs the method according to any one of claims 1 to 6.
9. An electronic device comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is configured to execute the method described in any one of claims 1 to 6 through the computer program.
Citation Information
Patent Citations
Man-machine interaction method and device based on artificial intelligence
CN106503156A
Text anaphora resolution method and device and medium
CN110968678A