Question and answer method, apparatus, device, medium, and product

By acquiring the granularity of the question text information and outputting corresponding granularity of response information in the multimodal text-image question answering model, the problem that large multimodal models cannot recognize items of different granularities is solved, achieving a balance between the accuracy of image recognition and user needs.

CN122154895APending Publication Date: 2026-06-05BEIJING CO WHEELS TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING CO WHEELS TECH CO LTD
Filing Date
2024-12-03
Publication Date
2026-06-05

Smart Images

  • Figure CN122154895A_ABST
    Figure CN122154895A_ABST
Patent Text Reader

Abstract

A question and answer method, device, equipment, medium and product are disclosed. The method comprises: obtaining question text information; in response to the question granularity of the question text information being a target granularity, determining reply information corresponding to the target granularity; and outputting the reply information. The present application solves the technical problem that the accuracy of image recognition is low due to the inability to identify different granularities of objects in the image in the prior art, and achieves the identification of reply information corresponding to question text information of different question granularities, thereby ensuring the balance between recognition accuracy and precision.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of human-computer interaction technology, and in particular to a question-and-answer method, apparatus, device, medium, and product. Background Technology

[0002] In open-source multimodal image-text question answering systems, an end-to-end multimodal large model is typically adopted: the acquired images are directly input into the multimodal large model, and the multimodal large model directly provides the answer.

[0003] However, existing multimodal large models have low accuracy and cannot identify objects of different granularities in images, resulting in low image recognition accuracy. Summary of the Invention

[0004] This invention provides a question-and-answer method, apparatus, device, medium, and product to solve the technical problem of low image recognition accuracy caused by the inability to identify items of different granularities in images in the prior art.

[0005] According to one aspect of the present invention, a question-and-answer method is provided, comprising:

[0006] Obtain the question text information;

[0007] In response to the question granularity of the question text information being the target granularity, the corresponding response information is determined;

[0008] Output the response information.

[0009] According to another aspect of the present invention, a question-and-answer device is provided, comprising:

[0010] The first acquisition module is used to acquire the question text information;

[0011] The determination module is used to determine the response information corresponding to the target granularity in response to the question granularity of the question text information being the target granularity;

[0012] The output module is used to output the response information.

[0013] According to another aspect of the present invention, a question-and-answer device is provided, the question-and-answer device comprising:

[0014] At least one processor;

[0015] A memory communicatively connected to the at least one processor; and a central control display screen communicatively connected to the at least one processor; wherein,

[0016] The central control display screen is used to display the vehicle information interface;

[0017] The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the question-and-answer method according to any embodiment of the present invention.

[0018] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the question-and-answer method according to any embodiment of the present invention.

[0019] According to another aspect of the present invention, a computer program product is provided, the computer program product comprising a computer program that, when executed by a processor, implements the question-and-answer method described in any embodiment of the present invention.

[0020] The technical solution of this invention solves the technical problem in the prior art that the inability to identify items of different granularities in an image leads to low image recognition accuracy. It obtains question text information and determines the question granularity of the question text information; based on the question granularity of the question text information, it determines the corresponding response information for question text information of different question granularities, thereby ensuring a balance between recognition precision and accuracy.

[0021] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0022] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0023] Figure 1 This is a flowchart of a question-and-answer method provided in an embodiment of the present invention;

[0024] Figure 2 This is a flowchart of another question-and-answer method provided in an embodiment of the present invention;

[0025] Figure 3 This is a flowchart of another question-and-answer method provided in an embodiment of the present invention;

[0026] Figure 4 This is a flowchart of another question-and-answer method provided in an embodiment of the present invention;

[0027] Figure 5 This is a schematic diagram illustrating a classification of question granularity provided in an embodiment of the present invention;

[0028] Figure 6 This is a schematic diagram of the structure of a question-and-answer device provided in an embodiment of the present invention;

[0029] Figure 7 This is a structural block diagram of a question-and-answer device provided in an embodiment of the present invention. Detailed Implementation

[0030] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0031] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0032] In one embodiment, Figure 1 This is a flowchart of a question-and-answer method provided by an embodiment of the present invention. This embodiment is applicable to situations where users ask questions in the form of images and text in voice interaction scenarios. The method can be executed by a question-and-answer device, which can be implemented in hardware and / or software. The device can be configured in an in-vehicle system or in a client application of a smart terminal. The client application can be a standalone application or a third-party mini-program. For example, the smart terminal can be a smartphone, tablet, iPad, or other smart device with a built-in camera. Figure 1 As shown, the method includes:

[0033] S110. Obtain the question text information.

[0034] In one example, the question text information refers to the question asked by the user in text form. In practice, the question text information can be obtained in ways including, but not limited to, at least one of the following: text box input; speech-to-text input. In one example, the question text information can be obtained through text box input; specifically, the relevant text information to be asked can be directly entered through the text input interface on the smart device's interactive interface. Alternatively, the question text information can be obtained through speech-to-text input; specifically, the user's voice information can be collected by the smart device's voice acquisition device and converted into text format to obtain the question text information.

[0035] In one example, the question-and-answer method can be applied to an in-vehicle system. In this case, voice input can be performed using a voice acquisition device within the vehicle, or relevant text information can be manually entered on the vehicle's central control screen to obtain the question text information. In another example, the question-and-answer method can be applied to a client-side application of a smart device. In this case, voice input can be performed using a voice acquisition device configured within the smart device, or relevant text information can be manually entered on the smart device's interactive interface to obtain the question text information. Exemplarily, the voice acquisition device may include, but is not limited to, at least one of the following: a microphone array, a standalone microphone, or a microphone integrated into the device (e.g., a smartphone microphone, a laptop microphone, or a wearable device microphone).

[0036] S120. In response to the question granularity of the question text information as the target granularity, determine the reply information corresponding to the target granularity.

[0037] In one example, question granularity characterizes the level of detail or breadth of the question posed by the user; target granularity characterizes the granularity of the question text. For example, target granularity can include first granularity or second granularity. In terms of answer precision, first granularity is more refined than second granularity. For instance, first granularity refers to fine granularity, which involves a more detailed and precise division and description of things, events, or concepts, such as the type of plant (what tree, what flower), or the name of a celebrity; while second granularity refers to coarse granularity, the opposite of fine granularity, typically referring to broad categories of items, animals, and plants, such as dog, elephant, rose, etc. For instance, if the question text is "What animal is this?", then the question granularity is coarse granularity; if the question text is "What specific breed of animal is this?", then the question granularity is fine granularity.

[0038] In one example, semantic analysis can be performed on the question text to determine its granularity. Specifically, keywords in the question text can be identified and extracted, and then analyzed. If the keywords are relatively abstract and broad, the granularity of the question text is determined to be second-level; if the keywords are relatively specific and clear, the granularity is determined to be first-level. By performing semantic analysis on the question text to determine its granularity, the accuracy of understanding the question text can be improved, thereby ensuring the accuracy of determining the granularity of the question itself.

[0039] In one example, a pre-created text classification model can be used to classify the granularity of the question text information. That is, the question text information is input into the pre-created text classification model to directly obtain the granularity of the question text information. By determining the granularity of the question text information through the pre-created target text classification model, the classification efficiency and accuracy of the granularity of the question text information can be improved.

[0040] In one example, pre-created text classification rules can be used to classify the granularity of the question text information. Specifically, the text classification rules can be judged from the following different dimensions: the scope of the topic to which the question text information belongs, and / or the specificity of the question. In one example, when judging the granularity of the question text information based on its scope of the topic, the granularity can be determined by whether the question text information involves a broad field. If the question text information involves a broad field, the granularity of the question text information is determined to be the second granularity; if the question text information involves specific content within a specific topic field, the granularity of the question text information is determined to be the first granularity. By only judging the scope of the topic to which the question text information belongs, the granularity of the question text information can be determined, thereby ensuring the accuracy of the judgment of the granularity of the question and reducing the amount of data computation. In one example, when judging the granularity of a question based on its specificity, we can consider whether the question is general or specific. If the question is general, the granularity is determined to be second-level; if it is specific, the granularity is determined to be first-level. This simple determination of the question's specificity reduces computational burden while ensuring accuracy. Alternatively, in another example, the granularity can be determined by combining the scope of the question's topic and its specificity to ensure accuracy.

[0041] In one example, the response information refers to the specific content of the answer to the question text. The level of detail in the response information is related to the granularity of the question text. When the question text has a first-level granularity, the corresponding response information is relatively more detailed; when the question text has a second-level granularity, the corresponding response information is relatively coarser. This can be understood as the response information associated with a first-level granularity question text containing more information than the response information associated with a second-level granularity question text.

[0042] In one example, the response information corresponding to the target granularity can be determined based on the question text information. This includes inputting the question text information into a pre-created multimodal graph-text question-answering model to obtain the response information corresponding to the target granularity. In one example, the target granularity can be either a first granularity or a second granularity.

[0043] In one example, when the target granularity is the first granularity, image recognition information of the original image can be determined using an image recognition method associated with the first granularity. The original image and its image recognition information, along with the question text, are then input into a pre-created multimodal graph-text question-answering model to obtain a more detailed response. In another example, when the target granularity is the second granularity, the question text can be directly input into the pre-created multimodal graph-text question-answering model to obtain a coarser response. This allows for a more accurate understanding of the user's question intent, enabling the output of responses with different levels of detail based on user needs, thus satisfying the user's personalized questioning requirements.

[0044] In one example, the question text can be input into a Large Language Model (LLM) to understand the question text and output a response that matches it. Since the input parameter for an LLM model needs to be text containing words, sentences, or paragraphs, image acquisition is unnecessary when using an LLM model to determine the response to a question. In other words, an LLM model can be used to determine the response when only the question text needs to be input.

[0045] S130, Output reply information.

[0046] In one example, the response information can be output in at least one of the following ways: text display and voice output. In one example, when using text display, the response information can be directly displayed on the smart device's interface, allowing users to view it intuitively. In another example, when using voice output, the response information can be output through the smart device's voice playback module, allowing users to receive the response through sound, eliminating the cumbersome process of viewing it on the interface and ensuring vehicle safety during driving. In yet another example, both text display and voice output can be used simultaneously, allowing users to view the response intuitively on the interface while also receiving it promptly through sound.

[0047] The technical solution of this embodiment obtains the question text information and determines the question granularity of the question text information; based on the question granularity of the question text information, it determines the corresponding response information for the question text information at different question granularities, thereby solving the technical problem in the prior art that the inability to identify items of different granularities in an image leads to low image recognition accuracy. It realizes the recognition of response information corresponding to question text information at different question granularities, thereby ensuring a balance between recognition precision and accuracy.

[0048] In an optional embodiment, during actual operation, since users may not fully understand the question-and-answer rules of the question-and-answer system, there may be situations where users want detailed reply information but cannot provide detailed question text information.

[0049] In one optional embodiment, the method for determining the granularity of the question text information includes one of the following:

[0050] Determine the granularity of the question text information based on the question text information;

[0051] The granularity of the question text information is determined based on the question text information and the associated scenario information;

[0052] The granularity of the question text information is determined based on the question text information and the associated user attribute information.

[0053] The granularity of the question text information is determined based on the question text information, associated user attribute information, and context information.

[0054] In one example, the granularity of the question text information can be determined based on the question text information and the context information. In this example, the context information refers to information related to the specific physical environment in which the question occurs and information related to the application scenario. The specific physical environment information can refer to the user's geographical location, the device information used, and the question time information. For example, a map location can include: country, city, and specific address. The device information used refers to the type of device the user is using, such as an in-vehicle system, or a terminal device such as a smartphone or tablet. The question time information refers to the specific time the question occurs, which can also be understood as the time the question text information is obtained. For example, the context information can include, but is not limited to, one of the following: a botanical garden, a museum, or an outdoor road. Here, an outdoor road can refer to a driveway, sidewalk, etc., that is not a specific location.

[0055] In one example, the methods for obtaining scene information may include, but are not limited to, device positioning, image recognition, and text recognition. For instance, in a question-answering method applied to an in-vehicle system, scene information associated with the question text can be obtained through the vehicle's Global Positioning System (GPS) positioning function; alternatively, scene information associated with the question text can be determined through environmental image information contained in the original images captured by the vehicle's image acquisition device; or, scene information associated with the question text can be determined through text information contained in the original images captured by the vehicle's image acquisition device. For example, assuming the captured original image contains text information such as road names (XX Street) or landmark buildings (XX Shopping Mall), this text information can be used to search on SD map information to obtain the scene information associated with the question text, which can be an outdoor road. For example, assuming the captured original image contains environmental image information of various plants, and each plant contains its name, the scene information associated with the question text can be determined to be a botanical garden.

[0056] In one example, the question text information and scenario information can be input into a pre-created target text classification model to directly obtain the question granularity of the question text information. By combining the question text information and scenario information to determine the question granularity of the question text information, it is ensured that even if the user cannot provide detailed question text information, relatively detailed response information can still be provided to the user. This avoids the situation where the user cannot fully understand the question and answer rules of the question-and-answer system, but still provides the necessary response information.

[0057] For example, in a scenario where the scene is a botanical garden, the user can generally be identified as a preschool child, elementary school student, or middle school student seeking knowledge. In this case, the question granularity can be directly set to the first granularity, meaning the output response will be a more detailed answer, allowing the user to obtain more detailed information and gain a better understanding. Similarly, in a scenario where the scene is XX road, the question granularity can be determined directly based on the question text. If the question text is "What exactly is this?", the question granularity can be set to the first granularity, and the final question granularity will be the first granularity. If the question text is "What is this?", the question granularity can be set to the second granularity, and the final question granularity will be the second granularity.

[0058] In one example, when the granularity of the question text information is determined based on the question text information and the context information, semantic analysis can be performed on the question text information and the context information simultaneously to determine the granularity of the question based on the question text information and the granularity of the question based on the context information, respectively; then, the granularity of the question determined by the two is comprehensively considered to determine the final granularity of the question.

[0059] In an optional embodiment, when the granularity of the question text information is determined based on the question text information and context information, the question text information and context information can be input into a pre-created target text classification model. This allows the target text classification model to directly output the final question granularity, ensuring the classification efficiency and accuracy of the question text information's granularity. It should be noted that the target text classification model can be iteratively trained on an initial text classification model using the question text information and context information as a training sample set to obtain the corresponding target text classification model.

[0060] In one example, if the granularity of the question determined based on the question text information is the first granularity, but the granularity of the question determined based on the scenario information is the second granularity, then the final granularity of the question text information is determined to be the first granularity.

[0061] In one example, if the granularity of the question determined based on the question text information is the second granularity, and the granularity of the question determined based on the scenario information is also the second granularity, then the final granularity of the question text information is determined to be the second granularity.

[0062] In one example, if the granularity of the question determined based on the question text information is the second granularity, but the granularity of the question determined based on the scenario information is the first granularity, then the final granularity of the question text information is determined to be the first granularity.

[0063] In one example, if the granularity of the question determined based on the question text information is the first granularity, and the granularity of the question determined based on the scenario information is also the first granularity, then the final granularity of the question text information is determined to be the first granularity.

[0064] In an optional embodiment, the granularity of the question text information can be determined based on user attribute information and question text information. This avoids situations where the user's age is too young, resulting in unclear question text information and an inability to determine the appropriate granularity to meet the user's needs, thus failing to provide a satisfactory response and improving the user's interactive experience. User attribute information refers to relevant information that characterizes an individual. For example, user attribute information may include, but is not limited to, age, education level, and occupation. In one example, when obtaining question text information via voice input, the user's age can be determined by analyzing relevant parameters of the voice associated with the question text information (e.g., pitch, timbre, speech rate and rhythm, and language expression style). In one example, a fingerprint recognition system or a facial recognition system can be installed in the device embedded in the question-and-answer system to identify and obtain user attribute information.

[0065] In one example, when the granularity of the question is determined based on the question text information and user attribute information, semantic analysis can be performed on both the question text information and the user attribute information simultaneously to determine the granularity of the question based on the question text information and the granularity of the question based on the user attribute information, respectively; then, the granularity of the question determined by both is comprehensively considered to determine the final granularity of the question.

[0066] In one example, when the granularity of the question text information is determined based on the question text information and user attribute information, these information can be input into a pre-created target text classification model. The target text classification model can then directly output the final question granularity, ensuring both the efficiency and accuracy of the question text information classification. It should be noted that the target text classification model can be iteratively trained on an initial text classification model using the question text information and user attribute information as a training sample set to obtain the corresponding target text classification model.

[0067] In one example, if the granularity of the question determined based on the question text information is the first granularity, but the granularity of the question determined based on the user attribute information is the second granularity, then the final granularity of the question text information is determined to be the first granularity.

[0068] In one example, if the granularity of the question determined based on the question text information is the second granularity, and the granularity of the question determined based on the user attribute information is also the second granularity, then the final granularity of the question text information is determined to be the second granularity.

[0069] In one example, if the granularity of the question determined based on the question text information is the second granularity, but the granularity of the question determined based on the user attribute information is the first granularity, then the final granularity of the question text information is determined to be the first granularity.

[0070] In one example, if the granularity of the question determined based on the question text information is the first granularity, and the granularity of the question determined based on the user attribute information is also the first granularity, then the final granularity of the question text information is determined to be the first granularity.

[0071] For example, if the user is 4 years old or has an education level below high school, the granularity of the question text can be determined as the first granularity; and if the question text is "What kind of plant is this?", the granularity of the question text can be determined as the first granularity, so as to provide the user with more detailed reply information such as "This is a rose, and it belongs to the Rosaceae family", which is easier for the user to understand.

[0072] In an optional embodiment, the granularity of the question text information can be determined based on user attribute information, scenario information, and question text information to determine the granularity of the question text information from different dimensions, so as to ensure that the level of detail of the response information associated with the corresponding question granularity is closer to the user's needs.

[0073] In one example, when the granularity of the question text information is determined based on the question text information, scenario information, and user attribute information, semantic analysis can be performed on both the question text information and the user attribute information simultaneously to determine the granularity of the question based on the question text information and the granularity of the question based on the user attribute information, respectively. Then, the granularity of the question determined by both is comprehensively considered to determine the final granularity of the question.

[0074] In one example, when the granularity of the question text information is determined based on the question text information, scene information, and user attribute information, the question text information and user attribute information can be input into a pre-created target text classification model. This allows the target text classification model to directly output the final question granularity, ensuring the classification efficiency and accuracy of the question text information's granularity. It should be noted that the target text classification model can be iteratively trained on an initial text classification model using the question text information, scene information, and user attribute information as a training sample set to obtain the corresponding target text classification model.

[0075] In one example, if the granularity of the question determined based on the question text information and the scenario information is the first granularity, but the granularity of the question determined based on the user attribute information is the second granularity, then the final granularity of the question text information is determined to be the first granularity.

[0076] In one example, if the granularity of the question determined based on the question text information, user attribute information, and scenario information is all second granularity, then the final granularity of the question text information is determined to be second granularity.

[0077] In one example, if the granularity of the question determined based on the question text information is the second granularity, but the granularity of the question determined based on the user attribute information and the scenario information is the first granularity, then the final granularity of the question text information is determined to be the first granularity.

[0078] In one example, if the granularity of the question determined based on the question text information, scenario information, and user attribute information is all of the first granularity, then the final granularity of the question text information is determined to be the first granularity.

[0079] In one example, if the granularity of the question determined based on the question text information and user attribute information is the first granularity, but the granularity of the question determined based on the scenario information is the second granularity, then the final granularity of the question text information is determined to be the first granularity.

[0080] For example, if the user is 4 years old or has an education level below high school, the granularity of the question text can be determined as first-level; if the question text is "What kind of plant is this?", the granularity of the question text can be determined as first-level; if the scene information is a botanical garden, the granularity of the question text can be determined as first-level; finally, based on the user attribute information, the question text information, and the scene information, the granularity of the question text information is determined as first-level. At this point, a more detailed reply can be provided to the user, such as "This is a rose, and it belongs to the Rosaceae family," to facilitate the user's understanding.

[0081] In one optional embodiment, determining the response information corresponding to the target granularity includes: determining the response information corresponding to the target granularity based on the question text information and associated scene information. In one example, the question granularity of the question text information is determined as a first granularity based on the question text information and scene information. An image recognition method associated with the first granularity can be used to determine the image recognition information of the original image. The original image, image recognition information, question text information, and scene information are then input into a pre-created multimodal graph-text question-answering model to obtain response information with higher detail. In another example, the question granularity of the question text information is determined as a second granularity based on the question text information and scene information. The question text information, scene information, and original image can be input into a pre-created multimodal graph-text question-answering model to obtain response information with lower detail. This allows for accurate understanding of the user's question intent and output of response information with a level of detail that meets the user's needs without increasing the amount of data computation.

[0082] In one optional embodiment, determining the response information corresponding to the target granularity includes: determining the response information corresponding to the target granularity based on the question text information and associated user attribute information. In one example, the question granularity of the question text information is determined as a first granularity based on the question text information and user attribute information. An image recognition method associated with the first granularity can be used to determine the image recognition information of the original image. The original image, image recognition information, question text information, and user attribute information are then input into a pre-created multimodal graph-text question-answering model to obtain response information with higher detail. In another example, the question granularity of the question text information is determined as a second granularity based on the question text information and user attribute information. The question text information, user attribute information, and original image can be input into a pre-created multimodal graph-text question-answering model to obtain response information with lower detail. This allows for accurate understanding of the user's question intent and accurate output of response information with a level of detail that meets the user's needs without increasing the amount of data computation.

[0083] In one optional embodiment, determining the response information corresponding to the target granularity includes: determining the response information corresponding to the target granularity based on the question text information and associated scene information and user attribute information. In one example, the question granularity of the question text information is determined as a first granularity based on the question text information and scene information. An image recognition method associated with the first granularity can be used to determine the image recognition information of the original image. The original image, image recognition information, question text information, scene information, and user attribute information are then input into a pre-created multimodal graph-text question-answering model to obtain response information with higher detail. In another example, the question granularity of the question text information is determined as a second granularity based on the question text information and scene information. The question text information, scene information, user attribute information, and original image can be input into a pre-created multimodal graph-text question-answering model to obtain response information with lower detail. This allows for accurate understanding of the user's question intent and accurate output of response information with a level of detail that meets the user's needs.

[0084] In an optional embodiment, when the target granularity is a first granularity, determining the response information corresponding to the target granularity includes: determining the image recognition information of the original image using an image recognition method associated with the first granularity; and inputting the original image, image recognition information, and question text information into a multimodal graph-text question-answering model to obtain the response information. Specifically, firstly, the question text information is obtained and input into a pre-created target text classification model to obtain the question granularity of the question text information; when the question granularity of the question text information is the first granularity, the image recognition information of the original image can be determined using an image recognition method associated with the first granularity, and the original image and image recognition information are used as input images, and the question text information is used as input text, and input into a pre-created multimodal graph-text question-answering model to obtain more detailed response information, ensuring that the level of detail of the response information is closer to the user's needs, and achieving the effect of accurately meeting the user's questioning intent and needs.

[0085] In an optional embodiment, when the target granularity is the second granularity, determining the response information corresponding to the target granularity based on the question text information includes: inputting the original image and question text information into a multimodal graph-text question-answering model to obtain the response information. In one example, the question text information can be obtained and input into a pre-created target text classification model to obtain the question granularity of the question text information; when the question granularity of the question text information is the second granularity, the question text information and the original image associated with the question text information can be directly used as input text and input into a pre-created multimodal graph-text question-answering model to obtain a relatively coarse response information, so as to accurately meet the user's questioning intent and needs, and avoid providing the user with response information containing too much information, which would lead to a poor user interaction experience.

[0086] In one embodiment, Figure 2 This is a flowchart of another question-and-answer method provided by an embodiment of the present invention. This embodiment further refines the process of determining the first-granularity and second-granularity response information based on the question text information, building upon the above embodiments. For example... Figure 2 As shown, the question-and-answer method in this embodiment includes:

[0087] S210. Obtain the question text information.

[0088] S220. In response to the question granularity of the question text information being the first granularity, the image recognition information of the original image is determined by using an image recognition method associated with the first granularity.

[0089] In one example, the raw image refers to an unprocessed image captured by an image acquisition device. This image acquisition device is integrated into a smart device or in-vehicle system. When the question-answering method is applied to an in-vehicle system, the image acquisition device can be a camera in the vehicle, such as a surround-view camera, an in-vehicle monitoring camera, a dashcam camera, or a streaming rearview camera. When the question-answering method is applied to a client-side smart device, the image acquisition device can be the smart device's front-facing camera and / or rear-facing camera.

[0090] In one example, when the question-answering method is applied to a vehicle's in-vehicle system, images can be acquired using the vehicle's onboard cameras to obtain raw images. In actual image acquisition, different types of cameras within the vehicle can be used. For instance, when recognizing objects in the vehicle's surroundings, images can be acquired using the vehicle's surround-view cameras; when recognizing objects inside the vehicle, images can be acquired using in-vehicle monitoring cameras; when recognizing objects in front of the vehicle, images can be acquired using dashcam cameras; and when recognizing objects behind the vehicle, images can be acquired using streaming rearview cameras—there is no limitation on this approach.

[0091] In one example, when the question-and-answer method is applied to a smart terminal, images can be captured using the front or rear camera of the smart terminal to obtain the original image.

[0092] In one example, image recognition information refers to the query result obtained by performing image recognition on items of the first granularity in the original image; image recognition information can be represented in text form. In practice, the original image may include one or more items, and may include multiple items of the same type. In this case, an image recognition method associated with the first granularity can be used to perform image recognition on one of the items of the same type in the original image to obtain the corresponding image recognition information.

[0093] In one example, image recognition methods may include, but are not limited to, one of the following: image recognition based on deep learning features; image recognition based on template matching; and image recognition based on a machine learning classifier. In one example, image recognition can be based on deep learning features. Specifically, the original image can be input into a pre-created image recognition model built using a convolutional neural network to obtain the corresponding image recognition information. In another example, image recognition can be based on template matching. Specifically, the original image to be recognized is matched with a predefined template image, and the similarity between the two images is determined by calculating the correlation of corresponding pixels. The image recognition information of the original image is then identified based on the similarity and a pre-configured similarity threshold. In yet another example, image recognition can be based on a machine learning classifier. Specifically, feature vectors of the image can be obtained through feature extraction methods, and these feature vectors are input into a machine learning classifier for classification and recognition to obtain the image recognition information of the original image.

[0094] S230. Input the original image, image recognition information, and question text information into the pre-created multimodal graph-text question-answering model to obtain the response information.

[0095] In one example, a multimodal graphical question-answering model refers to a question-answering system that supports asking questions about images. The input parameters of a multimodal graphical question-answering model can be images and text, and the output parameter can be text information. In an embodiment, the original image can be used as the image, and image recognition information and question text information can be used as text inputs into the multimodal graphical question-answering model to obtain the corresponding response information.

[0096] In one example, when the target-level response information is determined using the question text information, a multimodal text-image question answering model can be trained based on the question text information. Specifically, the training process of the multimodal text-image question answering model includes: acquiring an initial text dataset, an initial image dataset, and an initial text-image association dataset for training; then preprocessing the initial text dataset, initial image dataset, and initial text-image association dataset to obtain a target text dataset, a target image dataset, and a target text-image association dataset; then dividing the target text dataset, target image dataset, and target text-image association dataset into a training set, a validation set, and a test set according to a certain ratio; then loading the target text data, target image data, and target text-image association data from the training set into the initial text-image question answering model, and iteratively training the initial text-image question answering model based on the output results and in conjunction with a loss function to obtain the target text-image question answering model; finally, using relevant data from the validation set and the test set, evaluating the response accuracy of the target text-image question answering model until the response accuracy meets the pre-configured requirements, thus obtaining the multimodal text-image question answering model.

[0097] In one example, when using question text information and scene information to determine the response information corresponding to the target granularity, a multimodal graph-text question answering model can be trained based on the question text information and scene information. Specifically, the training process of the multimodal graph-text question answering model includes: acquiring an initial text dataset containing initial question text data and initial scene text data, an initial image dataset containing initial scene images, and an initial graph-text association dataset containing initial question graph-text association data and initial scene graph-text association data; then preprocessing the initial text dataset, initial image dataset, and initial graph-text association dataset to obtain a target text dataset containing target question text data and target scene text data, a target image dataset containing target scene images, and a target graph-text association dataset containing target question graph-text association data and initial scene graph-text association data; then, the target text dataset and target image dataset... The target image-text association dataset is divided into training, validation, and test sets according to a certain ratio. Then, the target text data (including target question text data and target scene text data), target image data (including target scene images), and target image-text association data (including target question image-text association data and initial scene image-text association data) in the training set are loaded into the initial image-text question answering model. Based on the output results and combined with the loss function, the initial image-text question answering model is iteratively trained to obtain the target image-text question answering model. Then, the correctness of the response of the target image-text question answering model is evaluated using relevant data in the validation and test sets until the correctness of the response meets the pre-configured requirements, thus obtaining the final multimodal image-text question answering model.

[0098] S240. Responding to the second granularity of the question text information, the original image and question text information are input into the pre-created multimodal graph-text question answering model to obtain the response information.

[0099] When the granularity of the question text information is at the second granularity, the original image and question message information can be directly input into the multimodal graph-text question answering model to obtain the response information.

[0100] S250, Output reply information.

[0101] The technical solution of this embodiment determines the granularity of the question text information. When the question granularity is the first granularity, it first uses an image recognition method that matches the question granularity to determine the image recognition information in the original image. Then, the original image, image recognition information, and question text information are input again into the multimodal graph-text question-answering model to obtain the response information. This effectively improves the recognition accuracy of fine-grained question text information and the accuracy of the response information. Alternatively, when the question granularity is the second granularity, the original image and question text information are directly input into the pre-created multimodal graph-text question-answering model to obtain the response information. This effectively ensures the accuracy of question text information recognition and response information response while improving the response efficiency and reducing the amount of data processing.

[0102] In one embodiment, Figure 3 This is a flowchart of another question-answering method provided in an embodiment of the present invention. This embodiment further refines the process of determining the question granularity and the process of determining image recognition information, based on the above embodiments. For example... Figure 3 As shown, the question-and-answer method in this embodiment includes:

[0103] S310. Obtain the question text information.

[0104] S320. Obtain the pre-created target text classification model.

[0105] In one example, the target text classification model refers to a deep learning model that can classify text to determine the granularity of the question text information. For instance, the target text classification model can be a Natural Language Understanding (NLU) model. NLU models process and understand natural language text or speech information used by humans, converting natural language into a structured representation that computers can understand and process. In practice, NLU models can be used to analyze and semantically understand question text information to accurately understand the user's intent in inputting the question text information, i.e., determining whether the user wants a relatively coarse or a more detailed response, thus determining the granularity of the question text information.

[0106] In one example, the creation process of the target text classification model includes: obtaining an initial question text sample set; if the number of samples in the initial question text sample set is less than a first sample number threshold, performing data augmentation on the question text samples in the initial question text sample set to obtain the target question text sample set; determining the text sample label associated with each question text sample in the target question text sample set; and iteratively training the initial text classification model using the target question text sample set and the text sample labels to obtain the corresponding target text classification model. In this example, the initial question text sample set refers to a set containing multiple question text samples; the target question text sample set refers to the sample set after data augmentation. Generally, the target question text sample set contains a larger number of samples than the initial question text sample set.

[0107] In practice, to ensure the accuracy of the target text classification model, more question text samples can be used to train the initial text classification model. Furthermore, to further enrich the diversity of questions, data augmentation can be performed on the question text samples in the initial question text sample set to generate more sample data. For example, data augmentation can be performed on the question text samples in the initial question text sample set using both manual generation and artificial intelligence (AI) rewriting: First, manual questioning of the items is required, targeting both broad and fine-grained categories to obtain corresponding text sample labels; then, GPT4 can be used to rewrite the question text samples in the initial question text sample set into sentences with the same meaning. For example, assuming a question text sample is "What is this?", it can be rewritten as "What item is this?". In one example, the first sample number threshold refers to a pre-configured threshold value for whether to perform data augmentation on the initial question text sample set. If the number of samples in the initial question text sample set is less than the first sample number threshold, then data augmentation is required to increase the number of samples in the initial question text sample set; if the number of samples in the initial question text sample set is greater than the first sample number threshold, then data augmentation is not required to increase the number of samples in the initial question text sample set.

[0108] In one example, text sample labels are used to characterize the granularity of each question text sample; that is, text sample labels include: first granularity and second granularity. For example, if the question text sample is "What is this?", then the text sample label for this question text sample is a second-granularity question; if the question text sample is "What exactly is this thing?", then the text sample label for this question text is a first-granularity question.

[0109] In one example, AI-generated data augmentation can be applied to the initial set of question text samples: various open-source LLM models can be used to generate question text samples for first and second granularities. Of course, to ensure the effectiveness of the question text samples, the augmented samples can be manually filtered to obtain only those that meet the requirements, thus yielding the target question text sample set.

[0110] After data augmentation of the initial text sample set, the target question text sample set is obtained. Then, the initial text classification model is trained using the target question text sample set as training data. The trained text classification model is then tested using question text test samples from the test set. If the question granularity output by the text classification model is consistent with the text sample label of the question text test sample, then the accuracy of the trained text classification model meets the requirements and can be used as the target text classification model.

[0111] S330. Input the question text information into the target text classification model to obtain the question granularity.

[0112] In one example, the question text is directly input into the target text classification model to determine whether the question granularity is first-level or second-level. For instance, the question text can be input into an NLU model to accurately identify the user's intent, i.e., to determine the level of detail required in the response. If the intent is to receive a relatively coarse response, the question granularity is determined to be second-level; if the intent is to receive a more detailed response, the question granularity is determined to be first-level.

[0113] S340, In response to the question granularity of the question text information being the first granularity and receiving an image cropping operation, a local image associated with the image cropping information is obtained from the original image.

[0114] In one example, a local image refers to a portion of the original image selected or segmented, containing specific local information from the original image; image cropping information is used to characterize which image regions in the original image are cropped. For example, assuming the question text is "What kind of plant is that in front of me?", the image in front is captured to obtain the original image; then, a target text classification model identifies the question granularity as the first granularity; upon detecting that the question granularity is the first granularity, an image cropping message is automatically output: "Which plant are you specifically asking about? Can you circle it for me?" Then, the user crops the original image on the interactive interface to obtain the local image associated with the image cropping information. The process of cropping the original image is equivalent to extracting a portion of the image region from the original image as a local image.

[0115] S350: Call the external image recognition interface to perform image recognition on a local image and obtain image recognition information.

[0116] In one example, image recognition information refers to the ability to accurately identify specific items in a local image; an external image recognition interface refers to an external tool specifically designed for image recognition. For example, if an external image recognition interface is a vehicle recognition interface, then in the case of a local image being a vehicle picture, the specific model of the vehicle can be directly obtained and used as the image recognition information.

[0117] S360. Input the original image, image recognition information, and question text information into the pre-created multimodal graph-text question-answering model to obtain the response information.

[0118] S370, Output reply information.

[0119] The technical solution of this embodiment, based on the above embodiments, utilizes a fine-grained external image recognition interface to recognize local images, which can more accurately recognize fine-grained local images. Furthermore, based on the acquisition of local images, the user's intent can be more clearly understood, thereby effectively improving the accuracy of image recognition.

[0120] In an optional embodiment, when the target granularity is a first granularity, the question-answering method further includes: outputting image cropping information; and obtaining a local image associated with the image cropping information in the original image based on the received image cropping operation.

[0121] Determine the response information corresponding to the target granularity, including: determining the response information corresponding to the target granularity based on the local image.

[0122] In one example, the question text information is first obtained. If the granularity of the question text information is determined to be the first granularity using the aforementioned method, image cropping information can be output to the user via text output. This allows the user to crop an image from the original image displayed on the interactive interface, obtaining a targeted local image. Then, image recognition is performed on the local image to obtain more accurate and detailed image recognition results, which serve as image recognition information. Finally, the image recognition information, the original image, and the question text information can be input into a multimodal graph-text question-answering model to obtain more detailed response information, corresponding to the finer granularity of the response. This achieves the goal of providing users with response information that meets their question's level of detail, effectively improving the user's interactive experience.

[0123] In one example, the question text information is first obtained. If the granularity of the question text information is determined to be the first granularity using the aforementioned method, image cropping information can be output to the user via text output. This allows the user to crop an image from the original image displayed on the interactive interface, obtaining a targeted local image. Then, image recognition is performed on the local image to obtain more accurate and detailed image recognition results, which serve as image recognition information. Finally, the image recognition information and the question text information can be input into a multimodal graph-text question-answering model to obtain more detailed response information, corresponding to the finer granularity of the response. This approach effectively reduces the computational load of the multimodal graph-text question-answering model while ensuring that the user receives a response that meets their required level of detail.

[0124] In one example, the question text information is first obtained. If the question granularity is determined to be the first granularity using the aforementioned method, image cropping information can be output to the user via text output. This allows the user to crop the image from the original image displayed on the interactive interface, obtaining a targeted local image. Then, image recognition is performed on the local image to obtain more accurate and detailed image recognition results, which serve as image recognition information. This image recognition information can then be converted into image-recognized text, and both the image-recognized text and the question text information are input into the LLM model to obtain more detailed response information, corresponding to the finer granularity. This approach effectively reduces the model's computational load while ensuring that the user receives a response that meets their question's required level of detail.

[0125] In an optional embodiment, the granularity of the question text information can be divided into two or more different granularity levels, and the target granularity corresponding to the question text information corresponds to one of the granularity levels. For example, suppose the granularity levels include a first granularity level and a second granularity level, where the first granularity level corresponds to fine granularity and the second granularity level corresponds to coarse granularity, that is, the question text information can be classified into coarse granularity or fine granularity. Alternatively, suppose the granularity levels include the following three types: a first granularity level, a second granularity level, and a third granularity level, where different granularity levels can be divided using pre-configured detailed threshold values. For example, two detailed threshold values ​​can be pre-configured: a first detailed threshold value and a second detailed threshold value, where the first detailed threshold value is greater than the second detailed threshold value. The granularity of the question can be classified according to the actual level of detail and the detail threshold of the question text information. If the actual level of detail is greater than the first detail threshold, it is classified into the first granularity level; if the actual level of detail is greater than the second detail threshold but less than the first detail threshold, it is classified into the second granularity level; if the actual level of detail is less than the second detail threshold, it is classified into the third granularity level. This is how the granularity of the question text information is determined.

[0126] Of course, the granularity of the question can be divided into more than three granularity levels. There is no limit to this, and it can be dynamically configured according to the actual needs of users to meet the different users' needs for different levels of detail in the reply information.

[0127] In one embodiment, the question-and-answer method may further include: in response to a received voice interaction command, acquiring an original image associated with the voice interaction command; and recognizing and extracting the question text information from the voice interaction command. In one example, the voice interaction command refers to a voice question command issued by a user to an image recognition device. This voice interaction command can be issued by a terminal device or by the user, without limitation. For example, assuming the voice interaction command is "What kind of plant is this?", after acquiring the voice interaction command, the image acquisition device in the image recognition device can acquire an image of the plant in front of it to obtain the corresponding original image. Furthermore, automatic speech recognition (ASR) technology can be used to recognize and extract the question text information from the voice interaction command.

[0128] In one embodiment, the question-and-answer method may further include: providing feedback to the user via voice playback. To improve the user's interactive experience, the image-recognized text can be played directly via voice playback, avoiding the process of the user viewing it on the image recognition device. This is especially important in scenarios where the user is identifying the model of a vehicle ahead while driving; the image-recognized text can be played directly through the vehicle's speakers, eliminating the need for the user to view it on the in-vehicle display screen and effectively ensuring driving safety.

[0129] In one example, Figure 4 This is a flowchart of another question-answering method provided in an embodiment of the present invention. This embodiment is a preferred embodiment, illustrating the image recognition process. In this embodiment, the image recognition process is illustrated using an NLU model as the target text classification model and an external API as the external image recognition interface. Figure 4 As shown, the image recognition process in this embodiment includes:

[0130] Step 1: Input the user's question text into the NLU model to obtain the question granularity.

[0131] In one example, Figure 5 This is a schematic diagram illustrating a classification of question granularity provided in an embodiment of the present invention. For example... Figure 5 As shown, the question text information is input into a pre-created NLU model to obtain the question granularity of the question text information.

[0132] Step 2: When the query granularity is fine, output an image cropping information.

[0133] Step 3: Crop the original image according to the image cropping information to obtain a partial image.

[0134] Step 4: Use an external image recognition API to recognize the local image and obtain image recognition information.

[0135] Step 5: Input the image recognition information, the original image, and the question text information into the multimodal graph-text question answering model to obtain the response information.

[0136] Step 6: When the question granularity is coarse, directly input the original image and question text information into the multimodal graph-text question answering model to obtain the response information.

[0137] The technical solution of this embodiment distinguishes different question granularities of question text information through an NLU model and performs image recognition using different image recognition methods, thereby achieving a balance between recognition accuracy and precision. Furthermore, by employing two-stage recognition of user intent, the user intent can be accurately understood, avoiding the occurrence of illusions. At the same time, by utilizing images and fine-grained external image recognition APIs, fine-grained images can be recognized more accurately.

[0138] In one embodiment, Figure 6 This is a schematic diagram of the structure of a question-and-answer device provided in an embodiment of the present invention. Figure 6 As shown, the question-and-answer device includes: a first acquisition module 610, a determination module 620, and an output module 630.

[0139] The first acquisition module 610 is used to acquire the question text information;

[0140] The determination module 620 is used to determine the corresponding response information in response to the question granularity of the question text information being the target granularity;

[0141] Output module 630 is used to output response information.

[0142] In one embodiment, determining the response information corresponding to the target granularity is specifically used for:

[0143] Determine the response information corresponding to the target granularity based on the question text information;

[0144] Alternatively, the response information corresponding to the target granularity can be determined based on the question text information and the associated scenario information.

[0145] In one embodiment, when the target granularity is the first granularity, the response information corresponding to the target granularity is determined, specifically for:

[0146] The image recognition information of the original image is determined using an image recognition method associated with the first granularity.

[0147] The original image, image recognition information, and question text information are input into a pre-created multimodal graph-text question-answering model to obtain the response information.

[0148] In one embodiment, the image recognition information of the original image is determined using an image recognition method associated with the first granularity, specifically for:

[0149] Extract a local image from the original image that is associated with the image cropping information;

[0150] Call an external image recognition interface to perform image recognition on a local image and obtain image recognition information.

[0151] In one embodiment, when the target granularity is a first granularity, the question-answering device further includes:

[0152] The output module is used to output image capture information;

[0153] The second acquisition module is used to acquire a local image associated with the image cropping information in the original image based on the received image cropping operation;

[0154] The step of determining the response information corresponding to the target granularity is specifically used for:

[0155] Based on the local image, the response information corresponding to the target granularity is determined.

[0156] In one embodiment, when the target granularity is the second granularity, the response information corresponding to the target granularity is determined based on the question text information, specifically for:

[0157] The original image and question text are input into a pre-created multimodal graph-text question-answering model to obtain the response information.

[0158] In one embodiment, the process of determining the granularity of the question text information includes:

[0159] Obtain a pre-created target text classification model;

[0160] The question text information is input into the target text classification model to obtain the question granularity.

[0161] In one embodiment, the training process of the target text classification model is specifically used for:

[0162] Obtain the initial question text sample set;

[0163] If the number of samples in the initial question text sample set is less than the first sample number threshold, data augmentation is used to augment the question text samples in the initial question text sample set to obtain the target question text sample set.

[0164] Determine the text sample labels associated with each question text sample in the target question text sample set;

[0165] The initial text classification model is iteratively trained using a target question text sample set and text sample labels to obtain the corresponding target text classification model.

[0166] In one embodiment, the question-and-answer device further includes:

[0167] The third acquisition module is used to acquire the original image associated with the received voice interaction command in response to the voice interaction command.

[0168] The recognition and extraction module is used to recognize and extract the question text information in the voice interaction commands.

[0169] In one embodiment, the question-and-answer device further includes:

[0170] The voice playback module is used to provide feedback and responses to users via voice playback.

[0171] The question-and-answer device provided in the embodiments of the present invention can execute the question-and-answer method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the method execution.

[0172] In one embodiment, Figure 7 This is a structural block diagram of a question-and-answer device provided in an embodiment of the present invention, such as... Figure 7The diagram illustrates a schematic representation of a question-and-answer device 10 that can be used to implement embodiments of the present invention. The in-vehicle system on the question-and-answer device is intended to represent various forms of digital computers, such as laptops, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. The question-and-answer device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0173] like Figure 7 As shown, the question-and-answer device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 can also store various programs and data required for the operation of the question-and-answer device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0174] Multiple components in the question-and-answer device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a disk, optical disk, etc.; and a communication unit 19, such as a network card, modem, wireless transceiver, etc. The communication unit 19 allows the question-and-answer device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0175] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as question-answering methods.

[0176] In some embodiments, the question-and-answer method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on the question-and-answer device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the question-and-answer method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to execute the question-and-answer method by any other suitable means (e.g., by means of firmware).

[0177] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0178] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0179] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0180] To provide interaction with a user, the systems and techniques described herein can be implemented on a question-and-answer device, which includes: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the question-and-answer device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including voice input, speech input, or tactile input).

[0181] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0182] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0183] This application also provides a computer program product, including a computer program that, when executed by a processor, can implement the question-and-answer method provided in any embodiment of this application.

[0184] In the implementation of the computer program product, computer program code for performing the operations of this application can be written in one or more programming languages ​​or a combination thereof. Programming languages ​​include object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0185] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0186] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A question-and-answer method, characterized in that, include: Obtain the question text information; In response to the question granularity of the question text information being the target granularity, the corresponding response information is determined; Output the response information.

2. The method according to claim 1, characterized in that, The method for determining the granularity of the question text information includes one of the following: Determine the granularity of the question text information based on the question text information; The granularity of the question text information is determined based on the question text information and the associated scene information; The granularity of the question is determined based on the question text information and the associated user attribute information. The granularity of the question is determined based on the question text information, associated user attribute information, and scenario information.

3. The method according to claim 1 or 2, characterized in that, When the target granularity is the first granularity, determining the response information corresponding to the target granularity includes: The image recognition information of the original image is determined using an image recognition method associated with the first granularity. The original image, the image recognition information, and the question text information are input into the multimodal graphic question answering model to obtain the response information.

4. The method according to claim 3, characterized in that, The step of determining the image recognition information of the original image using an image recognition method associated with the first granularity includes: Extract a local image from the original image that is associated with the image cropping information; An external image recognition interface is called to perform image recognition on the local image to obtain image recognition information.

5. The method according to any one of claims 1-4, characterized in that, When the target granularity is the first granularity, it further includes: Output image cropping information; Based on the received image cropping operation, a local image associated with the image cropping information is obtained from the original image; The step of determining the response information corresponding to the target granularity includes: Based on the local image, the response information corresponding to the target granularity is determined.

6. The method according to any one of claims 1-5, characterized in that, When the target granularity is the second granularity, determining the response information corresponding to the target granularity based on the question text information includes: The original image and the question text information are input into the multimodal graph-text question answering model to obtain the response information.

7. The method according to any one of claims 1-6, characterized in that, The process of determining the granularity of the question text information includes: Obtain the target text classification model; The question text information is input into the target text classification model to obtain the question granularity.

8. The method according to claim 7, characterized in that, The training process of the target text classification model includes: Obtain the initial question text sample set; If the number of samples in the initial question text sample set is less than the first sample number threshold, data augmentation is used to augment the question text samples in the initial question text sample set to obtain the target question text sample set. Determine the text sample label associated with each question text sample in the target question text sample set; The initial text classification model is iteratively trained using the target question text sample set and the text sample labels to obtain the corresponding target text classification model.

9. The method according to any one of claims 1-8, characterized in that, The method further includes: In response to a received voice interaction command, acquire the original image associated with the voice interaction command; Identify and extract the question text information from the voice interaction commands.

10. A question-and-answer device, characterized in that, include: The first acquisition module is used to acquire the question text information; The determination module is used to determine the response information corresponding to the target granularity in response to the question granularity of the question text information being the target granularity; The output module is used to output the response information.

11. A question-and-answer device, characterized in that, The question-and-answer device includes: At least one processor; A memory that is communicatively connected to the at least one processor; The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the question-and-answer method according to any one of claims 1-9.

12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the question-and-answer method according to any one of claims 1-9.

13. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the question-and-answer method according to any one of claims 1-9.