Question answering method and apparatus, device, and storage medium

By collecting image and voice data and using machine learning models to identify user intent and generate answers, the problem that existing question-answering systems are unable to understand visual information is solved, and the efficiency, accuracy and convenience of multimodal visual language question answering are achieved.

WO2024188242A9PCT designated stage expired Publication Date: 2025-09-25BEIJING YOUZHUJU NETWORK TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/081224
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-03-13
Filing Date
2024-03-12
Publication Date
2025-09-25

AI Technical Summary

Technical Problem

Existing question-answering systems are unable to effectively combine visual information, especially unable to understand users' visual-related intentions, resulting in the inability to provide accurate visual language question answers. Traditional visual recognition models are costly and closed, and cannot cope with open-world tasks.

Method used

By collecting image data and voice data, using machine learning models to identify user intent, determining the answer from the image data and outputting it in the form of voice, combined with visual language question answering technology, multimodal visual language question answering is achieved.

Benefits of technology

It improves the accuracy and efficiency of users in obtaining information about the external environment, especially the convenience for people with impaired vision, and can automatically and conveniently generate answers corresponding to user intentions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024081224_25092025_PF_FP_ABST
    Figure CN2024081224_25092025_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide a question answering method and apparatus, a device, and a storage medium. The question answering method comprises: in response to detecting a question answering initiation operation, capturing image data and voice data by using a device of a user, the voice data indicating an intention related to a question; according to the intention, determining, from the image data, an answer corresponding to the question; and outputting the answer at least in the form of a voice. Thus, the answer interaction corresponding to the image data and the voice data can be conveniently and quickly achieved, and the accuracy and convenience of the user acquiring information in the external environment can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Method, device, equipment and storage medium for question answering

[0001] This application claims priority to the Chinese invention patent application entitled “Method, apparatus, device and storage medium for question and answer” and application number 2023102385351, filed on March 13, 2023, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] Example embodiments of the present disclosure generally relate to the field of computers, and more particularly, to methods, devices, apparatuses, and computer-readable storage media for question answering. Background Art

[0003] With the rapid development of information technology, more and more applications are providing question-and-answer functions, which has brought many conveniences to users. Currently, applications with question-and-answer functions (such as voice assistants) can output corresponding answers based on the voice or text input by the user. For example, a voice assistant can play the corresponding answer audio based on the user's voice question. However, it is also expected to combine visual information to conveniently and quickly realize multimodal visual language question answering (VAQ).

[0004] Summary of the Invention

[0005] In a first aspect of the present disclosure, a question-and-answer method is provided. The method includes: in response to detecting a question-and-answer initiation operation, capturing image data and voice data using a user's device, the voice data indicating an intent associated with a question; determining an answer corresponding to the question from the image data based on the intent; and outputting the answer in at least voice form.

[0006] In a second aspect of the present disclosure, a device for question-and-answering is provided. The device includes: a data capture module configured to, in response to detecting a question-and-answer initiation operation, capture image data and voice data using a user's device, the voice data indicating an intent associated with the question; an answer determination module configured to determine an answer corresponding to the question from the image data based on the intent; and an answer output module configured to output the answer in at least voice form.

[0007] In a third aspect of the present disclosure, an electronic device is provided. The device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. When executed by the at least one processing unit, the instructions cause the electronic device to perform the method of the first aspect.

[0008] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided, wherein a computer program is stored on the medium, and when the computer program is executed by a processor, the method of the first aspect is implemented.

[0009] It should be understood that the content described in this section is not intended to limit the key features or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, wherein:

[0011] FIG1 shows a schematic diagram of an example environment in which embodiments of the present disclosure can be implemented;

[0012] FIG2 shows a flowchart of a question-and-answer process according to some embodiments of the present disclosure;

[0013] FIG3 shows a schematic diagram of a recording interface according to some embodiments of the present disclosure;

[0014] FIG4 shows a schematic diagram of a question-answering model according to some embodiments of the present disclosure;

[0015] FIG5 is a schematic diagram showing a question-answering process according to some embodiments of the present disclosure;

[0016] FIG6 shows a schematic diagram of an interface for prompting description information according to some embodiments of the present disclosure;

[0017] FIG7 shows a schematic diagram of an interface for prompting quality issues according to some embodiments of the present disclosure;

[0018] FIG8 shows a schematic diagram of question-answering model routing according to some embodiments of the present disclosure;

[0019] FIG9 shows a block diagram of an apparatus for applying question and answer according to some embodiments of the present disclosure; and

[0020] FIG10 shows a block diagram of a device capable of implementing various embodiments of the present disclosure. DETAILED DESCRIPTION

[0021] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.

[0022] In the description of the embodiments of the present disclosure, the term "including" and similar terms should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "based at least in part on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". The terms "first", "second", etc. may refer to different or the same objects. Other explicit and implicit definitions may also be included below. As used herein, the term "model" can represent the association relationship between various data. For example, the above-mentioned association relationship can be obtained based on a variety of technical solutions currently known and / or to be developed in the future.

[0023] Herein, unless explicitly stated otherwise, executing a step “in response to A” does not mean executing the step immediately after “A” but may include one or more intermediate steps.

[0024] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) must comply with the requirements of relevant laws, regulations and relevant provisions.

[0025] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0026] For example, in response to a user's active request, a prompt message is sent to the user to clearly inform the user that the operation requested will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operations of the disclosed technical solution based on the prompt message.

[0027] As an optional but non-limiting embodiment, in response to receiving a user's active request, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. Furthermore, the pop-up window may also include a target exploration control for the user to select "agree" or "disagree" to provide personal information to the electronic device.

[0028] It is understandable that the above notification and user authorization process are merely illustrative and do not limit the embodiments of the present disclosure. Other methods that meet relevant laws and regulations may also be applied to the embodiments of the present disclosure.

[0029] As used herein, the term "model" can learn the association between corresponding inputs and outputs from training data, so that after training is completed, corresponding outputs can be generated for given inputs. The generation of the model can be based on machine learning technology. Deep learning is a machine learning algorithm that processes inputs and provides corresponding outputs by using multiple layers of processing units. A neural network model is an example of a model based on deep learning. In this article, "model" may also be referred to as "machine learning model", "learning model", "machine learning network" or "learning network", and these terms are used interchangeably in this article.

[0030] A "neural network" is a machine learning network based on deep learning. A neural network is capable of processing inputs and providing corresponding outputs. It typically includes an input layer, an output layer, and one or more hidden layers between the input and output layers. Neural networks used in deep learning applications typically include many hidden layers, thereby increasing the depth of the network. The layers of a neural network are connected in sequence so that the output of the previous layer is provided as input to the next layer, where the input layer receives the input of the neural network and the output of the output layer serves as the final output of the neural network. Each layer of a neural network includes one or more nodes (also called processing nodes or neurons), each of which processes the input from the previous layer.

[0031] Generally speaking, machine learning can be roughly divided into three stages, namely the training stage, the testing stage, and the application stage (also known as the inference stage). In the training stage, a given model can be trained using a large amount of training data, and the parameter values ​​are continuously updated iteratively until the model can obtain consistent inferences that meet the expected goals from the training data. Through training, the model can be considered to be able to learn the association between input and output (also known as input-to-output mapping) from the training data. The parameter values ​​of the trained model are determined. In the testing stage, the test input is applied to the trained model to test whether the model can provide the correct output, thereby determining the performance of the model. In the application stage, the model can be used to process the actual input based on the parameter values ​​obtained through training to determine the corresponding output.

[0032] As briefly mentioned above, more and more applications that provide question-and-answer functions have brought many conveniences to users. Traditional applications with question-and-answer functions (such as voice assistants) can only output corresponding answers based on the voice or text input by the user. However, voice assistants cannot provide visual assistance to users. For example, traditional voice assistants can solve general inquiries that are not related to vision, such as weather forecasts, encyclopedia answers, and home appliance control, but cannot solve inquiries related to vision, such as clothing styles, colors, and product brands.

[0033] While some visual recognition models have been proposed, capable of outputting relevant recognition results for images (e.g., identifying the plant species and name contained in an image, or identifying product barcode information), these models are unable to understand specific user intent in different scenarios. Such models are passive (users cannot communicate intent to the application, but can only passively receive recognition results) and closed (the application can only complete predefined tasks and cannot handle open tasks in the open world, such as being unable to describe clothing styles in detail).

[0034] Visual Question Answering (VQA) is a multimodal understanding task that requires understanding visual content and then answering verbal questions. Traditionally, multimodal VQA can be achieved manually, for example by sharing the vision of volunteers or relevant staff with users via video calls (i.e., manually combining images and user questions to respond to users). However, this solution requires high labor costs, cannot guarantee volunteer working hours, and is subject to language limitations.

[0035] Embodiments of the present disclosure provide an improved solution for question-and-answering. This solution collects image data and voice data indicating a user's intent, determines a corresponding answer from the currently collected image data based on the intent, and outputs the answer in at least voice form. In this way, accurate answers corresponding to the user's intent can be automatically and conveniently generated from visual data, improving the accuracy and efficiency of users' acquisition of information about the external environment through electronic devices.

[0036] Furthermore, the question-and-answer solution proposed in this disclosure can effectively assist users, especially those with persistent or temporary vision impairment or obstruction, by enabling multimodal visual language question-and-answering. In some embodiments of this disclosure, it can also prompt and assist users in accurately capturing images of objects and provide users with more descriptive information related to the objects.

[0037] It should be understood that the solutions provided by the embodiments of the present disclosure can provide convenience for specific groups of people, but this does not imply any discrimination against specific groups of people.

[0038] FIG1 illustrates a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. In this example environment 100, an application 120 is installed on a terminal device 110. A user 140 can interact with the application 120 via the terminal device 110 and / or its attached devices. The application 120 is an application having at least a question-and-answer function.

[0039] In some embodiments, the terminal device 110 communicates with the server 130 to provide services for the application 120. The terminal device 110 can be any type of mobile terminal, fixed terminal or portable terminal, including a mobile phone, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a media computer, a multimedia tablet, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an e-book device, a gaming device or any combination thereof, including accessories and peripherals of these devices or any combination thereof. In some embodiments, the terminal device 110 can also support any type of interface for the user (such as "wearable" circuitry, etc.).

[0040] The terminal device 110 may, for example, include a sensor of an appropriate type for detecting user gestures. For example, the terminal device 110 may include a touch screen for detecting various types of gestures made by the user on the touch screen. Alternatively or additionally, the terminal device 110 may also include other appropriate types of sensing devices such as proximity sensors to detect various types of gestures made by the user within a predetermined distance above the screen. The terminal device 110 may also include, for example, a sound collection device (such as a microphone) for collecting user audio, a sound playback device (such as a speaker) for playing audio, an image collection device (such as a camera, a webcam, etc.) for collecting images, and a display screen for displaying an interface (the display screen may be a touch screen), etc.

[0041] Server 130 can be a standalone physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content distribution networks, and big data and artificial intelligence platforms. Server 130 can include, for example, computing systems / servers such as mainframes, edge computing nodes, computing devices in a cloud environment, and the like. Server 130 can provide backend services for application 120 in terminal device 110.

[0042] In some embodiments discussed below, the question-and-answer function can be implemented using multiple models with various functions. One or more of these models can be remotely deployed in the server 130, and the terminal device 110 can utilize these multiple models to implement the corresponding functions through communication with the server 130. This can save resources and power of the terminal device 110 and utilize the powerful resources of the server to improve computing efficiency. In some embodiments, one or more of these models can also be deployed locally on the terminal device 110. This can be selected based on actual circumstances.

[0043] In some embodiments, in the environment 100 of FIG. 1 , if the application 120 is active, the terminal device 110 may present an interface 150 of the application 120. Through the interface 150, the application 120 may provide the user 140 with one or more services related to the question-and-answer function, including voice capture, image capture, voice playback, text display, and the like.

[0044] It should be understood that the structure and functionality of environment 100 are described for exemplary purposes only and do not imply any limitation on the scope of the present disclosure.

[0045] Some example embodiments of the present disclosure will be described below with continued reference to the accompanying drawings.

[0046] 2 shows a flow chart of a question-answering process 200 according to some embodiments of the present disclosure. The process 200 may be implemented at the terminal device 110. For ease of discussion, the process 200 will be described with reference to the environment 100 of FIG. 1 .

[0047] In block 210 , the terminal device 110 captures image data and voice data in response to detecting a question-and-answer initiation operation, where the voice data indicates an intention associated with the question.

[0048] In some embodiments, the terminal device 110 may directly detect a Q&A initiation operation initiated by a user. For example, the terminal device 110 may determine that a Q&A initiation operation has been detected in response to detecting a Q&A initiation voice (e.g., "Open the Q&A function"). For another example, the terminal device 110 may determine that a Q&A initiation operation has been detected in response to detecting a preset operation on a hardware button (e.g., a press operation, a long press operation, etc.). In some embodiments, after detecting the Q&A initiation operation, the terminal device 110 runs the application 120 with the Q&A function and captures image data and voice data.

[0049] In some embodiments, after launching the application 120, the terminal device 110 presents a recording interface including at least a recording control through a display device. The terminal device 110 can detect a question-and-answer initiation operation by detecting a predetermined operation on the recording control, that is, the terminal device 110 can determine that a question-and-answer initiation operation is detected in response to detecting a predetermined operation on the recording control, and then capture image data and voice data. The predetermined operation on the recording control may include, for example, a click operation, a sliding operation, a long press operation, etc., which are not limited here. In some embodiments, the predetermined operation on the recording control can also be initiated by voice or other instructions.

[0050] In some embodiments, the terminal device 110 can, in response to detecting a question-and-answer initiation operation, capture voice data indicating the user's intention to ask a question through a sound acquisition device, and capture image data through an image acquisition device. The voice data can be in any language (e.g., Chinese, English, Japanese, etc.), of any duration (e.g., 3s, 5s, etc.), and of any timbre. The image data can be in any form (a still image, a video clip, etc.), of any resolution, and of any format (e.g., PNG, JPG, etc.). Alternatively or additionally, the image data can also be data pre-stored in the terminal device 110.

[0051] In some embodiments, during the process of capturing image data and voice data, the terminal device 110 may also stop capturing image data and voice data in response to receiving a capture end operation. Specifically, the terminal device 110 may determine that a capture end operation has been detected in response to detecting a voice such as "stop collecting data". The terminal device 110 may also determine that a capture end operation has been detected in response to detecting a preset operation on a hardware button (such as a press operation, a long press operation, etc.). The terminal device 110 may also determine that a capture end operation has been detected in response to detecting another predetermined operation on a recording control in the recording interface (such as a click operation, a release press operation, etc.).

[0052] Referring to Figure 3, a schematic diagram of a recording interface 300 according to some embodiments of the present disclosure is shown. Recording interface 300 may include a control display area 330, which displays at least a recording control 332. Terminal device 110 may determine that a Q&A initiation operation has been detected in response to detecting a predetermined operation on recording control 332. Terminal device 110 may then display text such as "Recording" in text prompt area 310 to indicate to the user that terminal device 110 is currently capturing data.

[0053] In some embodiments, in response to receiving a predetermined operation, the terminal device 110 changes the presentation of the recording control 332 (e.g., by changing the color or size of the recording control 332) to indicate that the terminal device 110 is capturing voice data and image data. Accordingly, in response to receiving a capture end operation, the terminal device 110 can stop capturing voice data and image data. In response to receiving a capture end operation, the terminal device 110 can switch the presentation of the recording control 332 back to the state before capture.

[0054] In some embodiments, terminal device 110 can convert captured voice data into text and display it in text prompt area 310. As shown in Figure 3, after terminal device 110 captures voice data, it displays the text corresponding to the voice data, "How many cups are there?", in text prompt area 310. Simultaneously, terminal device 110 can present the image data currently being captured by terminal device 110 in image display area 320. Speech-to-text conversion can be achieved using speech-to-text technology, which can be performed locally on terminal device 110 or on a remote server.

[0055] In some embodiments, to ensure the accuracy of the subsequently determined intention, the terminal device 110 may pre-process the captured voice data to eliminate noise in the voice data that is not related to the question (eg, ambient sound).

[0056] Referring back to FIG. 2 , at block 220 , the terminal device 110 determines an answer corresponding to the question from the image data according to the intention indicated by the voice data.

[0057] In some embodiments, the captured voice data may be subjected to intent recognition to determine the intent associated with the question from the voice data. The terminal device 110 then recognizes the image data based on the determined intent to obtain an answer corresponding to the intent.

[0058] Specifically, the terminal device 110 can identify intent based on a rule method of a dictionary and a template. Different intents will have different domain dictionaries, such as book titles, song titles, product names, etc. The terminal device 110 can make a judgment based on the degree of match or overlap between the user's intent and the dictionary. The terminal device 110 can also discriminate the user's intent based on a machine learning model. The terminal device 110 can train and learn the annotated domain corpus through machine learning and deep learning methods to obtain an intent recognition model (for example, a model based on fastText). The terminal device 110 then identifies the intent indicated by the input language data based on the model.

[0059] Continuing with Figure 3 , after receiving the voice data corresponding to "How many cups are there?", terminal device 110 recognizes the voice data and determines that the question's intent is to determine the number of cups in the image data. Terminal device 110 then recognizes the image data based on this intent and determines that the number of cups in the image data is 3, indicating that the answer corresponding to this intent is "3."

[0060] Referring back to Figure 2 , in some embodiments, terminal device 110 can use a trained question-answering model to determine the answer to a question. Specifically, after voice data is converted into a text sequence, the text sequence and image data can be input into the trained question-answering model, causing the question-answering model to output an answer corresponding to the indication in the voice data.

[0061] In some embodiments, the question-answering model includes at least four modules: an image encoder for encoding image data into image features, a language encoder for encoding text sequences into semantic features, a fusion module for obtaining combined features based on image features and semantic features, and a decoder for obtaining corresponding answers based on the combined features. In other embodiments, in order to simplify the model structure, the language encoder and the fusion module can be combined into a visual-language encoder, which is used to encode text sequences into semantic features and generate combined features based on the semantic features and the graphic feature encoding output by the image encoder. In some embodiments, the multiple encoders in the question-answering model can be Transformer-based encoders, and the decoder in the question-answering model can be a decoder with a structure similar to BERT, which can generate answers based on multimodal combined features.

[0062] In some embodiments, the trained question-answering model is deployed in the server 130, wherein the server 130 may be a remote server (e.g., a cloud server). The terminal device 110 may utilize the trained question-answering model to implement the question-answering function through communication with the server 130. Specifically, the terminal device 110 may send the captured image data and voice data to the server 130, and the server 130 may generate answers based on the image data and voice data using the trained question-answering model. The terminal device 110 may obtain answers from the server 130. Alternatively or additionally, in other embodiments, the trained question-answering model may also be deployed locally on the terminal device 110, and the terminal device 110 may directly utilize the locally deployed trained question-answering model to generate answers based on the captured image data and voice data.

[0063] In some embodiments, the terminal device 110 can convert voice data into a text sequence based on voice technology (e.g., automatic speech recognition (ASR) technology), and then provide image data and text sequences to the question-answering model in the server 130 or to a local question-answering model.

[0064] The question-answering model located on the server 130 or the terminal device 110 can generate questions and answers based on visual technologies and visual-language multimodal technologies. Examples of visual technologies include image quality control, optical character recognition (OCR), and image detection. Examples of visual-language multimodal technologies include visual question answering (VQA), visual captioning, and visual-language pre-training (VLP).

[0065] Referring to FIG4 , FIG4 shows a schematic diagram of a question-answering model 400 according to some embodiments of the present disclosure. The question-answering model 400 may include a visual encoder 410, a visual-language encoder 420, and a decoder 430. The visual encoder 410 includes a self-attention module 412 and a feedforward network module 414. The visual-language encoder 420 includes a self-attention module 422, a cross-attention module 424, and a feedforward network module 426. The decoder includes a self-attention module 432 and a feedforward network module 434.

[0066] In some implementations, the general processing of the attention module can be expressed as follows:

[0067] Where Q represents query input, K represents key input, V represents value input, and d k Represents the number of columns of Q and K, that is, the feature dimension. The above process can be understood as using the query input Q and the key input K to calculate the attention weight matrix, and using the attention weight matrix to perform weighted summation on the value input V.

[0068] In the question-answering model 400, the terminal device 110 can use the image data 401 as the query input (Q V ), key input (K V ) and value input (V V ). The self-attention module 412 in the visual encoder 410 further V , K V and V V The three inputs are extracted to obtain image features 403 .

[0069] The terminal device 110 can convert the speech data 402 into a text sequence and use the text sequence as a query input (Q L ), key input (K L ) and value input (V L ). The self-attention module 422 in the visual-language encoder 420 further L , K L and V LThe three inputs are extracted to obtain semantic features 404. In some embodiments, to facilitate the subsequent cross-attention module 424 to generate combined features 405, the image features 403 and the semantic features 404 are features of the same dimension.

[0070] Further, the semantic features 404 are provided as query input (Q L ), the image features 403 are provided as key inputs (K V ) and value input (V V ). The cross attention module 424 can be based on Q L , K V and V V The three inputs generate a combined feature 405 .

[0071] The feedforward network module 414 in the visual encoder 410 and the feedforward network module 426 in the visual-language encoder 420 can perform spatial changes on the input data, explore the nonlinear relationship between features, and enhance the expressiveness of features.

[0072] The combined features 405 are provided to a decoder 430, and a self-attention module 432 in the decoder 430 can determine an answer 406 based on the combined features 405. The feedforward network module 434 in the decoder 430 can perform spatial variation on the input data. In some embodiments, the feedforward network modules 414, 426, and 434 can include a feedforward neural network (FFN) using one or more fully connected layers.

[0073] It should be understood that the question-answering model 400 shown in FIG4 is merely exemplary and should not constitute any limitation on the functionality and structure of the question-answering model described herein. The number and type of each module in the encoder and decoder shown in FIG4 may vary. In other embodiments, various other models capable of processing multimodal data may be used to generate answers. The embodiments of the present disclosure are not limited in this regard.

[0074] 2. At block 230, the terminal device 110 outputs the answer at least in speech form.

[0075] After receiving the answer, the terminal device 110 can play the answer in voice form through the audio playback device. As shown in Figure 3, the terminal device 110 can play the answer audio through the speaker. In some embodiments, the answer can be in text form. The text can be converted into speech for output through speech synthesis (TTS). This makes it convenient for users, especially those with visual impairments, to quickly obtain the answer.

[0076] In some embodiments, the terminal device 110 may alternatively present the answer in text form via the display screen. In some embodiments, the terminal device 110 may also output the answer in an additional vibrational or visual form. Visual forms may include, for example, magnified or highlighted images. For example, if the user inputs voice data indicating the name of an object in the query image data, the terminal device 110 may, while playing the audio answer containing the object's name, magnify the image data on the display screen to highlight the object.

[0077] In this way, the user's intention to ask a question can be determined based on the voice data input by the user, the captured image data can be identified based on the intention to generate an answer, and the answer can be output at least in the form of voice. This can conveniently and quickly realize the answer interaction corresponding to the image data and voice data, and can improve the accuracy and convenience of the user in obtaining information in the external environment.

[0078] The above describes an embodiment of the question-answering process with reference to Figures 2 to 4. However, in some cases, the image data captured by the terminal device 110 through the image acquisition device cannot meet the question-answering requirements, that is, cannot generate a correct answer based on such image data. In such cases, the terminal device 110 will also perform other operations.

[0079] 5 , which shows a schematic diagram of a question-answering process 500 according to some embodiments of the present disclosure. The process 500 may be implemented at the terminal device 110. For ease of discussion, the process 500 will be described with reference to the environment 100 of FIG. 1 .

[0080] In some embodiments, during the process of capturing image data, users, especially visually impaired users, cannot accurately determine whether the device acquisition area is aligned with the actual target, and often need to find the target before asking questions. Therefore, in box 505 of process 500, a target exploration stage before asking questions is provided. In box 505, the user may be prompted with descriptive information about the image data captured by the terminal device 110. The descriptive information may generally describe the content in the picture captured by the terminal device. In some embodiments, the descriptive information includes at least one of the objects presented in the image data and the relative positional relationship between the objects. In this way, the user can adjust the posture of the image acquisition device based on the prompt to confirm that the target for the desired question can be captured.

[0081] In some embodiments, during the target exploration phase, a trained exploration model can be used to implement the target exploration function. The exploration model's input is image data, and its output is descriptive information. Similar to the question-answering model, the exploration model can be deployed on the server 130, and the terminal device 110 uses the exploration model to implement the target exploration function through communication with the server 130. The exploration model can also be deployed on the terminal device 110, and the terminal device 110 directly uses the model to implement the target exploration function.

[0082] In some embodiments, during the target exploration phase, the exploration model can be continuously used to identify and understand the captured image data, and output description information corresponding to the image data.

[0083] Referring to Figure 6, Figure 6 shows a schematic diagram of an interface 600 for prompting descriptive information according to some embodiments of the present disclosure. The terminal device 110 can present the captured image data in the image display area 620. In some embodiments, target exploration can directly respond to the detection of a question-and-answer initiation operation, using the image acquisition device in the terminal device 110 to capture an image of the external environment, and prompt the user with descriptive information about the image data captured by the terminal device 110. In some embodiments, in response to receiving a predetermined operation (such as a click operation, a slide operation, a long press operation, etc.) on the exploration control 634 in the control display area 630, it is determined that the target exploration operation is received, and the user is prompted with descriptive information about the image data captured by the terminal device 110. In some embodiments, the audio of the descriptive information can be played using a sound playback device. In some embodiments, the terminal device 110 can also display the descriptive information in text form at the text prompt area 610 in the interface 600.

[0084] Thus, the terminal device 110 can continuously provide descriptive information to the user through target exploration, assisting the user in confirming that the device can capture the target of the intended question so as to proceed to the next step of the question-and-answer dialogue. For example, if the user cannot obtain visual information of the external environment and wants to know relevant information about a kettle, the terminal device 110 can assist the user in obtaining visual information of the external environment by providing continuous descriptive information to the user, and when the image data of the kettle is captured, prompt the user with information about the location and color of the kettle. If the user wants to know information such as the brand and capacity of the kettle, the user can make a corresponding voice to proceed to the next step of the question-and-answer dialogue.

[0085] The control display area 630 may also include a switch control 631 for switching cameras, a text recognition control 632 for indicating the recognition of text in the image data, and a target selection control 633 for target selection. In some embodiments, the terminal device 110 may determine to enter the target selection mode in response to receiving a predetermined operation (such as a click operation, a sliding operation, a long press operation, etc.) on the target selection control 633. In the target selection mode, the terminal device 110 may prompt the user with descriptive information about the object in response to the user's touch operation on any object in the image data. Exemplarily, the terminal device 110 may prompt the user with descriptive information about the kettle in response to receiving a touch operation on the area where the kettle is located in the image display area 620.

[0086] Referring back to FIG. 5 , in some embodiments, after prompting the user with the description information, in response to a user action (e.g., a user triggering a Q&A initiation action or a collection confirmation action), terminal device 110 captures image data 510 for subsequent Q&A initiation. Furthermore, voice data of user 140 may also be captured via voice recording 525 .

[0087] In some embodiments, in order to ensure the visual quality of the image data as much as possible and to improve the accuracy of question and answer, the image data captured by the terminal device 110 through the image acquisition device can be image data in the form of video. In this case, the terminal device 110 can capture image data in frames from the video to obtain an image dataset, and perform image quality control on the captured image dataset in box 520. Specifically, the terminal device 110 can detect the visual quality of each image data in the image dataset. The terminal device 110 can select image data whose visual quality exceeds a quality threshold from the captured image data set, and provide it together with the text sequence obtained by converting the voice data through speech-to-text conversion 530 to the trained question and answer model 545.

[0088] Regarding the specific method of detecting visual quality, since the most common quality problems of image data are blur, overexposure (too bright), underexposure (darkness), improper framing (for example, incomplete framing of an object), occlusion, and rotation, in some embodiments, the terminal device 110 can score the image data in the captured image data set for these problems. The higher the score, the better the visual quality of the image data, and the lower the score, the worse the visual quality of the image data. In this way, the terminal device 110 can reduce blurring problems caused by jitter, etc. by capturing video and selecting image data with better visual quality from it, which helps to improve the visual quality of the image data and thus improve the accuracy of subsequent answers.

[0089] In some embodiments, during the image quality control stage, the terminal device 110 may also prompt the user to adjust the capture environment to recapture the image data if it is determined that a predetermined visual quality problem exists. In some embodiments, the predetermined visual quality problem includes at least an underexposure problem or an overexposure problem. In some embodiments, the predetermined visual quality problem may also include problems such as blur, improper framing, occlusion, and rotation. In some embodiments, blur problems can be solved by selecting high-quality image data, while problems such as improper framing, occlusion, and rotation can also be solved through target exploration. In this way, it can be ensured that the image data used by the user to initiate the question is of high quality and meets the user's question requirements.

[0090] Referring to Figure 7, Figure 7 shows a schematic diagram of an interface 700 for prompting quality problems according to some embodiments of the present disclosure. The terminal device 110 can present the captured image data in the image display area 720. During the image quality control stage, it can be determined whether the captured image data has a predetermined visual quality problem. If it is determined that there is an underexposure problem, the image quality control unit 520 can use a speaker to play a prompt audio to prompt the user "Insufficient light, please turn on the light". In some embodiments, during the image quality control stage, the terminal device 110 can also display a prompt text "Insufficient light, please turn on the light" in the text prompt area 710.

[0091] In this way, the terminal device 110 prompts the user to recapture image data that meets the visual quality requirements, which helps to improve the visual quality of the image data and further improve the accuracy of subsequent answers.

[0092] Return to reference Figure 5. In some embodiments, when the predetermined visual quality issue is rotation, the terminal device 110 may change the image angle using an image processing algorithm, correct the rotated image, and then provide the corrected image data together with a text sequence obtained by converting the speech data through speech-to-text 530 to the trained question-answering model.

[0093] In some embodiments, during the image quality control phase, the terminal device 110 can utilize a trained quality detection and processing model to detect the visual quality of the image data and process the image data accordingly. Similarly, the quality detection and processing model can be deployed on the server 130 or locally on the terminal device 110.

[0094] In some embodiments, the voice data input by the user 140 may not be the default language type of the question-answering model (for example, the default language type is Chinese, and the language type corresponding to the voice data input by the user is English). In this case, the terminal device 110 can also perform language translation 535 to convert the text sequence into a text sequence of the model's default language type, and perform language translation 555 again to translate the answer output by the model into a language type corresponding to the user input language.

[0095] In some embodiments, the question-answering models in different application scenarios can be the same, and the image data and text sequence will be directly provided to the question-answering model. In other embodiments, different question-answering models can be pre-selected and trained for different application scenarios. In this way, each question-answering model can give a more accurate answer that is more in line with user expectations for its respective application scenario. In such an implementation, the image data and text sequence will first be provided to the scene routing model 540, which will determine the application scenario based on the two, and provide the two to the question-answering model 545 corresponding to the determined application environment.

[0096] Referring to Figure 8, Figure 8 shows a schematic diagram of question-answering model routing according to some embodiments of the present disclosure. The scenario routing model 540 can determine multiple question-answering models 545 corresponding to multiple question-answering scenarios. After the scenario routing model 540 obtains the image data and the text sequence, it determines the question-answering scenario (i.e., application scenario) corresponding to the two. For example, when the image data contains an image of clothing and the text sequence contains a text sequence corresponding to "What style of clothing is this", the scenario routing model 540 determines that the corresponding question-answering scenario is a clothing dressing scenario, and provides the image data and text sequence to the clothing dressing question model 545-2 corresponding to the clothing dressing scenario.

[0097] In some embodiments, if there is no corresponding question-answering scenario for image data or text sequences, that is, if the scenario routing model 540 cannot determine the corresponding question-answering scenario for the two, the scenario routing model 540 will provide the two to the general multimodal question-answering model 545-1. The general multimodal question-answering model 545-1 is a question-answering model that is applicable to general scenarios.

[0098] In some embodiments, the model structures of different question-answering models 545 corresponding to different scenarios are the same; the difference between these multiple question-answering models is the training data. For example, the question-answering model 545 trained using image data and voice data associated with clothing questions and answers is the clothing outfit question-answering model 545-2 corresponding to the clothing outfit scenario, the question-answering model trained using product-related image data and voice data is the product packaging question-answering model 545-3 corresponding to the product packaging scenario, and so on. The general multimodal question-answering large model 545-1 can be a question-answering model trained using training data corresponding to multiple scenarios.

[0099] In some embodiments, the training data used to train the question-answering model 545 can be data obtained by the terminal device 110 or the server 130 from different scenarios. For example, the training data used to train the general multimodal large question-answering model 545-1 can be network image-text data associated with general life scenarios, online Chinese materials, general visual question-answering data, etc. The training data used to train specific scenarios such as clothing wearing scenarios, product packaging scenarios, and drug question-answering scenarios can be e-commerce, live streaming data, product information data, and health knowledge graphs associated with these scenarios. In addition, the training data for the question-answering model corresponding to the visually impaired life scenarios of visually impaired users can be visual question-answering data for the visually impaired and private data obtained after user authorization.

[0100] In this way, a general question-answering model can solve most question-answering problems. In certain scenarios, a specific question-answering model can be used to generate more accurate answers, ensuring the accuracy of answers in different scenarios. Furthermore, a scenario-based routing model is used to automatically select question-answering models based on the scenario. This allows for flexible switching of question-answering models as the scenario changes, and this switching process is automatic and imperceptible to the user, improving both the efficiency and accuracy of question-answering.

[0101] Return to reference Figure 5. In some embodiments, the scenario routing model 540 can also be deployed at the server 130, and the terminal device 110 can use the scenario routing model 540 to determine the question-and-answer scenario and select the corresponding question-and-answer model through communication with the server 130. In other embodiments, the scenario routing model 540 can also be deployed locally on the terminal device 110, and the terminal device 110 can directly use the scenario routing model to determine the question-and-answer scenario and select the corresponding question-and-answer model.

[0102] In some embodiments, the selected question-answering model 545 can obtain a corresponding answer based on the image data and the text sequence. The terminal device 110 can display visual information 550 based on the answer, for example, by presenting the answer in text form on a display screen. The terminal device 110 can also convert the answer into audio through text-to-speech (e.g., TTS technology) at block 560, and then perform voice playback 565 on a sound playback device to play the audio corresponding to the answer.

[0103] The example process of question-and-answering is described above with reference to Figures 2 to 8. This question-and-answer solution can be applied to a variety of question-and-answer scenarios and provide convenience to users in these various question-and-answer scenarios. These various question-and-answer scenarios can include, for example, barrier-free offline shopping, barrier-free home life, barrier-free live shopping, and so on.

[0104] For example, in a barrier-free offline shopping scenario, the terminal device 110 can provide convenience for users' offline shopping, enhance their shopping experience, and greatly reduce various uncertainties in offline shopping scenarios. Specifically, when purchasing clothes, the terminal device 110 can help users easily obtain visual information such as the style, color, size, price, etc. of the clothes, so that users can select and try on clothes more freely. When purchasing food in a supermarket, users can ask Lingtong questions such as taste, weight, and shelf life. Without the help of others, you can complete the shopping experience independently, and there is no need to listen to long paragraphs of irrelevant content read aloud by auxiliary tools.

[0105] For example, in barrier-free home life scenarios, the terminal device 110 can flexibly support open tasks and open scenarios through voice dialogue, assisting the user's daily life in the most natural way. For example, the terminal device 110 using the question-and-answer solution described in this solution can easily handle common challenges such as "matching the color of socks" and "whether clothes need to be washed." The terminal device 110 can help users quickly find key information such as food name, shelf life, and calorie content from product details. The terminal device 110 can also help users quickly obtain information such as the ingredients, specifications, usage, and dosage of a drug from a drug instruction sheet.

[0106] For example, in a barrier-free live shopping scenario, a terminal device 110 using the question-and-answer solution described in this solution can provide visual capabilities. Using basic system capabilities such as shortcut commands, the terminal device 110 can initiate conversations on any interface, helping users understand the visual content presented on the display. For example, a user can use the terminal device 110 to learn basic information about clothing styles and colors. While browsing multimedia content, a user can ask the terminal device 110 questions to understand the content of accompanying images.

[0107] It should be understood that the above question-and-answer scenario is merely exemplary and should not constitute any limitation on the scope of application of the embodiments described herein.

[0108] 9 shows a block diagram of a question-answering apparatus 900 according to some embodiments of the present disclosure. The apparatus 900 may be implemented in or included in the terminal device 110. Each module / component in the apparatus 900 may be implemented by hardware, software, firmware, or any combination thereof.

[0109] As shown, device 900 includes a data capture module 910 configured to, in response to detecting a question-and-answer initiation operation, capture image data and voice data using a user's device, where the voice data indicates an intent associated with the question. Device 900 also includes an answer determination module 920 configured to determine an answer corresponding to the question from the image data based on the intent. Device 900 also includes an answer output module 930 configured to output the answer in at least voice form.

[0110] In some embodiments, the device 900 also includes: an interface presentation module, configured to present a recording interface, the recording interface including at least a recording control; and an operation detection module, configured to detect a question-and-answer initiation operation by detecting a predetermined operation on the recording control.

[0111] In some embodiments, the data capture module 910 is further configured to select image data having a visual quality exceeding a quality threshold from the captured image data set.

[0112] In some embodiments, the data capture module 910 includes: a quality determination module configured to determine whether the captured image data has a predetermined visual quality problem; and a prompt module configured to prompt the user to adjust the capture environment to recapture the image data if it is determined that the predetermined visual quality problem exists.

[0113] In some embodiments, the predetermined visual quality problem includes at least an underexposure problem or an overexposure problem.

[0114] In some embodiments, the data capture module 910 includes: a description information prompt module, configured to prompt the user with description information about the image data captured by the device during the capture process; and an image data acquisition module, configured to acquire the image data captured by the device in response to detecting a capture confirmation operation.

[0115] In some embodiments, the description information indicates at least one of the following: objects presented in the image data, and relative positional relationships between objects.

[0116] In some embodiments, the description information prompting module is further configured to: prompt description information in response to detecting a target exploration operation.

[0117] In some embodiments, the answer is determined using a trained question-answering model, whose model input includes text sequences corresponding to image data and speech data.

[0118] In some embodiments, the question-answering model includes at least a cross-attention module, and the question-answering model determines the answer by: extracting image features of the image data; extracting semantic features of the text sequence; generating combined features by providing the semantic features as query inputs to the cross-attention module and providing the image features as key inputs and value inputs to the cross-attention module; and determining the answer based on the combined features.

[0119] In some embodiments, the question-answering model corresponds to a first question-answering scenario among a plurality of question-answering scenarios, and the question-answering model is selected based on the image data and the speech data being classified into the first question-answering scenario.

[0120] FIG10 shows a block diagram of an electronic device 1000 in which one or more embodiments of the present disclosure may be implemented. It should be understood that the electronic device 1000 shown in FIG10 is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein. The electronic device 1000 shown in FIG10 may be used to implement the terminal device 110 and / or the server 130 of FIG1 .

[0121] As shown in FIG10 , electronic device 1000 is a general-purpose electronic device. Components of electronic device 1000 may include, but are not limited to, one or more processors or processing units 1010, memory 1020, storage device 1030, one or more communication units 1040, one or more input devices 1050, and one or more output devices 1060. Processing unit 1010 may be a real or virtual processor and is capable of performing various processes according to programs stored in memory 1020. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to enhance the parallel processing capabilities of electronic device 1000.

[0122] The electronic device 1000 typically includes a plurality of computer storage media. Such media can be any accessible media that is accessible to the electronic device 1000, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 1020 can be a volatile memory (e.g., a register, a cache, a random access memory (RAM)), a non-volatile memory (e.g., a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 1030 can be a removable or non-removable medium and can include a machine-readable medium, such as a flash drive, a disk, or any other medium that can be used to store information and / or data (e.g., training data for training) and can be accessed within the electronic device 1000.

[0123] The electronic device 1000 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 10 , a disk drive for reading from or writing to a removable, non-volatile disk (e.g., a “floppy disk”) and an optical drive for reading from or writing to a removable, non-volatile optical disk may be provided. In these cases, each drive may be connected to a bus (not shown) by one or more data media interfaces. The memory 1020 may include a computer program product 1025 having one or more program modules configured to perform various methods or actions of various embodiments of the present disclosure.

[0124] The communication unit 1040 enables communication with other electronic devices via a communication medium. Additionally, the functions of the components of the electronic device 1000 can be implemented as a single computing cluster or multiple computing machines that can communicate via a communication connection. Thus, the electronic device 1000 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or other network nodes.

[0125] Input device 1050 may be one or more input devices, such as a mouse, keyboard, or trackball. Output device 1060 may be one or more output devices, such as a display, a speaker, or a printer. Electronic device 1000 may also communicate with one or more external devices (not shown) via communication unit 1040 as needed, such as storage devices, display devices, or the like, with one or more devices that allow a user to interact with electronic device 1000, or with any device that allows electronic device 1000 to communicate with one or more other electronic devices (e.g., a network card, a modem, etc.). Such communication may be performed via an input / output (I / O) interface (not shown).

[0126] According to an exemplary embodiment of the present disclosure, a computer-readable storage medium is provided, on which computer-executable instructions are stored, wherein the computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary embodiment of the present disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, and the computer-executable instructions are executed by a processor to implement the method described above.

[0127] Various aspects of the present disclosure are described herein with reference to flowcharts and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.

[0128] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, such that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.

[0129] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.

[0130] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer program products according to multiple implementations of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a part for a module, program segment or instruction, and a part for a module, program segment or instruction comprises one or more executable instructions for realizing the logical function of the specification. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two continuous boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be realized by a special hardware-based system that performs the function or action of the specification, or can be realized by a combination of special hardware and computer instructions.

[0131] While various implementations of the present disclosure have been described above, the above descriptions are exemplary, non-exhaustive, and not intended to be limiting of the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is selected to best explain the principles of the implementations, their practical applications, or improvements to existing technologies, or to enable others skilled in the art to understand the various embodiments disclosed herein.

Claims

1. A question-answering method, comprising: In response to detecting the question-and-answer initiation operation, capturing, using a device of the user, image data and voice data, the voice data indicating an intent related to asking the question; determining an answer corresponding to the question from the image data according to the intention; as well as The answer is output at least in speech form.

2. The method according to claim 1, further comprising: Presenting a recording interface, wherein the recording interface includes at least a recording control; as well as The question-and-answer initiation operation is detected by detecting a predetermined operation on the recording control.

3. The method of claim 1 , wherein capturing the image data comprises: The image data having a visual quality exceeding a quality threshold is selected from the set of captured image data.

4. The method of claim 1 , wherein capturing the image data comprises: determining whether the captured image data has a predetermined visual quality problem; as well as If it is determined that the predetermined visual quality problem exists, the user is prompted to adjust the capturing environment to recapture the image data. The method according to claim 4 , wherein the predetermined visual quality problem comprises at least an underexposure problem or an overexposure problem.

6. The method of claim 1 , wherein capturing the image data comprises: During the capture process, prompting the user with descriptive information about the image data captured by the device; as well as In response to detecting a capture confirmation operation, the image data captured by the device is acquired. 7 . The method according to claim 6 , wherein the description information indicates at least one of the following: objects presented in the image data, and relative positional relationships between the objects.

8. The method according to claim 6, wherein the description information indicates: In response to detecting a target exploration operation, the description information is prompted.

9. The method according to claim 1, wherein the answer is determined using a trained question-answering model, the model input of the question-answering model including a text sequence corresponding to the image data and the speech data.

10. The method according to claim 9, wherein the question-answering model comprises at least a cross-attention module, and the question-answering model determines the answer by: extracting image features of the image data; Extracting semantic features of the text sequence; generating a combined feature by providing the semantic feature as a query input to the cross-attention module and providing the image feature as a key input and a value input to the cross-attention module; and An answer is determined based on the combined features.

11. The method according to claim 9, wherein the question-answering model corresponds to a first question-answering scenario among a plurality of question-answering scenarios, and the question-answering model is selected based on the image data and the voice data being classified into the first question-answering scenario.

12. A device for question-answering, comprising: a data capture module configured to capture, in response to detecting a question-and-answer initiation operation, image data and voice data using a user's device, the voice data indicating an intention associated with the question; an answer determination module, configured to determine an answer corresponding to the question from the image data according to the intention; as well as The answer output module is configured to output the answer at least in voice form.

13. An electronic device comprising: at least one processing unit; as well as At least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions causing the electronic device to perform the method according to any one of claims 1 to 11 when executed by the at least one processing unit.

14. A computer-readable storage medium having a computer program stored thereon, wherein the computer program implements the method according to any one of claims 1 to 11 when executed by a processor.