Image-based question and answer method and related device
Patent Information
- Application Number
- CN202380094865.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-18
- Publication Date
- 2025-10-03
AI Technical Summary
When the existing question-and-answer system enters images, it is difficult for the user to understand the user's intention to ask questions, resulting in the inaccurate answers and the inability to meet the users' deep-seated question needs.
By searching the relevant Q&A pairs based on image features in the Q&A library, and combining the user's historical search behavior, select the most matching Q&A pair to return to the user, thereby understanding the user's interests and intentions and improving the accuracy of the answers.
It realizes that when the user enters an image, the Q&A system can more accurately understand the user's intentions and return the Q&A pairs that are of interest to the user, improving the answer accuracy and user experience of the Q&A system.
Smart Images

Figure CN120752630A_ABST
Abstract
Description
Image-based question answering method and related device Technical Field
[0001] The present application relates to the field of content search technology, and in particular to an image-based question-answering method and related devices. Background Art
[0002] Question-and-answer (Q&A) is one of the most successful AI applications currently, with specific applications in areas such as encyclopedias, intelligent customer service, and conversational systems. In these scenarios, the Q&A system needs to find the corresponding answer in a database for a user-entered question and return it to the user, satisfying their need for information. For example, in an encyclopedia, a Q&A system can find articles or entries within the encyclopedia that answer the user's question. Therefore, Q&A systems have become a crucial component of new information interaction methods.
[0003] Current question-answering systems rely heavily on user-entered question text. Only when users enter accurate question text can the system recognize their intent and provide the appropriate answer. However, with mobile devices now able to easily capture images (for example, by taking photos or screenshots on smartphones), users often prefer to input images into the question-answering system instead of question text and receive the system's answer based on the image.
[0004] Based on this, related art has proposed a technology called photo recognition. After a user captures an image, a question-and-answer system can identify the object in the image and return an encyclopedia description of the identified object. However, the answers returned by these related art question-and-answer systems are essentially general descriptions of the objects in the image, making it difficult to understand the user's actual question based on the image, making it difficult for the question-and-answer system to return accurate answers to the user.
[0005] Summary of the Invention
[0006] This application provides an image-based question-answering method that can understand the user's question intention based on the user's interests and hobbies, thereby returning question-answer pairs that the user is interested in, and improving the accuracy of the answers returned by the question-answering system.
[0007] The first aspect of the present application provides an image-based question-answering method that can be applied to a server. In this method, the server first obtains a first search request and image features of a first image. The first search request is used to search for question-answer pairs related to the first image, and the first search request comes from a terminal device, that is, the terminal device requests the server to search for question-answer pairs related to the first image. The image features of the first image are obtained by performing feature extraction on the first image. In addition, the image features can be sent by the terminal device to the server, or can be obtained by the server performing feature extraction on the first image after the terminal device sends the first image to the server.
[0008] Then, the server obtains multiple candidate question-answer pairs based on the image features, for example, determining multiple candidate question-answer pairs among the multiple question-answer pairs. The multiple question-answer pairs exist in a question-answer base database pre-built by the server, and each of the multiple question-answer pairs includes a question text and an answer text with a corresponding relationship. Among the multiple question-answer pairs, the text features of the multiple candidate question-answer pairs match the image features. For example, by calculating the similarity between the image features and the text features of each question-answer pair, it is possible to determine multiple candidate question-answer pairs whose similarity with the image features reaches a certain threshold, that is, the similarity between the multiple candidate question-answer pairs that match the image features and the image features reaches a certain threshold.
[0009] Next, based on the terminal device's historical search behavior, the server determines a first question-answer pair from multiple candidate question-answer pairs that matches the historical search behavior. This historical search behavior is used to indicate the terminal device's behavior in searching for answers during the historical period. The terminal device can search for answers in a variety of ways. For example, a user can search for answers by entering question text or voice on the terminal device; another example is a user searching for question-answer pairs that include answer text by entering an image on the terminal device; another example is a user searching for answers by entering an image and question text on the terminal device. In other words, any answer search behavior by the user on the terminal device with the intention of asking a question can be considered as the terminal device's behavior in searching for answers during the historical period.
[0010] Finally, the server sends a first search result to the terminal device. The first search result indicates the text in the first question-answer pair. For example, the first search result may indicate the entire text in the first question-answer pair; or the first search result may indicate a portion of the text in the first question-answer pair, such as the answer text in the first question-answer pair.
[0011] In this solution, once an image is acquired, a search is first performed in a pre-built Q&A database based on image features, resulting in multiple candidate Q&A pairs whose text features match the image features. Then, based on the terminal device's historical search behavior, a Q&A pair that matches the historical search behavior is selected from the multiple candidate Q&A pairs, and the search results corresponding to the image are returned to the terminal device. Because this solution first searches the Q&A database based on image features to obtain multiple candidate Q&A pairs, it ensures that candidate Q&A pairs related to the image content are obtained first, and then, based on historical search behavior, the final search results are further selected from multiple candidate Q&A pairs that match the image content. This allows the system to understand the user's question intent based on their interests and hobbies, thereby returning Q&A pairs that the user is interested in, improving the accuracy of the answers returned by the Q&A system.
[0012] In one possible implementation, during the process of sending the first search result to the terminal device, the server may further send instruction information to the terminal device, where the instruction information is used to instruct the terminal device to highlight first text in the first search result, where the first text is a portion of text in the first search result. The highlighting may refer to displaying the first text in a different display mode than other text in the first search result, such as highlighting or bolding the first text.
[0013] When the user clicks on the first text in the first search result, thereby triggering an instruction to send a search instruction to the terminal device for results related to the first text, the server can obtain a second search request from the terminal device, which is used to search for results related to the first text.
[0014] In this way, the server can search for a corresponding second question-answer pair in multiple question-answer pairs based on the first text and send a second search result to the terminal device. The second search result is used to indicate the text in the second question-answer pair related to the first text, for example, indicating all or part of the text in the second question-answer pair.
[0015] In this solution, by indicating the first text to be highlighted in the feedback search results in the indication information, the terminal device can highlight the first text in the search results, making it easier for users to quickly capture the key content in the search results. At the same time, it further guides users to click on the highlighted first text to ask new questions, triggering a new round of question-and-answer retrieval services, truly realizing what users think and facilitating users to quickly acquire new knowledge.
[0016] In one possible implementation, when the server searches for a second question-answer pair corresponding to the first text, the server may first generate a first question text based on the object included in the first image and the first text. The server then searches multiple question-answer pairs based on the first question text to obtain a second question-answer pair that matches the first question text. For example, assuming the object included in the first image is a cat and the first text is a balcony, the first question text generated based on "cat" and "balcony" may specifically be "Do I need to enclose my balcony if I have a cat?"
[0017] In this solution, by combining the first text in the question-answer pair with the object in the image and considering the relationship between the first text and the object in the image, the user's question intention is fully understood, and a question text related to the scene shown in the current image can be generated. Then, based on the generated question text, a search is performed to obtain a question-answer pair that meets the user's intention, making it easier for users to quickly acquire new knowledge of interest.
[0018] In one possible implementation, the first text may be the text with the highest search volume among all the texts in the first search result, that is, the first text is the text with the most user searches in the first search result. Alternatively, the first text may be the text corresponding to the question-answer pair with the highest confidence among all the texts in the first search result. In other words, when searching for question-answer pairs based on the various texts in the first search result, the question-answer pair retrieved based on the first text has the highest confidence.
[0019] In one possible implementation, if the server is able to obtain the first image sent by the terminal device, the server may also send a second image to the terminal device. The second image is used to mark and display the first object in the first image. That is, the second image is specifically obtained by marking the first object in the first image. If the first image includes multiple objects, the first object may be one or more objects in the first image. In other cases, the first object may also be one or more parts of an object in the first image.
[0020] After the terminal device displays the second image and the user clicks the first object in the second image, the server can receive a third search request from the terminal device, which searches for results related to the first object. The server then sends the third search result to the terminal device, which indicates the text in the question-answer pair related to the first object.
[0021] In this solution, by returning the second image marked with the first object to the terminal device, the terminal device can display the second image marked with the first object to the user, further guiding the user to pay attention to and click on the first object marked in the second image to ask new questions about the details in the image, triggering a new round of question-and-answer retrieval services, and facilitating the user to quickly acquire deeper new knowledge.
[0022] In one possible implementation, the first object is the object with the highest search volume among all objects included in the first image. Alternatively, the first object is the object corresponding to the question-answer pair with the highest confidence among all objects included in the first image. In general, the first object can be an object with high attention in the first image, or an object with abundant and authoritative answers.
[0023] In one possible implementation, during the process of building the Q&A database, the server first obtains a second question text, which may be manually provided or machine-generated. The server then performs an online answer search based on the second question text, obtaining at least one answer text that matches the second question text. For example, the server performs an answer search based on the second question text on various web pages, thereby obtaining web page content that has a high probability of matching the second question text. This web page content can then serve as the at least one answer text that matches the second question text.
[0024] Finally, the server determines a target answer text having the highest matching degree with the second question text in at least one answer text, wherein the second question text and the target answer text are used to constitute a question-answer pair among multiple question-answer pairs.
[0025] In this solution, by performing an answer search on the Internet including various public materials based on the pre-acquired question text, and further calculating the matching degree between the question text and the searched answer text, the answer text that best matches the question text is selected to form a question-answer pair. This can effectively ensure the diversity of the answer text sources and ensure that the constructed question-answer pairs can have a high degree of confidence.
[0026] In one possible implementation, after obtaining the question-answer pairs corresponding to the second question text, the server may generate multiple question texts based on the second question text. These multiple question texts and the second question text are used to indicate the same question, and the multiple question texts are expressed differently from the second question text. The server then constructs multiple question-answer pairs by combining each question text in the multiple question texts with the target answer text. In other words, because the multiple question texts generated based on the second question text have the same meaning as the second question text but are expressed differently, the same answer text can be assigned to these question texts, thereby quickly completing the construction of multiple question-answer pairs.
[0027] In this solution, more question texts with the same meaning but different expressions are generated based on the question texts in the constructed question-answer pairs, thereby expanding the number of question texts and improving the coverage of the question-answer base.
[0028] In one possible implementation, the method further includes: the server obtaining a third question text and performing an online answer search based on the third question text, obtaining multiple answer texts that match the third question text, wherein the multiple answer texts correspond to multiple different viewpoints. Thus, based on the viewpoints corresponding to each of the multiple answer texts, the server fuses the answer texts corresponding to the same viewpoint to obtain multiple fuse-processed answer texts; the third question text and the multiple fuse-processed answer texts are used to form one of the multiple question-answer pairs.
[0029] In this solution, when performing an answer search on a question text and obtaining answers corresponding to different viewpoints, multiple answers corresponding to the same viewpoint are integrated to achieve the goal of summarizing the answers according to viewpoints, so that the question-answer pair finally constructed can include answer texts corresponding to multiple viewpoints, ensuring the comprehensiveness of the answers.
[0030] The second aspect of the present application provides an image-based question-answering device, including: an acquisition module, used to obtain a first search request and image features of a first image, the first search request is used to search for results related to the first image, and the first search request comes from a terminal device; a processing module, used to obtain multiple candidate question-answer pairs based on the image features, each of the multiple candidate question-answer pairs includes a question text and an answer text with a corresponding relationship, and the text features of the multiple candidate question-answer pairs match the image features; the processing module is also used to determine, based on the historical search behavior of the terminal device, a first question-answer pair that matches the historical search behavior among multiple candidate question-answer pairs; a sending module, used to send a first search result to the terminal device, the first search result is used to indicate the text in the first question-answer pair.
[0031] In one possible implementation, the sending module is further used to send indication information to the terminal device, where the indication information is used to instruct the terminal device to highlight the first text in the first search result, where the first text is part of the text in the first search result; the acquiring module is further used to acquire a second search request from the terminal device, where the second search request is used to search for question-and-answer pairs related to the first text; the sending module is further used to send a second search result to the terminal device, where the second search result is used to indicate the text in a second question-and-answer pair related to the first text.
[0032] In one possible implementation, the processing module is further used to: generate a first question text based on the object included in the first image and the first text; perform a search in multiple question-answer pairs based on the first question text to obtain a second question-answer pair that matches the first question text.
[0033] In a possible implementation, the first text is the text with the highest search volume among all texts of the first search result; or, the first text is the text corresponding to the question-answer pair with the highest confidence among all texts of the first search result.
[0034] In one possible implementation, the sending module is further used to send a second image to the terminal device, where the second image is used to mark and display the first object in the first image; the acquiring module is further used to acquire a third search request from the terminal device, where the third search request is used to search for results related to the first object; the sending module is further used to send a third search result to the terminal device, where the third search result is used to indicate the text in the question-answer pair related to the first object.
[0035] In one possible implementation, the first object is the object with the highest search volume among all objects included in the first image; or, the first object is the object corresponding to the question-answer pair with the highest confidence among all objects included in the first image.
[0036] In one possible implementation, the acquisition module is further used to obtain a second question text; the processing module is further used to perform an answer search on the Internet based on the second question text to obtain at least one answer text that matches the second question text; the processing module is further used to determine a target answer text that has the highest match degree with the second question text in the at least one answer text, wherein the second question text and the target answer text are used to form a question-answer pair.
[0037] In one possible implementation, the processing module is further used to: generate multiple question texts based on the second question text, where the multiple question texts and the second question text are used to indicate the same question, and the multiple question texts and the second question text are expressed differently; and construct multiple question-answer pairs by combining each question text in the multiple question texts with the target answer text.
[0038] In one possible implementation, the acquisition module is further used to obtain a third question text; the processing module is further used to perform an answer search on the Internet based on the third question text to obtain multiple answer texts matching the third question text, and the multiple answer texts correspond to multiple different viewpoints; the processing module is further used to fuse the answer texts corresponding to the same viewpoint based on the viewpoint corresponding to each answer text in the multiple answer texts to obtain multiple fused answer texts; wherein the third question text and the multiple fused answer texts are used to form a question-answer pair.
[0039] In a possible implementation, the acquisition module is further configured to acquire a first search request and a first image from a terminal device; and the processing module is further configured to perform image feature extraction on the first image to obtain image features.
[0040] In a possible implementation, the image features come from a terminal device.
[0041] A third aspect of the present application provides an image-based question-answering device, which includes: a memory and a processor; the memory stores code, and the processor is configured to execute the code. When the code is executed, the image-based question-answering device performs a method as implemented in any one of the first aspects.
[0042] A fourth aspect of the present application provides a computer-readable storage medium, which stores a computer program. When the computer-readable storage medium is run on a computer, the computer executes the method of any one of the implementation modes in the first aspect.
[0043] A fifth aspect of the present application provides a computer program product, which, when executed on a computer, enables the computer to execute a method as in any one of the implementations of the first aspect.
[0044] In a sixth aspect, the present application provides a chip comprising one or more processors, wherein some or all of the processors are configured to read and execute a computer program stored in a memory to perform the method in any one of the implementations of the first aspect.
[0045] Optionally, the chip includes a memory, and the memory is connected to the processor via a circuit or wire. Optionally, the chip also includes a communication interface, and the processor is connected to the communication interface. The communication interface is used to receive data and / or information to be processed, and the processor obtains the data and / or information from the communication interface, processes the data and / or information, and outputs the processing results through the communication interface. The communication interface can be an input / output interface. The method provided in this application can be implemented by a single chip or by multiple chips working together.
[0046] Among them, the technical effects brought about by any design method in the second to sixth aspects can refer to the technical effects brought about by different implementation methods in the above-mentioned first aspect, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] FIG1 is a schematic diagram of a solution for photographing and recognizing images in the related art;
[0048] FIG2 is a schematic diagram of an application of an image-based question-answering method provided in an embodiment of the present application;
[0049] FIG3 is a schematic diagram of an application architecture of an image-based question-answering method provided in an embodiment of the present application;
[0050] FIG4 is a schematic structural diagram of a server 101 provided in an embodiment of the present application;
[0051] FIG5 is a flow chart of an image-based question-answering method provided in an embodiment of the present application;
[0052] FIG6 is a schematic diagram of highlighting a first text in a question-answer pair to trigger a new question-answer pair search, provided by an embodiment of the present application;
[0053] FIG7 is a schematic diagram of another method of highlighting the first text in a question-answer pair to trigger a new question-answer pair search according to an embodiment of the present application;
[0054] FIG8 is a schematic diagram of a method of marking and displaying a first object in an image to trigger a new question-answer pair retrieval according to an embodiment of the present application;
[0055] FIG9 is a schematic diagram of the structure of a server provided in an embodiment of the present application;
[0056] FIG10 is a schematic diagram of a process for a user to obtain a question-answer pair based on an image, provided by an embodiment of the present application;
[0057] FIG11 is a schematic diagram of a core process for a user to obtain a question-answer pair based on an image, provided by an embodiment of the present application;
[0058] FIG12 is a schematic diagram of simultaneously displaying multiple question-answer pairs to a user according to an embodiment of the present application;
[0059] FIG13 is a schematic diagram of a question-answer pair including a single-viewpoint answer provided in an embodiment of the present application;
[0060] FIG14 is a schematic diagram of a question-answer pair including answers from multiple viewpoints provided in an embodiment of the present application;
[0061] FIG15 is a schematic diagram of highlighting a first text in a question-answer pair and marking a first object in an image, provided by an embodiment of the present application;
[0062] FIG16 is a schematic diagram of generating a new question text based on a first text and a first object provided by an embodiment of the present application;
[0063] FIG17 is a schematic diagram of the structure of an image-based question-answering device provided in an embodiment of the present application;
[0064] FIG18 is a schematic diagram of the structure of an execution device provided in an embodiment of the present application;
[0065] FIG19 is a schematic diagram of the structure of a computer-readable storage medium provided in an embodiment of the present application. DETAILED DESCRIPTION
[0066] In order to make the purpose, technical solutions and advantages of this application more clear, the embodiments of this application are described below in conjunction with the accompanying drawings. Obviously, the described embodiments are only embodiments of a part of this application, rather than all embodiments. It is known to those skilled in the art that with the emergence of new application scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.
[0067] The terms "first", "second", etc. in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the descriptions used in this way can be interchangeable where appropriate so that the embodiments can be implemented in a sequence other than that illustrated or described in this application. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or modules is not necessarily limited to those steps or modules clearly listed, but may include other steps or modules that are not clearly listed or that are inherent to these processes, methods, products or devices. The naming or numbering of steps in this application does not mean that the steps in the method flow must be executed in the time / logical sequence indicated by the naming or numbering. The named or numbered process steps can change the execution order according to the technical purpose to be achieved, as long as the same or similar technical effects can be achieved.
[0068] The division of units in this application is a logical division. In actual application, there may be other division methods. For example, multiple units can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between each other shown or discussed can be through some interfaces, and the indirect coupling or communication connection between units can be electrical or other similar forms, which are not limited in this application. Moreover, the units or sub-units described as separate components may or may not be physically separated, may or may not be physical units, or may be distributed into multiple circuit units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this application.
[0069] To facilitate understanding, some technical terms involved in the embodiments of this application are first introduced below.
[0070] (1) Question Answering System (QA)
[0071] Question-answering systems are an advanced form of information retrieval system, capable of accurately and concisely answering natural language questions posed by users. The primary reason for the rise of question-answering system research is the demand for rapid and accurate information acquisition. In general, question-answering systems are a highly sought-after research area in the fields of artificial intelligence and natural language processing, with broad development prospects.
[0072] (2) Zero-Click Search
[0073] Zero-click search means that after a user searches for a question, the search engine can integrate the knowledge on the web page and directly return an answer that can answer the user's question, rather than a source web page. In this way, the user can complete a satisfactory search without clicking in.
[0074] (3) Multimodality
[0075] Multimodality refers to data in multiple modalities, such as text, images, videos, and audio.
[0076] (4) Neural Network
[0077] A neural network can be composed of neural units. A neural unit can refer to an operation unit that takes xs (i.e., input data) and intercept 1 as input. The output of the operation unit can be:
[0078] Where s = 1, 2, ... n, n is a natural number greater than 1, Ws is the weight of xs, and b is the bias of the neural unit. f is the activation function of the neural unit, which is used to introduce nonlinear characteristics into the neural network to convert the input signal of the neural unit into the output signal. The output signal of the activation function can be used as the input of the next convolutional layer, and the activation function can be a sigmoid function. A neural network is a network formed by connecting multiple single neural units mentioned above, that is, the output of one neural unit can be the input of another neural unit. The input of each neural unit can be connected to the local receptive field of the previous layer to extract the features of the local receptive field. The local receptive field can be an area composed of several neural units.
[0079] Generally speaking, most current artificial intelligence models are composed of neural networks.
[0080] (5) Encoder
[0081] In the embodiment of the present application, the encoder is essentially a neural network model that can convert data such as text or images into a vector in a coding space (i.e., convert text or images into text features or image features).
[0082] (6) Vector Engine
[0083] The Vector Engine is based on deep learning and vector retrieval. It uses a neural network model to represent the input object as a dense vector, thereby retrieving other vectors similar to the input object's vector, ultimately retrieving other objects similar to the input object. The distance between different vectors can indicate the similarity between two feature vectors (or two objects). This engine is typically used in online query scenarios, often requiring high throughput and low latency.
[0084] (7) Sequence to Sequence Model (Seq2Seq)
[0085] A sequence-to-sequence model is a model that generates another sequence from a given sequence using a specific generative method. The two sequences can be of unequal length. This structure, also known as the Encoder-Decoder model, is a variant of the Recurrent Neural Network (RNN) and addresses the problem of RNNs requiring equal-length sequences.
[0086] (8) Attention Network
[0087] Attention networks are network models that utilize the attention mechanism to accelerate model training. Currently, typical attention networks include the Transformer model. Models that utilize the attention mechanism assign different weights to each part of the input sequence, thereby extracting more important features from the input sequence and ultimately achieving more accurate output.
[0088] In deep learning, the attention mechanism can be implemented using a weight vector that describes importance: when predicting or inferring an element, the weight vector determines the correlation between that element and other elements. For example, for a pixel in an image or a word in a sentence, the attention vector can be used to quantitatively estimate the correlation between the target element and other elements, and the weighted sum of the attention vectors can be used as an approximation of the target value.
[0089] The attention mechanism in deep learning simulates the attention mechanism of the human brain. For example, when a person looks at a painting, although their eyes can see the entire painting, when they look closely, their eyes actually focus on only a small part of the pattern. At this time, the human brain focuses primarily on this small part. In other words, when a person observes an image carefully, the human brain's attention is not evenly distributed across the entire image; instead, it is weighted differently. This is the core idea of the attention mechanism.
[0090] Simply put, the human visual processing system tends to selectively focus on certain parts of an image and ignore other irrelevant information, which helps the human brain perceive. Similarly, in some problems involving language, speech, or vision, certain parts of the input may be more relevant than others in deep learning. Therefore, through the attention mechanism in the attention model, the attention model can perform different processing on different parts of the input data, so that the attention model dynamically focuses only on data relevant to the task.
[0091] (9) Bidirectional Encoder Representation from Transformers (BERT)
[0092] BERT is a pre-trained language representation model. BERT emphasizes that it no longer uses traditional unidirectional language models or shallow concatenation of two unidirectional language models for pre-training. Instead, it uses a new masked language model (MLM) to generate deep bidirectional language representations.
[0093] (10) Recurrent Neural Network (RNN)
[0094] A recurrent neural network is a type of recursive neural network that takes sequence data as input, performs recursion in the direction of sequence evolution, and all nodes (recurrent units) are connected in a chain-like manner.
[0095] Recurrent neural networks (RNNs) possess memory, parameter sharing, and Turing completeness, giving them advantages in learning nonlinear features of sequences. They have applications in natural language processing (NLP), such as speech recognition, language modeling, and machine translation, and are also used for various time series forecasting tasks. RNNs, constructed by incorporating convolutional neural networks, can handle computer vision problems involving sequential inputs.
[0096] RNNs are designed to process sequential data. In traditional neural network prediction models, each layer is fully connected, from the input layer to the hidden layer to the output layer, and the nodes within each layer are disconnected. However, these standard neural networks are inadequate for many problems. For example, to predict the next word in a sentence, you generally need to use the previous words, because the previous and next words in a sentence are not independent. RNNs are called recurrent neural networks because the current output of a sequence is dependent on the previous output. Specifically, the network remembers previous information and applies it to the calculation of the current output. In other words, the nodes between hidden layers are no longer disconnected but connected, and the input to a hidden layer includes not only the output of the input layer but also the output of the previous hidden layer. In theory, RNNs can process sequence data of any length. However, in practice, to reduce complexity, it is often assumed that the current state is only dependent on a few previous states.
[0097] Specifically, the core part of RNN is a directed graph. The chain-connected elements in the directed graph are called RNN cells. Generally, the chain connection composed of recurrent cells can be compared to the hidden layer in a feedforward neural network, but in different discussions, the "layer" of RNN may refer to the recurrent unit of a single time step or all recurrent units. Therefore, as a general introduction, the concept of "hidden layer" is avoided here. Given the learning data X = {X1, X2, ..., X τ}, the expansion length of the RNN is τ. The sequence to be processed is usually a time series, and the direction of sequence evolution is called the time step.
[0098] (11) User Profile
[0099] User portraits, also known as user personas, are an effective tool for outlining target users and linking user needs with design directions. They are widely used in various fields. In the era of big data, user information is flooding the internet. By abstracting each piece of user information into tags, these tags can be used to concretize the user's image and provide targeted services. Simply put, a user portrait can be understood as one or more tags assigned to a user based on various user information. These tags effectively describe the user's characteristics.
[0100] Currently, most question-answering systems rely heavily on user-entered question text. Only when users enter accurate question text can the system recognize their intent and provide the appropriate answer. However, with the current ease of accessing images on mobile devices (for example, by taking photos or screenshots on smartphones), users often prefer to input images into question-answering systems, avoiding the tedious process of manually entering question text and receiving the system's answer based on the image.
[0101] Based on this, related art has proposed a technology called photo recognition. After a user captures an image, a question-and-answer system can identify the object in the image and return an encyclopedia description of the identified object. However, the answers returned by these related art question-and-answer systems are essentially general descriptions of the objects in the image, making it difficult to understand the user's actual question based on the image, making it difficult for the question-and-answer system to return accurate answers to the user.
[0102] For example, please refer to Figure 1, which is a schematic diagram of a photo recognition solution in the related art. As shown in Figure 1, in the related art, after a user enters the photo recognition interface on a terminal device, a clickable "Photo Recognition" button is displayed, along with a text description of the photo recognition technology: Take a photo or select an image from an album to automatically identify the object in the image. Thus, by clicking the "Photo Recognition" button on the interface, the user enters the photo recognition interface and takes a photo for performing photo recognition, such as the photo of a cat crouching in a litter box shown in Figure 1. Furthermore, after the user confirms that the captured photo will be processed for image recognition, the related art returns a recognition result, which specifically includes an encyclopedia description of the "cat" recognized in the photo. As shown in Figure 1, the actual photo recognition result returned in the related art is: Cats, belonging to the family Felidae, are widely found as pets in households worldwide. Typical cats have round heads, short faces, five toes on their front limbs, four toes on their hind limbs, and sharp, curved claws at the ends. These claws are retractable and have night vision.
[0103] It can be seen that the photo recognition solution provided in the related art actually recognizes the object in the photo and then returns the encyclopedia introduction corresponding to the object. However, in most scenarios, the user does not want to ask what the object in the photo is, but wants to ask deeper questions. For example, taking the scene shown in Figure 1 as an example, the photo shown in Figure 1 is a kitten playing in a cat litter box. The user who took the photo is very likely a cat owner, so the user has a relatively deep understanding of cats. In this case, the encyclopedia introduction about cats returned by the photo recognition method is actually a simple introduction to cats, which is not suitable for push to cat owners who have a relatively deep understanding of cats.
[0104] Specifically, in the scenario shown in FIG1 , with respect to the photo taken by the user, the question the user wants to ask may be a deeper question, such as “Are cats and litter boxes suitable to be kept together?”.
[0105] In view of this, an embodiment of the present application provides an image-based question-answering method, which first searches the question-answer base based on image features to obtain multiple candidate question-answer pairs, which can ensure that candidate question-answer pairs related to the image content are obtained first, and then further select the question-answer pairs that are finally returned to the user from multiple candidate question-answer pairs that fit the image content based on the user's historical search behavior. It can understand the user's question intention based on the user's interests and hobbies, and thus return the question-answer pairs that the user is interested in, thereby improving the accuracy of the answers returned by the question-answering system.
[0106] For example, please refer to Figure 2, which is a schematic diagram of the application of an image-based question-answering method provided in an embodiment of the present application. As shown in Figure 2, after the user captures an image, by clicking the search button displayed above the image in the interface, the image-based question-answering process provided in this embodiment can be triggered, thereby obtaining a question-answer pair related to the captured image. Specifically, the question-answer pair includes a question text and an answer text. The question text is: "Are cats and litter boxes suitable for being kept together?" The answer text is: "It is not recommended to keep cats and litter boxes together. Litter boxes are usually not too close to where cats sleep or eat. It is best to place them in a location with a good view, such as the living room or a corner of the room. Because litter boxes are not clean and hygienic, keeping cats and litter boxes together can easily cause cats to contract skin diseases or other diseases."
[0107] Obviously, based on the above analysis and the content shown in Figure 2, it can be seen that the image-based question-answering method provided in this embodiment can understand the user's question intention based on the user's interests and hobbies, thereby returning the question-answer pairs that the user is interested in and improving the accuracy of the returned answers.
[0108] Please refer to Figure 3, which is a schematic diagram of the application architecture of an image-based question-answering method provided in an embodiment of the present application. As shown in Figure 3, the image-based question-answering method provided in an embodiment of the present application can be applied to a server. The server interacts with a terminal device, receives a search request for an image from the terminal device, and returns the searched question-answer pairs to the terminal device, thereby enabling the terminal device to display the question-answer pairs related to the image.
[0109] Exemplarily, the terminal device may be, for example, a smartphone, a personal computer (PC), a laptop, a tablet computer, a smart TV, a mobile internet device (MID), a wearable device, a virtual reality (VR) device, an augmented reality (AR) device, a wireless communication device in industrial control, a wireless communication device in self-driving, a wireless communication device in remote medical surgery, a wireless communication device in a smart grid, a wireless communication device in transportation safety, a wireless communication device in a smart city, a wireless communication device in a smart home, etc.
[0110] Please refer to Figure 4, which is a schematic diagram of the structure of a server 101 provided in an embodiment of the present application. As shown in Figure 4, server 101 includes a processor 103, which is coupled to a system bus 105. Processor 103 can be one or more processors, each of which can include one or more processor cores. A display adapter (video adapter) 107 can drive a display 109, which is coupled to system bus 105. System bus 105 is coupled to an input / output (I / O) bus via a bus bridge 111. An I / O interface 115 is coupled to the I / O bus. The I / O interface 115 communicates with various I / O devices, such as an input device 117 (e.g., a touch screen), an external memory 121 (e.g., a hard disk, floppy disk, optical disk, or USB flash drive), a multimedia interface, etc., a transceiver 123 (capable of sending and / or receiving radio communication signals), a camera 155 (capable of capturing still and dynamic digital video images), and an external USB port 125. Optionally, the interface connected to the I / O interface 115 may be a USB interface.
[0111] The processor 103 may be any conventional processor, including a reduced instruction set computing (RISC) processor, a complex instruction set computing (CISC) processor, or a combination thereof. Alternatively, the processor may be a dedicated device such as an ASIC.
[0112] Server 101 can communicate with software deployment server 149 via network interface 129. Exemplarily, network interface 129 is a hardware network interface, such as a network card. Network 127 can be an external network, such as the Internet, or an internal network, such as Ethernet or a virtual private network (VPN). Alternatively, network 127 can be a wireless network, such as a WiFi network or a cellular network.
[0113] The hard drive interface 131 is coupled to the system bus 105. The hard drive interface is connected to the hard drive 133. The internal memory 135 is coupled to the system bus 105. The data running in the internal memory 135 may include the operating system (OS) 137 of the server 101, the application program 143, and the scheduler.
[0114] The operating system consists of a shell 139 and a kernel 141. Shell 139 is an interface between the user and the operating system's kernel. The shell is the outermost layer of the operating system. The shell manages the interaction between the user and the operating system: it waits for user input, interprets user input to the operating system, and processes various operating system output.
[0115] The kernel 141 consists of the parts of the operating system that manage memory, files, peripherals, and system resources. The kernel 141 directly interacts with the hardware. The operating system kernel typically runs processes and provides inter-process communication, CPU time slice management, interrupts, memory management, and I / O management.
[0116] Please refer to Figure 5, which is a flow chart of an image-based question-answering method provided in an embodiment of the present application. As shown in Figure 5, the flow of the image-based question-answering method includes the following steps 501-504.
[0117] Step 501: Obtain a first search request and image features of a first image. The first search request is used to search for results related to the first image, and the first search request comes from a terminal device.
[0118] In this embodiment, after the user obtains the first image by taking a photo, taking a screenshot, or saving a web page image, the user can trigger a search for question-and-answer pairs related to the first image on the terminal device, that is, send a search instruction for the first image to the terminal device. For example, after the user obtains the first image by taking a photo, the user clicks a search button displayed on the first image, thereby triggering a search for question-and-answer pairs related to the first image. For another example, after the user obtains the first image by taking a screenshot, the user can trigger a search for question-and-answer pairs related to the first image by using a specific gesture (such as double-clicking the first image with two fingers).
[0119] In this way, after the terminal device obtains the search instruction from the user, the terminal device can send a first search request to the server to request to search for results related to the first image. That is, the server can obtain the first search request from the terminal device.
[0120] Optionally, the terminal device may directly send the first image to the server, i.e., the terminal device simultaneously sends the first image and the first search request to the server. Alternatively, the terminal device may perform image feature extraction on the first image, obtain image features corresponding to the first image, and then send the image features corresponding to the first image to the server.
[0121] That is to say, there are many ways for the server to obtain image features.
[0122] In one implementation, the server receives a first search request and a first image from a terminal device and performs image feature extraction on the first image to obtain image features corresponding to the first image. This allows the server to perform image feature extraction on the first image based on a locally deployed, high-performance image feature extraction model, enabling accurate image features to be extracted, facilitating subsequent image feature-based retrieval to obtain accurate question-and-answer pairs.
[0123] In another possible implementation, the server receives image features sent by the terminal device. Specifically, the image features are received from the terminal device. Specifically, what the server receives from the terminal device is not the original image, but rather the image features obtained by the terminal device after performing feature extraction on the original image. This allows the search for question-and-answer pairs to be performed by having the terminal device perform feature extraction on the original image before sending it to the server, thus avoiding the potential privacy leaks that could arise from directly sending the original image to the server.
[0124] In step 502, a plurality of candidate question-answer pairs are obtained based on the image features, each of the plurality of candidate question-answer pairs includes a question text and an answer text having a corresponding relationship, and the text features of the plurality of candidate question-answer pairs are matched with the image features.
[0125] In this embodiment, a pre-built question-and-answer database is deployed on the server, which includes multiple pre-built question-and-answer pairs. Furthermore, for each of the multiple question-and-answer pairs in the database, each pair includes a corresponding question text and answer text. That is, each pair includes a pair of mutually corresponding question text and answer text. Furthermore, in addition to the answer text, the question-and-answer pair may also include an image serving as the answer. In other words, a question-and-answer pair may include a question text, an answer text, and an answer image.
[0126] Therefore, after obtaining the image features, the server can perform a question-answer pair search based on the image features among multiple question-answer pairs in the question-answer base database, thereby obtaining multiple candidate question-answer pairs. Specifically, for the multiple question-answer pairs in the question-answer base database, the corresponding text features of each question-answer pair can be extracted using the same neural network model. In this way, by calculating the similarity between the image features and the text features of each question-answer pair, multiple candidate question-answer pairs whose similarity with the image features reaches a certain threshold can be determined. In other words, the similarity between the multiple candidate question-answer pairs that match the image features and the image features reaches a certain threshold.
[0127] Since text features and image features are essentially vectors, the similarity between text features and image features can actually be determined by calculating the distance between the text features and image features (e.g., the L1 distance). In addition, whether the text features and image features of a question-answer pair match can also be determined by other methods, which are not specifically limited in this embodiment.
[0128] It should be noted that the number of candidate question-answer pairs can be predefined, that is, the number of candidate question-answer pairs is a preset value, such as 10 candidate question-answer pairs or 15 candidate question-answer pairs. The preset value can be adjusted according to actual conditions and is not specifically limited in this embodiment.
[0129] Optionally, since text and images belong to data of different modalities, in order to facilitate the determination of the degree of match between text features and image features, a multimodal model can be used to extract image features and text features of question-answer pairs. Among them, a multimodal model refers to a model that can process data of different modalities such as images and text. During the training of the multimodal model, the multimodal model is trained together with data of different modalities such as images and text. Therefore, when the trained multimodal data processes images or text, the extracted image features and text are located in the same feature space.
[0130] Furthermore, when the server performs image feature extraction, the multimodal model used to extract the image features of the first image and the multimodal model used to extract the text features of the question-answer pair can be the same model; when the terminal device performs image feature extraction, the multimodal model used to extract image features and the multimodal model used to extract text features can be different models. For example, a multimodal model with fewer parameters can be used on the terminal device to extract image features to reduce resource consumption, while a multimodal model with a larger number of parameters can be used on the server to extract the text features of the question-answer pair to improve the accuracy of the extracted text features.
[0131] Step 503: Based on the historical search behavior of the terminal device, determine a first question-answer pair that matches the historical search behavior from multiple candidate question-answer pairs.
[0132] In this embodiment, historical search behavior refers to the behavior of users searching for answers through terminal devices during the historical period. There are many ways for users to search for answers through terminal devices. For example, users search for answers by inputting question text or voice on the terminal device; for another example, users search for question-answer pairs including answer text by inputting images on the terminal device; for another example, users search for answers by inputting images and question text on the terminal device. In general, any answer search behavior by users with the intention of asking questions can be regarded as the behavior of users searching for answers during the historical period. Therefore, by obtaining the user's historical search behavior, it is possible to obtain the questions asked by the user during the historical period, and then understand the user's interests and hobbies, so as to infer the user's intention to ask questions based on the current image. The above-mentioned historical search behavior can be the behavior generated by the user who provided the first image, that is, the user who provided the first image to trigger the first search request.
[0133] After obtaining the user's historical search behavior, the correlation between each candidate question-answer pair and the historical search behavior can be determined, and then the first question-answer pair with the highest correlation with the historical search behavior can be determined among the multiple candidate question-answer pairs. This first question-answer pair is used as the question-answer pair fed back to the user.
[0134] Specifically, since each candidate question-answer pair includes a question text and an answer text, and the user's historical search behavior actually also has a corresponding question text, the relevance between each candidate question-answer pair and the historical search behavior can be determined by determining the similarity between the question text in the candidate question-answer pair and the question text corresponding to the historical search behavior. In other words, the higher the similarity between the question text in the candidate question-answer pair and the question text corresponding to the historical search behavior, the higher the relevance between the candidate question-answer pair and the historical search behavior.
[0135] Optionally, when users are more active in asking questions, a user may also have more historical search behaviors. Therefore, the user profile of the user can be determined in advance based on all the user's historical search behaviors, thereby achieving a high degree of summary of the user's characteristics. Then, since the user profile corresponding to the user is actually some text labels corresponding to the user, the correlation between the question text and the user profile in the candidate question and answer pair can be determined (for example, determined by a pre-trained neural network model) to obtain the correlation between each candidate question and answer pair and the historical search behavior. Among them, the higher the correlation between the question text and the user portrait in the candidate question and answer pair, the higher the attention of the user with the current user portrait to the question text; the lower the correlation between the question text and the user portrait in the candidate question and answer pair, the lower the attention of the user with the current user portrait to the question text. Therefore, the correlation between the question text and the user portrait can also be obtained by obtaining the feedback results of other users on the question text after recommending the question text to other users with the same user portrait. For example, the more positive feedback other users give to the question text, the higher the relevance between the question text and the user profile; the more negative feedback other users give to the question text, the lower the relevance between the question text and the user profile.
[0136] For example, suppose there are currently two candidate question-answer pairs, and the question text 1 of one candidate question-answer pair is "What kind of animal is a cat?", and the question text 2 of the other candidate question-answer pair is "Can a cat and a litter box be put together?". Then, when the user's user profile is determined to be a "cat slave" (that is, a user who pays attention to information about raising cats) based on the user's historical search behavior, the user portrait can be combined with question text 1 and question text 2 respectively and input into the pre-trained neural network model to obtain the correlation between the user portrait and each question text. Specifically, the correlation between question text 1 "What kind of animal is a cat?" and the user portrait is higher than the correlation between question text 2 "Can a cat and a litter box be put together?" and the user portrait, so it can be determined that the candidate question-answer pair corresponding to question text 1 is the final output question-answer pair.
[0137] Among them, the method of determining the user portrait based on the user's historical search behavior can refer to the existing method of determining the user portrait, which will not be elaborated here.
[0138] Optionally, when users are not active in asking questions or when they have just started using the question-and-answer system, a user may have very little historical search behavior or even no historical search behavior. Therefore, in this case where it is difficult to obtain the search behavior of a specific user, the correlation between each candidate question-and-answer pair and the historical search behavior can be determined based on the historical search behavior of all other users. Specifically, for each candidate question-and-answer pair, the more historical search behaviors there are whose corresponding question texts are identical or similar to the question text in the candidate question-and-answer pair, the higher the correlation between the candidate question-and-answer pair and the historical search behavior. In this way, when a specific user has less historical search behavior, question-and-answer pairs that are of interest to the general public (i.e., frequently recommended question-and-answer pairs) can be recommended to the specific user based on the historical search behavior of the general public, thereby satisfying the user's demand for feedback on the question-and-answer pairs to a greater extent.
[0139] Step 504: Send a first search result to the terminal device, where the first search result is used to indicate the text in the first question-answer pair.
[0140] After determining that the first question-answer pair is the question-answer pair that responds to the first search request, the server may send the first search result to the terminal device so that the terminal device can display the first search result on the display interface. The first search result may indicate the entire text in the first question-answer pair; or the first search result may indicate part of the text in the first question-answer pair, such as the answer text in the first question-answer pair. In addition, after obtaining the first search result, the terminal device may display the first image and the first search result on the display interface at the same time, or may only display the first search result, which is not specifically limited here.
[0141] In this solution, we first search the question-and-answer database based on image features to obtain multiple candidate question-and-answer pairs, which can ensure that candidate question-and-answer pairs related to the image content are obtained first. Then, based on the user's historical search behavior, we further select the final search results from multiple candidate question-and-answer pairs that fit the image content. This solution can understand the user's question intention based on the user's interests and hobbies, and thus return the question-and-answer pairs that the user is interested in, thereby improving the accuracy of the answers returned by the question-and-answer system.
[0142] The above describes the process of searching for question-and-answer pairs based on a user-provided first image and returning corresponding question-and-answer pairs to the user. In some embodiments, to facilitate users in conveniently asking new questions based on the returned question-and-answer pairs, this embodiment may highlight some of the first text in the question-and-answer pairs during the process of returning the question-and-answer pairs to the user, thereby guiding the user to ask new questions based on the highlighted first text.
[0143] For example, in the above-mentioned step 504, while the server is sending the first question-answer pair to the terminal device, the server may also send an instruction message to the terminal device, which is used to instruct the terminal device to highlight the first text in the first search result. In this way, after receiving the first search result and the instruction message, the terminal device may highlight the first text in the first question-answer pair while displaying the first search result based on the content in the instruction message, so as to attract the user's attention to the key content in the first search result. The way in which the terminal device highlights the first text may include, but is not limited to, one or more of the following ways: highlighting the first text, underlining the first text below the first text, and displaying the first text in bold.
[0144] Optionally, the first text may be the text with the highest search volume among all the texts in the first search result, that is, the first text is the text with the most user searches in the first search result. Alternatively, the first text is the text corresponding to the question-answer pair with the highest confidence among all the texts in the first search result. That is, when retrieving question-answer pairs based on the various texts in the first search result, the question-answer pairs retrieved based on the first text are the ones with the highest confidence. Alternatively, the first text may be a key entity (i.e., a person's name, an organization's name, an animal's name, a place name, etc.) identified by an entity recognition method. In general, the first text may be a text with a high degree of attention in the first question-answer pair, or a text with rich and authoritative answers, or a text with a specific meaning. This embodiment does not limit the method for determining the first text.
[0145] While the user is browsing the first search results, if the user has further questions regarding the first text in the first search results, the user can click on the first text in the first search results, thereby triggering a command to search for answers related to the first text to be sent to the terminal device. In this way, in response to the user clicking on the first text, the terminal device can send a second search request to the server, so that the server can obtain the second search request from the terminal device, which is used to search for results related to the first text.
[0146] After receiving the second search request, the server can continue searching the Q&A database for question-answer pairs related to the first text based on the second search request. Since the first text is in text form, conventional text search methods can be used to search for question-answer pairs related to the first text, thereby retrieving the most relevant question-answer pairs.
[0147] Finally, the server sends a second search result to the terminal device, where the second search result indicates text in a second question-answer pair related to the first text, for example, indicating all or part of the text in the second question-answer pair. The second question-answer pair is a question-answer pair determined from multiple question-answer pairs in the question-answer database and is related to the first text.
[0148] For example, please refer to Figure 6, which is a schematic diagram of an embodiment of the present application, wherein the first text in a question-and-answer pair is highlighted to trigger the search for a new question-and-answer pair. As shown in Figure 6, after a user obtains an image by taking a photo or screenshot on a terminal device, a search button can be displayed above the image on the terminal device. After the user clicks the search button displayed above the image, a search for question-and-answer pairs related to the image is triggered.
[0149] In response to the user clicking the search button, the terminal device sends a search request and the image to the server, and receives and displays the question-answer pair returned by the server. Specifically, the question-answer pair displayed by the terminal device includes question text and answer text. The question text is specifically: "Is the bald eagle the national bird of the United States?" The answer text is specifically: "The bald eagle, also known as the American eagle. It is a large bird of prey. An adult eagle can reach 1 meter in length and has a wingspan of more than two meters. It mainly inhabits coasts, lakes and rivers. It is a species unique to North America and is the national bird of the United States." In addition, since the server also sends an indication message to the terminal device, the terminal device highlights the first text in the question-answer pair based on the content of the indication message: "bird of prey" and "national bird."
[0150] After the user further clicks on the first text "birds of prey" highlighted on the terminal device, the terminal device continues to send a search request to the server, requesting a search for question-answer pairs related to the first text "birds of prey." After the server further returns a new question-answer pair, the terminal device can then display the new question-answer pair. The new question-answer pair includes a question text, an answer text, and an answer image. In the new question-answer pair, the question text is specifically: "What are some common birds of prey?" The answer text is specifically: "Generally speaking, there are two major categories of birds of prey: one is falconiformes, such as eagles and vultures, and the other is owls. Birds of prey are carnivorous birds, some of which are scavengers. Birds of prey have downward-curved hooked beaks and very sharp claws. They have good eyesight and can find food at great heights."
[0151] In addition, after browsing the content of a new question-and-answer pair, the user can use the back button in the upper left corner of the terminal device's interface to return to the previous search result, allowing the user to continue clicking on other text. Furthermore, new text may be highlighted in the new question-and-answer pair, allowing the user to continue clicking on the newly highlighted text in the current question-and-answer pair to achieve continuous knowledge acquisition.
[0152] In this solution, by indicating the first text to be highlighted in the feedback search results in the indication information, the terminal device can highlight the first text in the search results, making it easier for users to quickly capture the key content in the search results. At the same time, it further guides users to click on the highlighted first text to ask new questions, triggering a new round of question-and-answer retrieval services, truly realizing what users think and facilitating users to quickly acquire new knowledge.
[0153] Optionally, in the process of obtaining the second search result based on the first text retrieval, the server may first generate a first question text based on the object included in the first image and the first text. For example, the server may input the name of the object included in the first image and the first text into a pre-trained text generation model, and the text generation model generates the first question text. The text generation model is able to recognize the name of the input object and the meaning of the first text, and then generate the corresponding first question text based on the semantic relationship between the two. For example, assuming that the object included in the first image is a cat and the first text is a balcony, the first question text generated based on "cat" and "balcony" can specifically be "Do I need to enclose the balcony if I keep a cat?"
[0154] After obtaining the first question text, the server then searches multiple question-answer pairs based on the first question text to obtain the second question-answer pair with the highest degree of match to the first question text. Specifically, since each question-answer pair includes a question text, the degree of match between a question-answer pair and the first question text can be understood as the similarity between the question texts. The server can calculate the similarity between the first question text and the question text of each of the multiple question-answer pairs, and then determine the second question-answer pair with the highest similarity to the first question text.
[0155] For example, please refer to Figure 7, which is a schematic diagram of another embodiment of the present application provided by highlighting the first text in a question-answer pair to trigger a new question-answer pair retrieval. As shown in Figure 7, in the first question-answer pair displayed on the terminal device, the question text is "Where is the best place to put the cat litter box?", and the answer text is "The cat litter box should be placed in a quiet place, which will make the cat feel more secure. Generally speaking, an enclosed balcony or the bathroom in the master bedroom are good choices, while next to a noisy washing machine or a wet faucet are not ideal locations for placing a cat litter box." Among them, in the answer text, the first text "balcony" and "washing machine" are highlighted.
[0156] In one possible operation mode, after the user clicks on the first text "balcony", the terminal device triggers a new round of question-answer retrieval based on the first text "balcony", and obtains a new question-answer pair returned by the server. The new question-answer pair returned by the server is retrieved based on the object "cat" included in the original image and the first text "balcony". Specifically, the new question-answer pair returned by the server includes a question text, an answer text, and an answer image. The question text is "Do I need to enclose the balcony if I have a cat?"; the answer text is "The balcony needs to be enclosed. Because cats are very curious animals! If there are flying insects or birds outside the window, they may not be able to resist catching them and cause them to fall! Another point is that cats playing on the balcony may fall because they don't hold on firmly."; the answer image is a schematic diagram of a cat playing on the balcony.
[0157] In another possible operation mode, after the user clicks on the first text "washing machine", the terminal device triggers a new round of question-answer retrieval based on the first text "washing machine" and obtains a new question-answer pair returned by the server. The new question-answer pair returned by the server is retrieved based on the object "cat" included in the original image and the first text "washing machine". Specifically, the new question-answer pair returned by the server includes a question text, an answer text, and an answer image. The question text is "What should I do if my cat likes to get into the washing machine?"; the answer text is "First, always close the washing machine door or the access door to the washing machine. Second, spray some smells that cats don't like on the washing machine, such as lemon or chili. Alternatively, put some items that cats don't like on the washing machine, such as orange peels or paper."; the answer image is a schematic diagram of a cat getting into a washing machine.
[0158] In this solution, by combining the first text in the question-answer pair with the object in the image and considering the relationship between the first text and the object in the image, the user's question intention is fully understood, and a question text related to the scene shown in the current image can be generated. Then, based on the generated question text, a search is performed to obtain a question-answer pair that meets the user's intention, making it easier for users to quickly acquire new knowledge of interest.
[0159] The above describes how, when returning a question-and-answer pair, the server instructs the user to highlight the first text in the question-and-answer pair, thereby guiding the user to click on the first text to trigger a new question-and-answer search, thereby expanding the scope of questions and answers. In some embodiments, the server may also identify the first image used to retrieve the question-and-answer pair, instruct the user to highlight the first object in the first image, and thereby guide the user to click on the first object in the first image to trigger a new question-and-answer search.
[0160] For example, in step 504 above, the server not only sends the first search result to the terminal device, but also sends a second image to the terminal device, where the second image is used to mark and display the first object in the first image. That is, the second image is obtained by marking the first object in the first image based on the first image. Therefore, the content in the second image is the same as that in the first image, with the only difference being that the first object is marked in the second image with a dotted box or highlighting. In this way, after receiving the second image and the first search result, the terminal device can display the second image and the first search result simultaneously on the interface.
[0161] In the case where the first image includes multiple objects, the first object may be one or more objects in the first image. For example, if the first image includes multiple objects such as a cat, a litter box, and a carpet, the first object may be the cat and / or the litter box in the first image. In other cases, the first object may also be one or more parts of an object in the first image. For example, if the first image includes an eagle, the first object may be a part of the eagle in the first image, such as its beak or claws.
[0162] Alternatively, the server may send marking information to the terminal device, where the marking information indicates the location of the first object in the first image, and the marking information is used to instruct the terminal device to mark the first object in the first image. In this way, the terminal device can mark the first object in the first image based on the marking information, thereby obtaining a second image.
[0163] Optionally, the first object is the object with the highest search volume among all objects included in the first image. Alternatively, the first object is the object corresponding to the question-answer pair with the highest confidence among all objects included in the first image. In general, the first object can be an object with high attention in the first image, or an object with abundant and authoritative answers. This embodiment does not limit the method for determining the first object.
[0164] When a terminal device displays a second image tagged with a first object, the user can click on the tagged first object in the second image to trigger a search for question-and-answer pairs based on the first object. In this case, the terminal device can send a third search request to the server. That is, the server can receive the third search request from the terminal device, which is used to search for question-and-answer pairs related to the first object in the second image.
[0165] After obtaining the third search request, the server performs a question-answer pair search in the question-answer base based on the first object indicated in the third search request, thereby determining and sending a third search result to the terminal device, where the third search result is used to indicate the text in the third question-answer pair related to the first object. The third question-answer pair is a question-answer pair determined from multiple question-answer pairs and related to the first object. For example, since the first object can be described in text, the server can retrieve the question text that matches the text of the first object in the question-answer base based on the text of the first object (such as the name of the first object) in combination with the user's historical search behavior, and then determine that the question-answer pair to which the question text belongs is the third question-answer pair that needs to be sent to the terminal device.
[0166] For example, please refer to Figure 8, which is a schematic diagram of an embodiment of the present application, providing a method for triggering the retrieval of new question-and-answer pairs by marking and displaying a first object in an image. As shown in Figure 8, after a user obtains an image by taking a photo or screenshot on a terminal device, a search button can be displayed above the image on the terminal device. After the user clicks the search button displayed above the image, a search for question-and-answer pairs related to the image is triggered.
[0167] In response to the user clicking the search button, the terminal device sends a search request and the image to the server, and receives and displays the question-answer pair and the new image returned by the server. In the new image displayed on the terminal device, the bald eagle's beak and claws are marked with a dotted box, indicating that the bald eagle's beak and claws are the first object.
[0168] After the user further clicks on the first object "bald eagle's beak" marked and displayed on the terminal device, the terminal device continues to send a search request to the server to request a search for question-and-answer pairs related to the first object "bald eagle's beak". After the server further returns a new question-and-answer pair, the terminal device can then display the new question-and-answer pair. The new question-and-answer pair includes a question text and an answer text. In the new question-and-answer pair, the question text is specifically: "Do birds have sensory organs?" The answer text is specifically: "Some birds have a soft skin connecting the beak to the front of the head, called cere. The cere is rich in tactile corpuscles, which are a type of sensory organ."
[0169] In this solution, by returning the second image marked with the first object to the terminal device, the terminal device can display the second image marked with the first object to the user, further guiding the user to pay attention to and click on the first object marked in the second image to ask new questions about the details in the image, triggering a new round of question-and-answer retrieval services, and facilitating the user to quickly acquire deeper new knowledge.
[0170] The above describes how the server, based on user interaction, continuously searches for and returns the question-and-answer pairs of interest from multiple question-and-answer pairs in the Q&A database. For easier understanding, the following describes how the server constructs the question-and-answer pairs in the Q&A database.
[0171] For example, before constructing the question-answer pair, the server may first obtain a second question text, which may be provided manually or automatically generated by a machine. This embodiment does not specifically limit the source of the second question text.
[0172] The server then searches for answers online based on the second question text, obtaining at least one answer text that matches the second question text. For example, the server searches for answers on various web pages based on the second question text, thereby obtaining web page content that has a high probability of matching the second question text. This web page content can then serve as the at least one answer text that matches the second question text. When searching for answers online, the server is not limited to encyclopedia sites; instead, it can search for answers across all publicly available online resources, such as web pages, posts, or question-and-answer records, thereby expanding the source of information.
[0173] Finally, the server determines the target answer text with the highest degree of match with the second question text in at least one answer text. In this way, the second question text and the target answer text can be used to form one question-answer pair among multiple question-answer pairs. For example, the server can use a pre-trained neural network model to determine the degree of match between the answer text and the second question text, thereby selecting the target answer text with the highest degree of match with the second question text in at least one answer text. The neural network model can be a natural language processing model that can realize the matching degree determination between question and answer texts. For example, the neural network model can be a question-answer matching degree determination model based on BERT.
[0174] In this solution, by performing an answer search on the Internet including various public materials based on the pre-acquired question text, and further calculating the matching degree between the question text and the searched answer text, the answer text that best matches the question text is selected to form a question-answer pair. This can effectively ensure the diversity of the answer text sources and ensure that the constructed question-answer pairs can have a high degree of confidence.
[0175] Optionally, after constructing the question-answer pair corresponding to the second question text, the server can generate multiple question texts based on the second question text. The multiple question texts and the second question text are used to indicate the same question, and the multiple question texts are expressed differently from the second question text. That is to say, based on the second question text, multiple question texts with the same meaning as the second question text but different expressions can be generated, thereby expanding the number of question texts. For example, assuming that the second question text is "What should I do if my cat likes to drill into the washing machine?", then the following multiple question texts can be generated: "How to solve the problem of cats like to drill into the washing machine" and "What to do if cats often drill into the washing machine?"
[0176] After generating multiple question texts, multiple question-answer pairs are constructed by combining each of the multiple question texts with the target answer text. In other words, since the multiple question texts generated based on the second question text have the same meaning as the second question text but are expressed differently, the same answer text can be assigned to these question texts, thereby quickly completing the construction of multiple question-answer pairs.
[0177] In this solution, more question texts with the same meaning but different expressions are generated based on the question texts in the constructed question-answer pairs, thereby expanding the number of question texts and improving the coverage of the question-answer base.
[0178] It is understandable that for some questions, the answers to these questions may not be unique. In other words, different people may hold different views on the same question, resulting in some questions having multiple answers corresponding to different viewpoints. In this case, when generating the question-answer pairs corresponding to these question texts, the server may retain the answers to the different viewpoints, so that the answer text of the question-answer pair simultaneously includes multiple answers corresponding to different viewpoints.
[0179] Exemplarily, the server first obtains a third question text, which can be manually provided or automatically generated by a machine. The server then performs an online answer search based on the third question text, obtaining multiple answer texts that match the third question text. The multiple answer texts obtained by the server correspond to multiple different viewpoints, meaning that not all of the multiple answer texts correspond to the same viewpoint. The multiple answer texts can be divided into multiple parts based on the number of viewpoints, with the answer text in each part corresponding to a specific viewpoint.
[0180] In this way, the server can fuse the answer texts corresponding to the same viewpoint in the multiple answer texts based on the viewpoints corresponding to each answer text, thereby obtaining multiple fused answer texts. The third question text and the fused answer texts are used to form one of the multiple question-answer pairs. In other words, the answer text of the question-answer pair corresponding to the third question text actually includes multiple answers corresponding to different viewpoints.
[0181] For example, suppose the server searches for the question "What fruit should I eat when I'm sick?" and obtains six answer texts. Three of these answer texts suggest eating apples, two suggest eating watermelon, and one suggests eating pears. The server can then merge the three answer texts that suggest eating apples to obtain a single answer text that also suggests eating apples. The server can also merge the two answer texts that suggest eating watermelon to obtain a single answer text that also suggests eating watermelon.
[0182] In this solution, when performing an answer search on a question text and obtaining answers corresponding to different viewpoints, multiple answers corresponding to the same viewpoint are integrated to achieve the goal of summarizing the answers according to viewpoints, so that the question-answer pair finally constructed can include answer texts corresponding to multiple viewpoints, ensuring the comprehensiveness of the answers.
[0183] For ease of understanding, the application process of the image-based question-answering method provided in the embodiments of the present application will be described in detail below with reference to specific examples.
[0184] Please refer to Figure 9, which is a schematic diagram of the structure of a server provided in an embodiment of the present application. As shown in Figure 9, the server, at the software level, includes the answer extraction / determination module, vector engine, and vectorized reasoning module of the machine learning platform. The program code used to implement the method of this embodiment resides in the answer extraction / determination module, vector engine, and vectorized reasoning module of the machine learning platform and runs in the server's host memory or image processor memory.
[0185] At the hardware level, the server includes a question-and-answer database deployed on a storage medium such as a hard drive or memory. The server's software module interacts with the storage medium in the hardware to access the question-and-answer database and search for corresponding question-and-answer pairs.
[0186] Specifically, the answer extraction / determination module is responsible for building a question and answer base, which is specifically used to extract answer texts that match the question text on the Internet, and determine the answer text that best matches the question text among the multiple answer texts extracted, thereby realizing the construction of question and answer pairs.
[0187] The vector engine is used to obtain image features and perform a search for question-answer pairs in the question-answer base based on the image features, thereby obtaining multiple candidate question-answer pairs.
[0188] The vectorized reasoning service is used to select a question-answer pair to be returned to the user from multiple candidate question-answer pairs based on the user's historical search behavior, and to determine the first text to be highlighted in the question-answer pair and the first object to be marked in the image.
[0189] Please refer to Figure 10, which is a schematic diagram of a process for a user to obtain question-answer pairs based on an image, provided by an embodiment of the present application. As shown in Figure 10, in the process of a user obtaining question-answer pairs based on an image, the user first inputs an image for retrieving question-answer pairs, and the terminal device or server extracts the image features. Then, the server retrieves the question-answer pairs in the question-answer base based on the image features (i.e., the question-answer retrieval indicated by step S1 in Figure 10), and obtains multiple candidate question-answer pairs. Among them, before executing the question-answer retrieval, the server first performs site knowledge extraction on the network, and realizes question-answer pair generation based on the extracted knowledge, and performs question-answer pair index construction for the generated question-answer pairs, thereby completing the construction of the question-answer base.
[0190] After retrieving multiple candidate question-answer pairs, the intent understanding process, that is, combining the user's historical search behavior to select the question-answer pair that best matches the user from multiple candidate question-answer pairs, and displaying the question-answer pairs to the user, realizes the display of graphical question-answer results.
[0191] In addition, in the process of displaying the question-answer pair to the user, the server can execute the question-answer recommendation indicated in step S3 in Figure 10, that is, highlighting the first text in the question-answer pair or the first object in the image to the user, thereby guiding the user to click on the first text or the first object to trigger a new round of question-answer retrieval service, and finally displaying the recommendation results obtained from the new round of question-answer retrieval service to the user.
[0192] In general, the method provided in this embodiment includes four core processes, namely, database construction, question-answer pair retrieval, intent understanding, and question-answer recommendation. These four core processes will be introduced in detail below.
[0193] Please refer to Figure 11, which is a schematic diagram of the core process of obtaining question-answer pairs based on images provided by an embodiment of the present application. The following will introduce the four core processes of base library construction, question-answer pair retrieval, intent understanding, and question-answer recommendation in conjunction with Figure 11.
[0194] (1) Base library construction
[0195] In the process of building the question-answer base, based on the pre-acquired question text, a search for answers is performed on the Internet, and the paragraphs on the Internet that may include the answers are extracted, thereby obtaining the paragraphs on the Internet that may include the answers. Specifically, the way to perform the search for answers on the Internet can be to search for information such as documents, images, and text on encyclopedia sites, vertical domain sites, and general web pages. Generally speaking, searching encyclopedia sites can obtain high-quality and high-frequency answers; searching vertical domain sites can obtain high-quality and professional answers; searching general web pages can obtain massive and general answers. In this way, by performing the search for answers on various websites and web pages, the source of the answer text can be enriched, which is conducive to answering various complex questions. After searching for web page content that matches the question text, since a web page usually includes a lot of content, and some of the content may not be highly relevant to the question text, in order to extract accurate answer text, the web page content can be input into a long answer extraction model based on machine reading comprehension (MRC), and the web page content is coarse-grained segmented to filter out paragraphs with a high relevance to the question text, i.e., the answer text. In this way, by combining the multiple answer texts obtained by screening with the question text, the generation of question-answer pairs can be preliminarily completed.
[0196] After performing paragraph extraction on the searched webpage content to obtain multiple answer texts, the server can input the extracted multiple answer texts and the question text into the BERT-based question-answer matching determination model. The encoder in the BERT-based question-answer matching determination model performs dense vector encoding on the multiple answer texts and the question text, respectively, to obtain vector features corresponding to each answer text and vector features corresponding to the question text. The vector features corresponding to the question text and the vector features corresponding to each answer text are then similarly measured to obtain the answer text with the highest similarity. Based on this answer text and the question text, an accurate question-answer pair is constructed.
[0197] After constructing a question-answer pair by searching the content on the web page, the original question text can be input into the Seq2Seq generation model to generate a new question text with the same meaning but different syntax. The new question text is then combined with the original answer text in the question-answer pair to form a new question-answer pair, thereby expanding the questions in the question-answer base database and improving the coverage of the question-answer base database.
[0198] In addition, in order to facilitate the rapid retrieval of corresponding question-answer pairs in the question-answer base, a corresponding index can be constructed for each question-answer pair based on the text features of each question-answer pair in the question-answer base, so as to facilitate the subsequent rapid retrieval of question-answer pairs based on the index of the question-answer pair.
[0199] (2) Question-Answer Pair Retrieval
[0200] After obtaining the user-provided image for question-answer pair retrieval, a multimodal model can be used to perform dense vector encoding on the image to obtain image features. Then, based on the image features, a vector engine is used to search for question-answer pairs in the question-answer base database, thereby retrieving multiple candidate question-answer pairs that have the highest degree of match with the image features. The vector engine can be, for example, the Facebook AI Similarity Search (FAISS). This embodiment does not specifically limit the vector engine used for question-answer pair retrieval.
[0201] 3. Understanding Intention
[0202] When the server understands the user's intent, it first analyzes the user's historical search behavior and constructs a corresponding user profile. Then, based on the user profile, it selects the question-answer pair with the highest relevance to the user profile from among multiple candidate question-answer pairs and returns that question-answer pair to the user.
[0203] In addition, in some embodiments, among multiple candidate question-answer pairs, if multiple candidate question-answer pairs have the highest relevance to the user portrait, then multiple question-answer pairs with the highest relevance can be simultaneously used as question-answer pairs returned to the user. For example, please refer to Figure 12, which is a schematic diagram of an embodiment of the present application for simultaneously displaying multiple question-answer pairs to a user. As shown in Figure 12, when a user provides an image about red wine, the server searches for three question-answer pairs with the highest match to the image based on the user's historical search behavior, and then returns these three question-answer pairs to the user. Among them, the question texts of these three question-answer pairs are: "Does red wine have a shelf life?", "What types of red wine are there?" and "What food is best to drink red wine with?" When the terminal device displays multiple question-answer pairs to the user, the user can select one of the question-answer pairs for in-depth browsing.
[0204] It should be noted that, for the question-answer pairs constructed by the server, due to the differences in the question texts, the viewpoints of the answers corresponding to some question texts are unique, so the answer texts included in the question-answer pairs corresponding to these question texts are actually single-viewpoint answers, that is, they do not have multiple different viewpoints. For example, please refer to Figure 13, which is a schematic diagram of a question-answer pair including a single-viewpoint answer provided in an embodiment of the present application. As shown in Figure 13, for the question text "What are the types of red wine?", the answer text in the question-answer pair is a single-viewpoint answer, specifically "1. Classification by the color of the wine: white wine, red wine, rosé wine; 2. Classification by the sugar content in the wine: dry wine, semi-dry wine, semi-sweet wine, sweet wine."
[0205] For other question texts, the viewpoints of the answers corresponding to these question texts may not be unique, that is, there may be answers with multiple different viewpoints. In this case, if the server searches for the answer text corresponding to a certain question text during the construction of the question-answer pair and finds that the answer text includes answers corresponding to different viewpoints, then the server can first perform viewpoint extraction on these answers to determine the viewpoint of each answer. Then, based on the arguments given in each answer (i.e., the specific content of the answer), the matching degree between the arguments and the viewpoint of each answer is determined, and the multiple answers with the highest viewpoint matching degree under each viewpoint are selected, and the multiple answers under each viewpoint are aggregated to obtain the aggregated answer under each viewpoint. In addition, for each viewpoint, the number of answers corresponding to each viewpoint can also be calculated, thereby calculating the support degree of each viewpoint. For example, please refer to Figure 14, which is a schematic diagram of a question-answer pair including multiple viewpoint answers provided in an embodiment of the present application. As shown in Figure 14, for the question text "Does red wine have a shelf life?", the answer text in the question-answer pair is a multiple viewpoint answer, one of which is "no shelf life" and the other is "yes". Moreover, the support level for the viewpoint “there is no shelf life” is 8, and the support level for the viewpoint “there is a shelf life” is 5.
[0206] Furthermore, when the server obtains answer texts corresponding to multiple different viewpoints for a question text, the server can sort these answer texts according to user click data and the number of supporting viewpoints, thereby giving priority to displaying high-frequency, high-support answer texts to users.
[0207] (IV) Q&A Recommendations
[0208] After the server performs intent understanding and determines the question and answer pairs that need to be displayed to the user, it can asynchronously perform two business calculations: expanded question and answer and fine-grained question and answer. Among them, expanded question and answer will highlight the first text in the answer text of the current question and answer pair; fine-grained question and answer will frame and display the first object (i.e., detail entity) in the input image. After the user triggers and clicks on the first text or the first object, the server can infer the user portrait of the current user based on the relationship between the first text or the first object and the image itself, and use the text generation model to generate question text related to the scene, thereby triggering a new round of question and answer retrieval business, allowing users to further obtain information from both depth and breadth. Among them, the first text and the first object can be identified by entity recognition technology, specifically a specific entity or an entity with high attention. In the case where the first text or the first object is an entity identified by an entity recognition method, using a text generation model to generate question text related to the scene is actually an entity-based query generation process.
[0209] For example, please refer to Figures 15 and 16. Figure 15 is a schematic diagram of an embodiment of the present application for highlighting the first text in a question-answer pair and marking the first object in an image; Figure 16 is a schematic diagram of an embodiment of the present application for generating a new question text based on the first text and the first object.
[0210] As shown in Figure 15, for an image containing the word "bald eagle," the returned question-answer pair highlights the first text "raptor" and "national bird." Furthermore, the first objects "mouth" and "claws" are marked with dashed boxes in the image. When a user clicks on either the first text "raptor" or the first object "mouth," a new question-answer pair is returned.
[0211] As shown in Figure 16, for the image with the content of "a cat squatting in the litter box", in the returned question-answer pair, the question text is "Where is the best place to put the litter box?", and the answer text is "The litter box should be placed in a quiet place, which will make the cat feel more secure. Generally speaking, an enclosed balcony or the bathroom in the master bedroom are good choices, while a noisy washing machine or a wet faucet are not ideal places to put the litter box." In addition, in the answer text, the first text "balcony" and "washing machine" are highlighted. When the user clicks on the first text "balcony", the server can generate a new question text "Do I need to enclose the balcony if I have a cat?" based on the user portrait, image content and the first text "balcony". When the user clicks on the first text "washing machine", the server can generate a new question text "What should I do if my cat likes to get into the washing machine?" based on the user portrait, image content and the first text "washing machine".
[0212] Furthermore, in the image displayed on the terminal device, the first objects "cat" and "litter box" are marked and displayed. When the user clicks the first object "cat," the server can generate new questions based on the first object "cat" and the user's profile, such as "How to raise a British Shorthair Blue cat?" or "What is the normal weight of a British Shorthair Blue cat?" When the user clicks the first object "litter box," the server can generate new questions based on the first object "litter box" and the user's profile, such as "How often should the litter box be cleaned?" or "Is an open or closed litter box better?"
[0213] The above embodiments introduce the image-based question-answering method provided by the embodiments of the present application. The following will introduce the device for executing the above image-based question-answering method.
[0214] Please refer to Figure 17, which is a structural diagram of an image-based question-answering device provided in an embodiment of the present application. As shown in Figure 17, the image-based question-answering device includes: an acquisition module 1701, which is used to obtain a first search request and image features of a first image, the first search request is used to search for results related to the first image, and the first search request comes from a terminal device; a processing module 1702, which is used to obtain multiple candidate question-answer pairs based on image features, each of the multiple candidate question-answer pairs includes a question text and an answer text with a corresponding relationship, and the text features of the multiple candidate question-answer pairs match the image features; the processing module 1702 is also used to determine the first question-answer pair that matches the historical search behavior among the multiple candidate question-answer pairs based on the historical search behavior of the terminal device; a sending module 1703 is used to send the first search result to the terminal device, and the first search result is used to indicate the text in the first question-answer pair.
[0215] In one possible implementation, the sending module 1703 is further used to send indication information to the terminal device, where the indication information is used to instruct the terminal device to highlight the first text in the first search result, where the first text is part of the text in the first search result; the obtaining module 1701 is further used to obtain a second search request from the terminal device, where the second search request is used to search for question-and-answer pairs related to the first text; the sending module 1703 is further used to send a second search result to the terminal device, where the second search result is used to indicate the text in the second question-and-answer pair related to the first text.
[0216] In one possible implementation, the processing module 1702 is further used to: generate a first question text based on the object included in the first image and the first text; perform a search in multiple question-answer pairs based on the first question text to obtain a second question-answer pair that matches the first question text.
[0217] In a possible implementation, the first text is the text with the highest search volume among all texts of the first search result; or, the first text is the text corresponding to the question-answer pair with the highest confidence among all texts of the first search result.
[0218] In one possible implementation, the sending module 1703 is further used to send a second image to the terminal device, where the second image is used to mark and display the first object in the first image; the acquiring module 1701 is further used to acquire a third search request from the terminal device, where the third search request is used to search for results related to the first object; the sending module 1703 is further used to send a third search result to the terminal device, where the third search result is used to indicate text in a question-and-answer pair related to the first object.
[0219] In one possible implementation, the first object is the object with the highest search volume among all objects included in the first image; or, the first object is the object corresponding to the question-answer pair with the highest confidence among all objects included in the first image.
[0220] In one possible implementation, the acquisition module 1701 is further used to obtain a second question text; the processing module 1702 is further used to perform an answer search on the Internet based on the second question text to obtain at least one answer text that matches the second question text; the processing module 1702 is further used to determine a target answer text that has the highest match with the second question text in the at least one answer text, wherein the second question text and the target answer text are used to constitute a question-answer pair.
[0221] In one possible implementation, the processing module 1702 is further used to: generate multiple question texts based on the second question text, where the multiple question texts and the second question text are used to indicate the same question, and the multiple question texts and the second question text are expressed in different ways; and construct multiple question-answer pairs by combining each question text in the multiple question texts with the target answer text.
[0222] In one possible implementation, the acquisition module 1701 is also used to obtain a third question text; the processing module 1702 is also used to perform an answer search on the Internet based on the third question text to obtain multiple answer texts matching the third question text, and the multiple answer texts correspond to multiple different viewpoints; the processing module 1702 is also used to fuse the answer texts corresponding to the same viewpoint based on the viewpoint corresponding to each answer text in the multiple answer texts to obtain multiple fused answer texts; wherein the third question text and the multiple fused answer texts are used to form a question-answer pair.
[0223] In a possible implementation, the acquisition module 1701 is further configured to acquire a first search request and a first image from a terminal device; and the processing module 1702 is further configured to perform image feature extraction on the first image to obtain image features.
[0224] In a possible implementation, the image features come from a terminal device.
[0225] Next, an execution device provided in an embodiment of the present application is introduced. Please refer to Figure 18. Figure 18 is a structural schematic diagram of the execution device provided in an embodiment of the present application. The execution device 1800 can be specifically manifested as a mobile phone, a tablet, a laptop computer, a smart wearable device, a server, etc., which is not limited here. Specifically, the execution device 1800 includes: a receiver 1801, a transmitter 1802, a processor 1803 and a memory 1804 (wherein the number of processors 1803 in the execution device 1800 can be one or more, and Figure 18 takes one processor as an example), wherein the processor 1803 may include an application processor 18031 and a communication processor 18032. In some embodiments of the present application, the receiver 1801, the transmitter 1802, the processor 1803 and the memory 1804 may be connected via a bus or other means.
[0226] Memory 1804 may include read-only memory and random access memory, and provides instructions and data to processor 1803. A portion of memory 1804 may also include non-volatile random access memory (NVRAM). Memory 1804 stores processor and operation instructions, executable modules, or data structures, or subsets or extended sets thereof. The operation instructions may include various operation instructions for implementing various operations.
[0227] Processor 1803 controls the operation of the execution device. In specific applications, the various components of the execution device are coupled together via a bus system. In addition to a data bus, the bus system may also include a power bus, a control bus, and a status signal bus. However, for clarity, all bus systems are referred to as a bus system in the figure.
[0228] The methods disclosed in the above embodiments of the present application can be applied to or implemented by processor 1803. Processor 1803 can be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by hardware integrated logic circuits or software instructions in processor 1803. The above processor 1803 can be a general-purpose processor, a digital signal processor (DSP), a microprocessor, or a microcontroller, and can further include an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The processor 1803 can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in conjunction with the embodiments of the present application can be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium well-known in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in memory 1804, and processor 1803 reads information from memory 1804 and, in conjunction with its hardware, completes the steps of the above method.
[0229] Receiver 1801 can be used to receive input digital or character information and generate signal input related to executing device-related settings and function control. Transmitter 1802 can be used to output digital or character information through the first interface. Transmitter 1802 can also be used to send instructions to the disk pack through the first interface to modify data in the disk pack. Transmitter 1802 can also include a display device such as a display screen.
[0230] In an embodiment of the present application, in one case, the processor 1803 is used to execute the method in the embodiment corresponding to Figure 5.
[0231] The electronic device provided in the embodiment of the present application may specifically be a chip, and the chip includes: a processing unit and a communication unit, wherein the processing unit may be, for example, a processor, and the communication unit may be, for example, an input / output interface, a pin or a circuit, etc. The processing unit may execute the computer execution instructions stored in the storage unit so that the chip in the execution device executes the rendering method described in the above embodiment. Optionally, the storage unit is a storage unit in the chip, such as a register, a cache, etc. The storage unit may also be a storage unit located outside the chip in the wireless access device, such as a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM), etc.
[0232] Please refer to Figure 19, which is a schematic diagram of the structure of a computer-readable storage medium provided in an embodiment of the present application. The present application also provides a computer-readable storage medium. In some embodiments, the method disclosed in Figure 5 above can be implemented as computer program instructions encoded in a machine-readable format on a computer-readable storage medium or on other non-transitory media or products.
[0233] 19 schematically illustrates a conceptual partial view of an example computer-readable storage medium including a computer program for executing a computer process on a computing device, arranged in accordance with at least some embodiments presented herein.
[0234] In one embodiment, computer-readable storage medium 1900 is provided using signal-bearing medium 1901. Signal-bearing medium 1901 may include one or more program instructions 1902 that, when executed by one or more processors, may provide the functionality or portions of the functionality described above with respect to FIG. 5. Furthermore, program instructions 1902 in FIG. 19 also depict example instructions.
[0235] In some examples, signal bearing medium 1901 may include computer readable medium 1903 such as, but not limited to, a hard drive, compact disk (CD), digital video disk (DVD), digital tape, memory, ROM or RAM, and the like.
[0236] In some embodiments, the signal-bearing medium 1901 may include a computer-recordable medium 1904, such as, but not limited to, a memory, a read / write (R / W) CD, a R / W DVD, or the like. In some embodiments, the signal-bearing medium 1901 may include a communication medium 1905, such as, but not limited to, a digital and / or analog communication medium (e.g., a fiber optic cable, a waveguide, a wired communication link, a wireless communication link, or the like). Thus, for example, the signal-bearing medium 1901 may be communicated via a wireless form of the communication medium 1905 (e.g., a wireless communication medium conforming to the IEEE 802 standard or other transmission protocol).
[0237] The one or more program instructions 1902 may be, for example, computer-executable instructions or logic-implemented instructions. In some examples, the computing device may be configured to provide various operations, functions, or actions in response to the program instructions 1902 communicated to the computing device via one or more of computer-readable media 1903, computer-recordable media 1904, and / or communication media 1905.
[0238] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware, or any combination thereof. When implemented by software, all or part of the embodiments may be implemented in the form of a computer program product.
[0239] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present invention is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that a computer can store or a data storage device such as a server or data center that includes one or more available media integrated. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)).
[0240] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0241] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0242] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0243] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0244] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0245] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. An image-based question answering method, It is characterized in that include: Acquire a first search request and an image feature of a first image, wherein the first search request is used to search for results related to the first image, and the first search request comes from a terminal device; Based on the image features, a plurality of candidate question-answer pairs are obtained, each of the plurality of candidate question-answer pairs includes a question text and an answer text having a corresponding relationship, and the text features of the plurality of candidate question-answer pairs match the image features; Based on the historical search behavior of the terminal device, determining a first question-answer pair matching the historical search behavior from among the plurality of candidate question-answer pairs; A first search result is sent to the terminal device, where the first search result is used to indicate text in the first question-answer pair.
2. The method according to claim 1, It is characterized in that The method further comprises: Sending instruction information to the terminal device, where the instruction information is used to instruct the terminal device to highlight a first text in the first search result, where the first text is a portion of text in the first search result; Acquire a second search request from the terminal device, where the second search request is used to search for results related to the first text; A second search result is sent to the terminal device, where the second search result is used to indicate text in a second question-answer pair related to the first text.
3. The method according to claim 2, It is characterized in that The method further comprises: Generate a first question text based on the object included in the first image and the first text; A search is performed among a plurality of question-answer pairs based on the first question text to obtain the second question-answer pair matching the first question text.
4. The method according to claim 2 or 3, It is characterized in that The first text is the text with the highest search volume among all the texts of the first search result; Alternatively, the first text is the text corresponding to the question-answer pair with the highest confidence among all the texts of the first search result.
5. The method according to any one of claims 1 to 4, It is characterized in that The method further comprises: Sending a second image to the terminal device, where the second image is used to mark and display a first object in the first image; Obtaining a third search request from the terminal device, where the third search request is used to search for results related to the first object; A third search result is sent to the terminal device, where the third search result is used to indicate text in a question-answer pair related to the first object.
6. The method according to claim 5, It is characterized in that The first object is an object having the highest search volume among all objects included in the first image; Alternatively, the first object is an object corresponding to a question-answer pair with the highest confidence among all objects included in the first image.
7. The method according to any one of claims 1 to 6, It is characterized in that The method further comprises: Get the second question text; Performing an answer search on the network based on the second question text to obtain at least one answer text matching the second question text; A target answer text having the highest matching degree with the second question text is determined in the at least one answer text, wherein the second question text and the target answer text are used to form a question-answer pair.
8. The method according to claim 7, It is characterized in that The method further comprises: Based on the second question text, generate a plurality of question texts, wherein the plurality of question texts and the second question text are used to indicate the same question, and the plurality of question texts and the second question text are expressed in a different manner; By combining each question text in the multiple question texts with the target answer text, multiple question-answer pairs are constructed.
9. The method according to any one of claims 1 to 8, It is characterized in that The method further comprises: Get the third question text; Performing an answer search on the network based on the third question text to obtain a plurality of answer texts matching the third question text, wherein the plurality of answer texts correspond to a plurality of different viewpoints; Based on the viewpoints corresponding to each answer text in the multiple answer texts, answer texts corresponding to the same viewpoint are merged to obtain multiple answer texts after merging; The third question text and the multiple answer texts after fusion processing are used to form a question-answer pair.
10. The method according to any one of claims 1 to 9, It is characterized in that The obtaining of the first search request and the image feature of the first image includes: Acquire the first search request and the first image from the terminal device; Perform image feature extraction on the first image to obtain the image features.
11. The method according to any one of claims 1 to 9, It is characterized in that The image features are from the terminal device.
12. An image-based question-answering device, It is characterized in that include: The acquisition module is used to acquire a first search request and an image feature of a first image, wherein the first search request is used to search searching for results related to a first image, wherein the first search request comes from a terminal device; A processing module, configured to obtain a plurality of candidate question-answer pairs based on the image features, each of the plurality of candidate question-answer pairs comprising a question text and an answer text having a corresponding relationship, and the text features of the plurality of candidate question-answer pairs match the image features; The processing module is further configured to determine, based on the historical search behavior of the terminal device, a first question-answer pair matching the historical search behavior from among the multiple candidate question-answer pairs; A sending module is used to send the first search result to the terminal device, where the first search result is used to indicate the text in the first question and answer pair.
13. The device according to claim 12, It is characterized in that The sending module is further used to send instruction information to the terminal device, where the instruction information is used to instruct the terminal device to highlight the first text in the first search result, where the first text is a part of the text in the first search result; The acquisition module is further used to acquire a second search request from the terminal device, where the second search request is used to search for results related to the first text; The sending module is further used to send a second search result to the terminal device, where the second search result is used to indicate text in a second question-answer pair related to the first text.
14. The device according to claim 13, It is characterized in that The processing module is further used for: Generate a first question text based on the object included in the first image and the first text; A search is performed among a plurality of question-answer pairs based on the first question text to obtain the second question-answer pair matching the first question text.
15. The device according to claim 13 or 14, It is characterized in that The first text is the text with the highest search volume among all the texts of the first search result; Alternatively, the first text is the text corresponding to the question-answer pair with the highest confidence among all the texts of the first search result.
16. The device according to any one of claims 12 to 15, It is characterized in that The sending module is further used to send a second image to the terminal device, where the second image is used to mark and display the first object in the first image; The acquisition module is further used to acquire a third search request from the terminal device, where the third search request is used to search for results related to the first object; The sending module is further used to send a third search result to the terminal device, where the third search result is used to indicate text in a question-answer pair related to the first object.
17. The device according to claim 16, It is characterized in that The first object is an object having the highest search volume among all objects included in the first image; Alternatively, the first object is an object corresponding to a question-answer pair with the highest confidence among all objects included in the first image.
18. The device according to any one of claims 12 to 17, It is characterized in that The acquisition module is further used to acquire the second question text; The processing module is further used to perform an answer search on the network based on the second question text to obtain at least one answer text matching the second question text; The processing module is further used to determine a target answer text having the highest matching degree with the second question text among the at least one answer text, wherein the second question text and the target answer text are used to form a question-answer pair.
19. The device according to claim 18, It is characterized in that The processing module is further used for: Based on the second question text, generate a plurality of question texts, wherein the plurality of question texts and the second question text are used to indicate the same question, and the plurality of question texts and the second question text are expressed in a different manner; By combining each question text in the multiple question texts with the target answer text, multiple question-answer pairs are constructed.
20. The device according to any one of claims 12 to 19, It is characterized in that The acquisition module is further used to acquire the third question text; The processing module is further used to perform an answer search on the network based on the third question text to obtain a plurality of answer texts matching the third question text, and the plurality of answer texts correspond to a plurality of different viewpoints; The processing module is further used to merge answer texts corresponding to the same viewpoint based on the viewpoint corresponding to each answer text in the multiple answer texts to obtain multiple answer texts after merging; The third question text and the multiple answer texts after fusion processing are used to form a question-answer pair. 21.An image-based question-answering device, It is characterized in that The device comprises a memory and a processor; the memory stores codes, the processor is configured to execute the codes, and when the codes are executed, the device executes the method according to any one of claims 1 to 11.
22. A computer storage medium, It is characterized in that The computer storage medium stores instructions, which, when executed by a computer, cause the computer to implement the method of any one of claims 1 to 11.
23. A computer program product, It is characterized in that The computer program product stores instructions, which, when executed by a computer, cause the computer to implement the method according to any one of claims 1 to 11.