Information acquisition method and device, electronic equipment and computer readable storage medium

By configuring multiple equivalent question templates and optimization question templates, and combining metadata to generate target question and answer information, the problem of single method of question and answer information generation in the visual question and answer system is solved, and the effectiveness of question and answer information and image content understanding ability are improved.

CN120472440APending Publication Date: 2025-08-12BEIJING ACAD OF ARTIFICIAL INTELLLIGENCE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510541635.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

In the prior art, the visual question and answer information generation method of the visual question and answer system is single, resulting in insufficient deep understanding and multi-angle analysis of image content, limiting the learning depth of the model and the effectiveness of the question and answer information.

Method used

By configuring multiple equivalent question templates, filter the target question templates from the preset question templates, combine metadata to generate target question and answer information for the original image, and optimize the question templates according to the content object type to improve the diversity and accuracy of question and answer information.

Benefits of technology

It realizes the generation of diverse contents based on different depths and angles, improves the effectiveness of question and answer information, and improves the understanding of image content by training the target visual model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120472440A_ABST
    Figure CN120472440A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an information acquisition method and device, electronic equipment and a computer readable storage medium, and the method comprises the steps: obtaining an original image and metadata corresponding to the original image, screening a target problem template from at least one equivalent preset problem template, and storing the target problem template in a database, and generating target question and answer information for the original image according to the target question template and the metadata. By configuring a plurality of equivalent question templates to flexibly select and generate the target question and answer information, the problem that the generation mode of the target question and answer information is single is solved, the question and answer information with various contents can be generated from different depths and different angles based on different question templates, and the effectiveness of the question and answer information is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of information processing technology, and specifically to an information acquisition method, device, electronic device, and computer-readable storage medium. Background Art

[0002] With the rapid development of artificial intelligence (AI), image recognition and understanding have become crucial components of intelligent systems. In particular, in the fields of optical character recognition (OCR) and visual question answering (VQA), precise understanding of image content and extraction of textual information enable a more intelligent human-computer interaction experience. Currently, image information acquisition technology is widely used in a variety of fields, including intelligent search, content analysis, and autonomous driving.

[0003] In existing technologies, visual question answering systems typically obtain an image and a question, analyze and process them using a model, and then provide an answer. For example, they can extract text from an image based on the question. Furthermore, the model is primarily built by constructing question-and-answer information (a dataset of question-and-answer pairs). Through training based on this information, the model can respond to certain image-based question tasks.

[0004] However, the target question and answer information generation method is still relatively simple. For example, it is limited to the simple task of "extracting text from an image", resulting in a lack of in-depth understanding and multi-angle analysis of the image content, making the generated question and answer information less effective. Furthermore, after the model is trained based on the question and answer information obtained in this single method, the learning depth of the model is limited. Summary of the Invention

[0005] The embodiments of the present application provide an information acquisition method, device, electronic device, and computer-readable storage medium, which can improve the effectiveness of question-and-answer information.

[0006] In a first aspect, an embodiment of the present application provides an information acquisition method, the method comprising:

[0007] Acquire an original image and metadata corresponding to the original image, wherein the metadata includes description information of the original image;

[0008] Filtering a target question template from at least one equivalent preset question template;

[0009] Target question-answering information for the original image is generated according to the target question template and the metadata.

[0010] In a second aspect, an embodiment of the present application further provides an information acquisition device, the device comprising:

[0011] An acquisition module, configured to acquire an original image and metadata corresponding to the original image, wherein the metadata includes description information of the original image;

[0012] A screening module, configured to screen a target question template from at least one equivalent preset question template;

[0013] A generation module is used to generate target question-answer information for the original image based on the target question template and the metadata.

[0014] Optionally, in some embodiments of the present application, generating target question-answer information for the original image according to the target question template and the metadata includes:

[0015] determining a content object type according to the original image or the metadata;

[0016] Optimizing the target question template according to the content object type to obtain an optimized question template;

[0017] The target question and answer information for the metadata is generated according to the optimized question template.

[0018] Optionally, in some embodiments of the present application, optimizing the target question template according to the content object type to obtain an optimized question template includes:

[0019] Determining prompt information of the target question in the target question template according to the content object type;

[0020] The target question is optimized according to the prompt information to obtain an optimized question template.

[0021] Optionally, in some embodiments of the present application, after generating target question-answer information for the original image according to the target question template and the metadata, the method further includes:

[0022] The original visual model is trained according to the target question and answer information and the original image corresponding to the target question and answer information to obtain a target visual model; wherein the target visual model is used to extract image content in the image to be identified.

[0023] Optionally, in some embodiments of the present application, after obtaining the original image and metadata corresponding to the original image, the method further includes:

[0024] Recognize the text record result of the original image by a text recognition model;

[0025] Generating target question-answer information for the original image according to the target question template and the metadata includes:

[0026] If the text record result indicates that text exists on the original image, target question-answer information for the original image is generated according to the target question template and the metadata.

[0027] Optionally, in some embodiments of the present application, obtaining the original image and metadata corresponding to the original image includes:

[0028] determining at least one image source according to a content object type, the content object type comprising at least one of a book, an album, a game, a movie, or music;

[0029] At least one original image of a content object type is obtained from the image source, and metadata corresponding to each original image is obtained from each image source, wherein the metadata is a description of the original image added at the corresponding image source.

[0030] Optionally, in some embodiments of the present application, the metadata includes an instance of at least one entity, and generating target question-answer information for the original image based on the target question template and the metadata includes:

[0031] extracting instances for each entity from the metadata;

[0032] The target question template is filled in according to the instances corresponding to each entity to obtain target question and answer information.

[0033] In a third aspect, an embodiment of the present application further provides an electronic device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, the steps in the above-mentioned information acquisition method are implemented.

[0034] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps in the above-mentioned information acquisition method are implemented.

[0035] In a fifth aspect, embodiments of the present application further provide a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in the various optional implementations described in the embodiments of the present application.

[0036] An embodiment of the present application obtains an original image and metadata corresponding to the original image, where the metadata includes descriptive information about the original image, filters a target question template from at least one equivalent preset question template, and generates target question-and-answer information for the original image based on the target question template and metadata.

[0037] Among them, the embodiment of the present application flexibly selects and generates target question and answer information by configuring multiple equivalent question templates, thereby solving the problem of a single method of generating target question and answer information, and helping to generate question and answer information with diverse content from different depths and angles based on different question templates, thereby improving the effectiveness of the question and answer information. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] In order to more clearly illustrate the technical solutions in this application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.

[0039] Figure 1 Schematic diagram of a scenario in which an electronic device according to an embodiment of the present application executes the information acquisition method;

[0040] Figure 2 Schematic diagram of the information acquisition method provided in the embodiment of the present application;

[0041] Figure 3 is a structural diagram of an information acquisition device provided in an embodiment of the present application;

[0042] Figure 4 It is a structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0043] The following will be combined with the drawings in this application to clearly and completely describe the technical solutions in this application. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making any creative efforts shall fall within the scope of protection of the present invention.

[0044] The embodiments of the present application provide an information acquisition method, device, electronic device and computer-readable storage medium. Specifically, the embodiments of the present application provide an information acquisition device suitable for electronic devices, which is used to improve the effectiveness of question and answer information. Specifically, the electronic device includes a terminal device or a server, and the terminal device includes but is not limited to desktop computers, laptops, mobile phones, tablet computers and other devices. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content distribution networks (CDN, Content Delivery Network), and basic cloud computing services such as big data and artificial intelligence platforms. The server can be directly or indirectly connected via wired or wireless communication.

[0045] See also Figure 1 , Figure 1 : is a schematic diagram of a scenario in which an electronic device according to an embodiment of the present application performs the information acquisition method. The specific execution process of the electronic device performing the information acquisition method is as follows:

[0046] The electronic device 10 obtains an original image and metadata corresponding to the original image, where the metadata includes description information about the original image, selects a target question template from at least one equivalent preset question template, and generates target question-answer information for the original image based on the target question template and the metadata.

[0047] For example, after the user obtains the original image and corresponding metadata through an electronic device, he operates the electronic device or the electronic device automatically triggers the process of generating target question and answer information. Specifically: the electronic device filters the target question template from multiple equivalent, pre-set question templates, and generates target question and answer information for the original image based on the target question template and metadata.

[0048] In summary, the embodiment of the present application solves the problem of a single method of generating target question and answer information by configuring multiple equivalent question templates, and helps to generate question and answer information with diverse content from different depths and angles based on different question templates, thereby improving the effectiveness of the question and answer information.

[0049] The following are detailed descriptions. It should be noted that the order of description of the following embodiments does not limit the priority order of the embodiments. Figure 2 , Figure 2This is a flowchart of an information acquisition method provided in an embodiment of the present application. Although a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in an order different from that shown in the flowchart. Specifically, the process of the information acquisition method specifically includes:

[0050] 101. Obtain an original image and metadata corresponding to the original image, where the metadata includes description information of the original image.

[0051] The original image can be of various types, such as a book cover, movie poster, or music album cover. Metadata refers to descriptive information associated with the original image, including the image's title, category, shooting time and location, and descriptions of the objects contained in the image. For example, for a book cover, metadata may include the title, author, publisher, publication date, and synopsis. For a movie poster, metadata may include the film's title, director, starring actors, release date, and plot summary.

[0052] There are many ways to obtain the original image and its corresponding metadata, such as obtaining it from an image database on the Internet, obtaining it from a locally stored image library, or intercepting the image and its related information from a specific website.

[0053] 102. Filter a target question template from at least one equivalent preset question template.

[0054] In the embodiments of the present application, the preset question templates refer to a series of pre-set question patterns that can be used to generate questions for different types of content. For example, for book-type content, the preset question templates may include "Who is the author of this book?", "What is the main content of this book?", "When was this book published?", etc.; for movie-type content, the preset question templates may include "Who is the director of this movie?", "Who are the main actors in this movie?", "When was this movie released?", etc.

[0055] The preset question template can be generated by the GPT-4o model. Equivalent preset question templates are question templates that are semantically similar but expressed differently. By setting multiple (for example, 10) equivalent question templates, the style, content, or expression of the target question and answer information generated subsequently can be more diverse, helping to improve the model's learning and comprehension capabilities.

[0056] The target question template is selected from a pool of pre-set question templates. This selection process can be based on various factors, such as the type of original image, metadata completeness, and the requirements of a specific application scenario. Based on these factors, the system selects the question template from the pre-set question template library that best suits the current original image and metadata as the target question template.

[0057] 103. Generate target question-answer information for the original image according to the target question template and the metadata.

[0058] It is understood that the target question-answer information refers to the question generated based on the target question template and the corresponding answer generated based on the metadata, that is, the question-answer pair. For example, if the target question template is "Who is the author of this book?" and the metadata contains the author information "Zhang San", the generated target question-answer information may be "Q: Who is the author of this book? A: Zhang San". It is understood that each original image can correspond to multiple sets of question-answer pairs, and the set of multiple question-answer pairs is the target question-answer information.

[0059] It can be understood that the process of generating target question and answer information may include: first, generating specific questions based on the target question template; then, extracting information related to the question from metadata as answers; and finally, combining the question and answer to form complete question and answer information.

[0060] In summary, the embodiment of the present application solves the problem of a single method of generating target question and answer information by configuring multiple equivalent question templates, and helps to generate question and answer information with diverse content from different depths and angles based on different question templates, thereby improving the effectiveness of the question and answer information.

[0061] Among them, since the preset question template is pre-set, in order to ensure that the template content matches the current question and answer information generation requirements, the question template can also be optimized based on the original image and metadata to be processed, thereby improving the accuracy and effectiveness of generating the target question and answer information. That is, optionally, in some embodiments of the present application, the step of "generating target question and answer information for the original image based on the target question template and the metadata" includes:

[0062] determining a content object type according to the original image or the metadata;

[0063] Optimizing the target question template according to the content object type to obtain an optimized question template;

[0064] The target question and answer information for the metadata is generated according to the optimized question template.

[0065] Specifically, the content object type refers to the category or theme of the original image, such as books, movies, music, games, landscapes, or people. Determining the content object type can be achieved by analyzing the visual features of the original image or by analyzing metadata such as keywords and tags. For example, image recognition technology can be used to identify the primary object in the image, or natural language processing technology can be used to analyze text descriptions in metadata to determine the content object type.

[0066] It's important to note that the optimization process for target question templates involves adjusting the question template's presentation, focus, depth, and content based on the content object type, making it more suitable for that specific type of content object. For example, for a book-type content object, the question template might be optimized to focus on aspects such as the author, publisher, and synopsis; for a movie-type content object, the question template might be optimized to focus on aspects such as the director, actors, and plot; and for an album-type content object, the question template might be optimized to distinguish between band names and album titles to avoid confusion between the two.

[0067] It should be noted that in some scenarios, the features between texts are relatively similar, which can easily lead to the extraction of incorrect answers from metadata, that is, the question and the answer do not match. For example, for content objects of the album type, some band names are consistent with the content or style of the album name, which causes questions about the album name to be incorrectly extracted as band names, and questions about the band name to be incorrectly extracted as album names. Therefore, in an embodiment of the present application, prompts for questions can also be generated based on the content object type to optimize the description of the question in the question template and improve the accuracy of the question and answer pair generation. That is, optionally, in some embodiments of the present application, the step of "optimizing the target question template according to the content object type to obtain an optimized question template" includes:

[0068] Determining prompt information of the target question in the target question template according to the content object type;

[0069] The target question is optimized according to the prompt information to obtain an optimized question template.

[0070] For example, for the album type, if the question in the target question template is "Can you identify the name of this album?", in order to avoid confusion between the album name and the band name, a prompt message for the question can be generated, such as "(Released by band XX)", then the optimized target question is "Can you identify the name of this album? (Released by band XX)", in this way, by prompting that "Band XX" is the band name rather than the album name, the incorrect extraction of the band name as the album name is avoided.

[0071] Similarly, this prompt information can also be applied to other types of questions. For example, when asking about the author, the prompt information of "first author" can be added to avoid extracting the second or third author instead of the first author.

[0072] Similarly, more general questions can be adjusted to be more targeted. For example, for the movie genre, the general question "Who created this?" can be optimized to "Who is the director of this movie?" For the book genre, the general question "What is this?" can be optimized to "What is the title of this book?" Or the question "What color is this?" can be optimized to "What color is the cover of this book?" and so on.

[0073] In summary, the embodiments of the present application improve the diversity and accuracy of the target question and answer information ultimately generated by pre-setting multiple question templates and optimizing the question templates based on the content object type.

[0074] Furthermore, after generating the target question and answer information, the original visual model can be trained based on the target question and answer information to obtain a target visual model capable of extracting information from the image. That is, optionally, in some embodiments of the present application, after the step of "generating the target question and answer information for the original image based on the target question template and the metadata", the method further includes:

[0075] The original visual model is trained according to the target question and answer information and the original image corresponding to the target question and answer information to obtain a target visual model; wherein the target visual model is used to extract image content in the image to be identified.

[0076] That is, the generated target question and answer information is used to train the visual model to improve the model's ability to understand image content. The original visual model can be a pre-trained image recognition model, which is fine-tuned using the target question and answer information and the corresponding original image to obtain a target visual model that can better understand the image content. The training process can adopt a supervised learning method, using the original image as input and the target question and answer information as labels. By optimizing the model parameters, the model can accurately extract relevant information from the image. After training, the target visual model can be used to extract image content from new images to be recognized, achieving understanding and analysis of the image. That is, the target visual model can be applied to image content understanding and question-answering systems.

[0077] Optionally, in an embodiment of the present application, after obtaining the original image and the corresponding metadata, a text check may be performed on the original image to determine whether the original image contains text information, and further determine whether the original image is suitable as a training sample for the original visual model. That is, optionally, in some embodiments of the present application, after the step of "obtaining the original image and the metadata corresponding to the original image", the method further includes:

[0078] Recognize the text record result of the original image by a text recognition model;

[0079] Generating target question-answer information for the original image according to the target question template and the metadata includes:

[0080] If the text record result indicates that text exists on the original image, target question-answer information for the original image is generated according to the target question template and the metadata.

[0081] The text recognition model is used to detect text content in raw images, thereby filtering out raw images that do not contain text information, thereby improving the sample quality for subsequent model training. The text record results include at least two types: containing text and not containing text.

[0082] In this embodiment of the present application, the text recognition model may include a GPT-4o model.

[0083] In the embodiment of the present application, the entire process of generating the target question and answer information can be automated. For example, a prompt word can be set to trigger the recognition of whether the original image contains text. For example, the prompt word includes: Please confirm whether the original image contains any text or words (excluding labels such as watermarks). A simple "yes" or "no" answer is provided, wherein the "yes" or "no" answer is the text record result.

[0084] In order to improve the target visual model's ability to process images of various content object types, original images of various content object types may be collected, and target question-answer information for various content object types may be generated. Specifically, optionally, in some embodiments of the present application, the step of "obtaining original images and metadata corresponding to the original images" includes:

[0085] determining at least one image source according to a content object type, the content object type comprising at least one of a book, an album, a game, a movie, or music;

[0086] At least one original image of a content object type is obtained from the image source, and metadata corresponding to each original image is obtained from each image source, wherein the metadata is a description of the original image added at the corresponding image source.

[0087] Image sources refer to sources that provide original images, such as e-commerce websites, social media platforms, digital libraries, and movie databases. Determining image sources based on content object type means selecting appropriate image sources for different content objects. For example, for books, book websites and digital libraries can be selected as image sources; for movies, movie databases and video platforms can be selected as image sources.

[0088] It should be noted that these metadata are usually added by managers or users of image sources to describe and label image content.

[0089] It is understandable that with the diversity of content presentation methods, these original images contain texts of different font types and different artistic levels. Therefore, when original images of more types of text are selected as sample data, the recognition performance of the trained target visual model in different fonts and different text recording methods can also be improved.

[0090] In an embodiment of the present application, target question-answer information may be generated by identifying entities and instances corresponding to the entities from metadata. That is, optionally, in some embodiments of the present application, the metadata includes at least one instance of an entity. The step of “generating target question-answer information for the original image based on the target question template and the metadata” includes:

[0091] extracting instances for each entity from the metadata;

[0092] The target question template is filled in according to the instances corresponding to each entity to obtain target question and answer information.

[0093] An entity refers to a specific object or concept described in metadata, such as a person, place, time, or event. An instance of an entity is a specific instance of a particular entity. For example, instances of a person entity might include specific names like "Liming" and "Zhang Hua." Extracting entity instances from metadata can be achieved using natural language processing techniques, such as named entity recognition and entity linking.

[0094] It can be understood that entity instances extracted from metadata are populated into the target question template to generate a specific answer to the question. For example, for a movie poster image, the metadata contains entity instances such as the movie title "Interstellar" and the director "John Nolan". These instances can be populated into the answer areas corresponding to the question templates "What is the name of this movie?" and "Who directed this movie?" to generate the corresponding answers.

[0095] In this embodiment of the present application, prompt words for extracting entity instances can also be constructed to extract each instance from the metadata. For example, for book types, prompt words include:

[0096] “Analyze the title of the work and structure it into its designated components.

[0097] Provides a JSON object with the following fields:

[0098] -**Main Title**: Identifies the main name of the work.

[0099] - **Subtitle**: A list of secondary phrases that provide additional context or details. If none, use an empty string.

[0100] -**Edition**: Specifies a specific edition or format, such as "Special Edition" or "Director's Cut". If none, use an empty string.

[0101] - **Publisher**: The entity responsible for producing or distributing the work. If none, use the empty string.

[0102] - **Additional Information**: Any additional details such as author, creator, series, or notable features. If none, use an empty string.

[0103] #Output format:

[0104] Must return a JSON object with the defined fields. ".

[0105] For example, the title in the website meta information contains not only information such as the book title, but also information such as the subtitle and publisher. Therefore, GPT-4o is used as a parser to structure the metadata and separate the main title, subtitle, etc. from the metadata containing auxiliary information.

[0106] In summary, the embodiment of the present application solves the problem of a single method of generating target question and answer information by configuring multiple equivalent question templates, and helps to generate question and answer information with diverse content from different depths and angles based on different question templates, thereby improving the effectiveness of the question and answer information.

[0107] Among them, by optimizing the question template based on the content object type corresponding to the original image or metadata, the accuracy and effectiveness of the target question and answer information output based on the optimized question template are improved.

[0108] Among them, by determining multiple image sources based on the content object type and obtaining original images of different content object types respectively, the target question and answer information is enriched, so that the target visual model trained based on the target question and answer information can better process images of different content object types, thereby improving the model's recognition performance in various types of images.

[0109] To facilitate better implementation of the information acquisition method of the present application, the present application also provides an information acquisition device based on the above information acquisition method. The meanings of the terms are the same as those in the above information acquisition method, and the specific implementation details can be referred to the description in the method embodiment.

[0110] See also Figure 3 , Figure 3 : is a schematic diagram of the structure of the information acquisition device provided in an embodiment of the present application. The information acquisition device may be specifically as follows:

[0111] An acquisition module 201 is configured to acquire an original image and metadata corresponding to the original image, wherein the metadata includes description information of the original image;

[0112] A screening module 202 is configured to screen a target question template from at least one equivalent preset question template;

[0113] The generating module 203 is configured to generate target question-answer information for the original image according to the target question template and the metadata.

[0114] Optionally, in some embodiments of the present application, generating target question-answer information for the original image according to the target question template and the metadata includes:

[0115] determining a content object type according to the original image or the metadata;

[0116] Optimizing the target question template according to the content object type to obtain an optimized question template;

[0117] The target question and answer information for the metadata is generated according to the optimized question template.

[0118] Optionally, in some embodiments of the present application, optimizing the target question template according to the content object type to obtain an optimized question template includes:

[0119] Determining prompt information of the target question in the target question template according to the content object type;

[0120] The target question is optimized according to the prompt information to obtain an optimized question template.

[0121] Optionally, in some embodiments of the present application, after generating target question-answer information for the original image according to the target question template and the metadata, the method further includes:

[0122] The original visual model is trained according to the target question and answer information and the original image corresponding to the target question and answer information to obtain a target visual model; wherein the target visual model is used to extract image content in the image to be identified.

[0123] Optionally, in some embodiments of the present application, after obtaining the original image and metadata corresponding to the original image, the method further includes:

[0124] Recognize the text record result of the original image by a text recognition model;

[0125] Generating target question-answer information for the original image according to the target question template and the metadata includes:

[0126] If the text record result indicates that text exists on the original image, target question-answer information for the original image is generated according to the target question template and the metadata.

[0127] Optionally, in some embodiments of the present application, obtaining the original image and metadata corresponding to the original image includes:

[0128] determining at least one image source according to a content object type, the content object type comprising at least one of a book, an album, a game, a movie, or music;

[0129] At least one original image of a content object type is obtained from the image source, and metadata corresponding to each original image is obtained from each image source, wherein the metadata is a description of the original image added at the corresponding image source.

[0130] Optionally, in some embodiments of the present application, the metadata includes an instance of at least one entity, and generating target question-answer information for the original image based on the target question template and the metadata includes:

[0131] extracting instances for each entity from the metadata;

[0132] The target question template is filled in according to the instances corresponding to each entity to obtain target question and answer information.

[0133] In the embodiment of the present application, the acquisition module 201 acquires the original image and the metadata corresponding to the original image, where the metadata includes description information of the original image. The screening module 202 screens the target question template from at least one equivalent preset question template, and the generation module 203 generates target question and answer information for the original image based on the target question template and the metadata.

[0134] Among them, the embodiment of the present application flexibly selects and generates target question and answer information by configuring multiple equivalent question templates, thereby solving the problem of a single method of generating target question and answer information, and helping to generate question and answer information with diverse content from different depths and angles based on different question templates, thereby improving the effectiveness of the question and answer information.

[0135] In addition, the present application also provides an electronic device, such as Figure 4 As shown, it shows a schematic diagram of the structure of the electronic device involved in this application, specifically:

[0136] The electronic device may include one or more processors 301 of processing cores, one or more computer-readable storage media memories 302, a power supply 303, an input unit 304 and other components. Those skilled in the art will appreciate that Figure 4 The electronic device structure shown in the figure does not constitute a limitation of the electronic device, and may include more or fewer components than shown in the figure, or combine certain components, or arrange components differently.

[0137] The processor 301 is the control center of the electronic device. It connects all parts of the electronic device using various interfaces and lines. By running or executing software programs and / or modules stored in the memory 302 and accessing data stored in the memory 302, it performs various functions of the electronic device and processes data, thereby monitoring the electronic device as a whole. Optionally, the processor 301 may include one or more processing cores; preferably, the processor 301 may integrate an application processor and a modem processor, wherein the application processor primarily processes the operating system, user interface, and application programs, while the modem processor primarily handles wireless communications. It is understood that the modem processor may not be integrated into the processor 301.

[0138] The memory 302 can be used to store software programs and modules. The processor 301 executes various functional applications and data processing by running the software programs and modules stored in the memory 302. The memory 302 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area may store data created according to the use of the electronic device, etc. In addition, the memory 302 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage device. Accordingly, the memory 302 may also include a memory controller to provide the processor 301 with access to the memory 302.

[0139] The electronic device also includes a power supply 303 for supplying power to various components. Preferably, the power supply 303 can be logically connected to the processor 301 via a power management system, thereby enabling the power management system to manage charging, discharging, and power consumption. The power supply 303 can also include one or more DC or AC power supplies, a recharging system, a power supply device debugging circuit, a power converter or inverter, a power status indicator, and other arbitrary components.

[0140] The electronic device may further include an input unit 304, which may be configured to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.

[0141] Although not shown, the electronic device may further include a display unit, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 301 in the electronic device will load the executable files corresponding to the processes of one or more application programs into the memory 302 according to the following instructions, and the processor 301 will run the application programs stored in the memory 302, thereby implementing the steps of any of the information acquisition methods provided in the embodiments of the present application.

[0142] An embodiment of the present application obtains an original image and metadata corresponding to the original image, where the metadata includes descriptive information about the original image, filters a target question template from at least one equivalent preset question template, and generates target question-and-answer information for the original image based on the target question template and metadata.

[0143] Among them, the embodiment of the present application flexibly selects and generates target question and answer information by configuring multiple equivalent question templates, thereby solving the problem of a single method of generating target question and answer information, and helping to generate question and answer information with diverse content from different depths and angles based on different question templates, thereby improving the effectiveness of the question and answer information.

[0144] The specific implementation of the above operations can be found in the previous embodiments and will not be repeated here.

[0145] Those skilled in the art will appreciate that all or part of the steps in the various methods of the above embodiments may be accomplished by instructions, or by controlling related hardware through instructions. The instructions may be stored in a computer-readable storage medium and loaded and executed by a processor.

[0146] To this end, the present application provides a computer-readable storage medium, on which a computer program is stored. The computer program can be loaded by a processor to execute the steps in any information acquisition method provided in the present application.

[0147] The specific implementation of the above operations can be found in the previous embodiments and will not be repeated here.

[0148] The computer-readable storage medium may include a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0149] Since the instructions stored in the computer-readable storage medium can execute the steps in any information acquisition method provided in the present application, the beneficial effects that can be achieved by any information acquisition method provided in the present application can be achieved. Please refer to the previous embodiments for details and will not be repeated here.

[0150] The above is a detailed introduction to the information acquisition method, device, electronic device and computer-readable storage medium provided by the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core ideas. At the same time, for those skilled in the art, according to the ideas of the present invention, there will be changes in the specific implementation methods and application scopes. In summary, the content of this specification should not be understood as limiting the present invention.

Claims

1. An information acquisition method, characterized in that: The method comprises: Acquire an original image and metadata corresponding to the original image, wherein the metadata includes description information of the original image; Filtering a target question template from at least one equivalent preset question template; Target question-answering information for the original image is generated according to the target question template and the metadata.

2. The information acquisition method according to claim 1, characterized in that Generating target question-answer information for the original image according to the target question template and the metadata includes: determining a content object type according to the original image or the metadata; Optimizing the target question template according to the content object type to obtain an optimized question template; The target question and answer information for the metadata is generated according to the optimized question template.

3. The information acquisition method according to claim 2, characterized in that: The step of optimizing the target question template according to the content object type to obtain an optimized question template includes: Determining prompt information of the target question in the target question template according to the content object type; The target question is optimized according to the prompt information to obtain an optimized question template.

4. The information acquisition method according to claim 1, wherein: After generating target question-answer information for the original image according to the target question template and the metadata, the method further includes: The original visual model is trained according to the target question and answer information and the original image corresponding to the target question and answer information to obtain a target visual model; wherein the target visual model is used to extract image content in the image to be identified.

5. The information acquisition method according to claim 1, characterized in that: After obtaining the original image and metadata corresponding to the original image, the method further includes: Recognize the text record result of the original image by a text recognition model; Generating target question-answer information for the original image according to the target question template and the metadata includes: If the text record result indicates that text exists on the original image, target question-answer information for the original image is generated according to the target question template and the metadata.

6. The information acquisition method according to claim 1, characterized in that: The obtaining of the original image and metadata corresponding to the original image includes: determining at least one image source according to a content object type, the content object type comprising at least one of a book, an album, a game, a movie, or music; At least one original image of a content object type is obtained from the image source, and metadata corresponding to each original image is obtained from each image source, wherein the metadata is a description of the original image added at the corresponding image source.

7. The information acquisition method according to claim 1, characterized in that: The metadata includes an instance of at least one entity, and generating target question-answer information for the original image according to the target question template and the metadata includes: extracting instances for each entity from the metadata; The target question template is filled in according to the instances corresponding to each entity to obtain target question and answer information.

8. An information acquisition device, characterized in that: The device comprises: An acquisition module, configured to acquire an original image and metadata corresponding to the original image, wherein the metadata includes description information of the original image; A screening module, configured to screen a target question template from at least one equivalent preset question template; A generation module is used to generate target question-answer information for the original image based on the target question template and the metadata.

9. An electronic device, characterized in that: The method comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the information acquisition method according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the information acquisition method according to any one of claims 1 to 7 are implemented.