Image information processing method in knowledge question-answering system based on multiple modes and electronic equipment

By extracting and filtering image information from multimodal files, and combining semantic analysis and text filtering, the problems of image information noise and redundancy in multimodal question answering systems are solved, thereby improving the quality of the information database and the accuracy and efficiency of the response system.

CN120407757APending Publication Date: 2025-08-01BEIJING SANKUAI ONLINE TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510503880.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

Existing multimodal knowledge question answering systems rely too heavily on text information, resulting in limited room for improvement in user satisfaction. Furthermore, noise and redundancy exist in image information processing, reducing the accuracy of responses.

Method used

By extracting target images from files to be processed, filtering images containing text information, performing semantic analysis using a visual-language big data model, determining the relevance of images to contextual information, retaining explanatory or supplementary information, performing paragraph and sentence-level text filtering, and matching preset questions to generate answers.

Benefits of technology

It improves the quality of the information base and the accuracy of responses in multimodal question-answering systems, reduces noise and redundancy, and enhances the richness of the information base and the efficiency of the response system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120407757A_ABST
    Figure CN120407757A_ABST
Patent Text Reader

Abstract

The invention provides an image information processing method in a multi-mode-based knowledge question-answering system and electronic equipment. The method comprises the steps that a target image is acquired from a to-be-processed file, the to-be-processed file comprises text data and an image, and the target image comprises text information; extracting text information in the target image as first text information, and extracting context information of the target image in the to-be-processed file according to the text data; selecting a reserved target image according to the first text information and the context information; according to the context information, determining reserved second text information in reserved first text information of the target image, the second text information comprising paragraphs; according to the matching result of the second text information and the multiple preset questions, third text information corresponding to the multiple preset questions is selected from the second text information, and the third text information comprises sentences. According to the embodiment of the invention, the processing efficiency of the multi-modal information and the effective information proportion of the information base can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0002] A multimodal knowledge Q&A system integrates various different types of information modalities (such as text, image, voice, video, etc.), and can comprehensively process and analyze this multi-source heterogeneous information to achieve an intelligent system that can accurately answer users' questions. In a multimodal knowledge Q&A system, users can ask questions in ways such as text, voice, and image. The system analyzes the question information and provides reply information based on the information database of the system.

[0003] In order to build the information database, it is usually necessary to collect various types of information related to this knowledge Q&A system. In related technologies, relevant information is usually sorted out and collected based on text. However, in many cases, the text information related to this knowledge Q&A system is not sufficient to provide sufficient information. Due to the limited available relevant information, there is still room for improvement in user satisfaction when the multimodal knowledge Q&A system generates reply information only relying on the text information of the relevant information.

[0004] It should be noted that the information disclosed in the above background art section is only used to strengthen the understanding of the background of the present disclosure, and thus may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention

[0005] The purpose of the present disclosure is to provide an image information processing method and an electronic device in a multimodal knowledge Q&A system, which are used to provide accurate and rich reference information for the database of the Q&A system.

[0006] According to the first aspect of the embodiments of the present disclosure, there is provided an image information processing method in a multimodal knowledge Q&A system, including: obtaining a target image from a file to be processed, where the file to be processed includes text data and images, and the target image includes text information; extracting the text information in the target image as first text information, and extracting the context information of the target image in the file to be processed according to the text data; selecting the target images to be retained according to the first text information and the context information; determining the second text information to be retained in the first text information of the retained target images according to the context information, where the second text information includes paragraphs; and selecting the third text information corresponding to the multiple preset questions in the second text information according to the matching results between the second text information and the multiple preset questions, where the third text information includes sentences.

[0007] According to the second aspect of the present disclosure, there is provided an electronic device, including: a memory; and a processor coupled to the memory, where the processor is configured to execute the method as described in any one of the above according to instructions stored in the memory.

[0008] According to a third aspect of the present disclosure, there is provided a computer-readable storage medium having a program stored thereon, which when executed by a processor, implements the image information processing method in the multimodal-based knowledge Q&A system described in any one of the above.

[0009] According to a fourth aspect of the present disclosure, there is provided a computer program product including a computer program, characterized in that the steps of the method described in any one of the above are implemented when the computer program is executed by a processor.

[0010] In the embodiments of the present disclosure, in a multimodal-based knowledge Q&A system, image extraction is performed on a file to be processed including text data and images to obtain a target image with text information, and the target image is screened based on the context information of the target image in the file to be processed, and further, second text information related to the context information is extracted from the target images that pass the screening, and third text information related to a preset question is retained in the second text information. This can fully utilize multimodal information including images and text to form an information library of the multimodal-based knowledge Q&A system during the construction of the Q&A system, and can effectively avoid irrelevant or redundant information from reducing the quality of answer generation, providing rich and accurate reference information for the multimodal-based knowledge Q&A system, and improving the quality and efficiency of the construction of the Q&A system.

[0011] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure. Obviously, the accompanying drawings in the following description are only some embodiments of the present disclosure, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.

[0013] Figure 1 is a flowchart of the image information processing method in the multimodal-based knowledge Q&A system in an exemplary embodiment of the present disclosure.

[0014] Figure 2 is a sub-flowchart of step S3 in an exemplary embodiment of the present disclosure.

[0015] Figure 3 is a sub-flowchart of step S4 in an exemplary embodiment of the present disclosure.

[0016] Figure 4 is a sub-flowchart of step S5 in an exemplary embodiment of the present disclosure.

[0017] Figure 5It is a schematic diagram of the effect of image processing in an exemplary embodiment of the present disclosure.

[0018] Figure 6 It is a schematic diagram of the working process of a question-answering system in an exemplary embodiment of the present disclosure.

[0019] Figure 7 It is a block diagram of an image information processing device in a multi-modal based knowledge question-answering system in an exemplary embodiment of the present disclosure.

[0020] Figure 8 It is a block diagram of an electronic device in an exemplary embodiment of the present disclosure. Detailed implementation manners

[0021] Now, example embodiments will be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be more thorough and complete, and will fully convey the concept of the example embodiments to those skilled in the art. The features, structures, or characteristics described may be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present disclosure. However, those skilled in the art will realize that the technical solutions of the present disclosure can be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. may be used. In other cases, well-known technical solutions are not shown or described in detail to avoid obscuring the various aspects of the present disclosure.

[0022] In addition, the accompanying drawings are only schematic illustrations of the present disclosure, and the same reference numerals in the drawings denote the same or similar parts, so repeated descriptions thereof will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0023] The following will describe the exemplary embodiments of the present disclosure in detail with reference to the accompanying drawings.

[0024] Figure 1 It is a flowchart of an image information processing method in a multi-modal based knowledge question-answering system in an exemplary embodiment of the present disclosure.

[0025] Refer to Figure 1 and the image information processing method 100 in the multi-modal based knowledge question-answering system may include:

[0026] Step S1, obtain a target image from a file to be processed, where the file to be processed includes text data and images, and the target image includes text information;

[0027] Step S2, extract the text information in the target image as the first text information, and extract the context information of the target image in the file to be processed according to the text data;

[0028] Step S3, select the target images to be retained according to the first text information and the context information;

[0029] Step S4, determine the second text information to be retained in the first text information of the target images that are retained according to the context information, where the second text information includes paragraphs;

[0030] Step S5, select the third text information corresponding to the multiple preset questions in the second text information according to the matching results between the second text information and the multiple preset questions, where the third text information includes sentences.

[0031] In the embodiments of the present disclosure, by performing image extraction on a file to be processed that contains text data and images in a multimodal knowledge question-and-answer system, a target image with text information is obtained, and the target image is screened based on the context information of the target image in the file to be processed. Further, second text information related to the context information is extracted from the target images that pass the screening, and third text information related to the preset questions is retained in the second text information. It is possible to fully utilize multimodal information including images and text to form an information library of the multimodal knowledge question-and-answer system during the construction of the question-and-answer system, and effectively avoid irrelevant or redundant information from reducing the quality of answer generation, providing rich and accurate reference information for the multimodal knowledge question-and-answer system, and improving the quality and efficiency of the construction of the question-and-answer system.

[0032] Next, each step of the image information processing method 100 in the multimodal knowledge question-and-answer system will be described in detail.

[0033] In step S1, a target image is obtained from a file to be processed, where the file to be processed includes text data and images, and the target image includes text information.

[0034] The method of the embodiments of the present disclosure can be used for constructing an information library of a question-and-answer system. The question-and-answer system is, for example, a system constructed by an enterprise or an industry and specifically used to answer certain types of questions, and can provide more accurate, more targeted, and more professional answers based on a dedicated information library. Even some question-and-answer systems with limited public access can provide internal information answers for users.

[0035] When building a question-answering system, you can first determine the scope of the information involved, such as the type of information, the source of information, etc., and then collect the information within this scope through various methods and store it as reference information for subsequent questions and answers. Generally speaking, the information available for building reference information is mostly text information, but in some cases, some files with images, or even directly image files, may also involve the knowledge required to build an information base. These files can be called multimodal information. Therefore, the construction of a question-answering system involves integrating information from different modalities (such as text, images) to improve the performance of the question-answering system and the retrieval system.

[0036] Due to the diverse sources and types of images, multimodal information often includes webpage screenshots, mobile phone screenshots, and other information. The conventional approach is to perform optical character recognition (OCR) on images to extract textual information, which is then stored as part of the information database. However, images often contain a significant amount of noise, and relying solely on OCR to extract text can be extremely lengthy. This not only fails to facilitate understanding of the original knowledge fragment, but can also dilute useful information, negatively impacting the recall process and easily contaminating the response with invalid information.

[0037] Therefore, the application of multimodal files sometimes not only fails to improve the accuracy of responses, but may even reduce the accuracy of responses.

[0038] Based on the discovered problem, the inventors of the embodiments of the present disclosure set up an image information processing method specifically for multimodal files to overcome the problem of decreased response accuracy caused by the introduction of multimodal files.

[0039] In the embodiment of the present disclosure, the files to be processed may include files from multiple sources used to construct an information library of the question-answering system, and the contents thereof may include both text information and images.

[0040] In the exemplary embodiment, only one file to be processed is used as an example for description. In actual application, the embodiment of the present disclosure is suitable for batch processing of massive files to be processed, so as to build and enrich the information library of the question-answering system through the massive files to be processed.

[0041] In step S1 , an image may be first extracted from the file to be processed, and then it is determined whether the image can provide information that is helpful for the answer, and the image that may be helpful for the answer is determined as the target image.

[0042] In an exemplary embodiment, the category of each image in the image file to be processed may be determined, and then the image corresponding to the preset category may be determined as the target image.

[0043] When determining the category of an image, the image category can be determined according to the characteristics corresponding to different categories of images. Image categories include, for example, QR codes, photos, screenshots, flowcharts, uncertain types (with text information), uncertain types (without text information), etc. Among them, a QR code is a small square black-and-white pattern containing encoded information that can be read by a scanning device to obtain data such as links and text; a photo is taken of an actual scene, person, or object, and is different from other synthetic or generated types, having a sense of reality; a flowchart is a visual representation that includes a sequence of steps, decision points, and the relationships between processes, and is usually used to show process flow and logical order; a screenshot is, for example, a screenshot of a web page or application. There can be multiple categories of images, and those skilled in the art can set the categories and define the characteristics of each category according to the actual information scope to determine the category of the image.

[0044] The categories of images can be initially divided into categories with text information and categories without text information. For example, categories with text information may include flowcharts, screenshots, uncertain types (with text information), etc.; categories without text information may include QR codes, photos, uncertain types (without text information), etc.

[0045] During the initial screening, images corresponding to categories without text information can be excluded first, and images corresponding to categories with text information can be determined as target images. In some embodiments, the images can also be directly classified into categories with text information and categories without text information for classification and screening. By initially screening the images in the file to be processed, noise information and redundant information can be reduced, and the efficiency of subsequent information processing and the accuracy of information extraction can be improved.

[0046] The process of extracting, classifying, and screening the above images can be carried out either through a pre-trained classification model or through a vision-language large model.

[0047] Next, information extraction is performed on the target images.

[0048] In step S2, the text information in the target image is extracted as the first text information, and the context information of the target image in the file to be processed is extracted according to the text data.

[0049] In images with text information, the text information can be extracted first, and the set of text information in the target image is called the first text information. The text information in the target image can be extracted using OCR (Optical Character Recognition) technology or other recognition technologies. The text information refers to text information including but not limited to characters, numbers, letters, etc.

[0050] Next, context information corresponding to the target image is extracted according to the text data of the document to be processed. The context information corresponding to the target image may include, for example, the text information corresponding to the paragraphs within a preset distance range before and after the target image (such as the 5 paragraphs before the target image and the 5 paragraphs after the target image), or may include all the text information corresponding to the document to be processed. If there are multiple target images in the document to be processed, context information can be determined for each target image, and in subsequent analysis, each target image is processed based on the context information corresponding to it.

[0051] In an exemplary embodiment, both the extraction of text information from the target image and the extraction of context information can be performed by a vision-language large model to improve the intelligence and accuracy of this process.

[0052] Since not all images are helpful for text information and can provide effective information, the embodiments of the present disclosure set to further screen the target images to select the target images that can provide effective information.

[0053] In step S3, the target images to be retained are selected according to the first text information and the context information.

[0054] Figure 2 It is a sub-flowchart of step S3 in the exemplary embodiment of the present disclosure.

[0055] Reference Figure 2 , in an exemplary embodiment, step S3 may include:

[0056] Step S31, determining the degree of relevance between the first text information and the context information;

[0057] Step S32, when the degree of relevance is greater than a first preset value, determining the types of relevance between the first text information and the context information, and the types of relevance at least include explanatory information and supplementary information;

[0058] Step S33, when the types of relevance are explanatory information and / or supplementary information, retaining the target image.

[0059] The degree of relevance between the first text information and the context information can be determined by means such as semantic analysis, keyword matching, topic model analysis, and context understanding. Semantic analysis refers to deeply understanding and analyzing the semantics of the text to judge the degree of semantic association between the first text and the context. For example, analyzing the subject-predicate-object structure of a sentence and the semantic roles of words to determine whether it conforms to the semantic logic of the context. Keyword matching refers to checking whether the keywords in the first text appear in the context, as well as the frequency and position of their appearance. If the keywords appear frequently and are closely related in position in the context, then their degree of relevance may be high. Topic model analysis refers to using topic model algorithms such as Latent Dirichlet Allocation (LDA) to analyze the topics involved in the first text and the context. If they belong to the same or similar topic categories, it indicates that the two have a high degree of relevance. Context understanding refers to considering the context information of the context, including the background, purpose, and audience of the text. By comprehensively understanding these factors, the relevance of the first text in this context is judged. For example, in an article about the development of technology, the specific technology products or technological innovations mentioned are relevant to the context of the overall article, while content unrelated to technology has a lower degree of relevance.

[0060] The above methods for analyzing the degree of relevance between the first text information and the context information can be implemented through various natural language processing methods. For example, it can be implemented through a pre-trained neural network model. In an exemplary embodiment, it can be implemented through a vision-language large model to improve accuracy and efficiency.

[0061] If the degree of relevance between the first text information and the context information of a target image is low (not greater than the first preset value), it indicates that this information is noise information or redundant information, and the target image can be removed, and the next target image or the next file to be processed can be continued (if there is no next target image). If the degree of relevance between the first text information and the context information of a target image is high (greater than the first preset value), it indicates that the target image can provide information related to the context information and is very likely to be information that supplements, explains, or interprets the context information. At this time, the target image is temporarily retained for further processing. Different methods for determining the degree of relevance correspond to different first preset values. When using a large model to determine the degree of relevance, the first preset value can also be automatically set by the large model.

[0062] If a target image is temporarily retained, it is further determined whether the corresponding first text information is explanatory or supplementary information for the context information. If so, the target image is retained; otherwise, the target image is removed, and the next target image or the next file to be processed is continued (if there is no next target image).

[0063] Among them, determining the first text information as the explanatory information of the context information can be achieved, for example, through methods such as semantic similarity analysis, keyword coverage analysis, text structure analysis, anaphora resolution verification, logical coherence verification, and positional relationship analysis.

[0064] Semantic similarity analysis is, for example, extracting paragraph vectors through pre-trained language models (such as BERT, RoBERTa), calculating cosine similarity or Jaccard similarity. When the similarity needs to be higher than a threshold (such as cosine similarity ≥ 0.7), determine whether the semantic range of the first text information covers the semantic range of the context information through vector space projection. If so, determine that the semantics of the two are similar.

[0065] Keyword coverage analysis refers to extracting the keywords of the context information (such as through TF-IDF, TextRank, etc.), verifying whether the first text information contains the keywords of the context information and introduces new related words. If the coverage rate of the related words in the context information in the first text information is greater than a certain value (such as 80%), and new terms related to the theme of the context information are added (such as the proportion of new terms ≥ 30%), then it is judged that the first text information may be the explanatory information of the context information.

[0066] Text structure analysis refers to detecting whether the first text information contains explanatory language patterns (such as marker words like "because", "for example", "that is", etc.), and analyzing whether its sentence structure presents explanatory features (for example, the context information is mostly concise conclusions, while the first text information is mostly causal chains, examples, or detailed descriptions). If the probability of the appearance of explanatory marker words in the first text information ≥ 50%, and its average sentence length is significantly higher than that of the context information (for example, the average sentence length of the context information ≤ 15 words, and the first text information ≥ 20 words), then it further supports the judgment as explanatory information.

[0067] Anaphora resolution verification refers to identifying pronouns or vague references in the context information (such as "this", "the method", etc.), and checking whether the first text information replaces them with specific nouns or clear expressions. If all the referential components in the context information are concretized in the first text information, the credibility of the explanatory relationship can be enhanced.

[0068] Logical coherence verification refers to analyzing whether the first text information forms logical support for the context information. For example, verifying whether the first text information contains an inference chain (such as cause and effect, definition, example, etc.) that directly supports the context information by constructing a text dependency graph (such as using Stanford CoreNLP). If there is at least one clear logical support relationship, it conforms to the characteristics of explanatory information.

[0069] The analysis of positional relationship refers to examining the physical positional relationship between the first text information and the context information (such as whether they are adjacent, whether there are typesetting marks such as indentation or numbering). If the first text information immediately follows the context information without an intervening paragraph, or there are obvious explanatory typesetting features (such as indentation, dash guidance), it can be used as an auxiliary basis for judgment.

[0070] Through one or more of the above methods, if multiple indicators all meet the preset thresholds (such as semantic similarity ≥ 0.7, keyword coverage rate ≥ 80%, clear logical support, etc.), it can be determined that the first text information is the explanatory information of the context information. In practical applications, weighted scoring (for example, after setting weights for each indicator, the scores of each indicator are weighted to obtain the weighted sum) or machine learning models (such as XGBoost) can be used for automated decision-making, or directly use large models to achieve the above judgment to improve the accuracy and adaptability of the judgment.

[0071] To determine that the first text information is supplementary information of the context information, for example, it can be achieved through methods such as semantic expansion analysis, information increment detection, topic consistency verification, context cohesion analysis, redundancy evaluation, etc.

[0072] Among them, semantic expansion analysis refers to verifying whether the first text information expands new semantics based on the context information. For example, use pre-trained language models (such as BERT, GPT) to extract the semantic vectors of two texts, and analyze whether the first text information introduces new topics or refines certain aspects of the context information. Then calculate the information entropy difference. If the entropy value of the first text information is significantly higher than that of the context information (such as an increase of more than 20%), it may contain supplementary content. If the semantic similarity between the first text information and the context information is between 0.5 and 0.8 (neither completely repeated nor irrelevant), and the first text information contains at least 30% new entities or concepts (such as new terms, examples, data), then the first text information may be supplementary information of the context information.

[0073] Information increment detection refers to identifying whether the first text information provides new content not covered by the context information. For example, use named entity recognition (NER) and keyword extraction to compare two texts. If the first text information contains specific details not mentioned in the context (such as time, location, data support), it is regarded as supplementary. If the proportion of new entities ≥ 40% (such as the context mentions "climate change", and the first text supplements "the average annual melting rate of the Arctic ice sheet reaches 12%") and the new content is strongly related to the context theme (verified by the LDA topic model, the topic relevance ≥ 70%), then the first text information may be supplementary information of the context information.

[0074] Topic consistency verification ensures that the primary text and the contextual information share core themes but provide additional perspectives. For example, using topic modeling (such as LDA or BERTopic) to analyze the topic distribution of two texts, if the main topic overlaps by 70% or more, but the primary text introduces different subtopics (e.g., the context discusses "technical advantages," while the primary text adds "potential ethical risks"), then the relationship is complementary.

[0075] Contextual cohesion analysis examines whether the first textual information naturally extends the contextual information through cohesive words or logical relationships. For example, it identifies cohesive words (e.g., "In addition," "On the other hand," "It is worth noting") and analyzes whether the sentence structure exhibits progressive or parallel structures (e.g., "Not only that, ...", "Specifically, ..."). If the first textual information contains multiple cohesive words that logically extend the context (e.g., provide additional contrast or detailed explanation), then its supplementary nature is supported.

[0076] Redundancy assessment involves eliminating instances where the first text information and the contextual information are highly repetitive. For example, if the text repetition rate is calculated (using Rouge-L or edit distance), and if the percentage of repeated content is less than 20%, then check whether the new content is implemented through repetition or example explanation (e.g., expanding "efficient algorithm" to "gradient descent-based optimization algorithm"). If the core information repetition rate of the two texts is ≤30% and the percentage of new content is ≥40%, then the supplement is considered valid.

[0077] Each of the above indicators can be weighted, for example, by assigning weights to each indicator (e.g., 25% for semantic extension, 30% for information increment, and 20% for redundancy). A total score of 0.7 or higher indicates that the first text information supplements the context. Furthermore, if the first text information meets the three core criteria of "topic consistency, information increment, and low redundancy" for the context information, the first text information is directly considered supplementary to the context information.

[0078] When the first text information is explanation information or supplementary information of the context information, the target image is retained; otherwise, the target image is discarded and the processing continues with the next target image or the next to-be-processed file (if there is no next target image).

[0079] The types of correlation between the first text information and the context information may be, in addition to explanatory information or supplementary information, an explanatory relationship, a corresponding relationship, an emphasis relationship, and the like.

[0080] The explanatory relationship means that the context information can be information that explains the target image. For example, in an image showing a city gate and a plaque on the city gate, the first text information extracted is the plaque, and the context information is the introduction to the city gate. At this time, the first text information may not be able to explain or supplement the context information in terms of text, and it does not belong to the supplementary information or explanatory information of the context information.

[0081] The corresponding relationship means that there is a one-to-one correspondence between the context information, the elements in the target image, and the first text information. For example, on a map, the first text information is the names of various cities on the map, and the context information is the introduction to the map. At this time, the first text information may not be able to explain or supplement the context information in terms of text, and it does not belong to the supplementary information or explanatory information of the context information.

[0082] The emphasizing relationship means that the first text information plays a role in emphasizing the key information or important parts in the image to attract the reader's attention. For example, in a product promotion picture, the core features or advantages of the product are highlighted with prominent text. Such first text information plays a role in emphasizing the context information, but it may not be able to explain or supplement the context information in terms of text, and it does not belong to the supplementary information or explanatory information of the context information.

[0083] The above judgments of the degree of relevance and the relationship of relevance can all be achieved through machine learning models, such as fine-tuning pre-trained models (such as BERT) or training classifiers (such as SVM), and automatically classified through feature vectors (similarity, number of new entities, density of cohesive words, etc.). Or, directly use large models to achieve the above judgments to improve the accuracy and adaptability of the judgments.

[0084] Based on the text information corresponding to the target image, that is, the first text information, to determine the association relationship between the target image and the context information, the target images that are strongly related to the context information can be screened out, improving the efficiency and accuracy of the subsequent processing process.

[0085] However, due to the accuracy problems of character recognition technologies such as OCR, the text information directly extracted from the target image is prone to typos, missing words, etc. In addition, step S3 makes a judgment based on the overall first text information of the target image. Since images usually carry a large amount of information, the first text information usually contains information that is irrelevant and redundant to the context information. Even if the first text information includes explanatory information or supplementary information for the context information, it may also contain some information that is useless for the context information. For example, the first text information may also include decorative information (such as the text information on the billboard in a city photo) and other information that is irrelevant to the context information. Such information may not be screened out during the overall relevance judgment in step S31 (as long as most of the content in the first text information is relevant to the context information).

[0086] Therefore, after determining to retain the target image, the embodiments of the present disclosure set to further screen and streamline the first text information of the target image.

[0087] In step S4, according to the context information, in the first text information of the retained target image, determine the retained second text information, where the second text information includes paragraphs.

[0088] Among them, a paragraph may be a natural paragraph. In some other embodiments of the present disclosure, the second text information may be either a paragraph or a sentence or a word.

[0089] Figure 3 It is a sub - flowchart of step S4 in the exemplary embodiments of the present disclosure.

[0090] Refer to Figure 3 , in the exemplary embodiment, step S4 may include:

[0091] Step S41, segment the first text information to form a plurality of text units;

[0092] Step S42, determine the relevant types of each text unit with respect to the context information;

[0093] Step S43, save the text units whose relevant types are the explanatory information and / or the supplementary information as the second text information.

[0094] In the exemplary embodiment, sentences or paragraphs can be used as text units to finely screen the first text information, so as to remove text units with less relevance or less supplementation to the context information in the first text information, and save text units that explain or supplement the context information, which can greatly improve the effectiveness of the saved information. In some embodiments, the removed text units may also include redundant information with a coincidence degree greater than a threshold (for example, greater than 90%) with the context information. For a piece of first text information, the division of text units can be the same or different. For example, each paragraph can be set as a text unit, that is, the text units are all paragraphs; or, each sentence can be set as a text unit, that is, the text units are all sentences; or, according to the semantic coherence between sentences, one or more sentences with a semantic coherence greater than a threshold can be set as a text unit. At this time, the text unit may be one sentence, multiple sentences, or one paragraph, multiple paragraphs, or may include multiple sentences across paragraphs. The segmentation method of text units can refer to natural language processing solutions, and the present disclosure does not make special restrictions on this.

[0095] In addition, before determining the relevant types in step S42, the degree of relevance between the text unit and the context information can also be determined first, and the relevant types are determined only when the degree of relevance is greater than the threshold, so as to eliminate decorative information irrelevant to the context information such as the billboard copy in the city photo in the above example.

[0096] The method for determining the relevant types between the text unit and the context information can be the same as the method mentioned above, and can also be implemented through a large model, which will not be elaborated here.

[0097] Since the screening is based on text units, even if there are some misspelled or missing words caused by character recognition, the impact on the final screening is relatively small, which can increase the probability of preserving effective information and reduce the probability of effective information being deleted by mistake. The effective information here refers to the information that can explain or supplement the context.

[0098] After screening out one or more text units whose relevant types are explanatory information and / or supplementary information, the text information set composed of the one or more text units is called the second text information.

[0099] Obtaining the second text information by further compressing and refining the first text information can greatly increase the effective information ratio of the information library and reduce invalid information and redundant information.

[0100] In some embodiments, candidate information can be directly formed based on the second text information and its corresponding context information, and the candidate information can be saved to the information library to form reference information for the question-answering system. In addition, text processing such as typo correction, missing word correction, semantic fusion, and grammar correction can also be performed based on the second text information and the context information to form a text with clear logic and organization (which can be implemented through a large model here), and saved to the information library as candidate information to improve the accuracy of the reference information in the information library. The candidate information can be unclassified and only needs to be complete.

[0101] Thus, the question-answering system can respond to a user's question to obtain the question information, retrieve the relevant information of the question information from at least one piece of candidate information, and generate the question-and-answer information corresponding to the question information based on the relevant information of the question information.

[0102] By generating candidate information based on the second text information and the context information, multiple pieces of candidate information can be formed based on a document to be processed, improving the processing efficiency of the document to be processed, supplementing the deficiencies of the text information of the document to be processed, and providing more accurate reference information for the construction of the information library.

[0103] In an exemplary embodiment, in order to further improve the accuracy of reference information, the second text information (paragraph fragment) may be further extracted and filtered based on the function of the question-and-answer system, so as to further retain valid information and remove redundant information, filter and retain high-quality sentences, and improve the information accuracy of the information library of the question-and-answer system.

[0104] In step S5, according to the matching result between the second text information and multiple preset questions, the third text information corresponding to the multiple preset questions is selected from the second text information, and the third text information includes sentences.

[0105] Wherein, a sentence refers to a string ending with a full stop. In the embodiments of the present disclosure, the third text information may also be certain words or combinations of words in a complete sentence.

[0106] Figure 4 It is a sub-flowchart of step S5 in the exemplary embodiment of the present disclosure.

[0107] Reference Figure 4 , in an exemplary embodiment, step S5 may include:

[0108] Step S51, determining the text relevance between the second text information and the preset question;

[0109] Step S52, when the text relevance is greater than a second preset value, determining the second text information as the third text information corresponding to the preset question.

[0110] The preset questions may be determined according to the setting objectives of the question-and-answer system, or according to the historical questions of the user in the usage record of the question-and-answer system. When initially constructing the question-and-answer system, multiple possible questions of the user may be formed according to the setting objectives of the question-and-answer system (such as solving database retrieval problems), and then these possible questions are saved as preset questions. When subsequently maintaining and updating the question-and-answer system, multiple question themes or multiple questions with the highest frequency of user questions can be summarized according to the historical questions of the user to the question-and-answer system. Multiple preset questions are generated according to the multiple question themes, or directly the multiple questions with the highest frequency of questions are set as preset questions. The above process can be implemented by constructing a question library corresponding to the question-and-answer model, and the process of constructing the question library can be implemented by collecting a large number of user questions (ensuring that various ways of asking are covered). Thus, in step S5, the questions in the question library can be used as preset questions for processing.

[0111] In Figure 5In the illustrated embodiments, information related to the preset questions in the question bank (with a text relevance greater than the second preset value) can be further screened out from the second text information and saved as the third information to form the reference information corresponding to the question bank. In step S5, candidate information can also be formed based on the second text information and its corresponding context information, and then the text relevance between the candidate information and the preset question text (the text relevance is greater than the second preset value) can be determined, so as to screen out the candidate information related to the preset question.

[0112] In some embodiments, the mean similarity between the above information (the second text information or the candidate information) and the preset question can be calculated through a text matching model (such as BERT) to determine the effectiveness of the information. If the mean similarity is greater than the set threshold (0.5), it is considered that this information may be frequently asked and belongs to valid information, and it is retained; otherwise, it is regarded as invalid information and removed.

[0113] After determining the third text information, in order to avoid problems with the fluency of the text of the third text information obtained by screening based on the original information and affecting subsequent responses, the refined text can be further polished to enhance its fluency, so as to ensure that the generated text can be used for high-quality responses, thereby improving the accuracy and relevance of the question-and-answer system.

[0114] The above processes can all be completed through a large model to improve efficiency and accuracy.

[0115] After obtaining the third text information or the second text information related to the image after refinement, the target image in the document to be processed can be replaced with the third text information or the second text information, and the document to be processed can be segmented according to the segmentation principles such as semantic coherence and then stored in the information library of the question-and-answer system.

[0116] By further screening out the third text information related to the preset questions of the question-and-answer system from the second text information, the proportion of valid information in the reference information of the question-and-answer system can be increased, while reducing the storage capacity of the information library and improving the information relevance and accuracy of the information library.

[0117] Figure 5 It is a schematic diagram of the effect of image processing in an exemplary embodiment of the present disclosure.

[0118] Reference Figure 5 , the initially screened target image is shown as A, including various image elements and text information; the first text information extracted from the target image is shown as B, and the information is fragmented and redundant, and there may also be errors and omissions. The finally formed third text information is shown as C, which is more refined and the key information is more prominent.

[0119] It can be seen that in the embodiments of the present disclosure, by adopting a triple information compression method, the images in the document to be processed are refined and filtered. First, image filtering is performed to determine whether the image contains information useful for understanding the segment. Then, paragraph filtering is carried out to perform a refined summary of the image and retain the effective paragraph information. Finally, text filtering is performed to compare and filter the paragraphs through text compression technology, and retain high-quality sentences for generating responses, which can completely solve the problem of introducing a large amount of noise due to the introduction of images in the process of multi-modal document processing.

[0120] In an exemplary embodiment, after determining the third text information, one or more pieces of third text information corresponding to a preset question can also be used to generate candidate answer information for the preset question in advance. Thus, when the user subsequently asks a question through the question-answering system, the preset question related to the user's question can be determined according to the user's question, and then according to the candidate answer information corresponding to the preset question, an answer to the user's question can be quickly and accurately formed, greatly improving the response efficiency and response accuracy of the question-answering system.

[0121] In practical applications, the method of the embodiments of the present disclosure can be integrated into an information retrieval and question-answering system. By performing triple filtering on the multi-modal information (document to be processed) input by the user, the system can quickly identify and extract useful information, thereby generating more accurate and relevant answers. For example, in a swarm intelligence question-answering system, the screenshots uploaded by the user can be processed by the method of the embodiments of the present disclosure, and the system can quickly extract the key information in the screenshots and give effective responses.

[0122] Figure 6 It is a schematic diagram of the working process of the question-answering system in the exemplary embodiments of the present disclosure.

[0123] Refer to Figure 6 , in step S601, a document to be processed is obtained from the multi-modal information library, and the document to be processed includes image and text information.

[0124] In step S602, the images are classified, and the images of the preset category are retained as target images.

[0125] In step S603, first text information is extracted from the target images, and the context information of the target images is extracted.

[0126] In step S604, image filtering is performed, and the target images are filtered according to the degree of relevance and types of relevance between the first text information and the context information to obtain target images that explain or supplement the context information.

[0127] In step S605, paragraph filtering is performed to further filter the first text information to retain the key paragraphs that explain or supplement the context information.

[0128] In step S606, text filtering is performed to retain text information related to the preset questions of the question-answering system, which can be referred to as image-compressed text.

[0129] In step S607, the information repository can be vectorized according to the context information and the image-compressed text.

[0130] In step S608, in response to a user's question, the user's question is vectorized.

[0131] In step S609, according to the results of question vectorization and information repository vectorization, segment recall is performed to retrieve reference information related to the question. The segments therein are, for example, text information.

[0132] In step S610, the segment most relevant to the question is screened from the recalled segments.

[0133] In step S611, an answer to the question is generated.

[0134] When the user asks a question, the question-answering system can find the segment most relevant to the user's question from the information repository through the "vector similarity calculation (BERT)-large model generation method", and then generate an answer to answer the user.

[0135] In summary, the triple-filtering method of the present disclosure embodiment can effectively reduce the noise in the image information, improve the quality and relevance of the information. Compared with the traditional OCR technology, it can extract the content useful for the user's understanding more accurately. On the one hand, the increase in effective information will improve the recall accuracy; on the other hand, the reduction of interference and redundant information will make the answers generated by the answering system more refined, thereby improving the overall performance of the multi-modal information processing system.

[0136] Corresponding to the above method embodiment, the present disclosure also provides an image information processing device in a multi-modal-based knowledge question-answering system, which can be used to execute the above method embodiment.

[0137] Figure 7 It is a block diagram of an image information processing device in a multi-modal-based knowledge question-answering system according to an exemplary embodiment of the present disclosure.

[0138] Reference Figure 7 FIG., the image information processing device 700 in the multi-modal-based knowledge question-answering system may include:

[0139] An image extraction module 71, configured to obtain a target image from a file to be processed, where the file to be processed includes text data and images, and the target image includes text information;

[0140] An information extraction module 72 is configured to extract text information in the target image as first text information, and extract context information of the target image in the file to be processed according to the text data;

[0141] An image screening module 73 is configured to select and retain the target image according to the first text information and the context information;

[0142] An information screening module 74 is configured to determine, according to the context information, second text information to be retained in the first text information of the retained target image, where the second text information includes paragraphs;

[0143] An association matching module 75 is configured to select third text information corresponding to the multiple preset questions in the second text information according to a matching result between the second text information and the multiple preset questions, where the third text information includes sentences.

[0144] In an exemplary embodiment of the present disclosure, the image extraction module 71 is configured to: determine the category of each image in the images of the file to be processed; and determine the image corresponding to the preset category as the target image.

[0145] In an exemplary embodiment of the present disclosure, the image screening module 73 is configured to: determine the degree of correlation between the first text information and the context information; when the degree of correlation is greater than a first preset value, determine the type of correlation between the first text information and the context information, where the type of correlation at least includes explanatory information and supplementary information; and when the type of correlation is the explanatory information and / or the supplementary information, retain the target image.

[0146] In an exemplary embodiment of the present disclosure, the information screening module 74 is configured to: segment the first text information to form multiple text units; determine the type of correlation between each text unit and the context information; and save the text units whose type of correlation is the explanatory information and / or the supplementary information as the second text information.

[0147] In an exemplary embodiment of the present disclosure, the association matching module 75 is configured to: determine the text correlation between the second text information and the preset questions; and when the text correlation is greater than a second preset value, determine the second text information as the third text information corresponding to the preset questions.

[0148] In an exemplary embodiment of the present disclosure, it further includes an answer prefabrication module, which is configured to: generate candidate answer information for the preset questions according to the third text information corresponding to the preset questions.

[0149] In an exemplary embodiment of the present disclosure, it further includes an information collection module, which is configured to: form candidate information according to the second text information and the context information; obtain question information in response to a user's question, and determine relevant information of the question information from at least one of the candidate information; generate answer information corresponding to the question information according to the relevant information of the question information.

[0150] Since the functions of device 700 have been described in detail in their corresponding method embodiments, the present disclosure will not repeat them here.

[0151] It should be noted that although several modules or units of a device for action execution are mentioned in the above detailed description, such a division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more of the above-described modules or units can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0152] In an exemplary embodiment of the present disclosure, there is also provided an electronic device capable of implementing the above method.

[0153] Those skilled in the art can understand that various aspects of the present invention can be implemented as a system, a method, or a program product. Therefore, various aspects of the present invention can be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software aspects, which can be collectively referred to as "circuit", "module", or "system" here.

[0154] Next, refer to Figure 8 to describe the electronic device 800 according to this embodiment of the present invention. Figure 8 The electronic device 800 shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present invention.

[0155] As Figure 8 shown, the electronic device 800 is presented in the form of a general-purpose computing device. The components of the electronic device 800 may include, but are not limited to: at least one of the above-mentioned processing units 810, at least one of the above-mentioned storage units 820, and a bus 830 connecting different system components (including the storage unit 820 and the processing unit 810).

[0156] Among them, the storage unit stores program code, and the program code can be executed by the processing unit 810, so that the processing unit 810 executes the steps according to various exemplary embodiments of the present invention described in the above "exemplary method" part of this specification. For example, the processing unit 810 can execute the method as shown in the embodiments of the present disclosure.

[0157] The storage unit 820 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 8201 and / or a cache storage unit 8202, and may further include a read-only storage unit (ROM) 8203.

[0158] The storage unit 820 may also include a program / utilities 8204 having a set (at least one) of program modules 8205. Such program modules 8205 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment.

[0159] The bus 830 may represent one or more of several types of bus structures, including a storage unit bus or storage unit controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any of a variety of bus structures.

[0160] The electronic device 800 may also communicate with one or more external devices 900 (such as a keyboard, a pointing device, a Bluetooth device, etc.), may also communicate with one or more devices that enable a user to interact with the electronic device 800, and / or may communicate with any device that enables the electronic device 800 to communicate with one or more other computing devices (such as a router, a modem, etc.). Such communication may be carried out through an input / output (I / O) interface 850. Moreover, the electronic device 800 may also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through a network adapter 860. As shown in the figure, the network adapter 860 communicates with other modules of the electronic device 800 through the bus 830. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with the electronic device 800, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.

[0161] Through the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software, or can be implemented by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.

[0162] In an exemplary embodiment of the present disclosure, there is also provided a computer-readable storage medium, on which a program product capable of implementing the above methods of this specification is stored. In some possible implementation manners, various aspects of the present invention can also be implemented in the form of a program product, which includes program code. When the program product runs on a terminal device, the program code is used to cause the terminal device to execute the steps according to various exemplary embodiments of the present invention described in the above "Exemplary Method" section of this specification.

[0163] The program product for implementing the above method according to an embodiment of the present invention can be a portable compact disc read-only memory (CD-ROM) and includes program code, and can run on a terminal device, such as a personal computer. However, the program product of the present invention is not limited thereto. In this document, the readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0164] The program product can adopt any combination of one or more readable media. The readable media can be a readable signal medium or a readable storage medium. The readable storage medium can, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0165] The computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries the readable program code. Such a propagated data signal can take various forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the above. The readable signal medium can also be any readable medium other than the readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0166] The program code contained on the readable medium can be transmitted by any appropriate medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination of the above.

[0167] Program code for performing the operations of the present invention may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's device, executed as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., through the Internet using an Internet service provider).

[0168] In addition, the above-mentioned drawings are only schematic illustrations of the processes included in the method according to the exemplary embodiments of the present invention, and are not for limiting purposes. It is easy to understand that the processes shown in the above-mentioned drawings do not indicate or limit the time sequence of these processes. Additionally, it is also easy to understand that these processes may be executed, for example, synchronously or asynchronously in multiple modules.

[0169] Other embodiments of the present disclosure will be readily apparent to those skilled in the art after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of the present disclosure are pointed out by the claims.

Claims

1. An image information processing method in a multimodal-based knowledge Q&A system, characterized in that Including: Obtain a target image from a file to be processed, where the file to be processed includes text data and images, and the target image includes text information; Extract the text information in the target image as first text information, and extract the context information of the target image in the file to be processed according to the text data; Select the target images to be retained according to the first text information and the context information; According to the context information, determine the second text information to be retained in the first text information of the retained target image, where the second text information includes paragraphs; According to the matching results of the second text information and multiple preset questions, select the third text information corresponding to the multiple preset questions in the second text information, where the third text information includes sentences.

2. The method for processing image information in the multimodal-based knowledge Q&A system according to claim 1, wherein Obtaining a target image from a file to be processed includes: Determine the category of each image in the images of the file to be processed; Determine the images corresponding to the preset category as the target images.

3. The image information processing method in the multimodal-based knowledge Q&A system according to claim 1, wherein Selecting the target images to be retained according to the first text information and the context information includes: Determine the degree of relevance between the first text information and the context information; When the degree of relevance is greater than a first preset value, determine the types of relevance between the first text information and the context information, where the types of relevance at least include explanatory information and supplementary information; When the type of relevance is the explanatory information and / or the supplementary information, retain the target image.

4. The image information processing method in the multimodal-based knowledge Q&A system according to claim 1, characterized in that, The determining the second text information to be retained in the first text information of the retained target image according to the context information includes: Segment the first text information to form multiple text units; Determine the type of relevance between each text unit and the context information; Save the text units whose type of relevance is the explanatory information and / or the supplementary information as the second text information.

5. The method for processing image information in the multimodal-based knowledge Q&A system according to claim 1, wherein According to the matching results of the second text information and multiple preset questions, selecting the third text information corresponding to the multiple preset questions in the second text information includes: Determine the text relevance between the second text information and the preset questions; When the text relevance is greater than a second preset value, determine the second text information as the third text information corresponding to the preset questions.

6. The image information processing method in the multimodal-based knowledge Q&A system according to claim 1, wherein It further includes: Generate candidate answer information for the preset questions according to the third text information corresponding to the preset questions.

7. The image information processing method in the multimodal-based knowledge Q&A system according to claim 1, wherein It further includes: Form candidate information according to the second text information and the context information; Respond to a user's question to obtain question information, and determine the relevant information of the question information in at least one of the candidate information; Generate the Q&A information corresponding to the question information according to the relevant information of the question information.

8. An electronic device, characterized in that, Including: A memory; And A processor coupled to the memory, where the processor is configured to execute the method according to any one of claims 1-7 based on instructions stored in the memory.

9. A computer-readable storage medium, on which a program is stored, and when the program is executed by a processor, it implements the method according to any one of claims 1-7.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1-7.

Citation Information

Cited By

  • Multi-mode question answering system, reply information generation method and electronic equipment

    CN121543716A