Generative search illustration method and device, storage medium and program product

By scoring and filtering candidate images in multimodal AI search image matching technology, and combining this with the large language model rewriting problem, the issue of invalid or misleading image search results in existing technologies has been resolved, resulting in a higher quality search experience.

CN120892589AActive Publication Date: 2025-11-04UC MOBILE CHINA CO LTD
View PDF 12 Cites 0 Cited by

Patent Information

Application Number
CN202510775484.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-11-04
Estimated Expiration
2045-06-10

AI Technical Summary

Technical Problem

In existing multimodal AI search image matching technologies, image search results offer limited help to users' questions and may even mislead users, resulting in a poor search experience.

Method used

Multiple candidate images are identified based on the original question from the user's search. A multimodal model is then used to evaluate the value, quality, and time of the candidate images to select those that can effectively answer the question and are of high quality. Finally, a large language model is used to rewrite the question to improve retrieval accuracy.

Benefits of technology

Improving the images in search results truly helps users understand the problem, enhances the user experience, and ensures the relevance and quality of the images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120892589A_ABST
    Figure CN120892589A_ABST
Patent Text Reader

Abstract

Embodiments of the invention provide a generative search illustration method and device, a storage medium and a program product. The method comprises the following steps: determining a plurality of candidate pictures based on an original question searched by a user; the original question and the candidate pictures are input into a multi-modal model, the value score, the quality score and the picture time of each candidate picture are obtained, the value score represents the contribution degree of the candidate pictures to answer the original question, and the quality score represents the quality score of the candidate pictures to answer the original question. The quality score represents the picture quality of the candidate picture on the premise that the candidate picture is used for answering the original question; and according to the value score, the quality score and the picture time of each candidate picture, determining a to-be-displayed picture in a search answer of the original question, the search answer of the original question including the to-be-displayed picture and a text answer of the original question. According to the method, the problem that illustrations in search answers are not helpful for answering user questions or even mislead users is solved, and the search experience of the users is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence, and in particular to a generative search image matching method, device, storage medium, and program product. Background Technology

[0002] Most current generative AI (Artificial Intelligence) searches are limited to text-based search results, which restricts the ability to describe complex concepts or entities. Multimodal AI search with accompanying images aims to explore how to integrate visual elements such as images and charts into search results, achieving a richer, more multimedia presentation of information.

[0003] Currently, multimodal AI search image matching technology involves vectorizing images and user queries, performing alignment training, and then building an image vector index. When a user enters a question, the similarity between the question vector and the image vectors in the index is calculated, and images with high similarity are used as search results. However, this approach results in a large number of images that are unhelpful or even misleading in answering user questions appearing in search results, severely degrading the user's search experience. Summary of the Invention

[0004] This application provides a generative search image matching method, device, storage medium, and program product to solve the problem that the images in the search results do not help answer the user's question or even mislead the user, seriously reducing the user's search experience.

[0005] In a first aspect, embodiments of this application provide a generative search method for image matching, including:

[0006] Multiple candidate images were identified based on the user's original search question;

[0007] The original question and the multiple candidate images are input into a multimodal model to obtain the value score, quality score, and image time of each candidate image. The value score represents the contribution of the candidate image to answering the original question, and the quality score represents the image quality of the candidate image when used to answer the original question.

[0008] The search answer to the original question is determined based on the value score, quality score, and image time of each candidate image. The search answer to the original question includes the image to be displayed and the text answer to the original question.

[0009] In some implementations, the determination of multiple candidate images based on the user's original search question includes:

[0010] The original question is input into a large language model to obtain the rewriting question corresponding to the original question and the text understanding information of the original question.

[0011] If the text understanding information determines that an image search is triggered, the multiple candidate images are retrieved from the image library based on the original question and the rewritten question.

[0012] In some implementations, the large language model includes a generative model and a classification model. The step of inputting the original question into the large language model to obtain the rewritten question corresponding to the original question and the text understanding information of the original question includes:

[0013] The original problem is input into the generative model to obtain the rewritten problem corresponding to the original problem;

[0014] The original question is input into the classification model to obtain the text understanding information.

[0015] In some implementations, the text understanding information includes industry classification, content classification, and trigger intensity, wherein the trigger intensity is used to characterize the intensity level of the original question triggering the image search;

[0016] The step of determining the trigger for image search based on the text understanding information includes:

[0017] If the industry category and / or the content category belong to the category that triggers image search, a trigger strength threshold shall be determined based on the industry category and / or the content category;

[0018] If the trigger intensity is greater than or equal to the trigger intensity threshold, a trigger search is determined.

[0019] Some implementations also include:

[0020] Input the original question into the large language model to obtain the entity names in the original question and / or the rewritten question;

[0021] The step of inputting the original problem and the multiple candidate images into the multimodal model includes:

[0022] The candidate images are filtered based on the entity name and the feature information of each candidate image, and the filtered candidate images and the original question are input into the multimodal model.

[0023] In some implementations, the step of filtering the multiple candidate images based on the entity name and the feature information of each candidate image includes:

[0024] If the feature information of any candidate image among the plurality of candidate images does not contain the entity name, then the candidate image is removed.

[0025] In some implementations, determining the images to be displayed in the search answer based on the value score, quality score, and image time of each candidate image includes:

[0026] The total score for each candidate image is determined based on its value score, quality score, and image time.

[0027] The candidate images are ranked based on their total scores, and a predetermined number of the top-ranked candidate images are selected as the images to be displayed.

[0028] In some implementations, the text understanding information includes timeliness information;

[0029] The process of determining the total score for each candidate image based on its value score, quality score, and image time includes:

[0030] The timeliness score of each candidate image is determined based on the image time of each candidate image, and the timeliness weight is determined based on the timeliness information.

[0031] The total score for each candidate image is determined based on its value score, quality score, image time, preset value weight, preset quality weight, and timeliness weight.

[0032] Some implementations also include:

[0033] For any candidate image among all candidate images, if the value score of the candidate image is lower than the first preset value score and / or the quality score is lower than the preset quality score, the candidate image is removed.

[0034] In some implementations, the text understanding information includes timeliness information; the method further includes:

[0035] If the image time and the timeliness information of any candidate image do not match, the candidate image will be removed.

[0036] In some implementations, the step of retrieving multiple candidate images from the image library based on the original problem and the rewritten problem includes:

[0037] The feature vectors of the original problem and the rewritten problem are determined. The feature vectors of the original problem and the rewritten problem are then matched with the image-text feature vectors corresponding to the images in the image library for similarity. Images with similarity exceeding a preset similarity are identified as candidate images. The image-text feature vector is a fusion feature vector of the image and the corresponding text information, and it is obtained based on a pre-trained multimodal alignment model.

[0038] Some implementations also include:

[0039] Obtain the sample question, sample image, and text information of the sample image;

[0040] A multimodal alignment model is trained based on the sample question, the sample image, and the text information of the sample image. The multimodal alignment model is used to align the sample question and the fusion information of the image and text, where the fusion information of the image and text is the fusion information of the sample image and the text information of the sample image.

[0041] The trained multimodal alignment model is used to extract the image and text feature vectors corresponding to the images in the image library.

[0042] Some implementations also include:

[0043] Generate a text answer based on the original question and / or the rewritten question;

[0044] The image to be displayed is placed before the text answer in the search results.

[0045] Some implementations also include:

[0046] If the number of images with a value score greater than or equal to the second preset value score is less than a preset value, the images to be displayed in the search answer will be folded.

[0047] Some implementations also include:

[0048] If it is determined that an image search is triggered, but the result of the candidate image is empty or the result of the image to be displayed is empty, the text answer of the original question is input into the image generation model to obtain the image generated by the image generation model.

[0049] Add the image generated by the image generation model to the search results.

[0050] Some implementations also include:

[0051] The initial search results for the original question and the image to be displayed are input into the generative large model to obtain the search answer for the original question. The image to be displayed is used to be displayed along with the text answer during the generation of the text answer in the search answer.

[0052] In some implementations, the determination of multiple candidate images based on the user's original search question includes:

[0053] If the original question is not in text form, the original question will be converted into text.

[0054] Based on the original question after textualization, multiple candidate images are identified.

[0055] Secondly, embodiments of this application provide a generative search image matching device, comprising:

[0056] The retrieval module is used to identify multiple candidate images based on the user's original search question;

[0057] The evaluation module is used to input the original question and the multiple candidate images into a multimodal model to obtain a value score, a quality score, and an image time for each candidate image. The value score represents the contribution of the candidate image to answering the original question, and the quality score represents the image quality of the candidate image under the premise of answering the original question.

[0058] The determination module is used to determine the image to be displayed in the search answer to the original question based on the value score, quality score and image time of each candidate image, wherein the search answer to the original question includes the image to be displayed and the text answer to the original question.

[0059] In some implementations, the retrieval module is used for:

[0060] The original question is input into a large language model to obtain the rewriting question corresponding to the original question and the text understanding information of the original question.

[0061] If the text understanding information determines that an image search is triggered, the multiple candidate images are retrieved from the image library based on the original question and the rewritten question.

[0062] In some implementations, the large language model includes a generative model and a classification model, and the retrieval module is used for:

[0063] The original problem is input into the generative model to obtain the rewritten problem corresponding to the original problem;

[0064] The original question is input into the classification model to obtain the text understanding information.

[0065] In some implementations, the text understanding information includes industry classification, content classification, and trigger intensity, wherein the trigger intensity is used to characterize the intensity level of the original question triggering the image search;

[0066] The retrieval module is used for:

[0067] If the industry category and / or the content category belong to the category that triggers image search, a trigger strength threshold shall be determined based on the industry category and / or the content category;

[0068] If the trigger intensity is greater than or equal to the trigger intensity threshold, a trigger search is determined.

[0069] In some implementations, the retrieval module is further used for:

[0070] Input the original question into the large language model to obtain the entity names in the original question and / or the rewritten question;

[0071] The evaluation module is used for:

[0072] The candidate images are filtered based on the entity name and the feature information of each candidate image, and the filtered candidate images and the original question are input into the multimodal model.

[0073] In some implementations, the evaluation module is used for:

[0074] If the feature information of any candidate image among the plurality of candidate images does not contain the entity name, then the candidate image is removed.

[0075] In some implementations, the determining module is used to:

[0076] The total score for each candidate image is determined based on its value score, quality score, and image time.

[0077] The candidate images are ranked based on their total scores, and a predetermined number of the top-ranked candidate images are selected as the images to be displayed.

[0078] In some implementations, the text understanding information includes timeliness information; the determining module is used for:

[0079] The timeliness score of each candidate image is determined based on the image time of each candidate image, and the timeliness weight is determined based on the timeliness information.

[0080] The total score for each candidate image is determined based on its value score, quality score, image time, preset value weight, preset quality weight, and timeliness weight.

[0081] In some implementations, the determining module is further configured to:

[0082] For any candidate image among all candidate images, if the value score of the candidate image is lower than the first preset value score and / or the quality score is lower than the preset quality score, the candidate image is removed.

[0083] In some implementations, the text understanding information includes timeliness information; the determining module is further configured to:

[0084] If the image time and the timeliness information of any candidate image do not match, the candidate image will be removed.

[0085] In some implementations, the retrieval module is used for:

[0086] The feature vectors of the original problem and the rewritten problem are determined. The feature vectors of the original problem and the rewritten problem are then matched with the image-text feature vectors corresponding to the images in the image library for similarity. Images with similarity exceeding a preset similarity are identified as candidate images. The image-text feature vector is a fusion feature vector of the image and the corresponding text information, and it is obtained based on a pre-trained multimodal alignment model.

[0087] Some implementations also include: a training module, used for:

[0088] Obtain the sample question, sample image, and text information of the sample image;

[0089] A multimodal alignment model is trained based on the sample question, the sample image, and the text information of the sample image. The multimodal alignment model is used to align the sample question and the fusion information of the image and text, where the fusion information of the image and text is the fusion information of the sample image and the text information of the sample image.

[0090] The trained multimodal alignment model is used to extract the image and text feature vectors corresponding to the images in the image library.

[0091] In some implementations, a first processing module is also included, used for

[0092] Generate a text answer based on the original question and / or the rewritten question;

[0093] The image to be displayed is placed before the text answer in the search results.

[0094] In some implementations, a second processing module is also included, used for:

[0095] If the number of images with a value score greater than or equal to the second preset value score is less than a preset value, the images to be displayed in the search answer will be folded.

[0096] Some implementations also include a third processing module, used for:

[0097] If it is determined that an image search is triggered, but the result of the candidate image is empty or the result of the image to be displayed is empty, the text answer of the original question is input into the image generation model to obtain the image generated by the image generation model.

[0098] Add the image generated by the image generation model to the search results.

[0099] In some implementations, a fourth processing module is also included, used for:

[0100] The initial search results for the original question and the image to be displayed are input into the generative large model to obtain the search answer for the original question. The image to be displayed is used to be displayed along with the text answer during the generation of the text answer in the search answer.

[0101] In some implementations, the retrieval module is used for:

[0102] If the original question is not in text form, the original question will be converted into text.

[0103] Based on the original question after textualization, multiple candidate images are identified.

[0104] Thirdly, embodiments of this application provide an electronic device, including: a memory and a processor;

[0105] The memory stores computer-executed instructions;

[0106] The processor executes computer execution instructions stored in the memory, causing the processor to perform the method described in any of the first aspects.

[0107] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the method described in any of the first aspects.

[0108] Fifthly, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the method shown in any of the first aspects.

[0109] This application provides a generative search image matching method, device, storage medium, and program product. The method determines multiple candidate images based on the user's original search question; it inputs the original question and multiple candidate images into a multimodal model to obtain a value score, quality score, and image time for each candidate image. The value score represents the candidate image's contribution to answering the original question, and the quality score represents the image quality of the candidate image when used to answer the original question. Based on the value score, quality score, and image time of each candidate image, it determines the images to be displayed in the search answer. By utilizing the matching of candidate images to the value, quality, and timeliness of answering the question, it selects images from the candidate images that can answer the user's question and are of good quality, thereby ensuring that the images in the search results truly help the user understand the question and improve the user experience. Attached Figure Description

[0110] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0111] Figure 1 A schematic diagram illustrating the application scenarios of the multimodal AI search and image matching technology provided in this application;

[0112] Figure 2 A flowchart illustrating a generative search image matching method provided for an exemplary embodiment of this application;

[0113] Figure 3 A schematic diagram of a search answer provided for an exemplary embodiment of this application;

[0114] Figure 4 Another schematic diagram of a search answer provided for an exemplary embodiment of this application;

[0115] Figure 5 Another schematic flowchart of a generative search image matching method provided for an exemplary embodiment of this application;

[0116] Figure 6 A schematic diagram of a generative search mapping device provided for an exemplary embodiment of this application;

[0117] Figure 7 This is a schematic diagram of the structure of an electronic device provided as an exemplary embodiment of this application. Detailed Implementation

[0118] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with relevant laws, regulations and standards, and corresponding operation entry points are provided for users to choose to authorize or refuse.

[0119] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0120] Combination Figure 1 The following example illustrates the application scenario of the multimodal AI search image matching technology involved in this application. When a user searches for the question "Explanation of the offside rule in football," the search results, in addition to outputting the explanation of the offside rule in text form, also output images to assist the user's understanding. The offside rule involves restrictions on the positional relationships of players, and the illustrations help users understand the rule more intuitively. In other words, the purpose of adding images to the search results in the multimodal AI search image matching technology is to answer the user's search question using images, in addition to textual answers. It is understood that the images in the search answer can be one or more pictures. Figure 1 The image shown is used for illustration.

[0121] However, because related technologies often only calculate vector similarity between user questions and images when creating images, the output images may not be helpful in answering user questions and could even mislead users. For example, regarding the question "Explanation of the offside rule in football," the output image might be used to explain the rules of football, but it could explain other rules that do not include the offside rule, resulting in a poor search experience.

[0122] To achieve more accurate image matching, in this embodiment, candidate images are retrieved based on the user's search question. These candidate images are then further selected using a multimodal model. The model evaluates the candidate images, considering their helpfulness in answering the user's question. The multimodal model outputs a value score, a quality score, and time information for each candidate image. By leveraging the match between the candidate images and the question's value, quality, and timeliness, images that effectively answer the user's question and are of good quality are selected. This ensures that the images in the search results truly help the user understand the question, thus improving the user experience.

[0123] The technical solutions shown in this application will now be described in detail through specific embodiments. It should be noted that the following embodiments may exist independently or in combination with each other; for identical or similar content, the description will not be repeated in different embodiments.

[0124] The execution subject in this application embodiment can be an electronic device or a generative search mapping device installed in an electronic device. The request processing device can be implemented by software or by a combination of software and hardware. The generative search mapping device can be a processor in an electronic device. For ease of understanding, the following description will use an electronic device as the execution subject.

[0125] Figure 2 This is a flowchart illustrating a generative search image matching method provided for an exemplary embodiment of this application. Please refer to... Figure 2 The methods may include:

[0126] S201. Determine multiple candidate images based on the original question searched by the user.

[0127] The original question in a user search refers to the content entered by the user when conducting an AI search. It should be noted that when the original question is used to represent the content entered by the user, the "question" here does not limit the content entered by the user to the form of a question, such as "What is the offside rule in a football match?". The original question can also be any form of a declarative sentence, imperative sentence, etc., such as "Please explain the offside rule in a football match".

[0128] Searching the image library based on the original question can yield multiple candidate images. Optionally, a web search can be performed based on the original question, combining the original question and the web search results for image retrieval. Alternatively, the original question can be rewritten, and the rewritten question can be used together with the original question for image retrieval.

[0129] The original question can be in text or non-text form; for example, a non-text question can be speech. Optionally, if the original question is in non-text form, it can be first processed into text, and multiple candidate images can be determined based on the text-processed original question. In subsequent embodiments, examples will be provided where the original question is in text form or a text-processed original question.

[0130] Taking the user question "Please explain the offside rule in a football match" as an example, candidate images might match some words in the original question but offer no help in answering it. For instance, a candidate image might be a picture of a football field, which wouldn't help answering the question. Candidate images might be related to the original question, offering little or no help in answering it. For example, a candidate image might be a real picture of a player being offside in a football match. While such an image shows offside, it doesn't explain the rule, making it difficult for the user to truly understand the rule and thus offering little help in answering the question. Candidate images might include answers to the original question. For example, a candidate image might be a diagram of a football field or part of a field, illustrating the positional relationships of players that constitute offside, and including textual explanations of the diagram. Such candidate images can directly answer the user's question. In other words, the candidate images in this step might have some correlation with the original question, but whether they truly help answer the original question requires further evaluation.

[0131] S202. Input the original question and multiple candidate images into the multimodal model to obtain the value score, quality score and image time of each candidate image. The value score represents the contribution of the candidate image to answering the original question, and the quality score represents the image quality of the candidate image under the premise of answering the original question.

[0132] Multimodal models can simultaneously understand and generate multiple types of data (such as text, images, audio, and video), and achieve more complex tasks through the correlation between different modalities. For example, multimodal models can perform cross-modal understanding and analyze the correlation between data from different modalities. In this embodiment, a multimodal model is used to perform cross-modal understanding of the original question and candidate images, thereby determining whether the candidate images are helpful in answering the original question.

[0133] Value scores characterize the contribution of candidate images to answering the original question, that is, the degree to which the candidate image can answer the original question. Value scores can be presented in the form of scores and / or tiers. For example, a value score of 4 indicates that the candidate image answers the question directly; the image's main content is the answer to the question, and the user can easily grasp the answer from the image. A value score of 3 indicates that the candidate image partially answers the question. The candidate image cannot completely answer the user's question, but it can partially answer it; the content shown in the image is only a part of the complete answer to the user's question. A value score of 2 indicates that the candidate image is relevant information to the original question. This candidate image does not answer the user's question; the answer to the question cannot be found in the image, but the image content is relevant information and can provide some help in answering the user's question. A value score of 1 indicates that the candidate image is an illustration for the original question. This candidate image cannot answer the user's question, but the image is aesthetically pleasing, has no watermark, and provides visual recognition of the main entities in the question, thus providing useful value to the user in understanding their question. A value score of 0 indicates that the candidate image is a worthless image. This candidate image cannot answer the user's question, does not provide any useful value to the user in understanding the question asked, and may even mislead the user in understanding the question.

[0134] The quality score characterizes the image quality of candidate images in answering the original question. This means the candidate image is rich in image and text, concise, and free of low-quality elements such as watermarks and obstructions. The quality score can be presented as a score and / or a tiered system. For example, a quality score of 5 indicates a high-quality candidate image that is valuable to the user question, of high quality, rich in image and text, concise, free of watermarks, and with an obvious answer. A quality score of 4 indicates a good candidate image that is valuable to the user question, but of lower quality than a high-quality image and less aesthetically pleasing. A quality score of 3 indicates a fair candidate image that is valuable to the user question, with no obvious quality issues, but its aesthetic appeal is insufficient. A quality score of 2 indicates that the candidate image is of poor quality. While valuable to the user's question, it has minor quality issues, such as a screenshot with relatively clear text, poorly designed text, a prominent watermark making it unattractive, the main text not being in simplified Chinese, the answer being obscure, cropped content, or other poor quality characteristics. A quality score of 1 indicates that the candidate image is of low quality. While valuable to the user's question, it has significant quality issues, such as a screenshot with unclear text, a page from a full-text Word document, an e-commerce product image, a large watermark affecting visual appeal, or other very poor visual quality. A quality score of 0 indicates that the candidate image is worthless. It fails to answer the user's question, provides no useful information, and may even mislead the user; the answer to the user's question is not found in the image.

[0135] Image time refers to the temporal information extracted from candidate images by the multimodal model. For example, it represents the time span from the current time when the main content of an image is extracted. This time span can be categorized as current month, within six months, current year, one year, two years or more, or when the time cannot be confirmed. In many cases, user searches are time-sensitive. For instance, a user might search for "Who is the CEO of XX company?", and the CEO of XX company rotates according to a certain time period. Therefore, temporal information needs to be considered to determine whether the content in the candidate image is a correct answer for the current time or outdated historical information.

[0136] S203. Determine the images to be displayed in the search response to the original question based on the value score, quality score, and image time of each candidate image. The search response to the original question includes the images to be displayed and the text response to the original question.

[0137] The value score, quality score, and image time of a candidate image collectively determine whether the candidate image adequately answers the user's question. Optionally, candidate images can be filtered using any one or more of the value score, quality score, and image time. Optionally, image time can be converted into a timeliness score and combined with the value score and quality score to determine the total score of each candidate image. Based on the total score of each candidate image, the candidate images are ranked, and a predetermined number of the top-ranked candidate images are selected as the images to be displayed. For example, if there are 10 candidate images, the top 5 candidate images with the highest total scores are selected as the images to be displayed. In this embodiment, the search answer to the original question includes a text answer and images to be displayed determined in this step. The images to be displayed are used to answer or partially answer the original question. The images to be displayed can supplement the text answer or present the content of the text answer in a more intuitive and concise way.

[0138] The image search method provided in this application determines multiple candidate images based on the user's original search question. The original question and the multiple candidate images are input into a multimodal model to obtain a value score, quality score, and image time for each candidate image. The value score represents the contribution of the candidate image to answering the original question, and the quality score represents the image quality of the candidate image when used to answer the original question. Based on the value score, quality score, and image time of each candidate image, the method determines the images to be displayed in the search results. This method retrieves candidate images based on the user's original search question, then evaluates the candidate images using a multimodal model. Considering whether the candidate images are helpful in answering the user's question, the multimodal model outputs the value score, quality score, and time information of the candidate images. By utilizing the matching of the candidate images to the value, quality, and timeliness of answering the question, images that can answer the user's question and are of good quality are selected from the candidate images, thereby ensuring that the images in the search results can truly help users understand the question and improve the user experience.

[0139] Based on the above embodiments, the determination of multiple candidate images based on the original question searched by the user in S201 will be explained.

[0140] In one implementation, the original question is input into a large language model to obtain the rewritten question corresponding to the original question and the text understanding information of the original question; when the text understanding information determines that an image search is triggered, multiple candidate images are retrieved from the image library based on the original question and the rewritten question.

[0141] The original question searched by a user may not be sufficiently detailed or precise, or may contain ambiguities, due to the user's own language habits. In this embodiment, a large language model is used to rewrite the user's original question, generating a rewritten question. The rewritten question still describes the problem in natural language, but can expand or modify the description based on the original question. Optionally, based on the original question, web search results can be combined, i.e., a web search is performed on the original question. The information from the web search results and the original question are used together as input information for the large language model, allowing the rewritten question generated by the large language model to describe the problem more precisely. Optionally, the large language model includes a generation model. The original question is input into the generation model, or the original question and web search results are input into the generation model to obtain the rewritten question corresponding to the original question. The rewritten question output by the large language model can be one or more. For example, the original question is "Who is the CEO of XX Company?", and the rewritten questions are "Introduction to the CEO of XX Company" and "New CEO of XX Company". By rewriting the original question, the user's true intent can be captured more accurately, improving the accuracy and relevance of subsequent searches, thereby improving image retrieval effects.

[0142] Furthermore, in some cases, certain questions are not suitable for presentation in the form of images. Therefore, in this embodiment of the application, before searching for candidate images, a judgment is made on whether to trigger image search. The large language model is used to output textual understanding information of the original question, and the judgment is made on whether to trigger image search based on the textual understanding information.

[0143] Optionally, the large language model includes a classification model. The original question is input into the classification model, or the original question and webpage search results are input into the classification model to obtain text understanding information. Optionally, the text understanding information includes industry classification, content classification, and trigger strength, where trigger strength is used to characterize the intensity level of the original question triggering image search.

[0144] Industry categories indicate the industry to which the user's original search question belongs. For example, industry categories include the following: Metaphysics / Folklore, Language / Literature, Educational Information, Electrical and Electronic Engineering, Accounting, Sports, Home Appliances and Furniture, Traditional Games, Video Games, Mechanical / Automation Engineering, Plumbing and Electrical Installation, 3C Digital Products, Other - Everyday Knowledge, Common Software, Other - Professional Knowledge, Economics, Tourism, Philosophy, Biology, Animation, Geology / Geography, Social and Political Affairs, Finance / Financial Management, Unclear Intent, General Food, Animals and Plants, General Commodities, Education, History / Cultural Relics, Workplace, Foreign Languages, Business / Management, General Entertainment, Mathematics, Physics and Chemistry, Music, Real Estate, General Apparel, Transportation, Legal Professionals, Novels, Civil Engineering, General Healthcare, Automobiles / Electric Vehicles, Medicine / Pharmacy, Film and Television, and Computer Science.

[0145] Content categories indicate what kind of content a user is searching for. For example, content categories include: TV series / chapter retrieval, formulas / chemical formulas, score lines, game locations / maps, pronunciation, etiquette, strategies / guides, salary and benefits, legal provisions / law retrieval, plot summaries / interpretations, contact information / social numbers, time retrieval, words / article retrieval, price retrieval, town locations, work inferences, term explanations / definitions, match scores, cause / principle explanations, fault codes, unclear intent, institution locations, institution retrieval, item dimensions, security checks / prohibited items regulations, character inferences, institution attributes / introductions, recommendation lists, classical Chinese translations, foreign language translations, weather retrieval, tourist attraction locations, other content, professional introductions, sports records, character introductions / attributes, traffic regulations / car insurance, numerical calculations / conversions, actor retrieval, game cheat codes, license plate locations, exam / competition regulations and standards, social statistics, distance retrieval, ranking lists, code programming, and internet memes.

[0146] Trigger strength is used to characterize the intensity level of the original question triggering the image search. Trigger strength can be presented as a score or a level. For example, trigger strength is divided into level 3, level 2, level 1, and level 0, with the intensity decreasing in that order.

[0147] For industry categories and content categories, you can pre-define which categories trigger image searches or do not. You can set separate categories for each category, or you can set categories for both categories. For example, if the industry category is "Mysticism" and the content category is "Character Introduction," then an image search will not be triggered.

[0148] If an industry category and / or content category falls under the category that triggers image search, a corresponding trigger strength threshold can be set. That is, if an industry category and / or content category falls under the category that triggers image search, the trigger strength threshold is determined based on the industry category and / or content category; if the trigger strength is greater than or equal to the trigger strength threshold, an image search is triggered. For example, if the industry category is "Civil Engineering" and the content category is "Methods and Strategies," the corresponding trigger strength threshold is level 3. Similarly, if the industry category is "Tourism" and the content category is "Scenic Spot Locations," the corresponding trigger strength threshold is level 0.

[0149] Once the text understanding information determines that an image search is triggered, multiple candidate images are retrieved from the image library based on the original question and the rewritten question.

[0150] The feature vectors of the original problem and the rewritten problem are determined. The feature vectors of the original problem and the rewritten problem are then matched with the image and text feature vectors corresponding to the images in the image library for similarity. Images with similarity exceeding a preset similarity are identified as candidate images. The image and text feature vector is a fusion feature vector of the image and the corresponding text information, and it is obtained based on a pre-trained multimodal alignment model.

[0151] The image library in this embodiment includes images and their image-text feature vectors. Optionally, it may also include feature information for each image. The image feature information refers to key elements and their attributes extracted from the image, image tags, image type, etc., such as the image type being a portrait image, and the name of the person in the image. The image feature information can provide richer and more detailed foundational support for subsequent retrieval or filtering, helping to improve accuracy.

[0152] Optionally, when retrieving candidates from the image library, the feature vector of the original question is compared with the image-text feature vector of the corresponding image, for example, by calculating cosine similarity. Images with a similarity exceeding a preset similarity are identified as candidate images. Similarly, the feature vector of the rewritten question is compared with the image-text feature vector of the corresponding image, for example, by calculating cosine similarity. Images with a similarity exceeding a preset similarity are identified as candidate images.

[0153] The above image and text feature vectors are obtained based on a pre-trained multimodal alignment model. The training of the multimodal alignment model will be introduced here.

[0154] Obtain sample questions, sample images, and text information of the sample images; train a multimodal alignment model based on the sample questions, sample images, and text information of the sample images. The multimodal alignment model is used to align the sample questions and the fused information of the image and text, which is the fused information of the sample images and the text information of the sample images.

[0155] The image-text feature vector corresponding to the sample image is a feature vector of the fused image-text information. Through multimodal alignment, the sample question is matched with the fused image-text information, enabling a better understanding and association between the sample question and the sample image. This overcomes the problem of information silos and greatly improves the system's ability to understand content in complex scenarios. In addition to images, the image library also contains corresponding text information. This text information can be feature information extracted from the image or other existing image description information. After model training, the trained multimodal alignment model is used to extract the image-text feature vectors corresponding to the images in the image library. In other words, the trained multimodal alignment model is used to extract the image-text feature vectors of the images and their corresponding text information in the image library.

[0156] In some embodiments, before training the multimodal alignment model using the aforementioned sample questions, sample images, and text information of the sample images, the model can also be trained using a large number of copyright-free images and their corresponding text information scraped from the internet. First, the collected copyright-free images and their corresponding text information are cleaned, removing duplicate, low-quality, and irrelevant data. The cleaned copyright-free images and their corresponding text information are then used for training. In this stage, the text information corresponding to the copyright-free images is used as the sample questions, the copyright-free images are used as the sample images, and the text information of the sample images is empty. Training is performed using a method similar to that described above.

[0157] After retrieving candidate images, it is necessary to further determine whether they are helpful in answering the user's question. In some embodiments, the candidate images can be initially screened to exclude images that are clearly unsuitable.

[0158] Optionally, when rewriting the original question using a large language model, in addition to generating the rewritten question, entity names can also be extracted from the original and / or rewritten questions using the large language model. That is, inputting the original question into the large language model yields not only the rewritten question but also the entity names from the original and / or rewritten questions. For example, entity names can include names of people, places, organizations, animals, plants, works, and products. After retrieving candidate images, multiple candidate images can be filtered based on the entity names in the original and / or rewritten questions, as well as the feature information of each candidate image. The filtered candidate images and the original question are then input into a multimodal model. When filtering candidate images, if the feature information of any candidate image does not contain an entity name, the candidate image is discarded.

[0159] For example, the original question is "Introduce the film and television works of actor Zhang". The entity name in the original question includes the name Zhang. Among the candidate images retrieved, the feature information of image 1 includes the name Zhang of the person in image 1, while the feature information of image 2 includes the image type as a person image but does not include the specific person's name. Therefore, image 1 will be retained and input into the multimodal model, while image 2 will be discarded.

[0160] After obtaining the value score, quality score, and image time of each candidate image through a multimodal large model, the method of determining the images to be displayed in the search answer based on the value score, quality score, and image time of each candidate image in S203 is explained.

[0161] The above embodiments introduced the inclusion of industry classification and content classification in text understanding information. In addition, text understanding information may also include timeliness information, which characterizes the timeliness of the original question searched by the user. For example, timeliness information can be categorized as updated annually, semi-annually, monthly, daily, fixed information, or ambiguous. For instance, the original question is "Exam dates for all subjects in the National College Entrance Examination," and its timeliness information is categorized as updated annually. For instance, the original question is "The offside rule in football," and its timeliness information is categorized as fixed information. For instance, the original question is "Jewelry suitable for middle-aged people," and its timeliness information is categorized as ambiguous.

[0162] Clearly, for each original question, if the image time of the retrieved candidate image does not match the timeliness information of the original question, then the candidate image is unlikely to truly answer the question. For example, if the original question is "Examination time for each subject in the National College Entrance Examination (NCEE)," and its timeliness information is categorized as updated annually, users often want to know the latest NCEE schedule. If a candidate image displays the NCEE time, but it is from ten years ago, then the candidate image is not very helpful in answering the user's question. Therefore, in this embodiment, the image time of the candidate image and the timeliness information of the original question can be comprehensively considered, combined with value scoring and quality scoring, to determine the total score of each candidate image.

[0163] In one implementation, the timeliness score of each candidate image is determined based on its image time, and the timeliness weight is determined based on the timeliness information; the total score of each candidate image is determined based on its value score, quality score, image time, and preset value weight, preset quality weight, and timeliness weight.

[0164] For example, a candidate image with a timeframe within the current month receives a timeliness score of 5; a candidate image with a timeframe within six months receives a timeliness score of 4; a candidate image with a timeframe within one year receives a timeliness score of 3; a candidate image with a timeframe within two years receives a timeliness score of 2; a candidate image with a timeframe more than two years receives a timeliness score of 0; and a candidate image with a timeframe in the future receives a timeliness score of 7.

[0165] For example, if the timeliness information of the original question is updated annually, semi-annually, or monthly, the corresponding timeliness weight is 1; if the timeliness information of the original question is updated irregularly, the corresponding timeliness weight is 0.3; if the timeliness information of the original question is fixed, the corresponding timeliness weight is 0.2; and if the timeliness information of the original question is unclear, the corresponding timeliness weight is 0.

[0166] The preset value weight and preset quality weight in this application embodiment can be set as needed. Since the value score reflects the contribution of the candidate image to answering the user's question, in some embodiments, the preset value weight can be set higher than the preset quality weight, meaning the value score has a relatively greater weight in the total score calculation. Based on the preset value weight, preset quality weight, and timeliness weight, the value score, quality score, and image score of each candidate image are weighted and summed to determine the total score of each candidate image. The candidate images are then sorted in descending order of their total scores, and the top-ranked candidate images are selected as the images to be displayed in the search results.

[0167] Based on the above embodiments, other truncation rules can also be used to remove candidate images.

[0168] In some implementations, for any candidate image among the candidate images, if the value score of the candidate image is lower than a first preset value score and / or the quality score is lower than a preset quality score, the candidate image is eliminated. This step can be performed before calculating the total score and sorting, or it can be performed after sorting; this application embodiment is not limited in this regard.

[0169] For example, if a candidate image has a value score of 0, it will be removed. Alternatively, if a candidate image has a value score of 1 and the first preset value score is 2, it can also be removed.

[0170] In some implementations, for any candidate image among the candidate images, if the image time and timeliness information of the candidate images do not match, the candidate image is eliminated. This step can be performed before calculating the total score and sorting, or it can be performed after sorting; this application embodiment is not limited in this regard.

[0171] The mismatch between the image time and timeliness information of the candidate images can be configured as needed. For example, if the timeliness information is updated monthly, the matching images should be from within the previous month. If the timeliness information of the original question is updated monthly, only candidate images from the current month and the previous month will be retained, and other candidate images can be discarded.

[0172] After determining the image to be displayed, different display methods can be set for the image. It is understood that in AI search, in addition to the images in the search results described in this embodiment, it is also necessary to generate a text answer to the user's original search question. In this embodiment, a text answer can be generated based on the original question and / or a rewritten question. The generation of the text answer can employ methods from related technologies, and this embodiment is not limited in this regard.

[0173] In one implementation, such as Figure 3As shown, to allow users to intuitively see the answer to their question, the image to be displayed in the search results is placed before the text answer. When browsing, users see the image first, followed by the text answer. This combination of text and images allows users to intuitively understand the answer, resulting in a better search experience. It should be noted that the image in the search results can be one or more images. Figure 3 An image is used for illustration. It is understood that the image to be displayed may also be displayed after or in the middle of the text answer, and this embodiment of the application is not limited in this respect.

[0174] In some implementations, the image to be displayed can be set to be presented in an expanded or collapsed form, which can also be operated by the user in the user interface. Optionally, if the number of images to be displayed with a value score greater than or equal to a second preset value score is less than a preset value, the image to be displayed in the search answer will be collapsed.

[0175] For example, after calculating the total score of candidate images, sorting them, and filtering and eliminating candidate images, there are 5 images to be displayed. However, only one of them has a value score of 4 or 3. In this case, the image to be displayed in the search results will be collapsed. Figure 4 As shown. Users can expand and view the images themselves when needed. This avoids displaying too many images in the search results that are not helpful in answering the question, allowing users to quickly obtain the information they need from the text answer.

[0176] The above embodiments illustrate schemes for acquiring the image to be displayed and combining it with the generated text answer in different forms. That is, acquiring the image to be displayed and generating the text answer are independent processes; they are combined and displayed only after acquisition and text answer generation. In other embodiments, the image to be displayed can be displayed simultaneously during the text answer generation process. The initial search results of the original question and the image to be displayed are input into the generative large-scale model to obtain the search answer for the original question. The image to be displayed is used to be displayed along with the text answer during the generation of the search answer. The initial search results of the original question can be based on the original question and / or a rewritten question. These initial search results may include web search results, etc. The initial search results and the image to be displayed are input together into the generative large-scale model. During the process of generating the text answer based on the initial search results, the generative large-scale model places the image to be displayed in a matching position, so that the image is displayed along with the text answer, thereby displaying the search answer for the original question in a visually appealing format, improving the user experience.

[0177] Figure 5Another schematic diagram of a generative search image matching method provided for an exemplary embodiment of this application.

[0178] Please see Figure 5 The methods may include:

[0179] S501. Input the original question and webpage search results into the generation model to obtain the rewritten question corresponding to the original question, as well as the entity names in the original question and / or the rewritten question. Input the original question and webpage search results into the classification model to obtain the text understanding information of the original question. The text understanding information includes industry classification, content classification and timeliness information.

[0180] S502. Determine whether to trigger image search based on text understanding information. If yes, execute S503; otherwise, end the image matching process.

[0181] S503. Determine the feature vector of the original problem and the feature vector of the rewritten problem. Perform similarity matching between the feature vector of the original problem and the feature vector of the rewritten problem and the image and text feature vectors corresponding to the images in the image library. Select images with similarity exceeding the preset similarity as candidate images.

[0182] S504. For any candidate image among multiple candidate images, if the feature information of the candidate image does not contain the entity names in the original question and / or the rewritten question, then the candidate image shall be removed.

[0183] S505. Input the original problem and multiple candidate images into the multimodal model to obtain the value score, quality score and image time of each candidate image.

[0184] S506. For any candidate image among the candidate images, if the value score of the candidate image is lower than the first preset value score and / or the quality score is lower than the preset quality score, and / or if the image time and timeliness information of the candidate image do not match, the candidate image shall be removed.

[0185] S507. Determine the timeliness score of each candidate image based on its image time, and determine the timeliness weight based on the timeliness information; determine the total score of each candidate image based on its value score, quality score, image time, and preset value weight, preset quality weight, and timeliness weight.

[0186] S508. Sort the candidate images according to their total scores, and select the top-ranked candidate images as the images to be displayed.

[0187] The image search method in this embodiment rewrites the user's original search question using a large language model, extracts entity names, and performs text understanding on the original question. Based on the text understanding results, it determines whether to trigger an image search. If an image search is triggered, candidate images are retrieved based on the original and rewritten questions. These candidate images are then filtered based on entity names. A multimodal model is used to evaluate the candidate images, considering whether they are helpful in answering the user's question. The multimodal model outputs a value score, a quality score, and time information for each candidate image. By utilizing the matching of the candidate images to the value, quality, and timeliness of answering the question, images that can answer the user's question and are of good quality are selected from the candidate images. This ensures that the images in the search results can truly help the user understand the question and improve the user experience.

[0188] The above embodiments illustrate how to retrieve candidate images, how to further determine the image to be displayed in the search answer based on the candidate images, and how to display the image to be displayed when an image search is triggered. However, it should also be noted that even when an image search is triggered, there may be cases where no matching image is found, that is, the result of determining candidate images is empty or the result of determining the image to be displayed is empty. In this case, the text answer of the original question is input into the image generation model to obtain the image generated by the image generation model; the image generated by the image generation model is then added to the search answer.

[0189] For example, the image search might be triggered based on the textual understanding of the original question, or based on user actions. In such cases, if no candidate images are subsequently retrieved, or if candidate images are retrieved but then removed in subsequent processing steps, the final result for the image to be displayed might be empty—meaning there are no images to display in the search answer. To improve the user experience, this embodiment utilizes an image generation model to generate images based on the textual answer to the original question. For instance, the image generation model could be a pre-trained text-to-image model. When the textual answer to the original question is input into the model, a description of the desired image, such as image style and color, can be combined. The textual answer and the description of the desired image are then used to create prompts, which are then input into the text-to-image model to obtain the image generated by the model. This satisfies the image matching requirements while ensuring that the generated image is helpful in answering the user's question.

[0190] Optionally, images generated by the image generation model can be added before the text answer in the search results, allowing users to see the answer to the question more intuitively and improving the user experience.

[0191] Figure 6This is a schematic diagram of a generative search mapping device provided for an exemplary embodiment of this application. Please refer to... Figure 6 The generative search mapping device 600 includes:

[0192] The retrieval module 601 is used to determine multiple candidate images based on the user's original search question;

[0193] Evaluation module 602 is used to input the original question and multiple candidate images into a multimodal model to obtain the value score, quality score and image time of each candidate image. The value score represents the contribution of the candidate image to answering the original question, and the quality score represents the image quality of the candidate image under the premise of answering the original question.

[0194] The determination module 603 is used to determine the images to be displayed in the search answer based on the value score, quality score and image time of each candidate image.

[0195] In some implementations, the retrieval module 601 is used for:

[0196] Input the original question into the large language model to obtain the rewritten question corresponding to the original question and the text understanding information of the original question.

[0197] Once the text understanding information determines that an image search is triggered, multiple candidate images are retrieved from the image library based on the original question and the rewritten question.

[0198] In some implementations, the large language model includes a generative model and a classification model, and the retrieval module 601 is used for:

[0199] Input the original problem into the generative model to obtain the rewritten problem corresponding to the original problem;

[0200] The original question is input into the classification model to obtain text understanding information.

[0201] In some implementations, text understanding information includes industry classification, content classification, and trigger intensity, where trigger intensity is used to characterize the strength level of the original question triggering the image search.

[0202] The retrieval module 601 is used for:

[0203] If the industry category and / or content category are categories that trigger image search, the trigger strength threshold shall be determined based on the industry category and / or content category.

[0204] If the trigger strength is greater than or equal to the trigger strength threshold, determine the trigger image search.

[0205] In some implementations, the retrieval module 601 is also used for:

[0206] Input the original question into the large language model to obtain the entity names in the original question and / or the rewritten question;

[0207] Evaluation module 602 is used for:

[0208] Multiple candidate images are filtered based on entity names and their respective feature information. The filtered candidate images and the original question are then input into a multimodal model.

[0209] In some implementations, the evaluation module 602 is used for:

[0210] If the feature information of any candidate image does not contain an entity name, then the candidate image is removed.

[0211] In some implementations, module 603 is used for:

[0212] The total score for each candidate image is determined based on its value score, quality score, and image time.

[0213] The candidate images are ranked based on their total scores, and a predetermined number of the top-ranked candidate images are selected as the images to be displayed.

[0214] In some implementations, the text understanding information includes timeliness information; the determination module 603 is used for:

[0215] The timeliness score of each candidate image is determined based on its image time, and the timeliness weight is determined based on the timeliness information.

[0216] The total score for each candidate image is determined based on its value score, quality score, image time, and preset value weight, preset quality weight, and timeliness weight.

[0217] In some implementations, the determining module 603 is also used for:

[0218] For any candidate image among all candidate images, if the value score of the candidate image is lower than the first preset value score and / or the quality score is lower than the preset quality score, the candidate image will be removed.

[0219] In some implementations, the text understanding information includes timeliness information; the determination module 603 is also used for:

[0220] For any candidate image among all candidate images, if the image time and timeliness information of the candidate image do not match, the candidate image will be removed.

[0221] In some implementations, the retrieval module 601 is used for:

[0222] The feature vectors of the original problem and the rewritten problem are determined. The feature vectors of the original problem and the rewritten problem are then matched with the image and text feature vectors corresponding to the images in the image library for similarity. Images with similarity exceeding a preset similarity are identified as candidate images. The image and text feature vector is a fusion feature vector of the image and the corresponding text information, and it is obtained based on a pre-trained multimodal alignment model.

[0223] Some implementations also include: a training module, used for:

[0224] Obtain the sample question, sample image, and text information of the sample image;

[0225] The multimodal alignment model is trained based on sample questions, sample images, and text information of sample images. The multimodal alignment model is used to align the sample questions and the fused information of images and text. The fused information of images and text is the fused information of sample images and text information of sample images.

[0226] The trained multimodal alignment model is used to extract the image and text feature vectors corresponding to the images in the image library.

[0227] In some implementations, a first processing module is also included, used for

[0228] Generate text answers based on the original question and / or a rewritten question;

[0229] In the search results, place the image to be displayed before the text answer.

[0230] In some implementations, a second processing module is also included, used for:

[0231] If the number of images with a value score greater than or equal to the second preset value score is less than a preset value, the images to be displayed in the search results will be collapsed.

[0232] Some implementations also include a third processing module, used for:

[0233] If the number of images with a value score greater than or equal to the second preset value score is less than a preset value, the images to be displayed in the search results will be collapsed.

[0234] If it is determined that an image search is triggered, but the result of the candidate image is empty or the result of the image to be displayed is empty, the text answer of the original question is input into the image generation model to obtain the image generated by the image generation model.

[0235] Add images generated by the image generation model to the search results.

[0236] In some implementations, a fourth processing module is also included, used for:

[0237] The initial search results for the original question and the image to be displayed are input into a generative large model to obtain the search answer for the original question. The image to be displayed is used to be displayed along with the text answer during the generation of the text answer for the search answer of the original question.

[0238] In some implementations, the retrieval module is used for:

[0239] If the original question is not in text form, the original question will be converted into text.

[0240] Based on the original question after textualization, multiple candidate images are identified.

[0241] The generative search mapping device provided in this application embodiment can execute the technical solution shown in the above method embodiment. Its implementation principle and beneficial effects are similar, and will not be described again here.

[0242] Figure 7 This is a schematic diagram of the structure of an electronic device provided for an exemplary embodiment of this application. Please refer to... Figure 7 The electronic device 700 may include at least one processor 701 for implementing the generative search mapping method provided in the embodiments of this application.

[0243] Optionally, the electronic device 700 further includes at least one memory 702 for storing program instructions and / or data. The memory 702 is coupled to the processor 701. The coupling in this embodiment is an indirect coupling or communication connection between devices, units, or modules, and can be electrical, mechanical, or other forms, used for information exchange between devices, units, or modules. The processor 701 may operate in conjunction with the memory 702. The processor 701 may execute program instructions stored in the memory 702. At least one of the at least one memory may be included in the processor.

[0244] Optionally, the electronic device 700 further includes a communication interface 703 for communicating with other devices via a transmission medium, thereby enabling the electronic device 700 to communicate with other devices. The communication interface 703 may be, for example, a transceiver, interface, bus, circuit, or a device capable of transmitting and receiving functions. The processor 701 can utilize the communication interface 703 to transmit and receive data and / or information, and to implement the methods provided in the embodiments of this application. For details, please refer to the detailed descriptions in the preceding embodiments; further elaboration is not repeated here.

[0245] This application embodiment does not limit the specific connection medium between the processor 701, memory 702, and communication interface 703. This application embodiment... Figure 7 The processor 701, memory 702, and communication interface 703 are connected via bus 704. Bus 704 is... Figure 7 The connections between other components are shown in thick lines only and are not intended to be limiting. This bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, Figure 7 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0246] It should be understood that the processor in the embodiments of this application can be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above method embodiments can be completed by the integrated logic circuits in the processor's hardware or by instructions in software form. The processor can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software modules can be located in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. This storage medium is located in memory, and the processor reads the information in the memory and, in conjunction with its hardware, completes the steps of the above method.

[0247] It should also be understood that the memory in the embodiments of this application can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous linked dynamic random access memory (SLDRAM), and direct rambus RAM (DR RAM). It should be noted that the memory used in the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.

[0248] Accordingly, this application provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the method described in the above-described method embodiments.

[0249] Accordingly, embodiments of this application may also provide a computer program product, including a computer program, which, when executed by a processor, can implement the methods shown in the above-described method embodiments.

[0250] The terms “unit”, “module”, etc., used in this specification may be used to refer to computer-related entities, hardware, firmware, combinations of hardware and software, software, or software in execution.

[0251] Those skilled in the art will recognize that the various illustrative logical blocks and steps described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application. In the several embodiments provided in this application, it should be understood that the disclosed apparatus, devices, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for example, the division of units is merely a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the shown or discussed mutual couplings or direct couplings or communication connections may be through some interfaces; indirect couplings or communication connections between devices or units may be electrical, mechanical, or other forms.

[0252] The unit described as a separate component may or may not be physically separate. The component shown as a unit may or may not be a physical unit; that is, it may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0253] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0254] In the above embodiments, the functions of each functional unit can be implemented entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. This computer program product includes one or more computer instructions (programs). When the computer program instructions (programs) are loaded and executed on a computer, all or part of the flow or function according to the embodiments of this application is generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., digital video discs (DVDs)), or semiconductor media (e.g., solid-state drives (SSDs)).

[0255] If this function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, ROM, RAM, magnetic disks, or optical disks.

[0256] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.

[0257] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A generative search method for image matching, characterized in that, include: Multiple candidate images were identified based on the user's original search question; The original question and the multiple candidate images are input into a multimodal model to obtain the value score, quality score, and image time of each candidate image. The value score represents the contribution of the candidate image to answering the original question, and the quality score represents the image quality of the candidate image when used to answer the original question. The search answer to the original question is determined based on the value score, quality score, and image time of each candidate image. The search answer to the original question includes the image to be displayed and the text answer to the original question.

2. The method according to claim 1, characterized in that, The original question based on the user's search determines multiple candidate images, including: The original question is input into a large language model to obtain the rewriting question corresponding to the original question and the text understanding information of the original question. If the text understanding information determines that an image search is triggered, the multiple candidate images are retrieved from the image library based on the original question and the rewritten question.

3. The method according to claim 2, characterized in that, The large language model includes a generative model and a classification model. The process of inputting the original question into the large language model yields a rewrite question corresponding to the original question and text understanding information of the original question, including: The original problem is input into the generative model to obtain the rewritten problem corresponding to the original problem; The original question is input into the classification model to obtain the text understanding information.

4. The method according to claim 3, characterized in that, The text understanding information includes industry classification, content classification, and trigger intensity, wherein the trigger intensity is used to characterize the intensity level of the original question triggering the image search; The step of determining the trigger for image search based on the text understanding information includes: If the industry category and / or the content category belong to the category that triggers image search, a trigger strength threshold shall be determined based on the industry category and / or the content category; If the trigger intensity is greater than or equal to the trigger intensity threshold, a trigger search is determined.

5. The method according to any one of claims 2-4, characterized in that, Also includes: Input the original question into the large language model to obtain the entity names in the original question and / or the rewritten question; The step of inputting the original problem and the multiple candidate images into the multimodal model includes: The candidate images are filtered based on the entity name and the feature information of each candidate image, and the filtered candidate images and the original question are input into the multimodal model.

6. The method according to claim 5, characterized in that, The step of filtering the multiple candidate images based on the entity name and the feature information of each candidate image includes: If the feature information of any candidate image among the plurality of candidate images does not contain the entity name, then the candidate image is removed.

7. The method according to any one of claims 2-4, characterized in that, The process of determining the images to be displayed in the search answer based on the value score, quality score, and image time of each candidate image includes: The total score for each candidate image is determined based on its value score, quality score, and image time. The candidate images are ranked based on their total scores, and a predetermined number of the top-ranked candidate images are selected as the images to be displayed.

8. The method according to claim 7, characterized in that, The text understanding information includes timeliness information; The process of determining the total score for each candidate image based on its value score, quality score, and image time includes: The timeliness score of each candidate image is determined based on the image time of each candidate image, and the timeliness weight is determined based on the timeliness information. The total score for each candidate image is determined based on its value score, quality score, image time, preset value weight, preset quality weight, and timeliness weight.

9. The method according to any one of claims 1-4, characterized in that, Also includes: For any candidate image among all candidate images, if the value score of the candidate image is lower than the first preset value score and / or the quality score is lower than the preset quality score, the candidate image is removed.

10. The method according to any one of claims 2-4, characterized in that, The text understanding information includes timeliness information; the method further includes: If the image time and the timeliness information of any candidate image do not match, the candidate image will be removed.

11. The method according to any one of claims 2-4, characterized in that, The process of retrieving multiple candidate images from the image library based on the original question and the rewritten question includes: The feature vectors of the original problem and the rewritten problem are determined. The feature vectors of the original problem and the rewritten problem are then matched with the image-text feature vectors corresponding to the images in the image library for similarity. Images with similarity exceeding a preset similarity are identified as candidate images. The image-text feature vector is a fusion feature vector of the image and the corresponding text information, and it is obtained based on a pre-trained multimodal alignment model.

12. The method according to claim 11, characterized in that, Also includes: Obtain the sample question, sample image, and text information of the sample image; A multimodal alignment model is trained based on the sample question, the sample image, and the text information of the sample image. The multimodal alignment model is used to align the sample question and the fusion information of the image and text, where the fusion information of the image and text is the fusion information of the sample image and the text information of the sample image. The trained multimodal alignment model is used to extract the image and text feature vectors corresponding to the images in the image library.

13. The method according to any one of claims 2-4, characterized in that, Also includes: Generate a text answer based on the original question and / or the rewritten question; The image to be displayed is placed before the text answer in the search results.

14. The method according to any one of claims 1-4, characterized in that, Also includes: If the number of images with a value score greater than or equal to the second preset value score is less than a preset value, the images to be displayed in the search answer will be folded.

15. The method according to any one of claims 2-4, characterized in that, Also includes: If it is determined that an image search is triggered, but the result of the candidate image is empty or the result of the image to be displayed is empty, the text answer of the original question is input into the image generation model to obtain the image generated by the image generation model. Add the image generated by the image generation model to the search results.

16. The method according to any one of claims 1-4, characterized in that, Also includes: The initial search results for the original question and the image to be displayed are input into the generative large model to obtain the search answer for the original question. The image to be displayed is used to be displayed along with the text answer during the generation of the text answer in the search answer.

17. The method according to any one of claims 1-4, characterized in that, The original question based on the user's search determines multiple candidate images, including: If the original question is not in text form, the original question will be converted into text. Based on the original question after textualization, multiple candidate images are identified.

18. An electronic device, characterized in that, include: At least one processor; as well as A memory that is communicatively connected to the at least one processor; The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, cause the electronic device to perform the method according to any one of claims 1-17.

19. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, implement the method as described in any one of claims 1-17.

20. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1-17.

Citation Information

Patent Citations

  • Automatic determination method and system for selected pictures

    CN111291829A

  • News illustration method and device, equipment and storage medium

    CN113343012A

  • Image display method and device, computer equipment and storage medium

    CN116150421A

  • Streaming question and answer illustration method and system

    CN118035416A

  • Method and system for enhancing RAG questions and answers through mixed retrieval method

    CN118627625A