Image acquisition method and electronic equipment

By generating key information in image search and analyzing and adjusting the searched images, the problem of matching error between text vectors and image vectors is solved, and the accuracy of image search is improved.

CN120144803APending Publication Date: 2025-06-13LENOVO (BEIJING) LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510364886.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

When searching for images through text content, there is an error in the matching between text vectors and image vectors in the prior art, resulting in low accuracy of the image.

Method used

By obtaining text requests, generating key information, and analyzing and adjusting the searched images, ensuring that the searched images contain key information, thereby improving the accuracy of the image.

Benefits of technology

Through the analysis and adjustment of key information, the accuracy of image search is improved to ensure that the finalized image matches the text request.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144803A_ABST
    Figure CN120144803A_ABST
Patent Text Reader

Abstract

The invention discloses an image acquisition method and electronic equipment, and the method comprises the steps: obtaining a text request which is used for requesting to obtain an image; obtaining a target number of retrieval images based on the text request; generating key information based on the text request, wherein the key information is key information included in a to-be-acquired image specified by the text request; the target number of retrieval images are analyzed, an analysis result is obtained, and the analysis result represents whether each retrieval image comprises key information or not; and adjusting the target number of retrieval images based on the analysis result to obtain an adjusted target image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of model technology, and in particular, to an image acquisition method and an electronic device. Background Art

[0002] When searching for images through text content, usually the text content is converted into a text vector, and the matching degree between the text vector and the image vector of the image is determined, so as to determine the image corresponding to the text content.

[0003] However, when searching for images in this way, the text content may specify a certain target for description, while the image vector of the image is usually a general description, rather than a description of the specified target, which will lead to errors in similarity calculation, thus affecting the accuracy of the finally determined image. Summary of the Invention

[0004] In view of this, this application provides an image acquisition method and an electronic device, and the specific scheme is as follows:

[0005] An image acquisition method, comprising:

[0006] Obtaining a text request for requesting to acquire an image;

[0007] Obtaining a target number of retrieved images based on the text request;

[0008] Generating key information based on the text request, where the key information is the key information included in the image to be acquired specified by the text request;

[0009] Analyzing the target number of retrieved images to obtain an analysis result, where the analysis result represents whether each of the retrieved images includes the key information;

[0010] Adjusting the target number of retrieved images based on the analysis result to obtain an adjusted target image.

[0011] Further, the adjusting the target number of retrieved images based on the analysis result to obtain an adjusted target image includes:

[0012] Reordering the target number of retrieved images based on the analysis result, and determining the reordered images as the target images;

[0013] Or,

[0014] If the analysis result indicates that there are retrieved images among the target number of retrieved images that do not include the key information, delete the retrieved images that do not include the key information from the target number of retrieved images, and determine the remaining retrieved images after deleting the retrieved images that do not include the key information from the target number of retrieved images as the target images.

[0015] Further, obtaining the target number of retrieved images based on the text request includes:

[0016] Determine the first text embedding vector of the text request;

[0017] Match the first text embedding vector with a second text embedding vector to obtain a matching result, so as to obtain the target number of retrieved images based on the matching result;

[0018] Wherein, the second text embedding vector is a vector corresponding to the text description information that pre-stores the text description of the image.

[0019] Further, it also includes:

[0020] Display the adjusted target image;

[0021] Obtain the correction information input by the user based on the displayed adjusted target image, and the correction information is used to correct the corresponding relationship between the first target image in the adjusted target image and the text request.

[0022] Further, it also includes:

[0023] Determine the text description information of the first target image based on the correction information;

[0024] Obtain the vector corresponding to the text description information of the first target image, and store the vector corresponding to the text description information of the first target image as the second text embedding vector corresponding to the first target image, and the second text embedding vector corresponding to the first target image does not match the first text embedding vector.

[0025] Further, determining the text description information of the first target image based on the correction information includes:

[0026] Obtain the first text description information corresponding to the part to be corrected in the first target image specified by the correction information;

[0027] Obtain the second text description information corresponding to the part of the first target image except the part to be corrected in the first analysis result corresponding to the first target image, and the first analysis result is the analysis result obtained when the first target image is analyzed as a retrieved image;

[0028] Determine the text description information of the first target image based on the first text description information and the second text description information.

[0029] Further, the matching of the first text embedding vector and the second text embedding vector to obtain a matching result, so as to obtain a target number of retrieved images based on the matching result, includes:

[0030] Match the first text embedding vector with the second text embedding vector to obtain a first matching result;

[0031] Match the first text embedding vector with an image embedding vector to obtain a second matching result, where the image embedding vector is the embedding vector of an image that can be obtained;

[0032] Obtain a target number of retrieved images based on the first matching result and the second matching result.

[0033] Further, the obtaining a target number of retrieved images based on the first matching result and the second matching result includes:

[0034] If it is determined that the first matching result meets the matching condition, determine the image corresponding to the first matching result as the retrieved image.

[0035] Further, the obtaining a target number of retrieved images based on the first matching result and the second matching result includes:

[0036] If it is determined that the second matching result meets the matching condition and the first matching result does not meet the matching condition, and if it is determined that the text request meets the condition, determine the image corresponding to the second matching result as the retrieved image;

[0037] Wherein, the text request meeting the condition at least includes: the text description information determined based on the correction information input by the user in the historical record is not included in the text request.

[0038] An image acquisition system, comprising:

[0039] A first acquisition unit, configured to acquire a text request for requesting to acquire an image;

[0040] A second acquisition unit, configured to acquire a target number of retrieved images based on the text request;

[0041] A generation unit, configured to generate key information based on the text request, where the key information is the key information included in the image to be acquired specified by the text request;

[0042] An analysis unit for analyzing the target number of retrieved images to obtain an analysis result, where the analysis result characterizes whether each of the retrieved images includes the key information;

[0043] An adjustment unit for adjusting the target number of retrieved images based on the analysis result to obtain adjusted target images.

[0044] An electronic device, comprising:

[0045] A processor for obtaining a text request for requesting to obtain an image; obtaining a target number of retrieved images based on the text request; generating key information based on the text request, where the key information is the key information included in the image to be obtained specified by the text request; analyzing the target number of retrieved images to obtain an analysis result, where the analysis result characterizes whether each of the retrieved images includes the key information; adjusting the target number of retrieved images based on the analysis result to obtain adjusted target images;

[0046] A memory for storing a program required for the processor to execute the above processing procedures.

[0047] A computer storage medium carrying one or more computer programs, which can enable the electronic device to implement the image acquisition method described in any one of the above when the one or more computer programs are executed by the electronic device. Description of the Drawings

[0048] To more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments or related technologies. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0049] Figure 1 It is a flowchart of an image acquisition method disclosed in an embodiment of the present application;

[0050] Figure 2 It is a flowchart of an image acquisition method disclosed in an embodiment of the present application;

[0051] Figure 3 It is a schematic diagram of an image acquisition method disclosed in an embodiment of the present application for performing two types of matching simultaneously to obtain retrieved images;

[0052] Figure 4 It is a flowchart of an image acquisition method disclosed in an embodiment of the present application;

[0053] Figure 5 Schematic diagrams of a first target image and its text description information before and after correction disclosed in an embodiment of the present application;

[0054] Figure 6 Schematic diagram of the structure of an image acquisition system disclosed in an embodiment of the present application;

[0055] Figure 7 Schematic diagram of the structure of an electronic device disclosed in an embodiment of the present application. Detailed implementation manners

[0056] The embodiments of the present application will be described below with reference to the accompanying drawings in the embodiments of the present application. The terms used in the implementation part of the present application are only used to explain the specific embodiments of the present application, rather than intended to limit the present application.

[0057] The embodiments of the present application will be described below with reference to the accompanying drawings. Those of ordinary skill in the art will know that with the development of technology and the emergence of new scenarios, the technical solutions provided in the embodiments of the present application are equally applicable to similar technical problems.

[0058] The terms "first", "second", etc. in the specification, claims and above-mentioned drawings of the present application are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances, which is only a way of distinguishing when describing objects with the same attributes in the embodiments of the present application. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, so that a process, method, system, product or device including a series of units does not have to be limited to those units, but may include other units not clearly listed or inherent to these processes, methods, products or devices.

[0059] The present application discloses an image acquisition method, and its flowchart is as Figure 1 shown, including:

[0060] Step S11: Obtain a text request for requesting to obtain an image;

[0061] Step S12: Obtain a target number of retrieved images based on the text request;

[0062] Step S13: Generate key information based on the text request, where the key information is the key information included in the image to be obtained specified by the text request;

[0063] Step S14: Analyze the target number of retrieved images to obtain an analysis result, where the analysis result indicates whether each retrieved image includes the key information;

[0064] Step S15: Adjust the target number of retrieved images based on the analysis result to obtain the adjusted target images.

[0065] When searching for images through text content, usually the text content is converted into a text vector, and the matching degree between the text vector and the image vector of the image is determined, so as to determine the image corresponding to the text content.

[0066] However, when searching for images in this way, a certain target may be specified in the text content for description, while the image vector of the image is usually a general description, rather than a description of the specified target. This will lead to errors in similarity calculation, thus affecting the accuracy of the finally determined image.

[0067] Based on this, in this solution, while obtaining the target number of retrieved images using the text request, key information is also generated using the text request, and the key information is used to adjust the target number of retrieved images to obtain the target images that conform to the key information, so as to ensure the accuracy of the finally determined image.

[0068] Obtain a text request, which is used to request to obtain an image, that is: search for an image by text, such as: the technology of searching for images by text. The obtained text request, such as: "Help me find a picture with two men and one woman in it."

[0069] After obtaining the text request, the target number of retrieved images can be obtained based on the text request, such as: using the technology of searching for images by text to obtain the target number of retrieved images corresponding to the text request.

[0070] Obtaining the target number of retrieved images based on the text request can be: converting the text request into a text vector, and determining the matching degree between the text vector and the image vector of the image, so as to determine the target number of retrieved images that match the text request.

[0071] In addition, after obtaining the text request, key information also needs to be generated based on the text request. The key information is the key information included in the image to be obtained specified by the text request. Since the text request is a text-based request for obtaining an image, therefore, by analyzing the text request, it can be determined what information the image to be obtained needs to include, and these information are determined as the key information.

[0072] For example: the text request is: "Help me find a picture with two men and one woman in it.", and the key information generated based on the text request can be: two men and one woman, three people.

[0073] After obtaining the retrieved images and key information, in order to verify the retrieved images, the retrieved images can be verified based on the key information, that is, the key information is used to analyze a target number of retrieved images to determine whether each retrieved image includes the key information. Each of the target number of retrieved images can be analyzed to determine whether each retrieved image includes the key information. Specifically, it can be: generating a verification request based on the key information and verifying the target number of retrieved images based on the verification request.

[0074] For example: the text request is: "Help me find a picture with two men and one woman", and the key information is: "two men and one woman, three people". It can directly determine whether each retrieved image is "two men and one woman, three people" based on the key information; or, the verification request generated based on the key information is: "Does the picture contain three people, and the three people are two men and one woman", and each retrieved image is verified based on this verification request.

[0075] After determining the analysis result of whether each retrieved image includes the key information, the target number of retrieved images is adjusted based on this analysis result to obtain the adjusted target images.

[0076] After obtaining the adjusted target images, the adjusted target images can be output to implement the response to the text request. Among them, the target number of retrieved images is adjusted based on the analysis result to obtain the adjusted target images, so that the adjusted target images are more in line with the text request and improve the accuracy of the finally determined images.

[0077] Among them, the adjustment of the retrieved images can be specifically:

[0078] Based on the analysis result, the target number of retrieved images is re-sorted, and the re-sorted images are determined as the target images; or, if the analysis result indicates that there are retrieved images among the target number of retrieved images that do not include the key information, the retrieved images that do not include the key information are deleted from the target number of retrieved images, and the remaining retrieved images after deleting the retrieved images that do not include the key information from the target number of retrieved images are determined as the target images.

[0079] Based on the analysis result, re-sorting the target number of retrieved images can be specifically: analyzing each of the target number of retrieved images to determine the matching degree of each retrieved image with the key information. The higher the matching degree, the higher the ranking, and the lower the matching degree, the lower the ranking.

[0080] Specifically, it can be: determine whether each retrieved image includes key information. For the retrieved images that include key information, move their sorting forward; for the retrieved images that do not include key information, move their sorting backward; for the retrieved images that include part of the key information, move their sorting to the middle position. Then, in the target images after re-sorting, the more key information an image includes, the higher its sorting position; the less key information an image includes, the lower its sorting position.

[0081] For example, the key information is: "two men and one woman, three people". The retrieved images determined based on the text request are 3, namely retrieved image 1, retrieved image 2, and retrieved image 3. Among them, retrieved image 1 includes 3 people, retrieved image 2 includes 4 people, among which there are three men and one woman, and retrieved image 3 includes 3 people, two men and one woman. Then, retrieved image 3 has the highest degree of matching with the key information and includes all the key information. The degrees of matching of retrieved image 2 and retrieved image 1 with the key information are lower than that of retrieved image 3, and they only include part of the key information. After re-sorting retrieved image 1, retrieved image 2, and retrieved image 3, their order is: retrieved image 3 → retrieved image 2 → retrieved image 1.

[0082] Or, it can also be:

[0083] Analyze each retrieved image among the target number of retrieved images to determine the degree of matching of each retrieved image with the key information. For the retrieved images with a matching degree lower than a certain set threshold, directly delete them and do not regard them as one of the target images. For the retrieved images with a matching degree not lower than this set threshold, they are determined as target images. Or, sort the retrieved images with a matching degree not lower than this set threshold according to the matching degree, and use these sorted images as target images.

[0084] Specifically, it can be: determine whether each retrieved image includes key information. For the retrieved images that do not include key information, directly delete them, and only retain the retrieved images that include key information. Use the retained retrieved images that include key information as target images.

[0085] The image acquisition method disclosed in this embodiment obtains a text request for requesting to acquire an image; obtains a target number of retrieved images based on the text request; generates key information based on the text request, where the key information is the key information included in the image to be acquired specified by the text request; analyzes the target number of retrieved images to obtain an analysis result, where the analysis result indicates whether each retrieved image includes the key information; and adjusts the target number of retrieved images based on the analysis result to obtain an adjusted target image. After obtaining the text request, this solution generates key information while obtaining retrieved images based on the text request, and uses the key information to analyze the retrieved images to obtain an analysis result, so as to adjust the retrieved images based on the analysis result to obtain a target image, ensuring that after obtaining the retrieved images based on the text request, the retrieved images can be adjusted based on the key information, so that the finally obtained target image meets the actual requirements and ensures the accuracy of the finally determined image.

[0086] This embodiment discloses an image acquisition method, and its flowchart is as Figure 2 shown, including:

[0087] Step S21: Obtain a text request for requesting to acquire an image;

[0088] Step S22: Determine the first text embedding vector of the text request;

[0089] Step S23: Match the first text embedding vector with the second text embedding vector to obtain a matching result, and obtain a target number of retrieved images based on the matching result. The second text embedding vector is a vector corresponding to the text description information that pre-stores the text description of the image;

[0090] Step S24: Generate key information based on the text request, where the key information is the key information included in the image to be acquired specified by the text request;

[0091] Step S25: Analyze the target number of retrieved images to obtain an analysis result, where the analysis result indicates whether each retrieved image includes the key information;

[0092] Step S26: Adjust the target number of retrieved images based on the analysis result to obtain an adjusted target image.

[0093] After obtaining the text request, a target number of retrieved images are obtained based on the text request. At the same time, key information is generated based on the text request, and the key information is used to analyze the target number of retrieved images to obtain an analysis result to determine whether each retrieved image includes the key information, and the target number of retrieved images is adjusted based on this analysis result to obtain a target image.

[0094] Among them, obtaining a target number of retrieved images based on a text request may be: determining a first text embedding vector corresponding to the text request, matching the first text embedding vector with an image embedding vector, and determining the retrieved image corresponding to the text request based on the matching result, where the image embedding vector is the embedding vector of the image that can be obtained.

[0095] In the image acquisition method disclosed in this embodiment, to obtain a target number of retrieved images, the method of using image embedding vectors is not adopted, but the method of matching text embedding vectors with text embedding vectors is used to determine.

[0096] Specifically, for the pre-stored images, or the images that can be obtained, the second text embedding vector of each image can be determined, that is, each pre-stored image, or each image that can be obtained is textually described to obtain text description information, and the text description information is converted into a vector to obtain the second text embedding vector. Each pre-stored image, or each image that can be obtained, has its corresponding second text embedding vector. The second text embedding vectors of each pre-stored image, or each image that can be obtained, are stored, and the association relationship between the second text embedding vector and its corresponding image is established.

[0097] When a text request is obtained and it is necessary to obtain a target number of retrieved images based on the text request, the text request can be first converted to obtain a first text embedding vector, the first text embedding vector is matched with each pre-stored second text embedding vector to obtain a matching result, the second text embedding vector that meets the matching requirements is determined based on the matching result, and then the retrieved image is determined from the pre-stored images, or the images that can be obtained, based on the association relationship between the second text embedding vector and the image.

[0098] For example, images 1, 2, 3, and 4 are pre-stored. Among them, the vector corresponding to the text description information of image 1 is the second text embedding vector 1, the vector corresponding to the text description information of image 2 is the second text embedding vector 2, the vector corresponding to the text description information of image 3 is the second text embedding vector 3, and the vector corresponding to the text description information of image 4 is the second text embedding vector 4. Establish an association relationship between the second text embedding vector and the corresponding image and store it. After obtaining a text request, determine the first text embedding vector corresponding to the text request, and match the first text embedding vector with each of the stored second text embedding vectors. If it is determined that the matching result between the first text embedding vector and the second text embedding vector 2 meets the matching requirement, then image 2 corresponding to the second text embedding vector 2 can be determined as one retrieved image based on the association relationship; if it is determined that the matching result between the first text embedding vector and the second text embedding vector 3 meets the matching requirement, then image 3 corresponding to the second text embedding vector 3 can be determined as one retrieved image based on the association relationship; if it is determined that the matching result between the first text embedding vector and the second text embedding vector 1 does not meet the matching requirement, then ignore image 1 and continue to determine whether other images can be determined as retrieved images.

[0099] In the image acquisition method disclosed in this embodiment, when obtaining a target number of retrieved images based on a text request, the first text embedding vector corresponding to the text request can be matched with the second text embedding vector to obtain a matching result, so as to determine a target number of retrieved images based on the matching result. Among them, the second text embedding vector is a vector corresponding to the text description information that pre-stores the text description of the image. In this solution, the vector corresponding to the text description information of the image is pre-stored and determined as the second text embedding vector. When a text request is obtained, the first text embedding vector corresponding to the text request is matched with the pre-stored second text embedding vector, so as to determine the retrieved image that can match the first text embedding vector. The vector of the text description information of the image is pre-stored, so that when there is a text request, the retrieved image can be accurately determined through the matching of the text embedding vector and the text embedding vector, improving the accuracy of determining the retrieved image. On the basis of ensuring the accuracy of the determined retrieved image, the accuracy of the target image determined based on the retrieved image is ensured.

[0100] Furthermore, the image acquisition method disclosed in this embodiment can be: while using the text embedding vector to match with the text embedding vector, the text embedding vector is also used to match with the image embedding vector. Through the matching of the text embedding vector and the text embedding vector, as well as the matching of the text embedding vector and the image embedding vector, the retrieved image of the target data is obtained.

[0101] Specifically, it can be:

[0102] Match the first text embedding vector with the second text embedding vector to obtain a first matching result; match the first text embedding vector with an image embedding vector to obtain a second matching result, where the image embedding vector is the embedding vector of an image that can be obtained; obtain a target number of retrieved images based on the first matching result and the second matching result.

[0103] Predetermine the image embedding vector of the image that can be obtained, and pre-store the text embedding vector of the image that can be obtained, that is, the second text embedding vector.

[0104] When a text request is obtained, determine the first text embedding vector corresponding to the text request. While matching the first text embedding vector with the second text embedding vector, match the first text embedding vector with the image embedding vector, so as to obtain a target number of retrieved images based on the matching results of these two matching methods.

[0105] Specifically, as Figure 3 shown, it is a schematic diagram of simultaneously performing two matches in the image acquisition method disclosed in this embodiment to obtain retrieved images. That is, first obtain a text request, and the text request can be the text content of the request to obtain an image input by the user; then convert the text request to obtain the first text embedding vector corresponding to the text request; there is a pre-stored association relationship between the first set of images and the image embedding vector corresponding to each image in the first set of images. In addition, there is a pre-stored association relationship between the second set of images and the second text embedding vector corresponding to each image in the second set of images. Match the first text embedding vector with each image embedding vector to obtain a second matching result. At the same time, match the first text embedding vector with each second text embedding vector to obtain a first matching result; based on the first matching result and the second matching result, the images in the first set of images and the images in the second set of images corresponding to the first text embedding vector can be determined, and a target number of retrieved images can be obtained based on the determined images in the first set of images and the images in the second set of images.

[0106] Among them, at least some of the images in the first set of images may overlap with the images in the second set of images. At this time, the images in the first set of images and the images in the second set of images determined based on the first matching result and the second matching result may have the same part; it is also possible that any image in the first set of images is different from the images in the second set of images.

[0107] Among them, matching the first text embedding vector with the image embedding vector can actually be calculating the similarity between the first text embedding vector and the image embedding vector, and matching the first text embedding vector with the second text embedding vector can also be calculating the similarity between the first text embedding vector and the second text embedding vector.

[0108] In this embodiment, the first text embedding vector is matched with the image embedding vector to obtain a second matching result, and at the same time, it is matched with the second text embedding vector to obtain a first matching result, and the retrieved image is jointly determined based on the two matching results, which improves the matching degree between the determined retrieved image and the text request.

[0109] Furthermore, in this embodiment, obtaining the target number of retrieved images based on the first matching result and the second matching result can be specifically:

[0110] If it is determined that the first matching result meets the matching condition, the image corresponding to the first matching result is determined as the retrieved image.

[0111] Or, if it is determined that the second matching result meets the matching condition and the first matching result does not meet the matching condition, and if the text request meets the condition, the image corresponding to the second matching result is determined as the retrieved image; wherein, the text request meeting the condition at least includes: the text description information determined based on the correction information input by the user in the historical record is not included in the text request.

[0112] When obtaining the target number of retrieved images based on the first matching result and the second matching result, images with the first matching degree meeting the threshold condition can be selected from the first group of images as the retrieved images based on the first matching degree between the first text embedding vector and each image embedding vector, and images with the second matching degree meeting the threshold condition can be selected from the second group of images as the retrieved images based on the second matching degree between the first text embedding vector and each second text embedding vector, so as to obtain the target number of retrieved images.

[0113] Specifically, by matching the first text embedding vector with each image embedding vector, a plurality of first matching degrees are obtained, and by matching the first text embedding vector with a plurality of second text embedding vectors, a plurality of second matching degrees are obtained. After mixing the plurality of first matching degrees and the plurality of second matching degrees, they are sorted in descending order, and the images corresponding to the first target number of matching degrees are selected from this sorting. Whether these images are from the first group of images or the second group of images, they can all be used as the retrieved images.

[0114] Or, when obtaining the target number of retrieved images based on the first matching result and the second matching result, after obtaining the first matching result and the second matching result, it is determined whether the first matching result meets the matching condition. When the first matching result meets the matching condition, the image corresponding to the first matching result is determined as the retrieved image.

[0115] Among them, determining whether the first matching result meets the matching condition can be specifically: as long as the matching degree between the first text embedding vector and a certain second text embedding vector reaches a certain specific matching degree, the image corresponding to the second text embedding vector is determined as the retrieved image. In this process, it is not necessary to consider the ranking of the matching degree between the first text embedding vector and the second text embedding vector among all matching degrees, nor to consider the matching degree between the first text embedding vector and the image embedding vector, that is, the first matching result between the first text embedding vector and the second text embedding vector takes precedence.

[0116] Among them, the second text embedding vector corresponding to each image in the second group of images is pre-stored, which can be determined and stored based on historical records, or can be determined and stored based on user input. Then, the second text embedding vector corresponding to each image in the second group of images has higher accuracy compared to the image embedding vector corresponding to each image in the first group of images. Therefore, the first matching result between the first text embedding vector and the second text embedding vector takes precedence.

[0117] If the first group of images includes Image 1, and Image 1 corresponds to Image Embedding Vector 1, and the second group of images includes Image 2, and Image 2 corresponds to Second Text Embedding Vector 2, where Image 1 and Image 2 are exactly the same image. When there is a text request, the first text embedding vector corresponding to the text request is matched with multiple second text embedding vectors to obtain a second matching result. If the matching degree between the first text embedding vector and Second Text Embedding Vector 2 in the first matching result meets the matching condition, then based on the fact that the matching degree between the first text embedding vector and Second Text Embedding Vector 2 meets the matching condition, Image 2 can be directly determined as the retrieved image, without matching the first text embedding vector with the image embedding vector 1 corresponding to this Image 2;

[0118] Or, while matching the first text embedding vector corresponding to the text request with multiple second text embedding vectors to obtain a second matching result, the first text embedding vector is also matched with multiple second text embedding vectors to obtain a first matching result. If the matching degree between the first text embedding vector and Second Text Embedding Vector 2 in the first matching result meets the matching condition, while the second matching result indicates that the matching degree between the first text embedding vector and Image Embedding Vector 1 does not meet the matching condition, then based on the fact that the matching degree between the first text embedding vector and Second Text Embedding Vector 2 meets the matching condition, Image 2 can be directly determined as the retrieved image. At this time, it is not necessary to consider whether the matching degree between the first text embedding vector and Image Embedding Vector 1 meets the matching condition.

[0119] Alternatively, when obtaining the target number of retrieved images based on the first matching result and the second matching result, after obtaining the first matching result and the second matching result, if the first matching result does not meet the matching condition, but the second matching result meets the matching condition, then when the text request meets the condition, the image corresponding to the second matching result is determined as the retrieved image.

[0120] Among them, determining whether the first matching result meets the matching condition can be specifically: determining the matching degree between the first text embedding vector and each second text embedding vector, determining whether the matching degree reaches a certain specific matching degree. If the matching degree reaches the specific matching degree, it can be determined that the first matching result meets the matching condition; if the matching degree does not reach the specific matching degree, it can be determined that the first matching result does not meet the matching condition. Correspondingly, for whether the second matching result meets the matching condition, it can also be: determining the matching degree between the first text embedding vector and each image embedding vector, determining whether the matching degree reaches a certain specific matching degree. If the matching degree reaches the specific matching degree, it can be determined that the second matching result meets the matching condition; if the matching degree does not reach the specific matching degree, it can be determined that the second matching result does not meet the matching condition.

[0121] When the first matching result meets the matching condition, the image corresponding to the first matching result can be directly determined as the retrieved image; if the first matching result does not meet the matching condition, it is necessary to further determine whether the second matching result meets the matching condition. If the second matching result also does not meet the matching condition, it is necessary to continue to match the first text embedding vector with other second text embedding vectors and image embedding vectors to determine the retrieved image; if the first matching result does not meet the matching condition, and at the same time, the second matching result meets the matching condition, then when determining the image corresponding to the second matching result as the retrieved image, it is necessary to further determine whether the text request meets the condition. Only when the text request meets the condition will the image corresponding to the second matching result be determined as the retrieved image. If the text request does not meet the condition, it is necessary to continue to match the first text embedding vector with other second text embedding vectors and image embedding vectors to determine the retrieved image.

[0122] Among them, the text request meeting the condition can at least include: the text request does not contain the text description information determined based on the correction information input by the user in the historical record.

[0123] The second text embedding vector is pre-stored based on historical records. To ensure the correctness of the second text embedding vector, after storing the second text embedding vector, the user may input correction information to correct the second text embedding vector to obtain the second text embedding vector corrected based on the user's correction information. Therefore, there will be a record of the user inputting correction information in the historical records. Similarly, if a certain second text embedding vector has been corrected based on the correction information input by the user, then this information can also be determined based on this second text embedding vector.

[0124] In addition, since the second text embedding vector is a vector corresponding to the text description information for textually describing an image that is pre-stored, when the second text embedding vector is corrected by the correction information input by the user, the text description information corresponding to the second text embedding vector also needs to be corrected by the correction information input by the user.

[0125] Therefore, when the text request does not meet the conditions, that is, when the text request contains the text information determined based on the correction information input by the user, it can be determined that the image required by the text request can be obtained from the second group of images. Moreover, the second text embedding vector corresponding to this image has been corrected by the user and has higher accuracy. The second group of images is the images pre-stored and having an associated relationship with the second text embedding vector. At this time, the retrieved image can be directly determined based on the first matching result;

[0126] When the text request meets the conditions, that is, when the text request does not contain the text information determined based on the correction information input by the user, it can be determined that the image required by the text request cannot be obtained from the second group of images, or even if it can be obtained from the second group of images, the second text embedding vector corresponding to the obtained image has not been corrected by the user and has relatively low accuracy. Therefore, when the first matching result does not meet the matching conditions, the second matching result meets the matching conditions, and the text request meets the conditions, the probability that the image corresponding to the first matching result matches the text request is relatively low, and the image corresponding to the second matching result can be directly determined as the retrieved image, that is, the retrieved image is obtained from the first group of images. The first group of images is the images pre-stored and having an associated relationship with the image embedding vector.

[0127] This embodiment discloses an image acquisition method, and its flowchart is as Figure 4 shown, including:

[0128] Step S41, obtain a text request, where the text request is used to request to obtain an image;

[0129] Step S42: Determine the first text embedding vector of the text request; match the first text embedding vector with the second text embedding vector to obtain a matching result, and obtain a target number of retrieved images based on the matching result. The second text embedding vector is the vector corresponding to the text description information that pre-stores the text description of the image.

[0130] Step S43: Generate key information based on the text request. The key information is the key information included in the image to be retrieved specified by the text request.

[0131] Step S44: Analyze the target number of retrieved images to obtain an analysis result, where the analysis result indicates whether each retrieved image includes the key information.

[0132] Step S45: Adjust the target number of retrieved images based on the analysis result to obtain and display the adjusted target images.

[0133] Step S46: Obtain the correction information input by the user based on the displayed adjusted target images. The correction information is used to correct the correspondence between the first target image in the adjusted target images and the text request.

[0134] Obtaining a target number of retrieved images based on the text request may specifically be: determining the first text embedding vector of the text request, matching the first text embedding vector with the second text embedding vector to obtain a matching result, so as to obtain a target number of retrieved images based on the matching result, where the second text embedding vector is the vector corresponding to the text description information that pre-stores the text description of the image;

[0135] Or, it may also be: determining the first text embedding vector of the text request, matching the first text embedding vector with the image embedding vector to obtain a matching result, and obtaining a target number of retrieved images based on the matching result;

[0136] Or, it may also be: determining the first text embedding vector of the text request, matching the first text embedding vector with the second text embedding vector to obtain a first matching result; matching the first text embedding vector with the image embedding vector to obtain a second matching result, where the image embedding vector is the embedding vector of the image that can be obtained; obtaining a target number of retrieved images based on the first matching result and the second matching result.

[0137] Regardless of which of the above methods is used to obtain a target number of retrieved images, when generating key information based on the text request and adjusting the target number of retrieved images based on the key information to obtain the adjusted target images, the target images can be corrected based on the correction information input by the user to adjust the second text embedding vector corresponding to the target images.

[0138] If the adjusted target image is obtained, and this target image is the image requested in the text request, then it can be determined that there is a corresponding relationship between the adjusted target image and the text request.

[0139] After obtaining the adjusted target image, the target image can be displayed to respond to the text request, display and output the target image obtained based on the text request, so that the user can see the target image corresponding to the text request.

[0140] After the target image is displayed and the user sees the displayed target image, the user can determine whether the displayed target image meets the text request. If the user determines that the first target image displayed does not meet the text request, the user can input correction information to correct the corresponding relationship between the first target image and the text request, that is, determine that the first target image does not correspond to the text request, that is, determine that the second text embedding vector of the first target image does not match the first text embedding vector of the text request.

[0141] At this time, the displayed target image can be corrected. That is, since the currently displayed target image includes the first target image that does not meet the text request, the first target image is no longer displayed, and only other target images that meet the text request and exclude the first target image are displayed to ensure that the finally displayed target image can match the initially obtained text request.

[0142] Among them, the correction information input by the user can be: a text input operation performed by the user at a certain target position on the displayed first target image. Through this text input operation, the correction information of the first target image is determined to determine the correct text description of the first target image, so as to generate the correct second text embedding vector of the first target image;

[0143] Or, a box selection operation and a text input operation performed by the user on the displayed first target image. By recognizing the box selection operation, the position of the first target image corresponding to the box selection operation is determined, and the correct description of this position is determined through the text input operation, so as to determine the correction of the information at this position on the first target image.

[0144] Further, in the solution disclosed in this embodiment, after the target image is displayed, a prompt message can be output. The prompt message is used to prompt the user to input correction information, and then the box selection operation or text input operation of the user on the displayed target image can be obtained.

[0145] For example: the text request is "Help me find a picture with two men and one woman". Based on this text request, the target image is determined. The target image includes the first target image. After the target image is displayed, the user performs a box selection operation and a text input operation on the first target image, such as Figure 5As shown, it is a schematic diagram of the first target image and its text description information before and after correction. Figure 5 The upper half shows the first target image and its text description information before correction. Figure 5 The lower half shows the first target image and its text description information after correction, and the correction of the first target image and its text description information is indicated by an arrow.

[0146] The first target image determined in response to the text request, its corresponding text description information is as Figure 5 shown in the upper half: "There are three people in the picture. The people on the left and right sides have short hair, and the person in the middle has long hair. Therefore, there are two men and a woman in the picture."

[0147] Obtain the correction information input by the user. The correction information may include the text content corresponding to the box selection operation and the text input operation. The text content may be as Figure 5 shown in the middle part: "The person in the middle has long hair, but looks like a man."

[0148] The user performs a box selection operation on the first target image, such as Figure 5 the border 51 in the lower half. According to the text content input by the user, it can be determined that the first target image does not correspond to the text request. Moreover, a second text embedding vector is added to the first target image. The text description information corresponding to the second text embedding vector needs to include at least "The person in the middle has long hair, but looks like a man." Then, the text description information of the first target image after correction may be as Figure 5 shown in the lower half: "There are three people in the picture. The people on the left and right sides have short hair. Although the person in the middle has long hair, he looks like a man. Therefore, there are three men in the picture." to clarify that the first target image is an image including three men.

[0149] After obtaining the correction information, the correction information can be analyzed to determine the text description information of the first target image. For example, the text description information is determined by the part of the first target image corresponding to the box selection operation and the text content corresponding to the text input operation. After obtaining the text description information, the vector corresponding to the text description information is obtained and determined as the second text embedding vector corresponding to the first target image.

[0150] Furthermore, in the image acquisition method disclosed in this embodiment, it may further include:

[0151] Determine the text description information of the first target image based on the correction information; obtain the vector corresponding to the text description information of the first target image, and store the vector corresponding to the text description information of the first target image as the second text embedding vector corresponding to the first target image, where the second text embedding vector corresponding to the first target image does not match the first text embedding vector.

[0152] That is, based on the correction information, not only can the displayed target image be corrected, but also the second text embedding vector corresponding to the first target image can be generated or corrected.

[0153] If the second text embedding vector of the first target image is pre-stored, then the second text embedding vector of the first target image needs to be corrected based on the correction information to make it match the actual information of the first target image; if the second text embedding vector of the first target image is not pre-stored, then the second text embedding vector of the first target image can be directly generated based on the correction information, and the association relationship between the second text embedding vector and the first target image is stored, so that when a text request is obtained later, the pre-stored second text embedding vector can be used to match the first text embedding vector of the obtained text request to enrich the pre-stored second text embedding vector and image.

[0154] It should be noted that if the second text embedding vector of the first target image is not pre-stored, then it may be that the image embedding vector of the first target image is pre-stored. By the matching degree between the first text embedding vector of the text request and the image embedding vector of the first target image, the first target image is determined as the retrieval image, and further based on the key information, the first target image is determined as the target image. When it is determined based on the user's correction information that there is no corresponding relationship between the first target image and the text request, the second text embedding vector of the first target image can be generated based on the correction information, and the association relationship between the first target image and the second text embedding vector is stored to increase the accuracy of the vector corresponding to the pre-stored first target image.

[0155] Furthermore, in the image acquisition method disclosed in this embodiment, determining the text description information of the first target image based on the correction information can be specifically:

[0156] Obtain the first text description information corresponding to the part to be corrected in the first target image specified by the correction information; obtain the second text description information corresponding to the part of the first target image other than the part to be corrected in the first analysis result of the first target image, where the first analysis result is the analysis result obtained when the first target image is used as the retrieval image; determine the text description information of the first target image based on the first text description information and the second text description information.

[0157] The correction information is the information for correcting at least a part of the first target image. Therefore, based on the correction information, the first text description information corresponding to the part to be corrected in the first target image can be determined. For example, Figure 5 the correction information of " Figure 5 " is "The person in the middle has long hair, but looks like a man", which is only for the correction of the person in the middle in Figure 5 Figure 5 . Then the part to be corrected is the person in the middle, and the first text description information corresponding to the part to be corrected can be "looks like a man".

[0158] For the text description information corresponding to the other parts of the first target image except the part to be corrected, it can be determined based on the analysis result obtained when the first target image is used as the retrieval image. That is, when the first target image is used as the retrieval image, the first target image is analyzed based on the key information to obtain the analysis result, and the content corresponding to the part of the first target image except the part to be corrected is extracted from the analysis result as the second text description information.

[0159] By combining the first text description information and the second text description information, the complete text description information of the first target image can be obtained. Then, the second text embedding vector of the first target image is obtained using the complete text description information of the first target image and stored for use in subsequent image retrieval processes.

[0160] In the image acquisition method disclosed in this embodiment, after using the first text embedding vector and the second text embedding vector for matching to obtain the retrieval image, the retrieval image is adjusted using the key information to obtain the target image. After the target image is displayed, the correction information input by the user based on the displayed target image can be obtained, and the corresponding relationship between the first target image and the text request is corrected based on the correction information, so that the target image determined after being corrected by the correction information is completely matched with the text request, ensuring the accuracy of target image acquisition; and further, the text description information of the first target image in the target image is determined, and the vector corresponding to the text expression information of the first target image is determined based on this and is determined as the second text embedding vector corresponding to the first target image. That is, this embodiment can adjust the second text embedding vector of the first target image based on the correction information input by the user to achieve the correction of the corresponding relationship between the first target image and the text request, realizing that after determining the target image corresponding to the text request, the corresponding relationship between the target image and the text request can be corrected based on the correction information input by the user, and at the same time, the second text embedding vector corresponding to the target image is adjusted, so that when a text request is received subsequently, it can be based on the adjusted second text embedding vector and the first text embedding vector corresponding to the text request for matching, thereby improving the accuracy of matching between the text embedding vectors and ensuring the accuracy of the finally determined target image.

[0161] This embodiment discloses an image acquisition system, and its structural schematic diagram is as Figure 6 shown, including:

[0162] A first acquisition unit 61, a second acquisition unit 62, a generation unit 63, an analysis unit 64 and an adjustment unit 65.

[0163] Among them, the first acquisition unit 61 is used to acquire a text request, and the text request is used to request to acquire an image;

[0164] The second acquisition unit 62 is used to acquire a target number of retrieved images based on the text request;

[0165] The generation unit 63 is used to generate key information based on the text request, and the key information is the key information included in the image to be acquired specified by the text request;

[0166] The analysis unit 64 is used to analyze the target number of retrieved images to obtain an analysis result, and the analysis result characterizes whether each retrieved image includes the key information;

[0167] The adjustment unit 65 is used to adjust the target number of retrieved images based on the analysis result to obtain the adjusted target image.

[0168] Furthermore, the adjustment unit is used for:

[0169] Reorder the target number of retrieved images based on the analysis result, and determine the reordered images as the target images;

[0170] Or,

[0171] If the analysis result characterizes that there are retrieved images among the target number of retrieved images that do not include the key information, delete the retrieved images that do not include the key information from the target number of retrieved images, and determine the remaining retrieved images after deleting the retrieved images that do not include the key information from the target number of retrieved images as the target images.

[0172] Furthermore, the second acquisition unit is used for:

[0173] Determine the first text embedding vector of the text request; match the first text embedding vector with the second text embedding vector to obtain a matching result, so as to obtain a target number of retrieved images based on the matching result; among them, the second text embedding vector is a vector corresponding to the text description information for textually describing the image stored in advance.

[0174] Furthermore, the image acquisition system disclosed in this embodiment may further include:

[0175] A correction unit, configured to display an adjusted target image; obtain correction information input by a user based on the displayed adjusted target image, where the correction information is used to correct the correspondence between a first target image in the adjusted target image and a text request.

[0176] Further, the correction unit may also be configured to:

[0177] Determine text description information of the first target image based on the correction information; obtain a vector corresponding to the text description information of the first target image, and store the vector corresponding to the text description information of the first target image as a second text embedding vector corresponding to the first target image, where the second text embedding vector corresponding to the first target image does not match the first text embedding vector.

[0178] Further, the correction unit is configured to:

[0179] Obtain first text description information corresponding to a part to be corrected in the first target image specified by the correction information; obtain second text description information corresponding to a part other than the part to be corrected in the first target image in a first analysis result corresponding to the first target image, where the first analysis result is an analysis result obtained when the first target image is used as a retrieval image for analysis; determine the text description information of the first target image based on the first text description information and the second text description information.

[0180] Further, the second obtaining unit is configured to:

[0181] Match the first text embedding vector with the second text embedding vector to obtain a first matching result; match the first text embedding vector with an image embedding vector to obtain a second matching result, where the image embedding vector is an embedding vector of an image that can be obtained; obtain a target number of retrieval images based on the first matching result and the second matching result.

[0182] Further, the second obtaining unit is configured to:

[0183] If it is determined that the first matching result meets the matching condition, determine the image corresponding to the first matching result as the retrieval image.

[0184] Further, the second obtaining unit is configured to:

[0185] If it is determined that the second matching result meets the matching condition and the first matching result does not meet the matching condition, and when it is determined that the text request meets the condition, determine the image corresponding to the second matching result as the retrieval image; where the text request meeting the condition at least includes: the text request does not include the text description information determined based on the correction information input by the user in the historical record.

[0186] The image acquisition system disclosed in this embodiment is implemented based on the image acquisition method disclosed in the above embodiment, and will not be elaborated here.

[0187] The image acquisition system disclosed in this embodiment obtains a text request for requesting to acquire an image, obtains a target number of retrieved images based on the text request, generates key information based on the text request, where the key information is the key information included in the image to be acquired specified by the text request, analyzes the target number of retrieved images to obtain an analysis result, where the analysis result represents whether each retrieved image includes the key information, and adjusts the target number of retrieved images based on the analysis result to obtain adjusted target images. After obtaining the text request, this solution generates key information while obtaining retrieved images based on the text request, and uses the key information to analyze the retrieved images to obtain an analysis result, so as to adjust the retrieved images based on the analysis result to obtain target images, ensuring that after obtaining the retrieved images based on the text request, the retrieved images can be adjusted based on the key information, so that the finally obtained target images meet the actual requirements and ensure the accuracy of the finally determined images.

[0188] This embodiment discloses an electronic device, and its structural schematic diagram is as Figure 7 shown, including:

[0189] A processor 71 and a memory 72.

[0190] Among them, the processor 71 is used to obtain a text request for requesting to acquire an image, obtain a target number of retrieved images based on the text request, generate key information based on the text request, where the key information is the key information included in the image to be acquired specified by the text request, analyze the target number of retrieved images to obtain an analysis result, where the analysis result represents whether each retrieved image includes the key information, and adjust the target number of retrieved images based on the analysis result to obtain adjusted target images;

[0191] The memory 72 is used to store the programs required for the processor to execute the above processing procedures.

[0192] The electronic device disclosed in this embodiment is implemented based on the image acquisition method disclosed in the above embodiment, and will not be elaborated here.

[0193] The electronic device disclosed in this embodiment obtains a text request for requesting to obtain an image; obtains a target number of retrieved images based on the text request; generates key information based on the text request, where the key information is the key information included in the image to be obtained specified by the text request; analyzes the target number of retrieved images to obtain an analysis result, where the analysis result indicates whether each retrieved image includes the key information; and adjusts the target number of retrieved images based on the analysis result to obtain an adjusted target image. After obtaining the text request, this solution generates key information while obtaining the retrieved images based on the text request, and uses the key information to analyze the retrieved images to obtain an analysis result, so as to adjust the retrieved images based on the analysis result to obtain the target image, ensuring that after obtaining the retrieved images based on the text request, the retrieved images can be adjusted based on the key information, so that the finally obtained target image meets the actual requirements and ensures the accuracy of the finally determined image.

[0194] A computer storage medium carrying one or more computer programs, which, when executed by an electronic device, can enable the electronic device to implement any one of the image acquisition methods provided in the embodiments of the present application.

[0195] In addition, it should be noted that the device embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the drawings of the device embodiments provided in the present application, the connection relationships between the modules indicate that they have communication connections, which can be specifically implemented as one or more communication buses or signal lines.

[0196] Through the description of the above embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general hardware. Of course, it can also be implemented by dedicated hardware including application-specific integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. Generally, functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structures for implementing the same function can also be diverse, such as analog circuits, digital circuits, or dedicated circuits. However, for this application, in more cases, software program implementation is a better embodiment. Based on such an understanding, the technical solution of this application, in essence, or the part that makes a contribution to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a computer floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk, or optical disc, etc., and includes several instructions to enable a computer device (which can be a personal computer, training device, or network device, etc.) to execute the methods described in various embodiments of this application.

[0197] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product.

[0198] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of this application are generated in whole or in part. The computer can be a general-purpose computer, a dedicated computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, training device, or data center to another website, computer, training device, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that a computer can store, or a data storage device such as a training device or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid-state drive (SSD)).

Claims

1. An image acquisition method, comprising: Obtaining a text request, wherein the text request is used to request to obtain an image; Obtaining a target number of retrieval images based on the text request; generating key information based on the text request, where the key information is key information included in the image to be acquired specified by the text request; Analyzing the target number of search images to obtain analysis results, wherein the analysis results indicate whether each of the search images includes the key information; The target number of retrieval images are adjusted based on the analysis result to obtain adjusted target images.

2. The method according to claim 1, wherein adjusting the target number of search images based on the analysis result to obtain an adjusted target image comprises: reordering the target number of retrieved images based on the analysis result, and determining the reordered images as the target images; or, If the analysis result indicates that there are retrieval images that do not include the key information in the target number of retrieval images, the retrieval images that do not include the key information are deleted from the target number of retrieval images, and the retrieval images remaining after deleting the retrieval images that do not include the key information from the target number of retrieval images are determined as the target images.

3. The method according to claim 1, wherein obtaining a target number of retrieved images based on the text request comprises: determining a first text embedding vector for the text request; Matching the first text embedding vector with the second text embedding vector to obtain a matching result, so as to obtain a target number of retrieval images based on the matching result; The second text embedding vector is a vector corresponding to pre-stored text description information for textually describing the image.

4. The method according to claim 3, further comprising: Displaying the adjusted target image; Correction information input by a user based on the displayed adjusted target image is obtained, where the correction information is used to correct a correspondence between a first target image in the adjusted target image and the text request.

5. The method according to claim 4, further comprising: Determining text description information of the first target image based on the correction information; A vector corresponding to the text description information of the first target image is obtained, and the vector corresponding to the text description information of the first target image is stored as a second text embedding vector corresponding to the first target image, and the second text embedding vector corresponding to the first target image does not match the first text embedding vector.

6. The method according to claim 5, wherein determining the text description information of the first target image based on the correction information comprises: Obtaining first text description information corresponding to the to-be-corrected portion of the first target image specified by the correction information; Obtaining second text description information corresponding to a portion of the first target image other than the portion to be corrected in a first analysis result corresponding to the first target image, wherein the first analysis result is an analysis result obtained when the first target image is analyzed as a search image; The text description information of the first target image is determined based on the first text description information and the second text description information.

7. The method according to claim 3, wherein matching the first text embedding vector with the second text embedding vector to obtain a matching result so as to obtain a target number of retrieval images based on the matching result comprises: Matching the first text embedding vector with the second text embedding vector to obtain a first matching result; Matching the first text embedding vector with an image embedding vector to obtain a second matching result, wherein the image embedding vector is an embedding vector of an image that can be obtained; A target number of retrieval images are obtained based on the first matching result and the second matching result.

8. The method according to claim 7, wherein obtaining a target number of retrieval images based on the first matching result and the second matching result comprises: If it is determined that the first matching result satisfies the matching condition, the image corresponding to the first matching result is determined as the search image.

9. The method according to claim 7, wherein obtaining a target number of retrieval images based on the first matching result and the second matching result comprises: If it is determined that the second matching result satisfies the matching condition, and the first matching result does not satisfy the matching condition, and if it is determined that the text request satisfies the condition, determining the image corresponding to the second matching result as the search image; The condition that the text request satisfies at least includes: the text request does not contain text description information determined based on the correction information input by the user in the historical record.

10. An electronic device, comprising: A processor, configured to obtain a text request, wherein the text request is used to request to obtain an image; Obtaining a target number of retrieval images based on the text request; Generate key information based on the text request, the key information being key information included in the image to be acquired specified by the text request; analyze the target number of search images to obtain analysis results, the analysis results indicating whether each of the search images includes the key information; adjust the target number of search images based on the analysis results to obtain adjusted target images; The memory is used to store the program required by the processor to execute the above processing process.