Image processing method and device, electronic equipment and computer readable storage medium

The multimodal big model generates description text and extracts the attribute values of the image content, which solves the problem of insufficient accuracy in image search technology, and achieves more efficient similarity calculation and error judgment reduction.

CN120296189APending Publication Date: 2025-07-11BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510337230.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing image search technology has insufficient accuracy, high manual testing costs, and the image similarity algorithm and target recognition algorithm cannot understand the image content, resulting in a high misjudgment rate.

Method used

A multimodal large model is used to process the reference image and the image to be compared, generate description text, extract keywords through natural language processing technology, calculate the image content attribute value, and determine the image similarity using the weighted average method.

Benefits of technology

It improves the accuracy of the image search function, improves the accuracy of similarity calculation, allows you to understand the image content more comprehensively, and reduces the misjudgment rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296189A_ABST
    Figure CN120296189A_ABST
Patent Text Reader

Abstract

The invention provides an image processing method and device, electronic equipment and a computer readable storage medium, and relates to the technical field of image processing, in particular to the technical fields of target recognition, large models, function testing and the like. According to the specific implementation scheme, a reference image is input into a pre-trained large model, and a reference description text of the reference image is obtained; inputting a to-be-compared image into a pre-trained large model, and obtaining a to-be-compared description text of the to-be-compared image; analyzing the reference description text to obtain at least one reference content attribute value corresponding to the reference image; analyzing the description text to be compared, and obtaining at least one content attribute value to be compared corresponding to the image to be compared; and determining the similarity between the reference image and the to-be-compared image according to the similarity between the reference content attribute value and the to-be-compared content attribute value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of image processing technologies, and particularly to technologies such as target recognition, large models, and functional testing. Specifically, the present disclosure relates to an image processing method, an apparatus, an electronic device, and a computer-readable storage medium. Background Art

[0002] With the development of intelligent technologies, image search technologies have also been greatly developed, making it possible to search for images by image. Searching for images by image supports uploading an image and retrieving images similar to the uploaded image through an algorithm.

[0003] With the popularization of the image search by image function, users' requirements for the accuracy of the image search by image function are also getting higher and higher. Testing the image search by image function and improving the algorithm according to the test results can improve the accuracy of the image search by image function and enhance the user experience. Summary of the Invention

[0004] The present disclosure provides an image processing method, an apparatus, an electronic device, and a computer-readable storage medium.

[0005] According to a first aspect of the present disclosure, there is provided an image processing method, the method comprising:

[0006] Inputting a reference image into a pre-trained large model to obtain a reference description text of the reference image; inputting a comparison image to be compared into the pre-trained large model to obtain a comparison description text of the comparison image to be compared;

[0007] Analyzing the reference description text to obtain at least one reference content attribute value corresponding to the reference image; analyzing the comparison description text to obtain at least one comparison content attribute value corresponding to the comparison image to be compared;

[0008] Determining the similarity between the reference image and the comparison image to be compared according to the similarity between the reference content attribute value and the comparison content attribute value.

[0009] According to a second aspect of the present disclosure, there is provided an image processing apparatus, the apparatus comprising:

[0010] A multimodal module configured to input a reference image into a pre-trained large model to obtain a reference description text of the reference image; input a comparison image to be compared into the pre-trained large model to obtain a comparison description text of the comparison image to be compared;

[0011] A text analysis module configured to analyze the reference description text to obtain at least one reference content attribute value corresponding to the reference image; analyze the comparison description text to obtain at least one comparison content attribute value corresponding to the comparison image to be compared;

[0012] A similarity calculation module, configured to determine the similarity between the reference image and the image to be compared according to the similarity between the reference content attribute value and the content attribute value to be compared.

[0013] According to a third aspect of the present disclosure, there is provided an electronic device, which includes:

[0014] At least one processor; and

[0015] A memory communicatively connected to the at least one processor; wherein,

[0016] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the above image processing method.

[0017] According to a fourth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the above image processing method.

[0018] According to a fifth aspect of the present disclosure, there is provided a computer program product, including a computer program, which implements the above image processing method when executed by a processor.

[0019] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0021] Figure 1 is a schematic flowchart of an image processing method provided by an embodiment of the present disclosure;

[0022] Figure 2 is a schematic flowchart of some steps of an image processing method provided by an embodiment of the present disclosure;

[0023] Figure 3 is a schematic flowchart of some steps of an image processing method provided by an embodiment of the present disclosure;

[0024] Figure 4 is a schematic flowchart of some steps of an image processing method provided by an embodiment of the present disclosure;

[0025] Figure 5 is a schematic diagram of the process of a specific embodiment of an image processing method provided by an embodiment of the present disclosure;

[0026] Figure 6 It is a schematic structural diagram of an image processing apparatus provided by an embodiment of the present disclosure;

[0027] Figure 7 It is a block diagram of an electronic device for implementing the image processing method of an embodiment of the present disclosure. Specific Embodiments

[0028] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to assist in understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted below.

[0029] In some related technologies, the accuracy of image search by image can be tested manually in batches, that is, the similarity between the result of image search by image and the uploaded image is judged by the naked eye, and the accuracy of image search by image is determined according to the judgment result.

[0030] In some related technologies, an image similarity algorithm (such as a cosine similarity algorithm) can be used to calculate the similarity between the result of image search by image and the uploaded image.

[0031] In some related technologies, a target recognition algorithm can be used to recognize the objects in the result of image search by image and the uploaded image, and the similarity between the result of image search by image and the uploaded image is determined by comparing the recognition results.

[0032] Manual testing has a high time cost. The image similarity algorithm cannot understand the objects and content in the image and is prone to misjudgment. The target recognition algorithm cannot understand factors such as the scene and weather in the image and there will also be certain misjudgments.

[0033] The image processing method, apparatus, electronic device, and computer-readable storage medium provided by the embodiments of the present disclosure are intended to solve at least one of the above technical problems in the prior art.

[0034] The image processing method provided by the embodiments of the present disclosure can be executed by an electronic device such as a terminal device or a server. The terminal device can be a vehicle-mounted device, a user equipment (UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, a vehicle-mounted device, a wearable device, etc. The method can be implemented by the processor calling computer-readable program instructions stored in the memory. Alternatively, the method can be executed by the server.

[0035] Figure 1 The flowchart shows the image processing method provided by the embodiments of the present disclosure. As Figure 1 shown, the image processing method provided by the embodiments of the present disclosure may include step S110, step S120, and step S130.

[0036] S110: Input the reference image into a pre-trained large model to obtain the reference description text of the reference image; input the image to be compared into the pre-trained large model to obtain the description text to be compared of the image to be compared.

[0037] Among them, the pre-trained large model may be a multimodal large model.

[0038] The multimodal large model, that is, the Multimodal Large Language Models (MLLM) or the Vision Language Models (VLM), is based on the Large Language Models (LLM) and the Large Vision Models (LVM). It can process various media data types including text, images, audio, and video, and learn the associations between data of different modalities through joint training to improve the performance and generalization ability of the model. The core lies in cross-modal information fusion and understanding, enabling the model to more comprehensively and accurately grasp the deep meaning behind the data.

[0039] The multimodal large model is a type of large model that jointly trains multimodal information such as text, images, videos, and audio. Based on the high performance of the large model, it can process data of other modalities in addition to text.

[0040] For example, the reference image can be filled into the prompt template of the preset multimodal large model to obtain the Prompt (hint) input into the multimodal large model. Among them, the prompt template may include: an instruction to prompt the multimodal large model to obtain information from the input image and generate a description text of the input image.

[0041] For example, the prompt template can be: Generate a detailed description of the picture {}. The Prompt input into the multimodal large model can be: Generate a detailed description of the picture {reference image}.

[0042] Similarly, the image to be compared can be filled into the prompt template of the preset multimodal large model to obtain the Prompt (hint) input into the multimodal large model. The Prompt input into the multimodal large model can be: Generate a detailed description of the picture {image to be compared}.

[0043] S120. Analyze the reference description text to obtain at least one reference content attribute value corresponding to the reference image; analyze the text to be compared to obtain at least one content attribute value to be compared corresponding to the image to be compared.

[0044] Analyzing the reference description text can be to use natural language processing techniques to split the reference description text into clauses, extract keywords, etc., to obtain the reference content attribute value according to the reference description text.

[0045] It can also be to use a pre-trained large language model to analyze the reference description text to obtain the reference content attribute value.

[0046] Among them, each reference content attribute value is the attribute value corresponding to an image content attribute of the reference image. Such as the image main body attribute in the reference image, etc.

[0047] Similarly, natural language processing techniques can be used to split the text to be compared into clauses, extract keywords, etc., to obtain the content attribute value to be compared according to the text to be compared.

[0048] It can also be to use a pre-trained large language model to analyze the text to be compared to obtain the content attribute value to be compared.

[0049] Among them, each content attribute value to be compared is the attribute value corresponding to an image content attribute of the image to be compared. Such as the image main body attribute in the image to be compared, etc.

[0050] S130. Determine the similarity between the reference image and the image to be compared according to the similarity between the reference content attribute value and the content attribute value to be compared.

[0051] For an image content attribute, calculate the similarity between the reference content attribute value and the content attribute value to be compared corresponding to this image content attribute.

[0052] Determine the similarity between the reference image and the image to be compared according to the similarities of multiple image content attributes.

[0053] In some possible implementation manners, a weighted average method can be used to determine the similarity between the reference image and the image to be compared. Specifically, the weighted weights corresponding to the image content attributes can be preset in advance, multiply the similarity score corresponding to this image content attribute by the weighted weight, and add the multiplication results corresponding to multiple image content attributes to determine the similarity score between the reference image and the image to be compared, and then determine the similarity between the reference image and the image to be compared.

[0054] In the image processing method provided by the embodiments of the present disclosure, based on a large model, the similarity between the reference image and the image to be compared is determined by the similarity between the reference content attribute value corresponding to the reference image and the content attribute value to be compared of the image to be compared. On the one hand, the high performance of the large model can be utilized to improve the accuracy of similarity calculation. On the other hand, the similarity of the image content of the images can be targeted to improve the accuracy of similarity calculation.

[0055] The following specifically introduces the image processing method provided by the embodiments of the present disclosure.

[0056] As described above, in some possible implementation manners, when obtaining the reference description text of the reference image and the description text to be compared of the image to be compared, the image content attributes to be obtained can be written into the Prompt to obtain the description text including the attribute values corresponding to the image content attributes.

[0057] For example, the reference image and the image content attributes can be filled into the prompt template of the preset multimodal large model to obtain the Prompt input to the multimodal large model. Among them, the prompt template can include: an instruction for prompting the multimodal large model to obtain the attribute value corresponding to the image content attribute from the input image and generate a description text of the input image according to the attribute value.

[0058] For example, the prompt template can be: generate a description of the picture {}, and the description should involve the following {} as much as possible, and the description should be detailed. The Prompt input to the multimodal large model can be: generate a description of the picture {reference image}, and the description should involve the following {image content attributes} as much as possible, and the description should be detailed.

[0059] Similarly, the image to be compared and the image content attributes can be filled into the prompt template of the preset multimodal large model to obtain the Prompt (prompt) input to the multimodal large model. The Prompt input to the multimodal large model can be: generate a description of the picture {image to be compared}, and the description should involve the following {image content attributes} as much as possible, and the description should be detailed.

[0060] After obtaining the description text including the attribute values corresponding to the image content attributes, based on natural language processing technology, the reference description text and the description text to be compared can be segmented and keyword extracted to obtain the attribute values corresponding to the image content attributes in the reference description text and the description text to be compared, that is, the reference content attribute value and the content attribute value to be compared.

[0061] For an image content attribute, calculate the similarity between the reference content attribute value corresponding to the image content attribute and the content attribute value to be compared, and determine the similarity between the reference image and the image to be compared according to the similarities of multiple image content attributes.

[0062] As described above, in some possible implementation manners, a pre-trained large language model can be used to analyze the reference description text to obtain the reference content attribute value, and a pre-trained large language model can be used to analyze the text to be compared description text to obtain the content attribute value to be compared.

[0063] Figure 2 The flowchart shows an implementation manner of using a large model to obtain content attribute values, as Figure 2 shown, it may include step S210 and step S220.

[0064] S210. Based on the reference description text and at least one preset image content attribute, construct a prompt text, input the prompt text into the pre-trained large model, and obtain the reference content attribute value corresponding to the image content attribute of the reference image.

[0065] For example, the reference description text and the image content attribute can be filled into the prompt template of the preset large model to obtain the Prompt input into the large model. Among them, the prompt template may include: prompting the large model to analyze the reference description text and obtain the attribute value corresponding to the image content attribute from the reference description text.

[0066] For example, the prompt template can be: You are a text analysis expert and need to analyze {}, and the analysis includes the following dimensions {}. The Prompt input into the large model can be: You are a text analysis expert and need to analyze {reference description text}, and the analysis includes the following dimensions {image content attribute}.

[0067] S220. Based on the text to be compared description text and at least one preset image content attribute, construct a prompt text, input the prompt text into the pre-trained large model, and obtain the content attribute value to be compared corresponding to the image content attribute of the image to be compared.

[0068] Similarly, the text to be compared description text and the image content attribute can be filled into the prompt template of the preset large model to obtain the Prompt, and the Prompt input into the multi-modal large model can be: You are a text analysis expert and need to analyze {text to be compared description text}, and the analysis includes the following dimensions {image content attribute}.

[0069] It should be emphasized that in step S210 and step S220, the large model used can be the multi-modal large model used in step S110, or other large models.

[0070] The present disclosure does not make any limitation on the execution order of step S210 and step S220. Step S210 can be executed first, step S220 can be executed first, or step S210 and step S220 can be executed simultaneously.

[0071] In some possible implementations, the image content attributes include at least one of the following:

[0072] Image main body attributes: The basic entities existing in the image, and the attributes of the basic entities, such as the color, shape, type, quantity, etc. of the basic entities.

[0073] Scene attributes of the scene where the image is located: The scene where the basic entity is located, such as artificial scenes like highways, ramps, urban roads, industrial parks, tunnels, rural roads, indoor scenes, in-vehicle scenes, etc., and natural scenes like forests, beaches, grasslands, snowfields, etc.

[0074] Environmental attributes of the environment where the image is located: Environmental attributes such as the weather and visibility of the scene where the basic entity is located, such as the weather conditions at that time, such as sunny, rainy, foggy, cloudy, snowy, etc.

[0075] Image shooting perspective: The perspective from which the basic entity is observed, such as front, side, rear, overhead, bottom, etc.

[0076] Attributes of other entities in the image except the image main body: Such as the quantity of other entities and the attributes of other entities, such as color, shape, type, etc.

[0077] Specifically, the image content attributes of the input Prompt as described above can be the following:

[0078] Basic entity: Clearly describe the basic entity existing in the picture. The basic item types can only be the following: sedan, truck, oil tanker, traffic cone, pedestrian, bus, covered truck, special-shaped vehicle, airplane, bicycle, boat, motorcycle, train, bottle, chair, dining table, potted plant, sofa, monitor / TV, bird, cat, cow, dog, horse, sheep, person. It must not exceed this range. Only describe the problems existing in the picture. For complex objects, they can be described with the specified simple objects. For example, a rider can be called a person riding a horse.

[0079] Basic entity attributes: Describe the color of the basic entity. The color should be one of the basic colors: red, orange, yellow, green, cyan, blue, purple, white, black. Describe the shape of the basic entity. The shape should be a regular shape, such as rectangle, square, circle, etc., or one of the irregular shapes.

[0080] Scene: Describe the scene where the object is located, such as highway, ramp, urban road, industrial park, tunnel, rural road, indoor scene, in-vehicle scene, etc.

[0081] Weather conditions: Describe the weather conditions at that time, such as sunny, rainy, foggy, cloudy, snowy, etc.

[0082] Perspective: Describe the perspective from which the object is observed, such as front, side, rear, overhead, bottom, etc.

[0083] Number of objects: Clearly describe the number of basic entities.

[0084] As described above, in some possible implementation manners, after obtaining the reference content attribute value and the content attribute value to be compared, the similarity between the reference image and the image to be compared is determined according to the similarity between the reference content attribute value and the content attribute value to be compared.

[0085] Figure 3 The flowchart shows an implementation manner of determining the similarity between the reference image and the image to be compared according to the similarity between the reference content attribute value and the content attribute value to be compared. As Figure 3 shown, it may include step S310.

[0086] S310. Construct a prompt text according to the reference content attribute value, the content attribute value to be compared, and the description text of the corresponding image content attribute, and input the prompt text into a pre-trained large model to determine the similarity between the reference image and the image to be compared.

[0087] For example, the reference content attribute value, the content attribute value to be compared, and the description text of the corresponding image content attribute can be filled into the prompt template of the preset large model to obtain the Prompt input into the large model. Among them, the prompt template may include: prompting the large model to analyze the reference content attribute value and the content attribute value to be compared to determine the similarity between the reference content attribute value and the content attribute value to be compared.

[0088] For example, the prompt template may be: You are a text analysis expert and need to perform a similarity score for {} and {}, and the score includes the following dimensions:

[0089] Basic entity score: The basic entity categories mainly described by {} and the main described objects of {}, and the object categories include sedan, truck, tanker, traffic cone, pedestrian, bus, tarpaulin truck, special-shaped vehicle, airplane, bicycle, ship, motorcycle, train, bottle, chair, dining table, potted plant, sofa, monitor / TV, bird, cat, cow, dog, horse, sheep, person. If they are the same, get 1 point, otherwise get 0 point.

[0090] Basic entity attribute score: Check whether the types and colors of the basic entities are the same. The colors include 9 types: red, orange, yellow, green, cyan, blue, purple, white, and black. If the types and colors are both the same, get 1 point. If only one item is the same, get 0.5 point. If they are different, get 0 point.

[0091] Scene score: Check whether the scenes described by {} and {} are the same. The scenes include highway, ramp, urban road, park, tunnel, rural road, indoor scene, in-vehicle scene, etc. If the scenes are the same, get 1 point. If they are different, get 0 point.

[0092] Environmental Score: Determine whether the weather mainly described in {} is consistent with the weather described in {}. The weather is divided into three categories. The first category is sunny and cloudy. The second category is rainy, foggy, and overcast. The third category is snowy. For the same category of weather, a score of 1 is given if they are consistent, and a score of 0 is given if they are different categories of weather.

[0093] Viewpoint Score: Check whether the viewpoints (front, side, back, overhead, bottom, etc.) in {} and {} are consistent. If they are consistent, a score of 1 is given; if not, a score of 0 is given.

[0094] Physical Quantity Score: Compare whether the number of objects described in {} and {} is consistent. If they are consistent, a score of 1 is given; if not, a score of 0 is given.

[0095] The Prompt input to the large model can be:

[0096] Basic Entity Score: {The reference image attribute value corresponding to the basic entity} Whether the category of the basic entity mainly described is consistent with the main described object in {the corresponding basic entity}. The object categories include cars, trucks, oil tankers, traffic cones, pedestrians, buses, tarp trucks, special-shaped vehicles, airplanes, bicycles, boats, motorcycles, trains, bottles, chairs, dining tables, potted plants, sofas, monitors / TVs, birds, cats, cows, dogs, horses, sheep, and people. If they are consistent, a score of 1 is given; otherwise, a score of 0 is given.

[0097] Basic Entity Attribute Score: Check whether the type and color of the basic entity described in {the reference content attribute value corresponding to the basic entity attribute} are consistent with the type and color of the basic entity described in {the content attribute value to be compared corresponding to the basic entity attribute}. The colors include 9 types: red, orange, yellow, green, cyan, blue, purple, white, and black. If both the type and color are consistent, a score of 1 is given; if only one is consistent, a score of 0.5 is given; if they are inconsistent, a score of 0 is given.

[0098] Scene Score: Check whether the scenes described in {the reference content attribute value corresponding to the scene} and {the content attribute value to be compared corresponding to the scene} are consistent. The scenes include highways, ramps, urban roads, parks, tunnels, rural roads, indoor scenes, in-vehicle scenes, etc. If the scenes are consistent, a score of 1 is given; if not, a score of 0 is given.

[0099] Environmental Score: Determine whether the weather mainly described in {the reference content attribute value corresponding to the weather condition} is consistent with the weather described in {the content attribute value to be compared corresponding to the weather condition}. The weather is divided into three categories. The first category is sunny and cloudy. The second category is rainy, foggy, and overcast. The third category is snowy. For the same category of weather, a score of 1 is given if they are consistent, and a score of 0 is given if they are different categories of weather.

[0100] Viewpoint score: Check whether the viewpoints (front, side, back, top-down, bottom, etc.) in the {reference content attribute value corresponding to the viewpoint} and the {content attribute value to be compared corresponding to the viewpoint} are the same. If they are the same, score 1 point; if not, score 0 point.

[0101] Physical quantity score: Compare whether the number of objects described in the {reference content attribute value corresponding to the number of objects} and the {content attribute value to be compared corresponding to the number of objects} is the same. If they are the same, score 1 point; if not, score 0 point.

[0102] After obtaining the similarity between the reference content attribute value and the content attribute value to be compared, determine the similarity between the reference image and the image to be compared according to the similarity between the reference content attribute value and the content attribute value to be compared.

[0103] Figure 4 The flowchart shows an implementation method for determining the similarity between the reference image and the image to be compared according to the similarity between the reference content attribute value and the content attribute value to be compared, as Figure 4 shown, which may include step S410 and step S420.

[0104] S410. When the correlation between the first image content attribute and the second image content attribute meets the preset condition, adjust the preset weighted weight corresponding to the second image content attribute according to the similarity between the reference content attribute value and the content attribute value to be compared corresponding to the first image content attribute.

[0105] Wherein, the first image content attribute and the second image content attribute are different image content attributes.

[0106] Preset the weighted weights of different image content attributes. For example, the designed weighted weights can be shown in the following table:

[0107]

[0108]

[0109] Calculate the correlation between different image content attributes (i.e., the first image attribute content and the second image attribute content). When the calculated correlation meets the preset condition (such as being greater than the preset value), obtain the similarity between the reference content attribute value and the content attribute value to be compared corresponding to the first image attribute content, and adjust the preset weighted weight corresponding to the second image content attribute according to the similarity value.

[0110] For example, if the content of the first image attribute is the scene and the content of the second image attribute is the weather, and the scenes of the reference image and the image to be compared are different, then it is very likely that their weathers are also different. Therefore, when the scene of the reference image is different from the scene of the image to be compared, the weighted weight corresponding to the weather can be reduced to prevent the special factor from having too much influence on the similarity calculation.

[0111] For another example, if the content of the first image attribute is the number of basic entities and the content of the second image attribute is the basic entity, if the number of basic entities is large, it means that the proportion of each basic entity in the image is small, and the probability of misidentifying each basic entity increases. Therefore, if the number of basic entities in the reference image is the same as that in the image to be compared and is greater than the set value, the weighted weights corresponding to the basic entity and the basic entity attribute can be reduced.

[0112] For another example, if the content of the first image attribute is the perspective and the content of the second image attribute is the basic entity attribute, if the perspective of the reference image is inconsistent with the perspective of the image to be compared, the same basic entity will present different forms, resulting in different basic entity attributes (such as color and shape). Therefore, when the perspectives are inconsistent, the weighted weight of the basic entity attribute can be reduced to avoid the influence of the different shapes of entities under different perspectives on the similarity calculation.

[0113] S420. Determine the similarity between the reference image and the image to be compared according to the similarity between the reference content attribute value and the content attribute value to be compared corresponding to the image content attribute, and the adjusted weighted weight corresponding to the image content attribute.

[0114] Multiply the similarity score corresponding to the image content attribute by the adjusted weighted weight, and add the multiplication results corresponding to multiple image content attributes to determine the similarity score between the reference image and the image to be compared, and further determine the similarity between the reference image and the image to be compared.

[0115] In some possible implementation manners, after obtaining the similarity score between the reference image and the image to be compared, normalization can be performed to make it between 0 and 1. The meanings of different scores are as follows:

[0116]

[0117] As described above, in some possible implementation manners, the reference image is an image input by the user; the image to be compared is an image retrieved by the application to be detected from the image library according to the reference image.

[0118] The image processing method provided by the embodiments of the present disclosure can be used to detect the image search function of an application.

[0119] Figure 5The process schematic diagram of a specific embodiment of the image processing method provided by the present disclosure is shown. As Figure 5 shown, in the image processing method provided by the embodiments of the present disclosure, the reference image is the image input by the user; the image to be compared is the image retrieved by the application to be detected from the image library according to the reference image.

[0120] As Figure 5 shown, the image input by the user is processed from image to text based on the large model, and the images retrieved by the application to be detected from the image library according to the image input by the user (i.e., search result 1, search result 2,..., search result 4) are processed from image to text, and the results of processing from image to text are processed based on the large model to obtain the similarity between the image input by the user and the retrieved images, and then the image search function of the application to be detected is evaluated.

[0121] Based on the same principle as the method shown in Figure 1 , Figure 6 The structural schematic diagram of an image processing device provided by the embodiments of the present disclosure is shown. As Figure 6 shown, the image processing device 60 may include:

[0122] A multimodal module 610, configured to input the reference image into a pre-trained large model to obtain a reference description text of the reference image; input the image to be compared into the pre-trained large model to obtain a comparison description text of the image to be compared;

[0123] A text analysis module 620, configured to analyze the reference description text to obtain at least one reference content attribute value corresponding to the reference image; analyze the comparison description text to obtain at least one comparison content attribute value corresponding to the image to be compared;

[0124] A similarity calculation module 630, configured to determine the similarity between the reference image and the image to be compared according to the similarity between the reference content attribute value and the comparison content attribute value.

[0125] In the image processing device provided by the embodiments of the present disclosure, based on the large model, the similarity between the reference image and the image to be compared is determined by the similarity between the reference content attribute value corresponding to the reference image and the comparison content attribute value of the image to be compared. On the one hand, the high performance of the large model can be used to improve the accuracy of similarity calculation, and on the other hand, the similarity of the image content of the image can be used to compare the similarity, thereby improving the accuracy of similarity calculation.

[0126] In some possible implementation manners, the text analysis module includes: a reference text analysis unit, configured to construct a prompt text based on a reference description text and at least one preset image content attribute, input the prompt text into a pre-trained large model, and obtain a reference content attribute value corresponding to the image content attribute of the reference image; a text to be compared analysis unit, configured to construct a prompt text based on a text to be compared description text and at least one preset image content attribute, input the prompt text into a pre-trained large model, and obtain a content attribute value to be compared corresponding to the image content attribute of the image to be compared.

[0127] In some possible implementation manners, the image content attribute includes at least one of the following: an image main body attribute; a scene attribute of the scene where the image is located; an environment attribute of the environment where the image is located; an image shooting perspective; other entity attributes in the image except the image main body.

[0128] In some possible implementation manners, the similarity calculation module includes: a large model calculation unit, configured to construct a prompt text according to the reference content attribute value, the content attribute value to be compared, and the description text of the corresponding image content attribute, input the prompt text into a pre-trained large model, and determine the similarity between the reference image and the image to be compared.

[0129] In some possible implementation manners, the similarity calculation module includes: a weighted average calculation unit, configured to adjust a preset weighted weight corresponding to a second image content attribute according to the similarity between the reference content attribute value and the content attribute value to be compared corresponding to a first image content attribute when the correlation between the first image content attribute and the second image content attribute meets a preset condition; the first image content attribute and the second image content attribute are different image content attributes; determine the similarity between the reference image and the image to be compared according to the similarity between the reference content attribute value and the content attribute value to be compared corresponding to the image content attribute, and the adjusted weighted weight corresponding to the image content attribute.

[0130] In some possible implementation manners, the reference image is an image input by a user; the image to be compared is an image retrieved by a to-be-detected application program from an image library according to the reference image.

[0131] It can be understood that each of the above modules of the image processing device in the embodiments of the present disclosure has the function of implementing the corresponding steps of the image processing method in the embodiments shown in Figure 1 The function can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions. The above modules can be software and / or hardware, and the above modules can be implemented separately or multiple modules can be integrated to implement. For the function descriptions of the respective modules of the above image processing device, reference can be specifically made to Figure 1The corresponding description of the image processing method in the embodiments shown herein will not be elaborated further here.

[0132] In the technical solutions of the present disclosure, the processing of the collection, storage, use, processing, transmission, provision, disclosure, and application of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0133] In the technical solutions of the present disclosure, before obtaining or collecting the user's personal information, the authorization or consent of the user has been obtained.

[0134] According to the embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0135] The electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the image processing method provided in the embodiments of the present disclosure.

[0136] Compared with the prior art, based on a large model, the similarity between the reference image and the image to be compared is determined by the similarity between the reference content attribute value corresponding to the reference image and the to-be-compared content attribute value of the image to be compared. On the one hand, the high performance of the large model can be utilized to improve the accuracy of similarity calculation, and on the other hand, the similarity of the image content of the image can be compared to improve the accuracy of similarity calculation.

[0137] The readable storage medium is a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the image processing method provided in the embodiments of the present disclosure.

[0138] Compared with the prior art, based on a large model, the similarity between the reference image and the image to be compared is determined by the similarity between the reference content attribute value corresponding to the reference image and the to-be-compared content attribute value of the image to be compared. On the one hand, the high performance of the large model can be utilized to improve the accuracy of similarity calculation, and on the other hand, the similarity of the image content of the image can be compared to improve the accuracy of similarity calculation.

[0139] The computer program product includes a computer program, and the computer program implements the image processing method provided in the embodiments of the present disclosure when executed by a processor.

[0140] Compared with the prior art, the computer program product determines the similarity between the reference image and the image to be compared based on a large model and by the similarity between the reference content attribute value corresponding to the reference image and the content attribute value to be compared of the image to be compared. On the one hand, the high performance of the large model can be utilized to improve the accuracy of similarity calculation. On the other hand, the similarity of the image content of the images can be targeted to improve the accuracy of similarity calculation.

[0141] Figure 7 FIG. shows a schematic block diagram of an exemplary electronic device 700 that can be used to implement the embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as, for example, a personal digital processor, a cellular phone, a smartphone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0142] As Figure 7 shown, the device 700 includes a computing unit 701 that can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 702 or a computer program loaded from a storage unit 708 into a random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the device 700 can also be stored. The computing unit 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0143] A plurality of components in the device 700 are connected to the I / O interface 705, including: an input unit 706, such as a keyboard, a mouse, etc.; an output unit 707, such as various types of displays, speakers, etc.; a storage unit 708, such as a magnetic disk, an optical disc, etc.; and a communication unit 709, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 709 allows the device 700 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0144] The computing unit 701 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 701 executes the various methods and processes described above, such as the image processing method. For example, in some embodiments, the image processing method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 708. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded into the RAM 703 and executed by the computing unit 701, one or more steps of the image processing method described above can be executed. Alternatively, in other embodiments, the computing unit 701 can be configured to execute the image processing method in any other suitable manner (e.g., by means of firmware).

[0145] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field-programmable gate arrays (FPGA), application-specific integrated circuits (ASIC), application-specific standard products (ASSP), systems-on-a-chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special or general-purpose programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0146] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to the processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program codes can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0147] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0148] In order to provide an interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide an interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and the input received from the user can be in any form (including acoustic input, speech input, or tactile input).

[0149] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0150] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The relationship of the client and the server is generated by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, or a server of a distributed system, or a server incorporating blockchain.

[0151] It should be understood that the various forms of processes shown above can be used, with steps reordered, added or deleted. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution disclosed in this disclosure can be achieved, and no limitations are imposed herein.

[0152] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub - combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.

Claims

1. An image processing method, comprising: Inputting a reference image into a pre-trained large model to obtain a reference description text of the reference image; Inputting a comparison image to be compared into a pre-trained large model to obtain a comparison description text of the comparison image to be compared; Analyzing the reference description text to obtain at least one reference content attribute value corresponding to the reference image; Analyzing the comparison description text to obtain at least one comparison content attribute value corresponding to the comparison image to be compared; Determining the similarity between the reference image and the comparison image to be compared according to the similarity between the reference content attribute value and the comparison content attribute value.

2. The method according to claim 1, wherein, The analyzing the reference description text to obtain at least one reference content attribute value corresponding to the reference image includes: Constructing a prompt text based on the reference description text and at least one preset image content attribute, inputting the prompt text into a pre-trained large model, and obtaining a reference content attribute value corresponding to the image content attribute of the reference image; The analyzing the comparison description text to obtain at least one comparison content attribute value corresponding to the comparison image to be compared includes: Constructing a prompt text based on the comparison description text and at least one preset image content attribute, inputting the prompt text into a pre-trained large model, and obtaining a comparison content attribute value corresponding to the image content attribute of the comparison image to be compared.

3. The method according to claim 2, wherein The image content attribute includes at least one of the following: Image main body attribute; scene attribute of the scene where the image is located; environment attribute of the environment where the image is located; image shooting perspective; other entity attributes in the image except the image main body.

4. The method according to claim 1, wherein The determining the similarity between the reference image and the comparison image to be compared according to the similarity between the reference content attribute value and the comparison content attribute value includes: Constructing a prompt text according to the reference content attribute value, the comparison content attribute value, and the description text of the corresponding image content attribute, inputting the prompt text into a pre-trained large model, and determining the similarity between the reference image and the comparison image to be compared.

5. The method according to claim 1, wherein The determining the similarity between the reference image and the comparison image to be compared according to the similarity between the reference content attribute value and the comparison content attribute value includes: When the correlation between the first image content attribute and the second image content attribute meets a preset condition, adjusting the preset weighted weight corresponding to the second image content attribute according to the similarity between the reference content attribute value and the comparison content attribute value corresponding to the first image content attribute; the first image content attribute and the second image content attribute are different image content attributes; Determining the similarity between the reference image and the comparison image to be compared according to the similarity between the reference content attribute value and the comparison content attribute value corresponding to the image content attribute, and the adjusted weighted weight corresponding to the image content attribute.

6. The method according to claim 1, wherein The reference image is an image input by a user; the comparison image to be compared is an image retrieved by a detection application program to be detected from an image library according to the reference image.

7. An image processing device, comprising: A multimodal module for inputting a reference image into a pre-trained large model to obtain a reference description text of the reference image; Inputting the image to be compared into a pre-trained large model to obtain a comparison description text of the image to be compared; A text analysis module for analyzing the reference description text to obtain at least one reference content attribute value corresponding to the reference image; Analyzing the comparison description text to obtain at least one comparison content attribute value corresponding to the image to be compared; A similarity calculation module for determining the similarity between the reference image and the image to be compared according to the similarity between the reference content attribute value and the comparison content attribute value.

8. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1-6.

9. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-6.

10. A computer program product, comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1-6.