Image processing method and device, computer equipment and storage medium

By determining the target area in the image and iteratively optimized based on the matching loss, the problem of the limited number of sample labels and training samples is solved, achieving more accurate target position recognition and broader applicability.

CN120580463APending Publication Date: 2025-09-02TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410240654.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-01
Publication Date
2025-09-02

AI Technical Summary

Technical Problem

The performance of existing object recognition models relies heavily on the accuracy of sample labels and the limited number of training samples, resulting in the inability to accurately identify the location of the target in the image.

Method used

By acquiring images and prompt text, determining the target area, and iterating based on matching losses, the target area is optimized to minimize semantic gaps and achieve precise positioning of the target object.

Benefits of technology

It improves the accuracy and generalization ability of target recognition, reduces dependence on training samples, and reduces costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120580463A_ABST
    Figure CN120580463A_ABST
Patent Text Reader

Abstract

The invention provides an image processing method and device, computer equipment and a storage medium, and belongs to the technical field of computers. The method comprises the steps that a to-be-processed image and a prompt text are acquired, the image comprises at least one object, and the prompt text is used for representing a target object to be found in the image; determining a target area in the image, wherein the target area is used for representing a position where the target object will appear in the image; based on the prompt text and the target area, determining a matching loss which is used for representing a semantic gap between the prompt text and the target area; and performing iteration on the target area by taking minimization of the matching loss as a target, the iterated target area being an area of the position of the target object in the image. According to the method, the region of the position of the target object in the image can be locked more accurately.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to an image processing method, apparatus, computer equipment, and storage medium. Background Art

[0002] With the development of computer technology, object recognition has been applied to a variety of fields, such as medicine and transportation. Object recognition involves identifying the target object in an image or video and determining its location. For example, identifying the location of lesions in medical images. Accurately determining the location of the target is a key research topic in this field.

[0003] Currently, object recognition models are typically trained using training samples. These training samples consist of images with sample labels. These labels accurately reflect the location of objects in the images. Using these labels as a reference, the object recognition model is trained to ensure that the object locations it determines from the images are closer to those indicated by the sample labels. This allows the object recognition model to accurately identify the location of objects.

[0004] However, in the above technical solution, the performance of the object recognition model is heavily dependent on the accuracy of the sample labels, and the number of training samples is limited, that is, the categories of the targets to be found are limited, which also limits the performance of the object recognition model and makes it impossible to accurately identify the position of the target in the image. Summary of the Invention

[0005] The embodiments of the present application provide an image processing method, apparatus, computer device, and storage medium that can accurately determine the area where a target object is located in an image. The technical solution is as follows:

[0006] In one aspect, an image processing method is provided, the method comprising:

[0007] Acquire an image to be processed and a prompt text, wherein the image contains at least one object, and the prompt text is used to indicate a target object to be found in the image;

[0008] determining a target area in the image, where the target area is used to indicate a location in the image where the target object will appear;

[0009] Determining a matching loss based on the prompt text and the target area, where the matching loss is used to represent a semantic gap between the prompt text and the target area;

[0010] With the goal of minimizing the matching loss, the target region is iterated, and the target region after iteration is the region where the target object is located in the image.

[0011] In another aspect, an image processing apparatus is provided, the apparatus comprising:

[0012] An acquisition module, configured to acquire an image to be processed and a prompt text, wherein the image contains at least one object and the prompt text is used to indicate a target object to be found in the image;

[0013] A first determining module is configured to determine a target area in the image, where the target area is used to indicate a location in the image where the target object will appear;

[0014] A second determination module is configured to determine a matching loss based on the prompt text and the target area, wherein the matching loss is used to represent a semantic gap between the prompt text and the target area;

[0015] The iterative module is used to iterate the target area with the goal of minimizing the matching loss, and the target area after iteration is the area where the target object is located in the image.

[0016] In some embodiments, the first determining module includes:

[0017] a first determining unit, configured to determine a plurality of candidate regions in the image, each candidate region being used to represent a location in the image where the target object will appear;

[0018] The second determining unit is configured to determine the target area from the multiple candidate areas based on the prompt text.

[0019] In some embodiments, the first determination unit is used to perform uniform sampling in the image to obtain multiple position points; for any position point, multiple candidate areas corresponding to the position point are determined with the position point as the center, and the position point is the center of each candidate area in the multiple candidate areas.

[0020] In some embodiments, the second determining unit includes:

[0021] A first determining subunit is configured to determine, for any candidate region, a prompt image where the candidate region is located based on the candidate region and the image;

[0022] a processing subunit, configured to process the prompt image and the prompt text using an object recognition model to obtain a correlation degree corresponding to a candidate region in the prompt image, wherein the correlation degree is used to indicate a degree of similarity between the candidate region in the prompt image and the prompt text;

[0023] The second determining subunit is configured to determine the target area from the multiple candidate areas based on the association degrees corresponding to the multiple candidate areas.

[0024] In some embodiments, the processing subunit is configured to perform at least one of the following:

[0025] Processing the prompt image and the prompt text by the object recognition model to obtain a similarity, wherein the similarity is used to represent the semantic similarity between the prompt image and the prompt text;

[0026] A contribution degree is determined based on the prompt image and the prompt text, where the contribution degree is used to indicate the degree of contribution of the candidate region in the prompt image to finding the target object.

[0027] In some embodiments, the processing sub-unit is used to process the prompt image and the prompt text using the Grad-CAM algorithm to obtain a contribution matrix of the prompt image, where each element in the contribution matrix is ​​used to represent the contribution of the corresponding pixel in the prompt image to determining the target object; based on the prompt image, a mask image of the prompt image is obtained, where the mask image is used to mask the image area outside the candidate area in the prompt image; and based on the contribution matrix and the mask image, the contribution is determined.

[0028] In some embodiments, the association degree includes the similarity degree and the contribution degree;

[0029] The second determination subunit is used to sort the multiple candidate areas in descending order of similarity; select a preset number of candidate areas with the highest ranking from the multiple candidate areas; and select the candidate area with the greatest contribution from the preset number of candidate areas as the target area.

[0030] In some embodiments, the iteration module includes:

[0031] an acquisition unit, configured to acquire a contribution matrix of the prompt image and a mask image of the prompt image, wherein each element in the contribution matrix represents a contribution of a corresponding pixel in the prompt image to determining the target object, and the mask image is configured to mask an image area outside a candidate area in the prompt image;

[0032] a third determining unit, configured to determine an expansion loss based on the contribution matrix, the mask image, and the size of the prompt image, wherein the expansion loss is used to represent a loss caused when the target area is too large during an iteration;

[0033] An adjusting unit is configured to adjust at least one of a position and a size of the target area with the goal of minimizing the matching loss and the expansion loss.

[0034] In some embodiments, the adjustment unit is used to determine the extrusion loss based on the contribution matrix and the mask image, and the extrusion loss is used to represent the loss caused by the target area being too small during the iteration process; with the goal of minimizing the matching loss, the expansion loss and the extrusion loss, at least one of the position and size of the target area is adjusted.

[0035] On the other hand, a computer device is provided, which includes a processor and a memory, wherein the memory is used to store at least one computer program, and the at least one computer program is loaded and executed by the processor to implement the image processing method in the embodiment of the present application.

[0036] On the other hand, a computer-readable storage medium is provided, in which at least one computer program is stored. The at least one computer program is loaded and executed by a processor to implement the image processing method in the embodiment of the present application.

[0037] On the other hand, a computer program product is provided, including a computer program, which is stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the computer device performs the image processing method provided in the above-mentioned various aspects or various optional implementations of various aspects.

[0038] An embodiment of the present application provides an image processing method. In the process of identifying the position of a target object in an image, a target area is first determined in the image, and the target area is used to indicate the position where the target object may appear. Then, the matching loss between the prompt text and the target area is determined. Since the matching loss can reflect the semantic gap between the prompt text and the target area, the presence of the target object in the target area can be accurately determined through the matching loss. Then, with the goal of minimizing the matching loss, the target area is iterated, so that the target area can not only accurately contain the entire target object, but also the target area will be reduced as much as possible on the basis of containing the entire target object as the iteration progresses, thereby more accurately locking the area where the target object is located in the image, that is, improving the accuracy of identifying the position of the target in the image. In addition, compared with the scheme of identifying the position of the target in the image through a trained object recognition model, this scheme does not require the help of samples for training, and can equally identify any target, which not only saves costs but also improves the generalization ability of target recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0040] Figure 1 is a schematic diagram of an implementation environment of an image processing method provided according to an embodiment of the present application;

[0041] Figure 2 is a flowchart of an image processing method provided according to an embodiment of the present application;

[0042] Figure 3 is a flowchart of another image processing method provided according to an embodiment of the present application;

[0043] Figure 4 is a schematic diagram of determining a candidate region according to an embodiment of the present application;

[0044] Figure 5 This is a flow chart for determining a target area according to an embodiment of the present application;

[0045] Figure 6 is a schematic diagram of an activation map provided according to an embodiment of the present application;

[0046] Figure 7 is a schematic diagram of a prompt image provided according to an embodiment of the present application;

[0047] Figure 8 This is a framework diagram of a multi-layer perceptron model provided according to an embodiment of the present application;

[0048] Figure 9 This is an effect diagram of object recognition provided according to an embodiment of the present application;

[0049] Figure 10 is a schematic diagram of object recognition provided according to an embodiment of the present application;

[0050] Figure 11 This is a comparison result diagram provided according to an embodiment of the present application;

[0051] Figure 12 is a block diagram of an image processing device provided according to an embodiment of the present application;

[0052] Figure 13 is a block diagram of another image processing device provided according to an embodiment of the present application;

[0053] Figure 14This is a structural block diagram of a terminal provided according to an embodiment of the present application;

[0054] Figure 15 It is a structural diagram of a server provided according to an embodiment of the present application. DETAILED DESCRIPTION

[0055] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.

[0056] In this application, the terms "first", "second", etc. are used to distinguish identical or similar items with substantially the same effects and functions. It should be understood that there is no logical or temporal dependency between "first", "second", and "nth", nor is there any limitation on the quantity and execution order.

[0057] In the present application, the term "at least one" means one or more, and the term "plurality" means two or more.

[0058] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, storage, and display, etc.), and signals involved in this application are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the images and prompt text involved in this application were obtained with full authorization.

[0059] For ease of understanding, the terms involved in this application are explained below.

[0060] Artificial Intelligence (AI) refers to the theories, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. AI also studies the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making.

[0061] Artificial intelligence (AI) technology is a comprehensive discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Foundational AI technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, big data processing, pre-trained models, operating / interaction systems, and mechatronics. Pre-trained models, also known as large models or basic models, can be fine-tuned and widely applied to downstream tasks across various AI disciplines. AI software technologies primarily encompass computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0062] Computer Vision (CV): Computer vision is the science of making machines "see." Specifically, it refers to using cameras and computers to replace the human eye in object recognition and measurement, and further processing the images to create images more suitable for human observation or transmission to instruments. As a scientific discipline, computer vision studies related theories and technologies, aiming to build artificial intelligence systems that can extract information from images or multidimensional data. Large model technology has brought significant changes to the development of computer vision technology. Pre-trained models in the field of vision, such as the swin-transformer, ViT (Vision Transformer), V-MOE (Vision-Mixture of Experts), and MAE (Masked Autoencoders), can be quickly and widely applied to specific downstream tasks through fine-tuning. Computer vision technology generally includes image processing, image recognition, image semantic understanding, image retrieval, OCR (Optical Character Recognition), video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D (Three Dimensions) technology, virtual reality, augmented reality and map construction, and also includes common biometric recognition technology. The image processing method provided in the embodiment of the present application can be applied to the field of computer vision technology. For example, the image processing method provided in the embodiment of the present application can identify the location of any object in any image.

[0063] Natural language processing (NLP) is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can enable effective communication between people and computers using natural language. Natural language processing involves natural language, that is, the language people use in daily life, and is closely related to linguistic research; it also involves important technologies for model training in computer science, mathematics, and artificial intelligence. The pre-trained model is developed from the Large Language Model in the field of NLP. After fine-tuning, the large language model can be widely used in downstream tasks. Natural language processing technology generally includes text processing, semantic understanding, machine translation, robot question answering, knowledge graphs and other technologies. In the image processing method provided in the embodiment of the present application, natural language processing technology can be used to determine what the target object to be found in the image indicated by the prompt text is.

[0064] Visual prompting: These are specific cues applied directly to the image to help the multimodal model correctly understand the target region or perform a specific task. These visual prompts can take different forms and styles, such as drawing a red circle around the target region, thickening the target bounding box, and highlighting key features.

[0065] The vision-language model CLIP (Cross-Language Image Pretraining) is a multimodal model for jointly processing image and text data. It is designed to enable computers to understand the semantic relationships between images and text. The object recognition model in this application is a vision-language model.

[0066] Referring Expression Comprehension (REC): In an image, referring expressions can refer to a target object through visual features or attributes described in language. For example, "the big dog," "the red car," "the apple in the upper left corner," etc. The referring expression comprehension task in computer vision requires the computer to understand these referring expressions and determine the location or bounding box of the target object they refer to. The prompt text in this application is called a referring expression and is used to indicate the target object to be found.

[0067] Gradient-weighted Class Activation Mapping (Grad-CAM): is a technique for visualizing class activations in deep convolutional neural networks. Grad-CAM calculates and visualizes image regions that respond to specific classes within a deep convolutional neural network, helping to understand the basis for the model's predictions for different classes.

[0068] The image processing method provided in the embodiment of the present application can be executed by a computer device. In some embodiments, the computer device is a terminal or a server. The following first takes the computer device as an example to introduce the implementation environment of the image processing method provided in the embodiment of the present application. Figure 1 Schematic diagram of an implementation environment of an image processing method provided according to an embodiment of the present application. Figure 1 The implementation environment includes a terminal 101 and a server 102. The terminal 101 and the server 102 can be directly or indirectly connected via wired or wireless communication, which is not limited in this application.

[0069] In some embodiments, terminal 101 is a smart phone, tablet computer, laptop computer, desktop computer, smart speaker, smart watch, intelligent voice interaction device, smart home appliance, vehicle-mounted terminal, etc., but is not limited to this. Terminal 101 installs and runs an application that supports image processing. The application can be a traffic application (for example, identifying the position of a vehicle in a traffic image), a game application (for example, interacting by identifying changes in the user's position), or a medical application (for example, identifying the position of a lesion in a medical image), etc., and the embodiments of the present application are not limited to this. Schematically, terminal 101 is a terminal used by a user. The user can input the image to be processed and the prompt text at terminal 101. Then, terminal 101 sends the image to be processed and the prompt text to server 102, and server 102 identifies whether there is a target object indicated by the prompt text in the image, and determines the position of the target object if there is a target object.

[0070] Those skilled in the art will appreciate that the number of the above-mentioned terminals may be more or less. For example, the above-mentioned terminal may be only one, or the above-mentioned terminals may be dozens or hundreds, or a larger number. The embodiments of the present application do not limit the number of terminals and device types.

[0071] In some embodiments, server 102 is an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), big data and artificial intelligence platforms. Server 102 is used to provide background services for applications that support image processing. In some embodiments, server 102 undertakes the main computing work and terminal 101 undertakes the secondary computing work; or, server 102 undertakes the secondary computing work and terminal 101 undertakes the main computing work; or, server 102 and terminal 101 adopt a distributed computing architecture for collaborative computing.

[0072] Figure 2 This is a flowchart of an image processing method provided according to an embodiment of the present application. Figure 2 In the embodiment of the present application, the image processing method is described as being executed by a server. The image processing method includes the following steps:

[0073] 201. The server obtains an image to be processed and a prompt text, where the image contains at least one object and the prompt text is used to indicate a target object to be found in the image.

[0074] In the embodiment of the present application, the image to be processed can be any image, and the embodiment of the present application does not limit this. The image contains at least one object, and the embodiment of the present application does not limit the number of objects in the image. The prompt text is used to indicate the target object to be found. That is, the server can determine the location of the target object in the image based on the prompt text. The target object can be a vehicle in a traffic image, a diseased tissue in a medical image, or the user's hand in a game image, etc., and the embodiment of the present application does not limit this. The server can obtain the image to be processed and the prompt text from other computer devices; or, the server can also obtain the image to be processed and the prompt text from a local database, and the embodiment of the present application does not limit this.

[0075] For example, consider a picture of an aquarium containing a yellow fish, a red fish, and a blue fish. The prompt text is "The yellow fish." The server can then locate the "yellow fish" in the image based on the prompt text.

[0076] 202. The server determines a target area in the image, where the target area is used to indicate a location where a target object will appear in the image.

[0077] In an embodiment of the present application, the server may randomly determine a target area in the image to indicate the location of the target object in the image. The embodiment of the present application does not limit the location and size of the target area. Alternatively, the server may first determine multiple candidate areas from the image, and then select the candidate area closest to the target object from the multiple candidate areas as the target area. The embodiment of the present application does not limit the method for determining the target area. Among them, the multiple candidate areas can be randomly determined or obtained by segmenting the image, and the embodiment of the present application does not limit this.

[0078] 203. The server determines a matching loss based on the prompt text and the target area, where the matching loss is used to represent the semantic gap between the prompt text and the target area.

[0079] In an embodiment of the present application, the server may perform feature extraction on the prompt text to obtain text features. The text features are used to represent the target object in the prompt text, that is, the semantics of the prompt text. The server may perform feature extraction on the target region to obtain region features. The region features are used to represent the image content within the target region, that is, the semantics of the target region. The server may then determine the matching loss based on the text features and the region features.

[0080] In addition to the above analysis from the perspective of the target area, the server can also analyze from the perspective of the entire image. That is, the server determines the matching loss based on the prompt text and the image where the target area is located. Accordingly, the server extracts features from the prompt text to obtain text features. The server extracts features from the image where the target area is located to obtain image features. Image features are used to represent the content within the image where the target area is located, that is, the semantics of the image where the target area is located. The server can then determine the matching loss based on the text features and image features. In the process of determining the matching loss based on image features, the weight of the features of the target area in the image features is higher than the weight of the features of the area outside the target area. That is, in the process of determining the matching loss, the server pays more attention to the target area in the image to calculate the matching loss. The matching loss can accurately reflect the presence of the target object in the target area.

[0081] 204. The server iterates the target area with the goal of minimizing the matching loss. The target area after the iteration is the area where the target object is located in the image.

[0082] In the embodiment of the present application, "iterating the target area" refers to adjusting the target area so that the target area meets the conditions, and the embodiment of the present application does not limit the number of adjustments. The server adjusts the target area with the goal of minimizing the matching loss. When the matching loss meets the matching condition, the server stops the iterative process of the target area. The matching condition can be that the matching loss is lower than a certain loss threshold, or it can be that the matching loss is in a certain loss interval, and the embodiment of the present application does not limit this. By iterating the target area, the target area can not only include the entire target object, but also be reduced as much as possible on the basis of including the entire target object as the iteration progresses, thereby more accurately locking the area where the target object is located in the image.

[0083] An embodiment of the present application provides an image processing method. In the process of identifying the position of a target object in an image, a target area is first determined in the image, and the target area is used to indicate the position where the target object may appear. Then, the matching loss between the prompt text and the target area is determined. Since the matching loss can reflect the semantic gap between the prompt text and the target area, the presence of the target object in the target area can be accurately determined through the matching loss. Then, with the goal of minimizing the matching loss, the target area is iterated, so that the target area can not only accurately contain the entire target object, but also the target area will be reduced as much as possible on the basis of containing the entire target object as the iteration progresses, thereby more accurately locking the area where the target object is located in the image, that is, improving the accuracy of identifying the position of the target in the image. In addition, compared with the scheme of identifying the position of the target in the image through a trained object recognition model, this scheme does not require the help of samples for training, and can equally identify any target, which not only saves costs but also improves the generalization ability of target recognition.

[0084] Figure 3 is a flowchart of another image processing method provided in accordance with an embodiment of the present application, see Figure 3 In the embodiment of the present application, the image processing method is described as being executed by a server. The image processing method includes the following steps:

[0085] 301. The server obtains an image to be processed and a prompt text, where the image contains at least one object and the prompt text is used to indicate a target object to be found in the image.

[0086] In the embodiment of the present application, the server obtains the image i to be processed and the prompt text t. H×W×3; R is used to represent an image set; H is used to represent the length of an image; W is used to represent the width of an image; 3 is used to represent the number of channels of an image, that is, the image is an RGB image, including data of three image channels: red, green, and blue. The data of these three image channels determine the pixels of the image. The prompt text t∈∑; ∑ is used to represent a letter set. The embodiment of the present application does not limit the image to be processed and the prompt text. Then, the server can determine the location of the target object in the image based on the prompt text. That is, the server continues to execute steps 302 to 305.

[0087] 302. The server determines a plurality of candidate regions in the image, each candidate region being used to represent a location where a target object may appear in the image.

[0088] In the embodiment of the present application, the multiple candidate regions can be randomly determined or obtained by uniformly sampling the image, which is not limited in the embodiment of the present application. The shape of the candidate regions can be elliptical, circular, or polygonal, etc., which is not limited in the embodiment of the present application.

[0089] In some embodiments, the server determines multiple candidate areas by uniformly sampling the image. Accordingly, the process of the server determining multiple candidate areas in the image includes: the server performs uniform sampling in the image to obtain multiple position points. For any position point, the server determines multiple candidate areas corresponding to the position point with the position point as the center. The position point is the center of each candidate area in the multiple candidate areas. The solution provided in the embodiment of the present application uniformly samples the image and determines the candidate area with the sampled position point as the center, thereby ensuring that each position in the image can be covered by the candidate areas of multiple position points, that is, ensuring that the candidate areas of the multiple position points can include the target object in the image, which is beneficial to the subsequent determination of the position of the target object and improves the recognition efficiency of the target object.

[0090] For example, the server uniformly samples N locations in the image. For each location, the server determines M candidate regions corresponding to that location, centered at that location. In this case, the server can determine a total of N*M candidate regions.

[0091] The sizes and rotation angles of the candidate regions corresponding to each location point are not exactly the same. The rotation angle refers to the angle by which the candidate region is rotated around the location point. Even if the candidate regions have the same size, the angles of the candidate regions corresponding to different rotation angles are different.

[0092] For example, Figure 4 Schematic diagram of determining a candidate region according to an embodiment of the present application. Figure 4The server evenly samples 24 locations in the image. Taking location 401 as an example, the server determines three candidate regions corresponding to location 401, centered on location 401. These three candidate regions are candidate region 402, candidate region 403, and candidate region 404. Candidate regions 402 and 403 are horizontal, with a rotation angle of 0 degrees. Candidate region 404 is vertical, with a rotation angle of 90 degrees.

[0093] 303. The server determines a target area from multiple candidate areas based on the prompt text.

[0094] In an embodiment of the present application, the server selects the candidate area closest to the target object from multiple candidate areas based on the prompt text as the target area. That is, the server can calculate the correlation between the prompt text and each candidate area. The correlation is used to indicate the degree of similarity between the candidate area and the prompt text. The higher the correlation, the closer the candidate area is to the target object; the lower the correlation, the farther the candidate area is from the target object. The server selects the candidate area with the highest correlation from multiple candidate areas as the target area of ​​the target object. Figure 5 This is a flow chart of determining a target area according to an embodiment of the present application. The specific process of determining the target area can be found in Figure 5 Steps 3031 to 3033 and the following content.

[0095] 3031. For any candidate region, the server determines a prompt image where the candidate region is located based on the candidate region and the image.

[0096] For any candidate region, the server can draw a prompt box in the image based on the candidate region, and use the prompt box to represent the candidate region, thereby obtaining a prompt image corresponding to the candidate region. If the server determines a total of N*M candidate regions, the server can obtain N*M prompt images, that is, {i′1,…,i′ NM}.

[0097] 3032. The server processes the prompt image and prompt text through the object recognition model to obtain the correlation degree corresponding to the candidate area in the prompt image.

[0098] In which, the object recognition model can be a CLIP (Cross-Language Image Pretraining) model or any other visual-language model, and the embodiments of the present application are not limited to this. The server processes the prompt image and prompt text through the object recognition model to obtain the correlation corresponding to the candidate area in the prompt image. The correlation is used to represent the similarity between the candidate area in the prompt image and the prompt text. That is, the server extracts features from the prompt image where the candidate area is located to obtain image features. The image features are used to represent the content in the prompt image where the candidate area is located, that is, the semantics of the image where the candidate area is located. Then, the server can calculate the correlation between the prompt image and the prompt text based on the text features and the image features. In which, in the process of calculating the correlation, the weight of the features of the candidate area in the image features is higher than the weight of the features of the area outside the candidate area. That is, in the process of calculating the correlation, the server pays more attention to the candidate area in the prompt image to calculate the correlation between it and the prompt text. The correlation can accurately reflect the presence of the target object in the target area.

[0099] In some embodiments, the relevance includes at least one of the similarity between the prompt image and the prompt text and the contribution of the prompt image to finding the target object. Accordingly, the process of the server determining the relevance between the prompt image and the prompt text includes at least one of the following methods.

[0100] Method 1: The server processes the prompt image and prompt text through the object recognition model to obtain the similarity. The similarity is used to represent the semantic similarity between the prompt image and the prompt text. The embodiment of the present application does not limit the method of calculating the similarity. The solution provided by the embodiment of the present application can accurately determine whether the candidate area in the prompt image contains the target object by calculating the semantic similarity between the prompt image and the prompt text, which is conducive to accurately locking the position of the target object in the image. If the server obtains N*M prompt images, the server can calculate N*M similarities, that is, {s(i′1, t),…, s(i′ NM , t)}. s(·) is used to represent the function for calculating similarity in the object recognition model. s(i′1,t) is used to represent the similarity between the first prompt image and the prompt text.

[0101] Method 2: The server determines the contribution based on the prompt image and prompt text. The contribution is used to indicate the contribution of the candidate area in the prompt image to the search for the target object. The contribution refers to the rationality (or possibility) of the target object in the candidate area in the prompt image. That is, the contribution corresponding to each candidate area is essentially a confidence score (probability) for identifying the target object in the candidate area. The greater the contribution, the greater the possibility of the target object appearing in the candidate area, and the more conducive it is to further determine the location of the target object; the smaller the contribution, the smaller the possibility of the target object appearing in the candidate area, and the less conducive it is to identify the location of the target object through the candidate area. The server can interpret the location of the target object by calculating the contribution. The solution provided in the embodiment of the present application calculates the contribution of the candidate area in the prompt image to the search for the target object, so that the importance of each position in the prompt image to the determination of the target object can be accurately determined, thereby accurately determining the situation where the candidate area in the prompt image contains the target object, which is conducive to accurately locking the location of the target object in the image.

[0102] The server can use the Grad-CAM algorithm to process the prompt image and prompt text to obtain the contribution matrix A∈R of the prompt image. H×W Each element in the contribution matrix is ​​used to represent the contribution of the corresponding pixel in the hint image to finding the target object. The contribution matrix can also be called an activation map, which is used to represent the importance of each position in the hint image for finding the target object. For example, Figure 6 is a schematic diagram of an activation map provided according to an embodiment of the present application. Figure 6 The darker the color, the greater the contribution to finding the target object, that is, the greater the possibility that the target object is located at that location.

[0103] The server then obtains a mask image for the prompt image based on the prompt image. The mask image is used to mask image regions outside the candidate region in the prompt image. The following example uses an elliptical candidate region as an example.

[0104] In some embodiments, the server may determine the candidate area using the following formula 1.

[0105] Formula 1:

[0106]

[0107] Among them, φ(x, y) is used to represent the candidate area, which is an elliptical area; (c x , c y) is used to represent the center of the candidate region; θ is used to represent the rotation angle of the candidate region; a is used to represent the major axis of the candidate region; b is used to represent the minor axis of the candidate region; (x, y) is used to represent the point on the boundary of the candidate region. The server substitutes the position of each pixel in the prompt image into (x, y) in the above formula 1 to calculate the Euclidean distance matrix D∈R of each pixel in the prompt image from the candidate region. H×W . That is, the elements in the Euclidean distance matrix are used to represent the distance from the corresponding pixel point to the candidate area. The elements (distance values) corresponding to the points on the boundary of the candidate area are all 0; the elements (distance values) corresponding to the points inside the candidate area are all negative numbers; the elements (distance values) corresponding to the points outside the candidate area are all positive numbers. Accordingly, the server can determine the positional relationship between the position obtained in the image and the candidate area based on the Euclidean distance matrix of the candidate area. The positional relationship includes three types of positional relationships: located within the candidate area, located on the boundary of the candidate area, and located outside the candidate area.

[0108] The server can then convert the Euclidean distance matrix into a matrix approximating an elliptic curve using a non-normalized Gaussian distribution f(·). For details, see Formula 2 below. The server can use Formula 2 to determine the matrix C that reflects the candidate region.

[0109] Formula 2:

[0110]

[0111] Among them, φ(x, y) is used to represent the candidate region; μ is used to represent the mean of the Gaussian distribution f(·); σ 2 It is used to represent the variance of Gaussian distribution f(·). By setting μ = 0 and σ, an elliptic curve can be approximated and used to provide visual prompts for the image to be processed to obtain a prompt image. For example, Figure 7 This is a schematic diagram of a prompt image provided according to an embodiment of the present application. Figure 7 (a) in FIG. 1 shows an elliptic curve (candidate region) as an example. The server determines a prompt image based on the elliptic curve (candidate region) and the image to be processed. Figure 7 In the hint image shown in (b), the elliptic curve (candidate region) contains a fish.

[0112] The server may also use the following formula 3 to process the Euclidean distance matrix to obtain a mask image of the prompt image.

[0113] Formula 3:

[0114]

[0115] Among them, φ(x, y) is used to represent the candidate area; g(·) is used to represent the inverse tangent function; ∈ is used to represent the blurriness, which is used to represent the blurriness of the boundary of the candidate area in the mask image; ∈ can be 50, and this is not limited in the embodiment of the present application. g(·) can also be replaced by any function that can approximate a step function (such as a Heaviside function), such as a Sigmoid function, and this is not limited in the embodiment of the present application. In the mask image, the value of the pixel point corresponding to the candidate area is 1; the value of the pixel point corresponding to the area outside the candidate area is 0.

[0116] The server then determines the contribution based on the contribution matrix and the mask image.

[0117] In some embodiments, the server may determine the contribution by using the following formula 4.

[0118] Formula 4:

[0119]

[0120] Among them, A1 is used to represent the contribution matrix corresponding to the first prompt image; M1 is used to represent the mask image (mask matrix) corresponding to the first prompt image; ∑(A1·M1) refers to the sum of the contribution matrix and the elements at the corresponding positions in the mask matrix. Since the elements at the corresponding positions outside the candidate area in the mask matrix are 0, ∑(A1·M1) calculates the sum of all contributions in the candidate area. Then, It is used to indicate the average contribution of each position in the first hint image to finding the target object.

[0121] 3033. The server determines a target area from the multiple candidate areas based on the correlations corresponding to the multiple candidate areas.

[0122] Among them, when the correlation includes similarity, the server can select the candidate area with the greatest similarity as the target area. Alternatively, when the correlation includes contribution, the server can select the candidate area with the greatest contribution as the target area. When the correlation includes similarity and contribution, the server comprehensively considers the similarity and contribution to select the target area. The embodiment of the present application does not limit the method for determining the target area. The solution provided by the embodiment of the present application, since the correlation can accurately reflect the degree of similarity between the candidate area in the prompt image and the prompt text, by selecting the candidate area with the greatest correlation as the target area, the target area can accurately cover the target object, that is, it can accurately determine the location of the target object.

[0123] In some embodiments, the degree of association includes similarity and contribution. Accordingly, the process of the server determining the target area from multiple candidate areas based on the degree of association corresponding to the multiple candidate areas includes: the server sorts the multiple candidate areas in order of similarity from high to low. Then, the server selects a preset number of candidate areas with the highest ranking from the multiple candidate areas. Then, the server selects the candidate area with the largest contribution from the preset number of candidate areas as the target area. The solution provided in the embodiment of the present application, in the process of selecting the target area from multiple candidate areas, not only considers the semantic similarity between the candidate area and the prompt text, but also considers the contribution of the candidate area to finding the target object, so that the selected target area can accurately cover the target object, that is, can accurately determine the location of the target object, which provides a guarantee for further precise location of the target object in the future.

[0124] The target area is an elliptical area. The parameters of the target area are The abscissa used to represent the center of the target area; The vertical coordinate used to represent the center of the target area; a * Used to represent the long axis of the target area; b * Used to represent the short axis of the target area; θ * Used to represent the rotation angle of the target area. Although the target area is the area closest to the target object among the above-mentioned multiple candidate areas, it is still not guaranteed that the target area accurately reflects the position of the target object. For example, the target area contains the target object, but the size of the target area is large and covers a large area in the image. In this case, it can only be determined that the target object is within the target area, and the specific position cannot be further determined. In order to improve this defect, the server can iterate based on the target area to further accurately determine the position of the target object. That is, the server continues to execute steps 304 to 305. The parameters of the target area can be adjusted in subsequent iterations.

[0125] 304. The server determines a matching loss based on the prompt text and the target area, where the matching loss is used to represent the semantic gap between the prompt text and the target area.

[0126] In an embodiment of the present application, the server can determine the matching loss based on the prompt text and the prompt image where the target area is located. The server performs feature extraction on the prompt text to obtain text features. The server performs feature extraction on the prompt image where the target area is located to obtain image features. Image features are used to represent the content within the prompt image where the target area is located, that is, the semantics of the prompt image where the target area is located. Then, the server can determine the matching loss based on the text features and the image features. In the process of determining the matching loss based on the image features, the weight of the features of the target area in the image features is higher than the weight of the features of the area outside the target area. That is, in the process of determining the matching loss, the server pays more attention to the target area in the prompt image to calculate the matching loss. The matching loss can accurately reflect the situation where the target object exists in the target area.

[0127] In some embodiments, the server may calculate the matching loss using the following formula 5.

[0128] Formula 5:

[0129]

[0130] in, Used to represent matching loss; i * The prompt image is used to indicate the target area; t is used to indicate the prompt text of the target object; s(i * , t) is used to represent the similarity (semantic similarity) between the prompt image where the target area is located and the prompt text of the target object; Used to indicate the hint text of the jth background, Is a background prompt collection The hint text of the jth background is shown in ; It is used to represent the similarity between the prompt image where the target area is located and the prompt text of the j-th background. It can maximize the semantic similarity between the prompt image and the prompt text of the target object, and minimize the semantic similarity between the prompt image and the prompt text of the background, so as to avoid the interference of the background in the prompt image on target recognition as much as possible, and make the target area and the area where the target object is located aligned as much as possible.

[0131] 305. The server iterates the target area with the goal of minimizing the matching loss. The target area after the iteration is the area where the target object is located in the image.

[0132] In an embodiment of the present application, the server adjusts at least one of the position and size of the target area with the goal of minimizing the matching loss, so that the target area can not only contain the entire target object, but also be reduced as much as possible on the basis of containing the entire target object as the iteration progresses, thereby more accurately locking the area where the target object is located in the image. The position of the target area includes the coordinates of the center of the target area and the rotation angle of the target area. The size of the target area includes the length and width of the target area. Taking the area where the target area is an ellipse as an example, the size of the target area includes the major axis and the minor axis of the target area. The server can predict the adjustment amount (t x , t y , t a , t b , t θ ). Among them, t x The adjustment amount of the horizontal coordinate used to represent the center of the target area; t y The adjustment amount of the vertical coordinate used to represent the center of the target area; t a Used to indicate the adjustment amount of the long axis of the target area; t b Used to indicate the adjustment amount of the minor axis of the target area; t θ It is used to indicate the adjustment amount of the rotation angle of the target area. The server calculates the adjustment amount (t x , t y , t a , t b , t θ ), adjust the target area and get the new target area (c x +t x , c y +t y , a+t a , b+t b ,θ+t θ ).For example, Figure 8 This is a framework diagram of a multi-layer perceptron model provided according to an embodiment of the present application. Figure 8 The multi-layer perceptron model is composed of a series of linear layers with Leaky ReLU activation functions and ends with a tanh activation function. The server can use the adjustable target region as input to the multi-layer perceptron model to determine the adjusted target region.

[0133] In the process of iterating the target area with the goal of minimizing the matching loss, the object recognition model can sometimes obtain relevant information aligned with the prompt text from the background of the prompt image. This means that there is a local optimal solution in the image space, which may distract the attention of the object recognition model. To this end, this application proposes a method for dynamically adjusting the target area to prevent the object recognition model from falling into the local optimal solution (background). That is, based on the matching loss, this application proposes two additional learning objectives, namely, the expansion loss and extrusion loss The expansion loss is used to represent the loss caused by the target area being too large during the iteration process. The squeeze loss is used to represent the loss caused by the target area being too small during the iteration process.

[0134] Accordingly, the server iterates the target region with the goal of minimizing the matching loss. The process includes: the server obtains a contribution matrix and a mask image of the prompt image. Each element in the contribution matrix represents the contribution of the corresponding pixel in the prompt image to determining the target object. The mask image is used to mask image areas outside the candidate region in the prompt image. The server determines a dilation loss based on the contribution matrix, the mask image, and the size of the prompt image. The server then adjusts at least one of the position and size of the target region with the goal of minimizing the matching loss and the dilation loss.

[0135] In some embodiments, the server may calculate the expansion loss using the following formula 6.

[0136] Formula 6:

[0137]

[0138] in, It is used to represent the expansion loss, which aims to expand the target area to include more contributions by minimizing the expansion loss. * The contribution matrix used to represent the prompt image where the target area is located; M * It is used to represent the mask image of the prompt image where the target area is located; H is used to represent the length of the prompt image where the target area is located; W is used to represent the width of the prompt image where the target area is located.

[0139] However, since the Grad-CAM algorithm usually focuses on areas that are irrelevant to the target object, the introduction of This may cause the target area to be over-expanded. To alleviate this problem, we introduce the squeeze loss Accordingly, the server determines the squeeze loss based on the contribution matrix and the mask image. This squeeze loss represents the loss caused by the target region being too small during the iteration. The server then adjusts the target region to minimize the matching loss, dilation loss, and squeeze loss.

[0140] In some embodiments, the server may calculate the squeeze loss using the following formula 7.

[0141] Formula 7:

[0142]

[0143] in, Used to represent the extrusion loss; minimizing the extrusion loss aims to maximize the average contribution of each position in the target area, which is equivalent to forcing the target area to closely cover (fit) the area where the target object is located. * The contribution matrix used to represent the prompt image where the target area is located; M * A mask image used to represent the hint image where the target area is located.

[0144] In some embodiments, the server may calculate the total loss using the following formula 8 so as to iterate the target area using the total loss.

[0145] Formula 8:

[0146]

[0147] in, Used to express total loss; Used to represent matching loss; Used to indicate expansion loss; Used to represent the squeezing loss. The weights of the three losses, matching loss, expansion loss, and squeezing loss, are all set to 1. The combination of these three losses (iteration targets) will produce a suitable target region (elliptic curve) that can cover the area where the target object is located. For example, Figure 9 This is an effect diagram of object recognition provided according to an embodiment of the present application. Figure 9 , the server is able to pinpoint the location of the left duck.

[0148] In order to more clearly describe the image processing method provided in the embodiment of the present application, the image processing method is further described below with reference to the accompanying drawings. Figure 10 is a schematic diagram of an object recognition provided according to an embodiment of the present application. Figure 10, for the image to be processed, the server performs uniform sampling in the image to obtain multiple candidate regions that are evenly distributed. For any candidate region, the server determines the prompt image where the candidate region is located based on the candidate region and the image. Then, the server inputs the prompt image and the prompt text provided by the user into the object recognition model. The server calculates the semantic similarity between the prompt image and the prompt text through the object recognition model. The server uses the Grad-CAM algorithm to calculate the contribution matrix of the prompt image. Then, the server determines the target region from multiple candidate regions based on the semantic similarity and contribution matrix. Then, the server determines the total loss of object recognition based on the semantic similarity and contribution matrix corresponding to the target region. The total loss includes matching loss, expansion loss, and extrusion loss. The server iterates the target region (coordinate transformation) through a multi-layer perceptron with the goal of minimizing the total loss. The target region after iteration is the region where the target object is located in the image. Among them, the process of calculating the loss based on the prompt image and prompt text where the target region is located is a forward propagation process ( Figure 10 The process of the server iterating the target region based on the loss is the back propagation process ( Figure 10 dashed line in the middle).

[0149] In an embodiment of the present application, the server can determine the effectiveness of the image processing method provided by this solution on multiple data sets such as RefCOCO, RefCOCO+, RefCOCOg, traffic image collection, game image collection, and medical image collection.

[0150] Dataset: The performance of the proposed image processing method is determined using the target referring expression comprehension (REC) task, which aims to find the image region (the region where the target object is located) that is most relevant to a given text (prompt text). The REC task is usually performed on RefCOCO, RefCOCO+, and RefCOCOg. Each image in the dataset is accompanied by multiple referring expressions, and each referring expression points to a unique object with bounding box information. In particular, the referring expressions in RefCOCO and RefCOCOg include relationship-based words such as left / bigger / closer, while RefCOCO+ only involves descriptions of the appearance of the object. For RefCOCO and RefCOCO+, the evaluation for humans and non-humans is divided into testA and testB.

[0151] Evaluation Metrics: The performance is evaluated using the percentage of accurate predictions. When the intersection-over-union ratio between the predicted box (the region determined by the image processing method) and the ground-truth box (the true region) exceeds 0.5, it is considered a correct prediction.

[0152] The server can perform ablation experiments on the three losses in the embodiments of this application. Each time, a partial loss is used to iterate the target area. Based on the intersection-over-union ratio of the target area after the iteration and the true area, the accuracy of the corresponding image processing method is determined. This results in the following Table 1. The accuracy rates in Table 1 demonstrate that the three losses in the embodiments of this application are meaningful and effective. They can more accurately locate the location of the target object.

[0153] Table 1

[0154]

[0155] The server can also compare the image processing method provided in the embodiment of the present application (hereinafter referred to as the present solution) with the existing REC method. Figure 11 , Figure 11 This is a comparison result diagram provided according to the embodiment of this application. Figure 11 It can be seen that when there is no precise target proposal box, the present invention outperforms other zero-shot image processing methods (target reference understanding methods) in multiple datasets.

[0156] A traffic image collection includes multiple traffic images. Each traffic image contains at least one vehicle. The server obtains a traffic image to be processed and a prompt text. The prompt text represents the target vehicle to be found in the traffic image. The server uniformly samples the traffic image to obtain multiple evenly distributed candidate regions. For each candidate region, the server determines the traffic warning image within which the candidate region resides based on the candidate region and the traffic image. The server then inputs the traffic warning image and prompt text into a vehicle recognition model. The server uses the vehicle recognition model to calculate the semantic similarity between the traffic warning image and the prompt text. The semantic similarity represents the semantic similarity between the traffic warning image and the prompt text. The semantic similarity here refers to the target vehicle. The server uses the Grad-CAM algorithm to calculate a contribution matrix for the traffic warning image. Each element in the contribution matrix represents the contribution of the corresponding pixel in the traffic warning image to finding the target vehicle. The server then determines a target region from the multiple candidate regions based on the semantic similarity and the contribution matrix. The server then determines the total object recognition loss based on the semantic similarity and contribution matrix corresponding to the target region. The total loss includes matching loss, expansion loss, and compression loss. Matching loss represents the semantic gap between the prompt text and the target region (target vehicle). Expansion loss represents the loss caused by the target region being too large compared to the target vehicle during iteration. Squeeze loss represents the loss caused by the target region being too small compared to the target vehicle during iteration. Using a multi-layer perceptron, the server adjusts at least one of the position and size of the target region to minimize the total loss, thereby determining the region in the traffic image where the target vehicle is located.

[0157] The traffic image may also include at least one license plate. The prompt text may be used to indicate the target license plate to be found in the traffic image. Accordingly, the server determines the area where the target license plate is located from the traffic image based on the prompt text. This will not be further described here.

[0158] A game image collection includes multiple game images. Each game image is captured while a user is interacting with the game and has been authorized by the user. Each game image includes the user's hand. During game interaction, the user can interact through hand movements. For example, the movement position of an object controlled by the user in the game can be determined based on the movement position of the user's hand. To determine the position of the hand, the server can employ the image processing methods provided in the embodiments of the present application to determine the position of the user's hand in the captured game image. Accordingly, the server obtains the game image to be processed and prompt text. The prompt text represents the hand to be found in the game image. The server uniformly samples the game image to obtain multiple evenly distributed candidate regions. For any candidate region, the server determines the game prompt image within the candidate region based on the candidate region and the game image. The server then inputs the game prompt image and prompt text into an object recognition model. Using the object recognition model, the server calculates the semantic similarity between the game prompt image and the prompt text. The semantic similarity represents the semantic similarity between the game prompt image and the prompt text. The semantic similarity here refers to the user's hand. The server employs the Grad-CAM algorithm to calculate the contribution matrix of the game prompt image. Each element in the contribution matrix is ​​used to represent the contribution of the corresponding pixel in the game prompt image to finding the hand. Then, the server determines the target area from multiple candidate areas based on the semantic similarity and the contribution matrix. Then, the server determines the total loss of object recognition based on the semantic similarity and contribution matrix corresponding to the target area. The total loss includes matching loss, expansion loss, and squeeze loss. Matching loss is used to represent the semantic gap (user's hand) between the prompt text and the target area. Expansion loss is used to represent the loss caused by the target area being too large than the user's hand during the iteration process. Squeeze loss is used to represent the loss caused by the target area being too small than the user's hand during the iteration process. The server uses a multi-layer perceptron to adjust at least one of the position and size of the target area with the goal of minimizing the total loss, thereby determining the area where the user's hand is located in the game image.

[0159] A medical image collection includes multiple medical images. Each medical image contains organs and tissues of a diseased subject (human or animal). The server obtains the medical image to be processed and a prompt text. The prompt text is used to indicate the diseased tissue to be found in the medical image. The embodiment of the present application does not restrict the type of diseased tissue or the organ in which it is located. The server uniformly samples the medical image to obtain multiple evenly distributed candidate regions. For any candidate region, the server determines the medical prompt image where the candidate region is located based on the candidate region and the medical image. The server then inputs the medical prompt image and prompt text into a lesion recognition model. Using the lesion recognition model, the server calculates the semantic similarity between the medical prompt image and the prompt text. The semantic similarity is used to indicate the semantic similarity between the medical prompt image and the prompt text. The semantics here refers to the diseased tissue. The server uses the Grad-CAM algorithm to calculate the contribution matrix of the medical prompt image. Each element in the contribution matrix represents the contribution of the corresponding pixel in the medical prompt image to the search for the diseased tissue. The server then determines the target region from the multiple candidate regions based on the semantic similarity and the contribution matrix. Then, the server determines the total loss of object recognition based on the semantic similarity and contribution matrix corresponding to the target area. The total loss includes matching loss, expansion loss, and squeeze loss. Matching loss is used to represent the semantic (lesion tissue) gap between the prompt text and the target area. Expansion loss is used to represent the loss caused when the target area is too large than the lesion tissue during the iteration process. Squeeze loss is used to represent the loss caused when the target area is too small than the lesion tissue during the iteration process. The server uses a multi-layer perceptron to adjust at least one of the position and size of the target area with the goal of minimizing the total loss, thereby determining the area where the lesion tissue is located in the medical image.

[0160] An embodiment of the present application provides an image processing method. In the process of identifying the position of a target object in an image, a target area is first determined in the image, and the target area is used to indicate the position where the target object may appear. Then, the matching loss between the prompt text and the target area is determined. Since the matching loss can reflect the semantic gap between the prompt text and the target area, the presence of the target object in the target area can be accurately determined through the matching loss. Then, with the goal of minimizing the matching loss, the target area is iterated, so that the target area can not only accurately contain the entire target object, but also the target area will be reduced as much as possible on the basis of containing the entire target object as the iteration progresses, thereby more accurately locking the area where the target object is located in the image, that is, improving the accuracy of identifying the position of the target in the image. In addition, compared with the scheme of identifying the position of the target in the image through a trained object recognition model, this scheme does not require the help of samples for training, and can equally identify any target, which not only saves costs but also improves the generalization ability of target recognition.

[0161] Figure 12 This is a block diagram of an image processing device provided according to an embodiment of the present application. The image processing device is used to execute the steps of the above-mentioned image processing method. Figure 12 The image processing device includes: an acquisition module 1201, a first determination module 1202, a second determination module 1203 and an iteration module 1204.

[0162] An acquisition module 1201 is configured to acquire an image to be processed and a prompt text, wherein the image contains at least one object and the prompt text is used to indicate a target object to be found in the image;

[0163] A first determining module 1202 is configured to determine a target region in an image, where the target region indicates a location where a target object will appear in the image;

[0164] A second determining module 1203 is configured to determine a matching loss based on the prompt text and the target region, where the matching loss is used to represent a semantic gap between the prompt text and the target region;

[0165] The iteration module 1204 is used to iterate the target area with the goal of minimizing the matching loss, and the area where the target object is located in the target area image after the iteration.

[0166] In some embodiments, Figure 13 is a block diagram of another image processing device provided according to an embodiment of the present application. Figure 13 The first determining module 1202 includes:

[0167] A first determining unit 12021 is configured to determine a plurality of candidate regions in the image, each candidate region being used to represent a location where a target object may appear in the image;

[0168] The second determining unit 12022 is configured to determine a target area from multiple candidate areas based on the prompt text.

[0169] In some embodiments, see Figure 13 The first determination unit 12021 is used to perform uniform sampling in the image to obtain multiple position points; for any position point, multiple candidate areas corresponding to the position point are determined with the position point as the center, and the position point is the center of each candidate area in the multiple candidate areas.

[0170] In some embodiments, see Figure 13 The second determining unit 12022 includes:

[0171] The first determining subunit 1301 is configured to determine, for any candidate region, a prompt image where the candidate region is located based on the candidate region and the image;

[0172] The processing subunit 1302 is configured to process the prompt image and the prompt text using an object recognition model to obtain a correlation degree corresponding to the candidate region in the prompt image. The correlation degree is used to indicate the degree of similarity between the candidate region in the prompt image and the prompt text.

[0173] The second determining subunit 1303 is configured to determine a target region from the multiple candidate regions based on the correlation degrees corresponding to the multiple candidate regions.

[0174] In some embodiments, see Figure 13 The processing subunit 1302 is configured to perform at least one of the following:

[0175] The prompt image and prompt text are processed by the object recognition model to obtain similarity, which is used to represent the semantic similarity between the prompt image and the prompt text;

[0176] Based on the prompt image and the prompt text, a contribution degree is determined, and the contribution degree is used to indicate the contribution degree of the candidate region in the prompt image to finding the target object.

[0177] In some embodiments, see Figure 13The processing sub-unit 1302 is used to process the prompt image and prompt text using the Grad-CAM algorithm to obtain a contribution matrix of the prompt image, where each element in the contribution matrix is ​​used to represent the contribution of the corresponding pixel in the prompt image to the determination of the target object; based on the prompt image, a mask image of the prompt image is obtained, where the mask image is used to cover the image area outside the candidate area in the prompt image; and based on the contribution matrix and the mask image, the contribution is determined.

[0178] In some embodiments, the relevance includes similarity and contribution;

[0179] Continue to see Figure 13 The second determining subunit 1303 is used to sort the multiple candidate regions in descending order of similarity; select a preset number of candidate regions with the highest ranking from the multiple candidate regions; and select the candidate region with the greatest contribution from the preset number of candidate regions as the target region.

[0180] In some embodiments, see Figure 13 , iterative module 1204, including:

[0181] An acquisition unit 12041 is configured to acquire a contribution matrix of the prompt image and a mask image of the prompt image, wherein each element in the contribution matrix represents the contribution of a corresponding pixel in the prompt image to determining the target object, and the mask image is configured to mask image areas outside the candidate area in the prompt image.

[0182] The third determining unit 12042 is configured to determine a dilation loss based on the contribution matrix, the mask image, and the size of the hint image. The dilation loss represents a loss caused by an excessively large target region during the iteration process.

[0183] The adjusting unit 12043 is configured to adjust at least one of the position and size of the target area with the goal of minimizing the matching loss and the expansion loss.

[0184] In some embodiments, see Figure 13 The adjustment unit 12043 is used to determine the extrusion loss based on the contribution matrix and the mask image. The extrusion loss is used to represent the loss caused by the target area being too small during the iteration process; the target area is adjusted with the goal of minimizing the matching loss, the expansion loss and the extrusion loss.

[0185] An embodiment of the present application provides an image processing device. In the process of identifying the position of a target object in an image, a target area is first determined in the image, and the target area is used to indicate the position where the target object may appear. Then, the matching loss between the prompt text and the target area is determined. Since the matching loss can reflect the semantic gap between the prompt text and the target area, the presence of the target object in the target area can be accurately determined through the matching loss. Then, with the goal of minimizing the matching loss, the target area is iterated, so that the target area can not only accurately contain the entire target object, but also the target area will be reduced as much as possible on the basis of containing the entire target object as the iteration progresses, thereby more accurately locking the area where the target object is located in the image, that is, improving the accuracy of identifying the position of the target in the image. In addition, compared with the scheme of identifying the position of the target in the image through a trained object recognition model, this scheme does not require the use of samples for training, and can equally identify any target, which not only saves costs but also improves the generalization ability of target recognition.

[0186] It should be noted that the image processing device provided in the above embodiment, when identifying the location of a target object in an image, is illustrated only by the division of the above functional modules. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the image processing device and the image processing method provided in the above embodiment are based on the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.

[0187] In the embodiments of the present application, the computer device can be configured as a terminal or a server. When the computer device is configured as a terminal, the terminal can be used as the execution subject to implement the technical solution provided in the embodiments of the present application. When the computer device is configured as a server, the server can be used as the execution subject to implement the technical solution provided in the embodiments of the present application. The technical solution provided in the present application can also be implemented through interaction between the terminal and the server. The embodiments of the present application do not limit this.

[0188] Figure 14This is a block diagram of a terminal 1400 according to an embodiment of the present application. Terminal 1400 may be a portable mobile terminal, such as a smartphone, tablet computer, MP3 player (Moving Picture Experts Group Audio Layer III), MP4 player (Moving Picture Experts Group Audio Layer IV), laptop computer, or desktop computer. Terminal 1400 may also be referred to as user equipment, portable terminal, laptop terminal, desktop terminal, or other similar names.

[0189] Typically, the terminal 1400 includes a processor 1401 and a memory 1402 .

[0190] The processor 1401 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 1401 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), or PLA (Programmable Logic Array). The processor 1401 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 1401 may be integrated with a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 1401 may also include an AI (Artificial Intelligence) processor, which is used to process computing operations related to machine learning.

[0191] Memory 1402 may include one or more computer-readable storage media, which may be non-transitory. Memory 1402 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash memory storage devices. In some embodiments, the non-transitory computer-readable storage medium in memory 1402 is used to store at least one computer program, which is executed by processor 1401 to implement the image processing method provided in the method embodiment of the present application.

[0192] In some embodiments, terminal 1400 may optionally include a peripheral device interface 1403 and at least one peripheral device. Processor 1401, memory 1402, and peripheral device interface 1403 may be connected via a bus or signal lines. Each peripheral device may be connected to peripheral device interface 1403 via a bus, signal lines, or circuit boards. Specifically, the peripheral device may include at least one of a radio frequency circuit 1404, a display screen 1405, a camera assembly 1406, an audio circuit 1407, and a power supply 1408.

[0193] The peripheral device interface 1403 can be used to connect at least one I / O (Input / Output)-related peripheral device to the processor 1401 and the memory 1402. In some embodiments, the processor 1401, the memory 1402, and the peripheral device interface 1403 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 1401, the memory 1402, and the peripheral device interface 1403 can be implemented on separate chips or circuit boards, which is not limited in this embodiment.

[0194] RF circuit 1404 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. RF circuit 1404 communicates with communication networks and other communication devices via electromagnetic signals. RF circuit 1404 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals into electrical signals. In some embodiments, RF circuit 1404 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, and the like. RF circuit 1404 can communicate with other terminals via at least one wireless communication protocol. Such wireless communication protocols include, but are not limited to, the World Wide Web, metropolitan area networks, intranets, various generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area networks, and / or WiFi (Wireless Fidelity) networks. In some embodiments, RF circuit 1404 may also include circuitry related to NFC (Near Field Communication), which is not limited in this application.

[0195] The display screen 1405 is used to display a UI (User Interface). The UI may include graphics, text, icons, videos, and any combination thereof. When the display screen 1405 is a touch screen display, the display screen 1405 also has the ability to collect touch signals on the surface or above the surface of the display screen 1405. The touch signal can be input as a control signal to the processor 1401 for processing. At this time, the display screen 1405 can also be used to provide virtual buttons and / or virtual keyboards, also known as soft buttons and / or soft keyboards. In some embodiments, there can be one display screen 1405, which is set on the front panel of the terminal 1400; in other embodiments, there can be at least two display screens 1405, which are respectively set on different surfaces of the terminal 1400 or in a folding design; in other embodiments, the display screen 1405 can be a flexible display screen, which is set on the curved surface or folding surface of the terminal 1400. Even more, the display screen 1405 can be set to a non-rectangular irregular shape, that is, a special-shaped screen. The display screen 1405 can be made of materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).

[0196] The camera assembly 1406 is used to capture images or videos. In some embodiments, the camera assembly 1406 includes a front camera and a rear camera. Typically, the front camera is set on the front panel of the terminal, and the rear camera is set on the back of the terminal. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth of field camera, a wide-angle camera, and a telephoto camera, so as to realize the fusion of the main camera and the depth of field camera to realize the background blur function, the fusion of the main camera and the wide-angle camera to realize panoramic shooting and VR (Virtual Reality) shooting function or other fusion shooting functions. In some embodiments, the camera assembly 1406 may also include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. A dual-color temperature flash refers to a combination of a warm light flash and a cold light flash, which can be used for light compensation at different color temperatures.

[0197] The audio circuit 1407 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, and convert the sound waves into electrical signals that are input into the processor 1401 for processing, or input into the radio frequency circuit 1404 to achieve voice communication. For the purpose of stereo sound collection or noise reduction, there may be multiple microphones, each located in different parts of the terminal 1400. The microphone may also be an array microphone or an omnidirectional collection microphone. The speaker is used to convert electrical signals from the processor 1401 or the radio frequency circuit 1404 into sound waves. The speaker may be a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert electrical signals into sound waves audible to humans, but also convert electrical signals into sound waves inaudible to humans for purposes such as ranging. In some embodiments, the audio circuit 1407 may also include a headphone jack.

[0198] Power supply 1408 is used to power various components in terminal 1400. Power supply 1408 can be AC ​​power, DC power, a disposable battery, or a rechargeable battery. When power supply 1408 includes a rechargeable battery, the rechargeable battery can be wired or wirelessly rechargeable. A wired rechargeable battery is charged via a wired line, while a wireless rechargeable battery is charged via a wireless coil. The rechargeable battery can also support fast charging technology.

[0199] In some embodiments, the terminal 1400 further includes one or more sensors 1409 , including but not limited to: an acceleration sensor 1410 , a gyroscope sensor 1411 , a pressure sensor 1412 , an optical sensor 1413 , and a proximity sensor 1414 .

[0200] Accelerometer 1410 can detect the magnitude of acceleration along the three coordinate axes of the coordinate system established by terminal 1400. For example, accelerometer 1410 can be used to detect the components of gravity acceleration along the three coordinate axes. Processor 1401 can control display screen 1405 to display the user interface in either a landscape or portrait view based on the gravity acceleration signal collected by accelerometer 1410. Accelerometer 1410 can also be used to collect game or user motion data.

[0201] The gyroscope sensor 1411 can detect the orientation and rotation angle of the terminal 1400. It can also work with the accelerometer 1410 to collect the user's 3D movements on the terminal 1400. Based on the data collected by the gyroscope sensor 1411, the processor 1401 can implement the following functions: motion sensing (for example, changing the UI based on the user's tilt), image stabilization during shooting, game control, and inertial navigation.

[0202] The pressure sensor 1412 can be provided on the side frame of the terminal 1400 and / or below the display screen 1405. When the pressure sensor 1412 is provided on the side frame of the terminal 1400, it can detect the user's gripping signal of the terminal 1400, and the processor 1401 can perform left-hand or right-hand recognition or shortcut operations based on the gripping signal collected by the pressure sensor 1412. When the pressure sensor 1412 is provided below the display screen 1405, the processor 1401 controls the operable controls on the UI interface based on the user's pressure operation on the display screen 1405. Operable controls include at least one of a button control, a scroll bar control, an icon control, and a menu control.

[0203] Optical sensor 1413 is used to detect ambient light intensity. In one embodiment, processor 1401 can control the display brightness of display screen 1405 based on the ambient light intensity detected by optical sensor 1413. Specifically, when the ambient light intensity is high, the display brightness of display screen 1405 is increased; when the ambient light intensity is low, the display brightness of display screen 1405 is decreased. In another embodiment, processor 1401 can also dynamically adjust the shooting parameters of camera assembly 1406 based on the ambient light intensity detected by optical sensor 1413.

[0204] Proximity sensor 1414, also known as a distance sensor, is typically located on the front panel of terminal 1400. Proximity sensor 1414 is used to detect the distance between the user and the front of terminal 1400. In one embodiment, when proximity sensor 1414 detects that the distance between the user and the front of terminal 1400 is gradually decreasing, processor 1401 controls display screen 1405 to switch from the screen-on state to the screen-off state. When proximity sensor 1414 detects that the distance between the user and the front of terminal 1400 is gradually increasing, processor 1401 controls display screen 1405 to switch from the screen-off state to the screen-on state.

[0205] Those skilled in the art will understand that Figure 14 The structure shown in the figure does not constitute a limitation on the terminal 1400, and the terminal 1400 may include more or fewer components than shown in the figure, or combine certain components, or adopt a different component arrangement.

[0206] Figure 151 is a schematic diagram of the structure of a server provided in accordance with an embodiment of the present application. The server 1500 may vary significantly due to different configurations or performances, and may include one or more processors (Central Processing Units, CPUs) 1501 and one or more memories 1502, wherein the memories 1502 store at least one computer program, which is loaded and executed by the processor 1501 to implement the image processing methods provided in the above-mentioned various method embodiments. Of course, the server 1500 may also have components such as a wired or wireless network interface, a keyboard, and an input / output interface for input and output. The server 1500 may also include other components for implementing device functions, which will not be described in detail here.

[0207] The present application also provides a computer-readable storage medium that stores at least one computer program. The at least one computer program is loaded and executed by a processor of a computer device to implement the operations performed by the computer device in the image processing method of the above embodiment. For example, the computer-readable storage medium can be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), a magnetic tape, a floppy disk, an optical data storage device, etc.

[0208] The present application also provides a computer program product, including a computer program stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium and executes the computer program, causing the computer device to perform the image processing methods provided in the various optional implementations described above.

[0209] Those skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware, or by a program to instruct the relevant hardware, and the program may be stored in a computer-readable storage medium, which may be a read-only memory, a disk, or an optical disk, etc.

[0210] The above description is merely an optional embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.

Claims

1. An image processing method, characterized in that: The method comprises: Acquire an image to be processed and a prompt text, wherein the image contains at least one object, and the prompt text is used to indicate a target object to be found in the image; determining a target area in the image, where the target area is used to indicate a location in the image where the target object will appear; Determining a matching loss based on the prompt text and the target area, where the matching loss is used to represent a semantic gap between the prompt text and the target area; With the goal of minimizing the matching loss, the target region is iterated, and the target region after iteration is the region where the target object is located in the image.

2. The method according to claim 1, characterized in that Determining the target area in the image includes: Determining a plurality of candidate regions in the image, each candidate region being used to represent a location in the image where the target object will appear; Based on the prompt text, the target area is determined from the multiple candidate areas.

3. The method according to claim 2, characterized in that The determining of a plurality of candidate regions in the image comprises: Performing uniform sampling in the image to obtain a plurality of position points; For any position point, multiple candidate areas corresponding to the position point are determined with the position point as the center, and the position point is the center of each candidate area in the multiple candidate areas.

4. The method according to claim 2, characterized in that The determining the target area from the plurality of candidate areas based on the prompt text includes: For any candidate region, determining a prompt image where the candidate region is located based on the candidate region and the image; Processing the prompt image and the prompt text through an object recognition model to obtain a correlation degree corresponding to a candidate region in the prompt image, wherein the correlation degree is used to represent a degree of similarity between the candidate region in the prompt image and the prompt text; The target area is determined from the multiple candidate areas based on the association degrees corresponding to the multiple candidate areas.

5. The method according to claim 4, characterized in that The determining of the correlation between the prompt image and the prompt text by using an object recognition model includes at least one of the following: Processing the prompt image and the prompt text by the object recognition model to obtain a similarity, wherein the similarity is used to represent the semantic similarity between the prompt image and the prompt text; A contribution degree is determined based on the prompt image and the prompt text, where the contribution degree is used to indicate the degree of contribution of the candidate region in the prompt image to finding the target object.

6. The method according to claim 5, characterized in that The determining of the contribution based on the prompt image and the prompt text includes: Using the Grad-CAM algorithm, the prompt image and the prompt text are processed to obtain a contribution matrix of the prompt image, where each element in the contribution matrix is ​​used to represent the contribution of the corresponding pixel in the prompt image to finding the target object; Based on the prompt image, obtaining a mask image of the prompt image, wherein the mask image is used to cover an image area outside the candidate area in the prompt image; The contribution is determined based on the contribution matrix and the mask image.

7. The method according to claim 5, characterized in that The association degree includes the similarity and the contribution; The determining the target area from the multiple candidate areas based on the correlation degrees corresponding to the multiple candidate areas includes: Sorting the multiple candidate regions in descending order of similarity; Selecting a preset number of candidate regions ranked top from the multiple candidate regions; From the preset number of candidate regions, the candidate region with the greatest contribution is selected as the target region.

8. The method according to claim 1, characterized in that The iterating of the target region with the goal of minimizing the matching loss includes: Obtaining a contribution matrix of the prompt image and a mask image of the prompt image, wherein each element in the contribution matrix is ​​used to represent the contribution of a corresponding pixel in the prompt image to determining the target object, and the mask image is used to mask an image area outside the candidate area in the prompt image; Determining an expansion loss based on the contribution matrix, the mask image, and the size of the prompt image, wherein the expansion loss is used to represent a loss caused when the target area is too large during an iteration; With the goal of minimizing the matching loss and the dilation loss, at least one of the position and the size of the target area is adjusted.

9. The method according to claim 8, characterized in that The adjusting at least one of the position and the size of the target area with the goal of minimizing the matching loss and the expansion loss includes: Determine a squeeze loss based on the contribution matrix and the mask image, where the squeeze loss represents a loss caused by the target area being too small during an iteration; With the goal of minimizing the matching loss, the expansion loss, and the squeezing loss, at least one of the position and the size of the target area is adjusted.

10. An image processing device, characterized in that: The device comprises: An acquisition module, configured to acquire an image to be processed and a prompt text, wherein the image contains at least one object and the prompt text is used to indicate a target object to be found in the image; A first determining module is configured to determine a target area in the image, where the target area is used to indicate a location in the image where the target object will appear; A second determination module is configured to determine a matching loss based on the prompt text and the target area, wherein the matching loss is used to represent a semantic gap between the prompt text and the target area; The iterative module is used to iterate the target area with the goal of minimizing the matching loss, and the target area after iteration is the area where the target object is located in the image.

11. A computer device, characterized in that: The computer device includes a processor and a memory, the memory is used to store at least one computer program, and the at least one computer program is loaded by the processor to execute the image processing method according to any one of claims 1 to 9.

12. A computer-readable storage medium, characterized in that The computer-readable storage medium is used to store at least one computer program, and the at least one computer program is used to execute the image processing method according to any one of claims 1 to 9.

13. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the image processing method according to any one of claims 1 to 9 is implemented.