Image Processing Method, Apparatus, Device, and Medium
By generating object feature vectors and object mask diagrams, determining the category of the target object belongs, the problem of low accuracy of object segmentation in the prior art is solved, and flexible identification and segmentation of multiple objects is realized.
Patent Information
- Application Number
- CN202310274692.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-17
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2043-03-17
AI Technical Summary
The prior art can only recognize specific objects learned when object segmentation in images, resulting in low accuracy of object segmentation.
By obtaining the image feature vector and object query vector of the target image, the object feature vector and object mask diagram are generated, and the category of the target object is determined, thereby realizing the identification and segmentation of multiple objects.
Improves the accuracy and flexibility of object segmentation in the image, allowing the identification of the categories of more types of objects.
Smart Images

Figure CN117746029B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of image processing, and particularly to an image processing method, apparatus, device, and medium. Background Art
[0002] Currently, object segmentation (also known as object detection) of images is an important task in the field of image processing. Among them, "segmentation" refers to identifying a target object (such as a vehicle, an airplane, etc.) from an image and determining the position of the target object in the image. The existing object segmentation method usually is to input an image into a pre-trained segmentation model, and the segmentation model detects and identifies the corresponding target object in the input image according to the image features of the sample objects learned during the training process, determines the image region of the target object, and outputs it. It can be seen that when the above method segments the target object in an image, it can only segment the object that has learned the corresponding image features. For example, if the segmentation model only learns the image features of a cat, then the segmentation model only has the ability to segment a cat from an image. However, since the objects included in an image can be diverse, this will result in only a limited number of objects being segmented from the image, leading to low accuracy of object segmentation in the image. Summary of the Invention
[0003] Embodiments of the present application provide an image processing method, apparatus, device, and medium, which can improve the accuracy of object segmentation in an image.
[0004] On the one hand, embodiments of the present application provide an image processing method, which includes:
[0005] Obtain a target image including a target object, and obtain an image feature vector of the target image;
[0006] Obtain an object query vector for querying the target object from the target image, and generate an object feature vector associated with the target image based on the object query vector and the image feature vector;
[0007] Generate an object mask map corresponding to the object feature vector based on the object feature vector and the image feature vector; the object mask map is used to mark the target object;
[0008] Based on the object mask map, obtain a category encoding vector corresponding to each object category in the object category set, and determine the similarity between the object feature vector and the category encoding vector corresponding to each object category, so as to determine the object category to which the target object belongs from the object category set.
[0009] On the one hand, embodiments of the present application provide an image processing method, which includes:
[0010] Obtain a first target image including a first target object and a second target image including a second target object, and obtain a first image feature vector of the first target image and a second image feature vector of the second target image;
[0011] Obtain an object query vector; the object query vector is used to query the first target object from the first target image and to query the second target object from the second target image;
[0012] Generate a first object feature vector associated with the first target image based on the object query vector and the first image feature vector, and generate a second object feature vector associated with the second target image based on the object query vector and the second image feature vector;
[0013] Generate a first object mask map corresponding to the first object feature vector based on the first object feature vector and the first image feature vector, and generate a second object mask map corresponding to the second object feature vector based on the second object feature vector and the second image feature vector; the first object mask map is used to mark the first target object, and the second object mask map is used to mark the second target object;
[0014] Determine an associated target object that matches the first target object from the second target object based on the first object feature vector and the second object feature vector, and generate a movement trajectory of the associated target object based on the first object mask map and the second object mask map.
[0015] On the one hand, an embodiment of the present application provides an image processing apparatus, and the apparatus includes:
[0016] A first acquisition module, configured to acquire a target image including a target object, and acquire an image feature vector of the target image;
[0017] A first processing module, configured to acquire an object query vector for querying the target object from the target image, and generate an object feature vector associated with the target image based on the object query vector and the image feature vector;
[0018] The first processing module is further configured to generate an object mask map corresponding to the object feature vector based on the object feature vector and the image feature vector; the object mask map is used to mark the target object;
[0019] The first processing module is further configured to, based on the object mask map, acquire a category encoding vector corresponding to each object category in the object category set, and determine the similarity between the object feature vector and the category encoding vector corresponding to each object category, so as to determine the object category to which the target object belongs from the object category set.
[0020] On the one hand, an embodiment of the present application provides an image processing apparatus, and the apparatus includes:
[0021] A second acquisition module, configured to acquire a first target image including a first target object and a second target image including a second target object, and acquire an image feature vector of the first target image and an image feature vector of the second target image;
[0022] A second processing module, configured to acquire an object query vector; the object query vector is used to query the first target object from the first target image and the second target object from the second target image;
[0023] The second processing module is further configured to generate a first object feature vector associated with the first target image based on the object query vector and the first image feature vector, and generate a second object feature vector associated with the second target image based on the object query vector and the second image feature vector;
[0024] The second processing module is further configured to generate a first object mask map corresponding to the first object feature vector based on the first object feature vector and the first image feature vector, and generate a second object mask map corresponding to the second object feature vector based on the second object feature vector and the second image feature vector; the first object mask map is used to mark the first target object, and the second object mask map is used to mark the second target object;
[0025] The second processing module is further configured to determine an associated target object that matches the first target object from the second target object based on the first object feature vector and the second object feature vector, and generate a movement trajectory of the associated target object based on the first object mask map and the second object mask map.
[0026] On the one hand, an embodiment of the present application provides an electronic device, which includes a processor and a memory. The memory is used to store a computer program, and the computer program includes program instructions. The processor is configured to call the program instructions to execute some or all of the steps in the above method.
[0027] On the one hand, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. The computer program includes program instructions. When the program instructions are executed by a processor, they are used to execute some or all of the steps in the above method.
[0028] Correspondingly, according to an aspect of the present application, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions. When the computer instructions are executed by a processor, some or all of the steps in the above method can be implemented.
[0029] In the embodiments of the present application, a target image including a target object can be obtained, and an image feature vector of the target image can be obtained. An object query vector for querying the target object from the target image can be obtained, and an object feature vector associated with the target image can be generated based on the object query vector and the image feature vector. The object feature vector can represent object query information determined from the target image through the object query vector, and the object query information can indicate that the object query vector queries for one object or does not query for any object. An object mask map corresponding to the object feature vector can be generated based on the object feature vector and the image feature vector. The object mask map is used to mark the target object. The object mask map can be used to mark the target object queried by the object query vector, and this can be used to identify the position where the target object is located in the target image. Based on the object mask map, a category encoding vector corresponding to each object category in the object category set can be obtained, and the similarity between the object feature vector and the category encoding vector corresponding to each object category can be determined to determine the object category to which the target object belongs from the object category set. It can be understood that this method can decouple the recognition logics of the object position and object category of a target object. The target object in the target image can be queried from the target image first through the object query vector, and a corresponding object feature vector can be generated, and the object position and object category can be determined respectively through the object feature vector. In addition, the object categories in the object category set can also be flexibly adjusted, such as adding or deleting some object categories, so that the object categories to which more types of objects belong can be identified in the target image, and the accuracy and flexibility of object segmentation in the image can be improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] To more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0031] Figure 1 A schematic diagram of an image processing scenario provided by an embodiment of the present application;
[0032] Figure 2 A flowchart of an image processing method provided by an embodiment of the present application;
[0033] Figure 3 A schematic diagram of an object mask map provided by an embodiment of the present application;
[0034] Figure 4 A schematic diagram of a scenario for determining an object mask map provided by an embodiment of the present application;
[0035] Figure 5A schematic diagram of a scenario for highlighting a target image provided by an embodiment of the present application;
[0036] Figure 6 A schematic flowchart of an image processing method provided by an embodiment of the present application;
[0037] Figure 7 A schematic diagram of a scenario for target object segmentation provided by an embodiment of the present application;
[0038] Figure 8 A schematic flowchart of an image processing method provided by an embodiment of the present application;
[0039] Figure 9 A schematic diagram of a scenario for target object segmentation provided by an embodiment of the present application;
[0040] Figure 10 A schematic structural diagram of an image processing apparatus provided by an embodiment of the present application;
[0041] Figure 11 A schematic structural diagram of an image processing apparatus provided by an embodiment of the present application;
[0042] Figure 12 A schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0043] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0044] The image processing method proposed in the embodiments of the present application is implemented on an electronic device, which may be a server or a terminal. Among them, the server may be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, and big data and artificial intelligence platforms. The terminal may be a smart phone, a tablet computer, a notebook computer, a desktop computer, etc., but is not limited thereto.
[0045] A schematic diagram of an image processing scenario proposed based on the image processing method may be as Figure 1 shown Figure 1A network architecture is proposed. The network architecture may include a service server and a user terminal cluster. Among them, the user terminal cluster may include one or more user terminals, and the number of user terminals in the user terminal cluster will not be limited here. There may be a communication connection between the user terminals in the user terminal cluster. At the same time, any user terminal in the user terminal cluster may have a communication connection with the service server, so that each user terminal in the user terminal cluster can perform data interaction with the service server through this communication connection. Among them, the above communication connection does not limit the connection method, and can be directly or indirectly connected through a wired communication method, or can be directly or indirectly connected through a wireless communication method, or can also be connected through other methods, which are not limited in this application. In addition, it can be understood that the electronic devices involved in the embodiments of this application may be Figure 1 the service server shown, or may be Figure 1 any one of the user terminals in the user terminal cluster shown.
[0046] For example, in the embodiments of this application, the server can obtain a target image uploaded by any user terminal and implement object segmentation of the image through the image processing method proposed in this application. For example, when the server obtains a target image containing a target object, it can obtain the image feature vector of the target image and obtain an object query vector for querying the target object from the target image. Based on the object query vector and the image feature vector, an object feature vector associated with the target image is generated. Based on the object feature vector and the image feature vector, an object mask map corresponding to the object feature vector is generated. The object mask map can be used to mark the target object in the target image. Based on the object mask map, the category coding vector corresponding to each object category in the object category set is obtained, and the similarity between the object feature vector and the category coding vector corresponding to each object category is determined. Based on the similarity, the object category to which the target object belongs is determined from the object category set (that is, the specific object meaning of the target object segmented from the target image is determined). Among them, it can be understood that the object query vector may or may not query out the target object from the target image. Therefore, if it is determined based on the object feature vector and the image feature vector that the target object in the target image is queried out through the object query vector (this can be understood as segmenting the target object from the target image), for example, an object mask map is generated based on the object feature vector and the image feature vector, and the target object is determined to be queried out through the object mask map, then the object category set is obtained, and the category coding vector corresponding to each object category is obtained. Subsequently, the target object and the object category to which it belongs can be associated and displayed in the target image.
[0047] It can be understood that if the object mask map includes a target mask value (such as 1) for marking the target object, it means that the target object is determined to be queried, and the image area formed by the target mask value is the object position of the target object. If the object mask map does not include a target mask value for marking the target object (such as only 0 in the object mask value), it means that no object is determined to be queried. Among them, the image size of the object mask map is consistent with the image size of the target image.
[0048] It can be understood that the above object category set may include one or more object categories. It should be understood that the object categories in the object category set can be adjusted, such as adding or deleting object categories. Therefore, the technical solution of this application does not limit the types of objects that can be recognized from the image, which can improve the object segmentation efficiency and flexibility of the image.
[0049] Optionally, in some embodiments, the electronic device can execute the image processing method according to actual business requirements to achieve efficient object segmentation. The technical solution of this application can be applied to any object segmentation scenario. For example, in an object movement detection scenario, the electronic device can obtain the video to be segmented, extract the video frames in the video to be segmented, and perform object segmentation on each video frame in turn according to the technical solution of this application, and generate corresponding object movement trajectories based on the object positions and object categories of the objects recognized in each video frame. Another example is that in a vehicle recognition scenario, the electronic device can obtain a driving image and perform object segmentation on the driving image according to the technical solution of this application to determine vehicles, pedestrians, buildings, etc. included in the driving image. The application scenario is not limited here.
[0050] Optionally, the data involved in this application, such as the target image, the segmented target object, etc., can be stored in a database or can be stored in a blockchain, such as stored through a blockchain distributed system, which is not limited in this application.
[0051] It should be noted that in the specific implementation of this application, when it comes to scenarios involving obtaining relevant data such as user information, such as obtaining the target image uploaded by the user, user permission or consent is required. That is, when the embodiments of this application are applied to specific products or technologies, the collection, use, and processing of relevant user data comply with the relevant laws, regulations, and standards of relevant countries and regions. For example, prompt information can be sent in the form of an interactive interface to prompt specifically what data will be collected or obtained. Specifically, the types, contents, etc. of these data can be prompted to the user through a list or the like. Only after receiving a confirmation operation or instruction to allow data collection on the interactive interface will further data collection and processing be carried out.
[0052] It can be understood that the above scenarios are only examples and do not constitute limitations on the application scenarios of the technical solutions provided by the embodiments of the present application. The technical solutions of the present application can also be applied to other scenarios. For example, as is known to those of ordinary skill in the art, with the evolution of the system architecture and the emergence of new business scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.
[0053] Based on the above description, an embodiment of the present application proposes an image processing method, which can be executed by the above-mentioned electronic device. Please refer to Figure 2 , Figure 2 which is a schematic flowchart of an image processing method provided by an embodiment of the present application.
[0054] As Figure 2 shown, the process of the image processing method according to an embodiment of the present application may include the following:
[0055] S101. Obtain a target image including a target object, and obtain an image feature vector of the target image.
[0056] Among them, the target image may be an image uploaded by a user terminal, or may be a video frame intercepted from a video, or may be an image downloaded from the network, or may be an image obtained from a business database (such as a database constructed in a vehicle recognition scenario). The source of the target image is not limited herein.
[0057] Among them, the target image may include a target object to be queried and recognized. The target object may also be referred to as an object. There may be one or more target objects in a target image. The number of target objects in the target image is not limited. It can be understood that the image part other than the target object in the target image can be referred to as the image background.
[0058] In some embodiments, obtaining the image feature vector of the target image may be to divide the target image into R image patches, obtain the embedding vector of each image patch, and input the embedding vectors of each image patch into an image encoder to sequentially output the image feature sub-vectors of each image patch, and use the R image feature sub-vectors as the image feature vector. It can be understood that an image feature sub-vector is used to characterize the image features of an image region in the target image, that is, to characterize the image features of the corresponding image patch. Among them, the division of the target image can be performed at the pixel granularity. For example, the target image is divided at a 1x1 pixel granularity, and each image patch obtained at this time corresponds to a pixel point. Or, the target image is divided at a 2x1 pixel granularity, and each image patch obtained at this time corresponds to two pixel points. For example, if the pixels of the target image are 4x4, when the target image is divided at a 1x1 pixel granularity, the R image patches obtained are 16 image patches; when the target image is divided at a 2x2 pixel granularity, the R image patches obtained are 4 image patches. Here, the pixel granularity used for dividing the target image is not limited. That is to say, R is related to the pixels of the target image (i.e., the image size).
[0059] Among them, for the sake of easy understanding, hereinafter, taking the division of the target image at a 1x1 pixel granularity as an example, for example, the image size of the target image is HxW, that is, the R image patches are also HW image patches. The image encoder can be the encoder in a transformer (a neural network model). When the image encoder processes an image based on the embedding vector of any image patch to output the corresponding image feature sub-vector, it can fuse the embedding vectors of the image patches associated with the any image patch (such as the adjacent image patches of the any image patch, etc.), so that the image feature sub-vector of any image patch can not only contain the image information of the any image patch, but also include the image information of the associated image patches.
[0060] Among them, if the image feature sub-vector is a C-dimensional vector, the image feature vector includes R C-dimensional vectors (which can also be expressed as HxW C-dimensional vectors). It can be understood that the R C-dimensional vectors are obtained by arranging the HxW C-dimensional vectors in sequence. Optionally, the image encoder can belong to an image segmentation model, and the embedding vector associated with the pixel value can be preset or obtained by training simultaneously with the image segmentation model.
[0061] S102. Obtain an object query vector for querying a target object from the target image, and generate an object feature vector associated with the target image based on the object query vector and the image feature vector.
[0062] Among them, the object query vector is used to query the target object (i.e., the object) contained in the target image. There can be N object query vectors, where N is a positive integer. Each object query vector can query the object contained in the target image. The query process and principle of each object query vector are the same. It can be understood that the object query vector can be obtained through simultaneous training with the image segmentation model. During training, N initial object query vectors can be obtained, and these N initial object query vectors are trained during the training of the image segmentation model to obtain the N object query vectors here. It can be understood that N can be a preset parameter and is not limited here.
[0063] In some embodiments, generating the object feature vector associated with the target image can be to input the object query vector and the image feature vector into the image decoder. The image decoder performs feature interaction on the object query vector and the image feature vector to obtain the object feature vector corresponding to the object query vector. Among them, the image decoder belongs to the image segmentation model, and this image decoder can be the encoder in the transformer.
[0064] Optionally, one object query vector can correspond to one object feature vector. This object feature vector is used to represent the object query information determined from the target image. This object query information can indicate whether the object query vector queries the target object from the image feature vector. It can be understood that one object query vector realizes querying the target object from the target image through feature interaction with the image feature vector. When an object query vector fails to query any target object, the object query information represented by the corresponding object feature vector indicates that no target object is queried. That is to say, the object query vector may identify from the target image not the target object but the image background. When an object query vector queries a target object, the object query information represented by the corresponding object feature vector indicates that the target object is queried.
[0065] Among them, the process of performing feature interaction on the object query vector and the image feature vector in the image decoder can be as follows:
[0066] Q 0 = softmax(Q 1 *F T )F + Q 1
[0067] Among them, Q 1 ∈R NxC represents N C-dimensional object query vectors processed by the image decoder; F ∈ R RxC represents R C-dimensional image feature sub-vectors processed by the image decoder. softmax(Q 1 *F T)A similarity matrix of size NxR can be obtained, which is used to measure the similarity between each object query vector and each image feature sub-vector. The higher the similarity, the more likely the image patch corresponding to the image feature sub-vector belongs to the object region of the target object queried by the corresponding object query vector, that is, the more likely the image patch represents the target object. softmax(Q 1 *F T )F should be N vectors of dimension C, which are used to represent the image features of the target image are sequentially assigned to each object query vector according to the similarity matrix. The resulting Q 0 ∈R NxC represents the N object feature vectors of dimension C after vector addition and fusion. The N object query vectors of dimension C correspond one-to-one with the N object feature vectors of dimension C. The object feature vector can be understood as a new object query vector centered on the object queried by the object query vector.
[0068] That is to say, the target object contained in the target image can be queried through the object query vector, the feature information of the queried target object can be extracted from the target image, and the corresponding object feature vector can be generated based on the feature information and the object query vector. When an object query vector queries out the contained target object, the feature information of the specifically queried target object can be extracted through the object query vector, and the corresponding object feature vector is generated. At this time, the object query information represented by the object feature vector indicates that the target object in the target image is queried through the object query vector. When an object query vector does not query out the contained target object, the object query information represented by the corresponding generated object feature vector indicates that no target object in the target image is queried through the object query vector, that is, the queried target object is empty.
[0069] S103. Generate an object mask map corresponding to the object feature vector based on the object feature vector and the image feature vector.
[0070] In some embodiments, generating the object mask map can specifically be to perform convolution processing on each image feature sub-vector through the object feature vector to obtain the convolution result corresponding to each image feature sub-vector, determine the mask value corresponding to the image region indicated by each image feature sub-vector based on the convolution result corresponding to each image feature sub-vector, and generate the object mask map through the mask value corresponding to the image region indicated by each image feature sub-vector. It can be understood that the image size of the object mask map is consistent with the image size of the target image. It can be understood that the object mask map can be used to mark the target object. Among them, the object mask map can include a target mask value (such as 1) and a non-target mask value (such as 0). The target mask value is used to mark the target object queried by the object query vector, and the non-target mask value is the mask value other than the target mask value.
[0071] In some embodiments, each image feature sub-vector is subjected to convolution processing by the object feature vector. The specific convolution result corresponding to each image feature sub-vector can be obtained by inputting the object feature vector into a convolution kernel generator (also known as a multi-layer perceptron) to obtain a dynamic convolution kernel corresponding to the object feature vector, and then performing convolution processing on each image feature sub-vector by the dynamic convolution kernel to obtain the convolution result corresponding to each image feature sub-vector. Among them, the convolution kernel generator can belong to the image segmentation model and is obtained during the training process of the image segmentation model. That is to say, an object query vector can obtain an object mask map.
[0072] It can be understood that determining the mask value corresponding to the image region indicated by each image feature sub-vector based on the convolution result corresponding to each image feature sub-vector can be as follows: if the convolution result corresponding to any image feature sub-vector is greater than or equal to a preset threshold, then it is determined that the mask value corresponding to the image region indicated by any image feature sub-vector is the target mask value, that is, the image region indicated by any image feature sub-vector belongs to the region where the target object is located; if the convolution result corresponding to any image feature sub-vector is less than the preset threshold, then it is determined that the mask value corresponding to the image region indicated by any image feature sub-vector is a non-target mask value, that is, the image region indicated by any image feature sub-vector does not belong to the region where the target object is located.
[0073] It can be understood that each object feature vector will generate an object mask map. The object mask map can represent the position and shape of the target object queried by the corresponding object query vector in the target image. When an object query vector queries an object, the corresponding output object mask map consists of 0s and 1s. In the object mask map, the region where the target object is located is represented by 1s, and the remaining image regions are represented by 0s. When an object query vector does not query a target object, the corresponding output object mask map consists of 0s. It can be understood that each object query vector can query a target object from the target image, and the target objects queried by different object query vectors can be different. Among them, it can be understood that when the target image is divided at a 1x1 pixel granularity, one mask value in the obtained object mask map corresponds to one pixel point. When the target image is divided at a 2x2 pixel granularity, one mask value in the obtained object mask map corresponds to four pixel points.
[0074] For example, the target image includes object A, and it is determined that object A is queried from the target image through the object query vector, then the corresponding generated object mask map can be as Figure 3 shown Figure 3Schematic diagram of an object mask map provided by an embodiment of the present application; wherein, a mask value is recorded for the corresponding pixel points in the target image in the object mask map. If any pixel point (i.e., the image area indicated by the image feature sub-vector) belongs to the area where object A is located, the mask value recorded for this any pixel point in the object mask map is the target mask value (i.e., 1). If any pixel point does not belong to the area where object A is located, the mask value recorded for this any pixel point in the object mask map is a non-target mask value (i.e., 0).
[0075] For example, as Figure 4 shown Figure 4 Schematic diagram of a scenario for determining an object mask map provided by an embodiment of the present application; wherein, the image feature vector of the target image includes R C-dimensional image feature vectors 40 (such as denoted as V 1 C 、V 2 C 、...、V R C ), the object query vector includes N C-dimensional object query vectors 41 (such as denoted as U 1 C 、U 2 C 、...、U N C ), each object query vector is successively multiplied by each image feature sub-vector and then passed through the softmax function to output an NxR similarity matrix 42 (S 11 、S 12 、...、S RN ), and the object feature vector 43 (such as denoted as W 1 C 、W 2 C 、...、W N C ) is determined through the similarity matrix 42, the image feature vector 40, and the object query vector 41; taking the object query vector U 2 C as an example, the object query vector U 2 C is successively multiplied by the image feature sub-vectors V 1 C 、V 2 C 、...、V R C and then passed through the softmax function to output the similarity between the object query vector U 2 C and the image feature vector V 1 C 、V2 C ,..., V R C The similarity (S 21 , S 22 ,..., S 2N ) between them. Multiply the similarities (S 21 , S 22 ,..., S 2N ) with the corresponding image feature sub-vectors (V 1 C , V 2 C ,..., V R C ). After that, add them to the object query vector U 2 C sequentially to obtain the object feature vector W 2 C corresponding to the object query vector U 2 C .
[0076] For another example, based on the above object query vector U 2 C corresponding to the object feature vector W 2 C to generate the object mask graph 44. It can be done by inputting the object feature vector W 2 C into the convolutional kernel generator 45 to obtain the dynamic convolutional kernel 46 (a C-dimensional vector). Then, use the dynamic convolutional kernel 46 to perform convolutional processing on the image feature vectors 40 (V 1 C , V 2 C ,..., V R C ) sequentially to obtain the convolutional results 47 (Q 1 C , V 2 C ,..., V R C ) corresponding to the image feature vectors 40 (V 1 C , Q 2 C ,..., Q R C ). Based on this convolutional result 47 (Q 1 C , Q 2 C ,..., Q R C ) to determine the corresponding image feature vectors 40 (V1 C , V 2 C ,..., V R C ) The mask value 48 corresponding to the image region (i.e., pixel points) indicated, that is, the mask values corresponding to each pixel point in the target image can be obtained, and the object mask map 44 is generated based on the mask value 48 corresponding to each pixel point.
[0077] S104. Based on the object mask map, obtain the category encoding vector corresponding to each object category in the object category set, and determine the similarity between the object feature vector and the category encoding vector corresponding to each object category, so as to determine the object category to which the target object belongs from the object category set.
[0078] Among them, obtaining the category encoding vector corresponding to each object category in the object category set based on the object mask map can be that if the object mask map includes the target mask value for marking the target object, obtain the category encoding vector corresponding to each object category in the object category set; if the object mask map does not include the target mask value for marking the target object, do not obtain the category encoding vector corresponding to each object category in the object category set.
[0079] It can be understood that if the object mask map includes the target mask value for marking the target object, it is determined that the object query information indicates that the target object is queried through the object query vector. Or, if the object mask map includes the target mask value and the number of target mask values is greater than or equal to the preset number, it is determined that the object query information indicates that the target object is queried through the object query vector. Correspondingly, if the object mask map does not include the target mask value for marking the target object, it is determined that the object query information indicates that the target object is not queried through the object query vector. Or, if the object mask map includes the target mask value and the number of target mask values is less than the preset number, it is determined that the object query information indicates that the target object is not queried through the object query vector. That is to say, an object query feature obtains an object feature vector, and an object feature vector determines an object mask map. Step S104 is executed for each object mask map.
[0080] In some embodiments, the set of object categories may include a preset plurality of object categories, and the number of object categories may be preset, which is not limited herein. For example, the object categories are vehicles, airplanes, cats, dogs, etc.; each object category has an associated category text. For example, if the object category is a vehicle, the category text is "vehicle"; the category text associated with the object category can be subjected to text encoding processing (such as using an encoder in a transformer model or any text encoder) to obtain a text encoding vector corresponding to the category text associated with the object category, and this text encoding vector can be used as the category encoding vector corresponding to the object category. One object category corresponds to one category encoding vector. It can be understood that the object categories in the set of object categories can be arbitrarily configured, which can improve the flexibility of category recognition of the objects included in the image.
[0081] Among them, the object feature vector can be multiplied by the category encoding vector corresponding to each object category respectively, and the result of the vector multiplication can be used as the similarity between the object feature vector and the category encoding vector corresponding to each object category respectively.
[0082] It can be understood that the object category to which the target object belongs can be determined from the set of object categories based on the similarity. For example, the object category with the highest similarity in the set of object categories is determined as the object category to which the target object belongs. Subsequently, the target object and the object category to which it belongs can be associated and displayed in the target image.
[0083] In some embodiments, the associated display of the target object and the object category to which it belongs can be to determine the object region indicated by the object mask map in the target image, and the target object and the object category to which it belongs are associated and displayed in the object region. It can be understood that the target image is masked based on the target mask value in the object mask map, and the mask region where the target mask value is located in the masked target image is used as the object region indicated by the object mask map in the target image. It can be understood that the object region indicated by the object mask map is the object region of the identified target object in the target image.
[0084] For example, the associated display of the target object and its object category in the object region can be to highlight the target object in the object region and generate a category label carrying the object category. For example, the image data in the object region is subjected to hue adjustment processing to obtain an adjusted object region, and the adjusted object region is superimposed on the target image to obtain a superimposed target image containing the highlighted target object, and a category label corresponding to the object category to which the highlighted target object belongs is generated in the superimposed target image. Among them, the hue adjustment processing can be to adjust the color values of the image data in the object region. For example, the color values of the image data are all rendered as purple, gray, etc. It can be understood that when multiple object query vectors are used to query multiple objects from the target object, that is, when multiple object mask maps are generated, the above method can be used to obtain an adjusted object region based on each object mask map respectively, and the multiple adjusted object regions can be superimposed on the target image, so that multiple segmented objects can be viewed in the target image.
[0085] For example, as Figure 5 shown, Figure 5 FIG. 3 is a schematic diagram of a scenario for highlighting a target image provided by an embodiment of the present application; the target image 51 is masked based on the object mask map 50 to obtain a masked region indicated by the object mask map 50, and the masked region can be used as the object region 53 where the object 52 in the target image is located. The image data in the object region 53 is subjected to hue adjustment processing to obtain an adjusted object region 54, and the adjusted object region 54 is superimposed on the target image 51; at the same time, the target image 51 is masked based on the object mask map 55 to obtain a masked region indicated by the object mask map 55, and the masked region can be used as the object region 57 where the object 56 in the target image is located. The image data in the object region 57 is subjected to hue adjustment processing to obtain an adjusted object region 58, and the adjusted object region 58 is superimposed on the target image 51 to obtain a superimposed target image 59, and category labels corresponding to the object categories to which the highlighted objects 52 and 56 belong are generated in the superimposed target image 59.
[0086] Through the above method, object segmentation in an image can be achieved. The target object contained in the target image can be queried from the target image through the object query vector (at this time, the physical meaning of the target object is not determined), and an object feature vector containing the feature information of the target object can be output. An object mask map corresponding to the queried target object can be generated through the object feature vector. The object region of the target object in the target image can be determined through the object mask map, and the object category to which the target object belongs can be determined through the object feature vector. It can be understood that this method first separately queries the object regions of one or more target objects contained in the target image from the target image, and then separately determines the object categories to which the one or more queried target objects belong. Through the way of logical decoupling, various objects contained in the target image can be queried, and the actual object category of the queried object can be determined based on the object category set. The efficiency and flexibility of object segmentation can be improved, and at the same time, the robustness and usability of the image segmentation model can be improved.
[0087] In the embodiment of the present application, the recognition logics of the object position and object category of an object can be decoupled. First, the target object in the target image can be queried from the target image through the object query vector, and a corresponding object feature vector can be generated, and the object position and object category can be separately determined through the object feature vector. In addition, the object categories in the object category set can also be flexibly adjusted, such as adding or deleting some object categories, so that the object categories to which more types of objects belong can be recognized in the target image, and the accuracy and flexibility of object segmentation in the image can be improved.
[0088] Please refer to Figure 6 , Figure 6 which is a schematic flowchart of an image processing method provided by an embodiment of the present application. This method can be executed by the above-mentioned electronic device. As Figure 6 shown, the process of the image processing method in the embodiment of the present application can include the following:
[0089] S201. Obtain a target image containing a target object, and obtain an image feature vector of the target image. Among them, the specific implementation manner of step S201 can refer to the relevant description of the above embodiment, and will not be elaborated here.
[0090] S202. Obtain an object query vector for querying the target object from the target image, and generate an object feature vector associated with the target image based on the object query vector and the image feature vector.
[0091] Among them, the image feature vector includes R image feature sub-vectors, and one image feature sub-vector is used to characterize the image features of an image region in the target image. Therefore, determining the object feature vector associated with the target object can be based on the object query vector and each image feature sub-vector to respectively determine the region probability that the image region indicated by each image feature sub-vector belongs to the object region, and the object region is used to characterize the region of the target object in the target image; based on the region probabilities indicated by each image feature sub-vector, perform a fusion process on the R image feature sub-vectors and the object query vector to obtain the object feature vector. Among them, the object feature vector is used to characterize the object query information determined from the target image. It can be understood that the region probability that the image region indicated by any one image feature sub-vector belongs to the object region is the similarity between the object query vector and any one image feature sub-vector.
[0092] In some embodiments, performing a fusion process on the R image feature sub-vectors and the object query vector to obtain the object feature vector can specifically be based on the region probabilities indicated by each image feature sub-vector to respectively determine the to-be-fused feature vector corresponding to each image feature sub-vector from each image feature sub-vector, and perform a fusion process on the to-be-fused feature vector corresponding to each image feature sub-vector and the object query vector to obtain the object feature vector. It can be understood that determining the to-be-fused feature vector corresponding to any one image feature sub-vector can be the product vector obtained by multiplying the region probability indicated by any one image feature sub-vector with any one image feature sub-vector. In addition, the vector sum of the to-be-fused feature vector corresponding to each image feature sub-vector and the object query vector can be used as the object feature vector.
[0093] S203. Generate an object mask graph corresponding to the object feature vector based on the object feature vector and the image feature vector.
[0094] S204. If the object mask graph includes a target mask value for marking the target object, obtain the category encoding vector corresponding to each object category in the object category set, and determine the similarity between the object feature vector and the category encoding vector corresponding to each object category.
[0095] S205. Determine the object category to which the target object belongs from the object category set based on the similarity. Among them, the specific implementation manners of steps S203 - S205 can refer to the relevant descriptions in the above embodiments and will not be elaborated here.
[0096] S206. Highlight the target object in the target image according to a preset display manner, and perform an associated display on the highlighted target object and the object category to which it belongs.
[0097] In some embodiments, highlighting a target object in a target image may be to perform masking processing on the target image based on the target mask value in the object mask map, and use the mask area where the target mask value is located in the masked target image as the object area of the target object in the target image, and highlight the target object on the object area in the target image according to a preset display method.
[0098] Among them, highlighting the target object on the object area in the target image according to a preset display method may be to perform hue adjustment processing on the target object on the object area to obtain an adjusted object area, and superimpose the adjusted object area on the target image to obtain a target image including the highlighted target object. Optionally, the specific method for associatively displaying the highlighted target object and its object category can refer to the relevant descriptions in the above embodiments.
[0099] For example, as Figure 7 shown, Figure 7 is a schematic diagram of a scenario for target object segmentation provided by an embodiment of the present application; among them, the object segmentation process of the image may be: obtaining a target image 70, inputting the target image 70 into an image encoder 71 to obtain an image feature vector 72 (HxW image feature sub-vectors) of the target image 70, obtaining N object query vectors 73, and inputting the N object query vectors 73 and the image feature vector 72 into an image decoder 74 to obtain an object feature vector 75 corresponding to each object query vector 73; generating an object mask map 76 corresponding to each object feature vector 75 based on the object feature vector 75 corresponding to each object query vector 73 and the image feature vector 72 through a convolution kernel generator; if the object feature vector 1 associated with the target image 70 is determined from the image feature vector 72 based on the object query vector 1, and the object mask Figure 1 does not contain the target mask value (for example, all mask values are 0), it means that the object query vector 1 fails to query any object from the target image 70; or, if the object feature vector 2 associated with the target image 70 is determined from the image feature vector 72 based on the object query vector 2, and the object mask Figure 2If the target mask value is included, it indicates that the object query vector 2 queries the target object 77 from the target image 70. Therefore, the object feature vector 75 for querying any object can be selected from the N object feature vectors 75 based on the object mask map 76 corresponding to each object feature vector 75 (for example, denoted as the reference object feature vector 78); the object category set 79 is obtained, and the category text 710 associated with each object category in the object category set 79 (such as people, vehicles, cats, dogs, etc.) is input into the text encoder to obtain the category encoding vector 711 corresponding to each object category. The similarity 712 between each reference object feature vector 78 and the category encoding vector 771 corresponding to each object category is determined, and the object category 713 to which the target object 77 indicated by each reference object feature vector 78 belongs is determined based on this similarity 712. The target object 77 and the object category 713 to which it belongs can be associated and displayed in the target image 70; for example, the target image 70 is masked based on the target mask value in the object mask map 76 corresponding to the reference object feature vector 78 to obtain the object area where the target object 77 is located, and the target object 77 in the object area is subjected to hue adjustment processing. The adjusted object area is superimposed on the target image 70 to obtain the highlighted target object 77, and the highlighted target object 77 and the object category 713 to which it belongs are associated and displayed in the target image 70.
[0100] Among them, the above-mentioned object query vector, image encoder, image decoder, and convolution kernel generator can be obtained through model training of an image segmentation model. The training process can be as follows: Obtain a training sample image and a preset object category set. The training sample image carries an object mask map label and an object category label to which it belongs. N initial object query vectors are generated, and the text encoder is used to encode the category text associated with each object category in the object category set to obtain the category encoding vector corresponding to each object category. The training sample image is input into the image encoder in the image segmentation model to obtain an image feature vector. The image feature vector and the N initial object query vectors are input into the image decoder, and the object feature vector corresponding to the N initial object query vectors is output through the image decoder. The object feature vector is input into the convolution kernel generator to obtain a dynamic convolution kernel. An object mask map is generated through the dynamic convolution kernel and the object feature vector, and the object category to which it belongs is determined based on the category encoding vector and the object feature vector corresponding to each object category. The N initial object query vectors, as well as the image encoder, image decoder, and convolution kernel generator in the image segmentation model, are trained based on the object mask map, the corresponding object category to which it belongs, the object mask map label, and the object category label, so as to obtain the trained image segmentation model and the N object query vectors.
[0101] Therefore, by combining the object segmentation of the image segmentation model guided by the text encoder, that is, the location and shape of the objects contained in the target image can be queried through the object query vector, and the object category to which the object queried by the object query vector belongs can be determined by combining the category encoding vector, realizing the recognition and segmentation of the objects in the image. That is to say, this solution first proposes to use the category encoding vector obtained by the text encoder to guide the classification and segmentation of objects of any category, solving the problem that the traditional method can only classify a fixed set of categories. It can be understood that the preset object categories in the object category set can be updated according to different scenarios. For example, the object category set used in the training phase may only include 10 object categories, but in the application phase, object categories can be added to the object category set. For example, at this time, it can include 20 object categories. Therefore, through the above method, the object categories that can be recognized can be extended, and the ability to recognize objects of new categories is possessed, thereby improving the flexibility and accuracy of object segmentation.
[0102] In the embodiments of the present application, the recognition logics of the object location and object category of an object can be decoupled. First, the target object in the target image can be queried from the target image through the object query vector, and the corresponding object feature vector can be generated, and the object location and object category can be determined respectively through the object feature vector. In addition, the object categories in the object category set can also be flexibly adjusted, such as adding or deleting some object categories, so that the object categories to which more types of objects in the target image belong can be recognized, and the accuracy and flexibility of object segmentation in the image can be improved.
[0103] Please refer to Figure 8 , Figure 8 which is a schematic flowchart of an image processing method provided by the embodiments of the present application. This method can be executed by the above-mentioned electronic device. As Figure 8 shown, the process of the image processing method in the embodiments of the present application can include the following:
[0104] S301. Obtain a first target image including a first target object and a second target image including a second target object, and obtain a first image feature vector of the first target image and a second image feature vector of the second target image.
[0105] Among them, there can be one or more first target objects in the first target image. There can be one or more second target objects in the second target image. For example, the first target image and the second target image can be video frames intercepted from a video. An object movement trajectory can be generated based on the objects segmented from each video frame. Here, two adjacent video frames are taken as an example, that is, the first target image and the second target image.
[0106] S302. Obtain an object query vector.
[0107] S303. Generate a first object feature vector associated with the first target image based on the object query vector and the first image feature vector, and generate a second object feature vector associated with the second target image based on the object query vector and the second image feature vector.
[0108] S304. Generate a first object mask map corresponding to the first object feature vector based on the first object feature vector and the first image feature vector, and generate a second object mask map corresponding to the second object feature vector based on the second object feature vector and the second image feature vector. For the specific implementation manners of steps S302 - S304, reference can be made to the relevant descriptions in the above embodiments.
[0109] S305. Determine an associated target object that matches the first target object from the second target object based on the first object feature vector and the second object feature vector, and generate a movement trajectory of the associated target object based on the first object mask map and the second object mask map.
[0110] Among them, the first object feature vector associated with the first target image includes the first object feature vector corresponding to the first target object, and the second object feature vector associated with the second target image includes the second object feature vector corresponding to the second target object. It can be understood that the first object feature associated with the first target image includes the first object feature vector for querying and identifying the first target object and the first object feature vector for querying and identifying the image background. Therefore, the first object feature vector for querying and identifying the first target object can be determined from the first object feature vector associated with the first target image, that is, it can be determined whether the first object feature vector queries and identifies the first target object in the first target image through the first object mask map corresponding to the first object feature vector. It can be understood that when it is determined that a first object feature vector queries and identifies a first target object in the first target image, it means that the first object feature vector is the first object feature vector corresponding to a first target object. For example, the first object feature vector associated with the first target image includes the first object feature vectors corresponding to U first target objects, and the second object feature vector associated with the second target image includes the second object feature vectors corresponding to V second target objects; both U and V are positive integers. Correspondingly, each first target object corresponds to a first object mask map, and each second target object corresponds to a second object mask map. That is to say, for example, the number of first target objects queried and identified from the first target image through the object query vector is U, and the number of second target objects queried and identified from the second target image through the object query vector is V. Both U and V are positive integers. For example, there are N object query vectors, corresponding to N first object feature vectors, and N first object mask maps are obtained. U first object mask maps are determined from the N first object mask maps to identify the first target object from the first target image. Therefore, the first object feature vectors corresponding to the U first object mask maps are used as the first object feature vectors corresponding to the U first target objects.
[0111] In some embodiments, determining an associated target object may be based on the first object feature vectors corresponding to U first target objects and the second object feature vectors corresponding to V second target objects to determine the object relevance between each first target object and each second target object; based on the object relevance, determine a second target object that matches any one of the U first target objects from the V second target objects, and use the determined second target object as the associated target object. Among them, the vector similarity between the object feature vector corresponding to any one first object and the object feature vector corresponding to any one second object may be used as the object relevance between any one first object and any one second object. It can be understood that one object query vector obtains one first object feature vector, and one first object feature vector corresponds to one first object mask map. When the first object mask map indicates that the first target object is queried and recognized from the first target image, determine the first object feature vector corresponding to the first object mask map as the first object feature vector corresponding to one first target object.
[0112] Among them, determining a second target object that matches any one of the U first target objects from the V second target objects based on the object relevance may be to determine a second target object that matches each first target object from the V second target objects, and based on the Hungarian algorithm, determine a second target object that matches any one of the first target objects according to the second target objects that match each first target object. Among them, the second target object that matches any one of the first target objects refers to the second target object with the highest object relevance to any one of the first target objects among the V second target objects. It can be understood that if a first target object has a matching second target object, it means that the first target object and the matching second target object are the same target object, that is, the position of the first target object in the first target image moves to the position in the second target image. That is to say, the movement trajectory of each target object in the video can be obtained based on the position information of the same target object segmented in any two adjacent video frames.
[0113] In some embodiments, generating the movement trajectory of the associated target object based on the first object mask map and the second object mask map may be to determine the first position information of the first target object that matches the associated target object in the first target image based on the first object mask map, and determine the second position information of the associated target object in the second target image based on the second object mask map; generate the movement trajectory of the associated target object through the first position information and the second position information. The first object mask map corresponding to the first object feature vector includes the first object mask maps corresponding to U first target objects; the second object mask map corresponding to the second object feature vector includes the second object mask maps corresponding to U second target objects.
[0114] That is, the position information indicated by the first object mask graph corresponding to the first target object that matches the associated target object in the first target image is determined as the position information of the associated target object in the first target image, and the position information indicated by the second object mask graph corresponding to the associated target object in the second target image is determined as the position information of the associated target object in the second target image; the movement trajectory of the associated target object is generated based on the position information of the associated target object in the first target image and its position information in the second target image. It can be understood that the object area of the first target object that matches the associated target object in the first target image can be determined through the first object mask graph corresponding to the first target object that matches the associated target object to obtain the position information in the first target image. The object area of the second target object that matches the associated target object in the second target image can be determined through the second object mask graph corresponding to the associated target object to obtain the position information in the second target image. Thus, the movement trajectory of the associated target object in the first target image and the second target image can be generated. At the same time, when generating the movement trajectory, the object category to which the associated target object belongs can also be determined in the above manner, and the associated target object can be associated and displayed in the first target image and the second target image.
[0115] For example, as Figure 9 , Figure 9A schematic diagram of a scenario for target object segmentation provided by an embodiment of this application; wherein, a first target image and a second target image are obtained, a first image feature vector of the first target image and a second image feature vector of the second target image are obtained; N object query features are obtained; based on the N object query vectors and the first image feature vector, N first object feature vectors associated with the first target image are generated, and based on the N object query vectors and the second image feature vector, N second object feature vectors associated with the second target image are generated, based on the N first object feature vectors and the first image feature vector, N first object mask graphs corresponding to the first object feature vectors are generated, and based on the N second object feature vectors and the second image feature vector, N second object mask graphs corresponding to the second object feature vectors are generated, U first object mask graphs that identify the first target object are obtained from the N first object mask graphs, and the first object feature vectors corresponding to the U first object mask graphs are used as the first object feature vectors corresponding to the U first target objects (1, 2,..., U), V second object mask graphs that identify the second target object are obtained from the N second object mask graphs, and the second object feature vectors corresponding to the V second object mask graphs are used as the second object feature vectors corresponding to the V second target objects (1, 2,..., V); that is, U first target objects are segmented from the first target image and V second target objects are segmented from the second target image through the N object query vectors; an object correlation matrix between the U first target objects and the V second target objects can be constructed based on the similarity between the first object feature vector corresponding to each first target object and the second object feature vector corresponding to each second target object (wherein, Figure 9 in the object correlation matrix shown, "11" represents the object correlation between the first target object 1 and the second target object 2, "12" represents the object correlation between the first target object and the second target object; and so on for the rest), and through the Hungarian algorithm, the second target object that matches any first target object is determined based on this object correlation matrix. For example, the first target object 1 matches the second target object 2, and the first target object 2 matches the second target object 3. Therefore, the position information of the first target object 1 in the first target image can be determined based on the first object mask graph corresponding to the first target object 1, and the position information of the second target object 2 in the second target image can be determined based on the second object mask graph corresponding to the second target object 2 to generate the movement trajectory of the second target object 2. The position information of the first target object 2 in the first target image can be determined based on the first object mask graph corresponding to the first target object 2, and the position information of the second target object 3 in the second target image can be determined based on the second object mask graph corresponding to the second target object 3 to generate the movement trajectory of the second target object 3.
[0116] In the embodiments of the present application, the recognition logics of the object position and object category of an object can be decoupled. First, a target object in the target image can be queried from the target image through an object query vector, and a corresponding object feature vector can be generated. Moreover, the object position and the same target object can be determined through the object feature vector, so as to generate the movement trajectory of the target object in the first target image and the second target image.
[0117] Please refer to Figure 10 , Figure 10 which is a schematic structural diagram of an image processing device provided by the present application. It should be noted that Figure 10 the shown image processing device is used to execute the methods of the embodiments of the present application Figure 2 and Figure 6 shown. For the sake of convenience of description, only the parts related to the embodiments of the present application are shown, and the specific technical details are not disclosed. Please refer to the embodiments Figure 2 and Figure 6 shown by the present application. The image processing device 1000 may include: a first acquisition module 1001 and a first processing module 1002. Among them:
[0118] The first acquisition module 1001 is configured to acquire a target image including a target object and acquire an image feature vector of the target image;
[0119] The first processing module 1002 is configured to acquire an object query vector for querying the target object from the target image and generate an object feature vector associated with the target image based on the object query vector and the image feature vector;
[0120] The first processing module 1002 is further configured to generate an object mask map corresponding to the object feature vector based on the object feature vector and the image feature vector; the object mask map is used to mark the target object;
[0121] The first processing module 1002 is further configured to, based on the object mask map, acquire a category coding vector corresponding to each object category in the object category set and determine the similarity between the object feature vector and the category coding vector corresponding to each object category, so as to determine the object category to which the target object belongs from the object category set.
[0122] In some embodiments, the image feature vector includes R image feature sub-vectors, and one image feature sub-vector is used to characterize the image feature of an image region in the target image;
[0123] When the first processing module 1002 is configured to generate an object feature vector associated with the target image based on the object query vector and the image feature vector, it is specifically configured to:
[0124] Based on the object query vector and each image feature sub-vector, respectively determine the regional probability that the image region indicated by each image feature sub-vector belongs to the object region; the object region is used to characterize the region of the target object in the target image.
[0125] Based on the regional probabilities indicated by each image feature sub-vector, perform a fusion process on the R image feature sub-vectors and the object query vector to obtain an object feature vector.
[0126] In some embodiments, when the first processing module 1002 is used to perform a fusion process on the R image feature sub-vectors and the object query vector based on the regional probabilities indicated by each image feature sub-vector to obtain an object feature vector, it is specifically used for:
[0127] Based on the regional probabilities indicated by each image feature sub-vector, respectively determine the feature vector to be fused corresponding to each image feature sub-vector from each image feature sub-vector;
[0128] Perform a fusion process on the feature vector to be fused corresponding to each image feature sub-vector and the object query vector to obtain an object feature vector.
[0129] In some embodiments, the image feature vector includes R image feature sub-vectors, and one image feature sub-vector is used to characterize the image features of an image region in the target image;
[0130] When the first processing module 1002 is used to generate an object mask graph corresponding to the object feature vector based on the object feature vector and the image feature vector, it is specifically used for:
[0131] Perform a convolution process on each image feature sub-vector through the object feature vector to obtain a convolution result corresponding to each image feature sub-vector;
[0132] Based on the convolution result corresponding to each image feature sub-vector, determine the mask value corresponding to the image region indicated by each image feature sub-vector;
[0133] Generate an object mask graph through the mask value corresponding to the image region indicated by each image feature sub-vector.
[0134] In some embodiments, when the first processing module 1002 is used to obtain the category encoding vector corresponding to each object category in the object category set based on the object mask graph, it is specifically used for:
[0135] If the object mask graph includes a target mask value for marking the target object, then obtain the category encoding vector corresponding to each object category in the object category set;
[0136] If the target mask value for marking the target object is not included in the object mask map, the category encoding vectors corresponding to each object category in the object category set are not obtained.
[0137] In some embodiments, the first processing module 1002 is further configured to:
[0138] Based on the target mask value in the object mask map, perform mask processing on the target image, and use the mask area where the target mask value is located in the masked target image as the object area of the target object in the target image;
[0139] Highlight the target object on the object area in the target image according to a preset display method, and perform associated display on the highlighted target object and its belonging object category.
[0140] Wherein, for the specific implementation manners of the first acquisition module and the first processing module, reference may be made to the descriptions of the above embodiments, and details will not be elaborated here. It should be understood that the descriptions of the beneficial effects obtained by using the same method will not be elaborated either.
[0141] Please refer to Figure 11 , Figure 11 which is a schematic structural diagram of an image processing device provided by this application. It should be noted that Figure 11 the illustrated image processing device is used to execute the method of the embodiment shown in this application. For the sake of convenience of description, only the parts related to the embodiments of this application are shown, and the specific technical details are not disclosed. Please refer to the embodiments shown in this application Figure 8 Figure 8 Figure 8 The image processing device 1100 may include: a second acquisition module 1101 and a second processing module 1102. Wherein:
[0142] The second acquisition module 1101 is configured to acquire a first target image including a first target object and a second target image including a second target object, and acquire the image feature vector of the first target image and the image feature vector of the second target image;
[0143] The second processing module 1102 is configured to acquire an object query vector; the object query vector is used to query the first target object from the first target image and query the second target object from the second target image;
[0144] The second processing module 1102 is further configured to generate a first object feature vector associated with the first target image based on the object query vector and the first image feature vector, and generate a second object feature vector associated with the second target image based on the object query vector and the second image feature vector;
[0145] The second processing module 1102 is further configured to generate a first object mask map corresponding to the first object feature vector based on the first object feature vector and the first image feature vector, and generate a second object mask map corresponding to the second object feature vector based on the second object feature vector and the second image feature vector; the first object mask map is used to mark the first target object, and the second object mask map is used to mark the second target object;
[0146] The second processing module 1102 is further configured to determine, from the second target objects, associated target objects that match the first target object based on the first object feature vector and the second object feature vector, and generate a movement trajectory of the associated target object based on the first object mask map and the second object mask map.
[0147] Wherein, the first object feature vector associated with the first target image includes first object feature vectors corresponding to U first target objects, and the second object feature vector associated with the second target image includes second object feature vectors corresponding to V second target objects; both U and V are positive integers;
[0148] When the second processing module 1102 is configured to determine, from the second target objects, associated target objects that match the first target object based on the first object feature vector and the second object feature vector, it is specifically configured to:
[0149] Determine the object correlation between each first target object and each second target object based on the first object feature vectors corresponding to U first target objects and the second object feature vectors corresponding to V second target objects;
[0150] Based on the object correlation, determine, from the V second target objects, second target objects that match any one of the U first target objects, and use the determined second target objects as the associated target objects.
[0151] Wherein, when the second processing module 1102 is configured to generate a movement trajectory of the associated target object based on the first object mask map and the second object mask map, it is specifically configured to:
[0152] Determine first position information of the first target object that matches the associated target object in the first target image based on the first object mask map, and determine second position information of the associated target object in the second target image based on the second object mask map;
[0153] Generate a movement trajectory of the associated target object through the first position information and the second position information.
[0154] Wherein, for the specific implementation manners of the second acquisition module and the second processing module, reference may be made to the descriptions of the above embodiments, and details will not be elaborated herein. It should be understood that the descriptions of the beneficial effects obtained by using the same method will not be elaborated either.
[0155] Please refer to Figure 12 , Figure 12 , which is a schematic structural diagram of an electronic device provided by an embodiment of the present application. As Figure 12 shown, the electronic device 1200 includes: at least one processor 1201 and a memory 1202. Optionally, the electronic device may further include a network interface. Among them, data can be exchanged between the processor 1201, the memory 1202, and the network interface. The network interface is controlled by the processor 1201 for sending and receiving messages. The memory 1202 is used to store computer programs, and the computer programs include program instructions. The processor 1201 is used to execute the program instructions stored in the memory 1202. Among them, the processor 1201 is configured to call the program instructions to execute the above method.
[0156] The memory 1202 may include a volatile memory, such as a random-access memory (RAM); the memory 1202 may also include a non-volatile memory, such as a flash memory, a solid-state drive (SSD), etc.; the memory 1202 may further include a combination of the above types of memories.
[0157] The processor 1201 may be a central processing unit (CPU). In one embodiment, the processor 1201 may also be a Graphics Processing Unit (GPU). The processor 1201 may also be a combination of a CPU and a GPU. The processor 1201 may be used to call the device control application program stored in the memory 1202 to execute the above Figure 2 , Figure 6 and Figure 8 descriptions of the image processing method in the corresponding embodiments, and may also execute the descriptions of the image processing device in the corresponding embodiments of the foregoing Figure 10 and Figure 11 , which will not be elaborated here. In addition, the description of the beneficial effects of adopting the same method will not be elaborated either.
[0158] In specific implementation, the devices, processors, memories, etc. described in the embodiments of the present application may execute the implementation manners described in the above method embodiments, and may also execute the implementation manners described in the embodiments of the present application, which will not be elaborated here.
[0159] In an embodiment of the present application, a computer (readable) storage medium is further provided. The computer storage medium stores a computer program, and the computer program includes program instructions. When the program instructions are executed by a processor, the processor can execute some or all of the steps executed in the above method embodiment. Optionally, the computer storage medium can be volatile or non-volatile. The computer-readable storage medium mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function, etc.; the data storage area can store data created according to the use of the blockchain node, etc.
[0160] An embodiment of the present application provides a computer program product, which may include a computer program. When the computer program is executed by a processor, some or all of the steps in the above method can be implemented, and details are not described here.
[0161] As used herein, "a plurality of" means two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.
[0162] Those of ordinary skill in the art can understand that all or part of the processes of implementing the above method embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer storage medium, and the computer storage medium can be a computer-readable storage medium. When the program is executed, it can include the processes of the above method embodiments. Among them, the storage medium can be a magnetic disk, an optical disc, a read-only memory (ROM), or a random access memory (RAM), etc.
[0163] The above-disclosed are only some embodiments of the present application. Of course, the scope of rights of the present application cannot be limited thereby. Those of ordinary skill in the art can understand all or part of the processes of implementing the above embodiments, and the equivalent changes made according to the claims of the present application still fall within the scope covered by the present application.
Claims
1. An image processing method, characterized in that, the method includes: obtaining a target image including a target object, and obtaining an image feature vector of the target image; the image feature vector includes R image feature sub-vectors, and one image feature sub-vector is used to characterize the image feature of an image region in the target image; obtaining an object query vector for querying the target object from the target image, and respectively determining, based on the object query vector and each image feature sub-vector, the region probability that the image region indicated by each image feature sub-vector belongs to the object region, and performing a fusion process on the R image feature sub-vectors and the object query vector based on the region probability indicated by each image feature sub-vector to obtain an object feature vector associated with the target image; the object region is used to characterize the region of the target object in the target image; performing a convolution process on the image feature vector by using the object feature vector to obtain a convolution result, and determining a mask value corresponding to the image region indicated by the image feature vector in the target image based on the convolution result, so as to generate an object mask map corresponding to the object feature vector; the object mask map is used to mark the target object; based on the object mask map, obtaining a category coding vector corresponding to each object category in the object category set, and determining the similarity between the object feature vector and the category coding vector corresponding to each object category, so as to determine the object category to which the target object belongs from the object category set; the category coding vector is obtained by performing text coding on the category text associated with the object category.
2. The method according to claim 1, wherein the performing a fusion process on the R image feature sub-vectors and the object query vector based on the region probability indicated by each image feature sub-vector to obtain the object feature vector, includes: respectively determining, based on the region probability indicated by each image feature sub-vector, a to-be-fused feature vector corresponding to each image feature sub-vector from each image feature sub-vector; performing a fusion process on the to-be-fused feature vector corresponding to each image feature sub-vector and the object query vector to obtain the object feature vector.
3. The method according to claim 1, characterized in that, the image feature vector includes R image feature sub-vectors, and one image feature sub-vector is used to characterize the image feature of an image region in the target image; the performing a convolution process on the image feature vector by using the object feature vector to obtain a convolution result, and determining a mask value corresponding to the image region indicated by the image feature vector in the target image based on the convolution result, so as to generate an object mask map corresponding to the object feature vector, includes: performing a convolution process on each image feature sub-vector by using the object feature vector to obtain a convolution result corresponding to each image feature sub-vector; determining a mask value corresponding to the image region indicated by each image feature sub-vector based on the convolution result corresponding to each image feature sub-vector; Generate the object mask map based on the mask values corresponding to the image regions indicated by each of the image feature sub-vectors.
4. The method according to claim 1, wherein, obtaining, based on the object mask map, a category encoding vector corresponding to each object category in the object category set includes: if the mask values in the object mask map include target mask values for marking the target object, obtaining a category encoding vector corresponding to each object category in the object category set; if the mask values in the object mask map do not include target mask values for marking the target object, not obtaining a category encoding vector corresponding to each object category in the object category set.
5. The method according to claim 4, wherein, the method further includes: performing mask processing on the target image based on the target mask values in the object mask map, and taking the mask region where the target mask values are located in the masked target image as the object region of the target object in the target image; highlighting the target object on the object region in the target image according to a preset display manner, and associatively displaying the highlighted target object and the object category to which it belongs.
6. An image processing method, wherein, the method includes: obtaining a first target image including a first target object and a second target image including a second target object, and obtaining a first image feature vector of the first target image and a second image feature vector of the second target image; the first image feature vector includes at least one image feature sub-vector, and one image feature sub-vector included in the first image feature vector is used to characterize the image feature of an image region in the first target image, the second image feature vector includes at least one image feature sub-vector, and one image feature sub-vector included in the second image feature vector is used to characterize the image feature of an image region in the second target image; obtaining an object query vector; the object query vector is used to query the first target object from the first target image and to query the second target object from the second target image; determining, based on the object query vector and each image feature sub-vector included in the first image feature vector respectively, the regional probability that the image region indicated by each image feature sub-vector included in the first image feature vector belongs to the object region for the first target object, and performing fusion processing on at least one image feature sub-vector included in the first image feature vector and the object query vector based on the regional probability indicated by each image feature sub-vector included in the first image feature vector, to obtain a first object feature vector associated with the first target image; the object region for the first target object is used to characterize the region of the first target object in the first target image; Based on the object query vector and each image feature sub-vector included in the second image feature vector, respectively determine the region probability that the image region indicated by each image feature sub-vector included in the second image feature vector belongs to the object region for the second target object. Based on the region probabilities indicated by each image feature sub-vector included in the second image feature vector, perform a fusion process on at least one image feature sub-vector included in the second image feature vector and the object query vector to obtain the second object feature vector associated with the second target image; the object region for the second target object is used to characterize the region of the second target object in the second target image; Perform a convolution process on the first image feature vector through the first object feature vector to obtain a convolution result for the first object feature vector. Based on the convolution result for the first object feature vector, determine the mask value corresponding to the image region indicated by the first object feature vector in the first target image, so as to generate a first object mask map; the first object mask map is used to mark the first target object; Perform a convolution process on the second image feature vector through the second object feature vector to obtain a convolution result for the second object feature vector. Based on the convolution result for the second object feature vector, determine the mask value corresponding to the image region indicated by the second object feature vector in the second target image, so as to generate a second object mask map; the second object mask map is used to mark the second target object; Based on the first object feature vector and the second object feature vector, determine an associated target object that matches the first target object from the second target object, and generate a movement trajectory of the associated target object based on the first object mask map and the second object mask map.
7. The method according to claim 6, wherein, The first object feature vector associated with the first target image includes first object feature vectors corresponding to U first target objects, and the second object feature vector associated with the second target image includes second object feature vectors corresponding to V second target objects; both U and V are positive integers; The determining, from the second target objects, an associated target object that matches the first target object based on the first object feature vector and the second object feature vector includes: Based on the first object feature vectors corresponding to the U first target objects and the second object feature vectors corresponding to the V second target objects, determine the object correlation degree between each first target object and each second target object; Based on the object correlation degree, determine a second target object that matches any one of the U first target objects from the V second target objects, and use the determined second target object as the associated target object.
8. The method according to claim 6, wherein, The generating the movement trajectory of the associated target object based on the first object mask map and the second object mask map includes: Determine the first position information of the first target object matching the associated target object in the first target image based on the first object mask map, and determine the second position information of the associated target object in the second target image based on the second object mask map; Generate the movement trajectory of the associated target object based on the first position information and the second position information.
9. An image processing device, characterized in that, the device includes: A first acquisition module, configured to acquire a target image including a target object, and acquire an image feature vector of the target image; the image feature vector includes R image feature sub-vectors, and one image feature sub-vector is used to characterize the image feature of an image region in the target image; A first processing module, configured to acquire an object query vector for querying the target object from the target image, and respectively determine the region probability that the image region indicated by each image feature sub-vector belongs to the object region based on the object query vector and each image feature sub-vector, and perform a fusion process on the R image feature sub-vectors and the object query vector based on the region probability indicated by each image feature sub-vector to obtain an object feature vector associated with the target image; the object region is used to characterize the region of the target object in the target image; The first processing module is further configured to perform a convolution process on the image feature vector through the object feature vector to obtain a convolution result, and determine a mask value corresponding to the image region indicated by the image feature vector in the target image based on the convolution result, so as to generate an object mask map corresponding to the object feature vector; the object mask map is used to mark the target object; The first processing module is further configured to, based on the object mask map, acquire a category encoding vector corresponding to each object category in the object category set, and determine the similarity between the object feature vector and the category encoding vector corresponding to each object category, so as to determine the object category to which the target object belongs from the object category set; the category encoding vector is obtained by performing text encoding on the category text associated with the object category.
10. An image processing device, characterized in that, the device includes: A second acquisition module, configured to acquire a first target image including a first target object and a second target image including a second target object, and acquire a first image feature vector of the first target image and a second image feature vector of the second target image; the first image feature vector includes at least one image feature sub-vector, and one image feature sub-vector included in the first image feature vector is used to characterize the image feature of an image region in the first target image, the second image feature vector includes at least one image feature sub-vector, and one image feature sub-vector included in the second image feature vector is used to characterize the image feature of an image region in the second target image; A second processing module, configured to obtain an object query vector; the object query vector is used to query the first target object from the first target image and to query the second target object from the second target image; The second processing module is further configured to respectively determine, based on the object query vector and each image feature sub-vector included in the first image feature vector, a region probability that the image region indicated by each image feature sub-vector included in the first image feature vector belongs to an object region for the first target object, and perform a fusion process on at least one image feature sub-vector included in the first image feature vector and the object query vector based on the region probabilities indicated by each image feature sub-vector included in the first image feature vector, to obtain a first object feature vector associated with the first target image; the object region for the first target object is used to characterize the region of the first target object in the first target image; The second processing module is further configured to respectively determine, based on the object query vector and each image feature sub-vector included in the second image feature vector, a region probability that the image region indicated by each image feature sub-vector included in the second image feature vector belongs to an object region for the second target object, and perform a fusion process on at least one image feature sub-vector included in the second image feature vector and the object query vector based on the region probabilities indicated by each image feature sub-vector included in the second image feature vector, to obtain a second object feature vector associated with the second target image; the object region for the second target object is used to characterize the region of the second target object in the second target image; The second processing module is further configured to perform a convolution process on the first image feature vector through the first object feature vector, obtain a convolution result for the first object feature vector, and determine a mask value corresponding to the image region indicated by the first object feature vector in the first target image based on the convolution result for the first object feature vector, so as to generate a first object mask map; the first object mask map is used to mark the first target object; The second processing module is further configured to perform a convolution process on the second image feature vector through the second object feature vector, obtain a convolution result for the second object feature vector, and determine a mask value corresponding to the image region indicated by the second object feature vector in the second target image based on the convolution result for the second object feature vector, so as to generate a second object mask map; the second object mask map is used to mark the second target object; The second processing module is further configured to determine, based on the first object feature vector and the second object feature vector, an associated target object that matches the first target object from the second target object, and generate a movement trajectory of the associated target object based on the first object mask map and the second object mask map.
11. An electronic device Characterized in that It includes a processor and a memory. Among them, the memory is used to store a computer program, and the computer program includes program instructions. The processor is configured to call the program instructions to execute the method according to any one of claims 1-8.
12. A computer-readable storage medium, characterized in that, the computer-readable storage medium stores a computer program, and the computer program includes program instructions. When the program instructions are executed by a processor, the processor is caused to execute the method according to any one of claims 1-8.
13. A computer program product, characterized in that, the computer program product includes computer instructions, and the computer instructions are used to execute the method according to any one of claims 1-8 when executed by a processor.
Citation Information
Patent Citations
Target detection segmentation method and device, intelligent terminal and storage medium
CN115147598A