Image labeling method and device, electronic equipment and computer program product
By acquiring the image features and text prompt information of the target image under multiple preset image parameters, and combining the cross attention mechanism for feature fusion, the problem of low image labeling efficiency in the prior art is solved, and high accuracy and high efficiency image labeling is achieved.
Patent Information
- Application Number
- CN202411943079.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-26
- Publication Date
- 2025-05-16
AI Technical Summary
In the prior art, the image annotation method has low efficiency, the manual annotation is complex and the semi-automatic annotation-dependent model has low accuracy, resulting in the need of manual proofreading.
By obtaining the corresponding image features and text prompt information of the target image to be marked under multiple preset image parameters, and combining features with the cross attention mechanism to generate multimodal features for image segmentation and labeling processing.
The accuracy and efficiency of image annotation are improved without manual proofreading, and the output image annotation results are more accurate.
Smart Images

Figure CN120014644A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of image processing technology, and in particular relates to an image annotation method, device, electronic equipment and computer program product. Background Art
[0002] There are two main image annotation methods currently: one is manual annotation, and the other is semi-automatic annotation combined with deep learning. The manual annotation method relies on manual use of annotation tools to annotate the image to be annotated. Since the manual annotation process is relatively complicated, the image annotation efficiency is low; the semi-automatic annotation method can improve the annotation efficiency to a certain extent, but if the accuracy of the model used is not high, the semi-automatic annotation results need to be manually proofread, so the image annotation efficiency is also low.
[0003] It can be seen that the efficiency of the image annotation method in the prior art is low. Summary of the invention
[0004] In view of this, embodiments of the present application provide an image annotation method, device, electronic device, and computer program product to solve the technical problem of low efficiency of existing image annotation methods.
[0005] In a first aspect, an embodiment of the present application provides an image annotation method, comprising:
[0006] Obtaining image features corresponding to the target image to be annotated under multiple preset image parameters;
[0007] Acquire text prompt information of the target image, where the text prompt information is used to describe the features of the object in the target image;
[0008] The target image is segmented and annotated according to the image features and the text prompt information.
[0009] Optionally, the obtaining of image features corresponding to the target image to be labeled under a plurality of preset image parameters includes:
[0010] Performing downsampling processing and / or upsampling processing on the target image to obtain the target image under the plurality of preset image parameters;
[0011] According to the target image under the multiple preset image parameters, the image features corresponding to each of the multiple preset image parameters are determined.
[0012] Optionally, the obtaining of image features corresponding to the target image to be labeled under a plurality of preset image parameters includes:
[0013] Changing the resolution of the target image to obtain the target image under the plurality of preset image parameters;
[0014] According to the target image under the multiple preset image parameters, the image features corresponding to each of the multiple preset image parameters are determined.
[0015] Optionally, determining, according to the target image under the multiple preset image parameters, the image features corresponding to each of the multiple preset image parameters includes:
[0016] For each of the preset image parameters, the target image under the preset image parameters is input into a trained image encoder for processing, and the image encoder outputs the corresponding image features under the preset image parameters.
[0017] Optionally, the segmenting and labeling the target image according to the image features and the text prompt information includes:
[0018] Performing a first feature fusion process on the image features corresponding to each of the plurality of preset image parameters to obtain an overall image feature;
[0019] The target image is segmented and labeled according to the overall image features and the text prompt information.
[0020] Optionally, after acquiring the text prompt information of the target image, the method further includes:
[0021] Obtain input user demand information;
[0022] Generate semantic coding according to the text prompt information and the user demand information;
[0023] The segmenting and labeling of the target image according to the image features and the text prompt information includes:
[0024] Performing a second feature fusion process on the total image feature and the semantic code through a cross attention mechanism to obtain a multimodal feature;
[0025] Inputting the multimodal features into a trained segmentation decoder for processing, and outputting masks of a plurality of preset areas through the segmentation decoder;
[0026] The target image is segmented and labeled according to each of the masks.
[0027] Optionally, the image features include spatial information and semantic information of the target image under each of the preset image parameters; and the second feature fusion processing is performed on the image features and the semantic encoding by the cross attention mechanism to obtain the multimodal features, including:
[0028] The spatial information, the semantic information and the semantic coding are subjected to a second feature fusion process through a cross-attention mechanism to obtain the multimodal feature.
[0029] In a second aspect, an embodiment of the present application provides an image annotation device, comprising:
[0030] An image feature acquisition unit, used to acquire image features corresponding to the target image to be annotated under a plurality of preset image parameters;
[0031] A text prompt information acquisition unit, used to acquire text prompt information of the target image, wherein the text prompt information is used to describe the features of the object in the target image;
[0032] The labeling unit is used to segment and label the target image according to the image features and the text prompt information.
[0033] In a third aspect, an embodiment of the present application provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, each step of the image annotation method as described in any one of the first aspects above is implemented.
[0034] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, each step of the image annotation method described in any one of the first aspects above is implemented.
[0035] In a fifth aspect, an embodiment of the present application provides a computer program product. When the computer program product is executed on a terminal device, the terminal device executes each step of the image annotation method as described in any one of the first aspects above.
[0036] The image annotation method, device, electronic device, and computer program product provided by the embodiments of the present application have the following beneficial effects:
[0037] In the image annotation method provided in the embodiment of the present application, firstly, the image features corresponding to the target image to be annotated under multiple preset image parameters are obtained, and then the text prompt information of the target image is obtained, wherein the text prompt information is used to describe the object features in the target image, and finally, the target image is segmented and annotated according to the image features and the text prompt information. The image annotation method of the present application enables the electronic device to segment and annotate the target image according to the extracted image features corresponding to the target image under multiple preset image parameters and the text prompt information of the target image, so the accuracy of the image annotation result output by the electronic device is high, and no manual proofreading is required, thereby improving the efficiency of image annotation. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0039] Figure 1 A flowchart of an implementation method of an image annotation method provided in an embodiment of the present application;
[0040] Figure 2 A schematic diagram of the structure of an image annotation device provided in an embodiment of the present application;
[0041] Figure 3 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0042] It should be noted that the terms used in the embodiments of the present application are only used to explain the specific embodiments of the present application, and are not intended to limit the present application. In the description of the embodiments of the present application, unless otherwise specified, "multiple" refers to two or more than two, and "at least one", "one or more" refers to one, two or more. The terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the number of indicated technical features. Thus, it is defined that the "first" and "second" features can explicitly or implicitly include one or more of the features.
[0043] References to "one embodiment" or "some embodiments" etc. described in this specification mean that a particular feature, structure or characteristic described in conjunction with the embodiment is included in one or more embodiments of the present application. Thus, the phrases "in one embodiment", "in some embodiments", "in some other embodiments", "in some other embodiments", etc. that appear at different places in this specification do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "including", "comprising", "having" and their variations all mean "including but not limited to", unless otherwise specifically emphasized in other ways.
[0044] The image annotation method provided in the embodiment of the present application can be applied to electronic devices, wherein the electronic devices may include but are not limited to image annotation devices, mobile phones, tablet computers, laptop computers, desktop computers, cloud servers and other devices.
[0045] The image annotation method provided in the embodiment of the present application can be applied to various scenarios where target images need to be annotated. When a user needs to annotate a target image, the various steps of the image annotation method provided in the embodiment of the present application can be executed by an electronic device, so that the electronic device can accurately output the annotation result of the target image, thereby eliminating the need for manual proofreading of the annotation result of the target image output by the electronic device, and ultimately improving the image annotation efficiency.
[0046] See also Figure 1 , Figure 1 The present invention provides an implementation flow chart of an image annotation method according to an embodiment of the present invention. The image annotation method may include S101 to S103, which are described in detail as follows:
[0047] In S101, image features corresponding to the target image to be annotated under a plurality of preset image parameters are obtained.
[0048] In an embodiment of the present application, when it is necessary to annotate a target image, the electronic device may first obtain image features corresponding to the target image to be annotated under a plurality of preset image parameters.
[0049] The electronic device may acquire the target image through a photographing device, or the user may input the target image into the electronic device so that the electronic device may acquire the target image.
[0050] The generation method and specific composition of the preset image parameters can be set by the user according to actual needs and input into the electronic device so that the electronic device can obtain various preset image parameters.
[0051] In one possible implementation, after acquiring the target image, the electronic device may downsample and / or upsample the target image to obtain the target image under multiple preset image parameters. After obtaining the target image under multiple preset image parameters, the electronic device may determine the image features corresponding to each of the multiple preset image parameters based on the target image under the multiple preset image parameters.
[0052] In another possible implementation, after acquiring the target image, the electronic device may change the resolution of the target image, thereby obtaining the target image under multiple preset image parameters. After obtaining the target image under multiple preset image parameters, the electronic device may determine the image features corresponding to each of the multiple preset image parameters based on the target image under the multiple preset image parameters.
[0053] It should be noted that the electronic device can also perform any image processing on the target image to obtain the target image under multiple preset image parameters, which will not be described one by one here.
[0054] In a possible implementation, after obtaining the target images under multiple preset image parameters, the electronic device may determine the image features corresponding to each of the multiple preset image parameters according to the target images under the multiple preset image parameters in the following manner, as described in detail below:
[0055] For each preset image parameter, the electronic device can input the target image under the preset image parameter into the trained image encoder for processing, and output the corresponding image features under the preset image through the image encoder.
[0056] The image encoder can be used to generate and output image features corresponding to the target image under any preset image parameters based on the received target image under the preset image parameters after receiving the target image under the preset image parameters.
[0057] In a possible implementation, the image features corresponding to the target image under the preset image parameters output by the image encoder may include spatial information and semantic information of the target image under the preset image parameters. The spatial information of the target image can be used to describe the position information of the target object in the target image, and can be used to describe the relative position information between each target object in the target image. The semantic information of the target image can be used to describe the object category of the target object in the target image, and / or the semantic information of the target image can be used to describe the object category corresponding to each pixel in the target image.
[0058] Among them, the image encoder can be obtained by pre-acquiring training data by the electronic device, and training the initial image encoder according to the training data, wherein the electronic device can obtain training data through a variety of channels and methods to ensure the diversity and comprehensiveness of the training data. Exemplarily, the channels and methods for obtaining training data may include but are not limited to the following: a method of collecting training data by oneself, a method of obtaining a public training data set, and a method of a web crawler, etc. After obtaining training data from a variety of channels and methods, the electronic device can perform data cleaning processing on the training data to improve the quality of the training data. Among them, data cleaning and combing are specifically used to eliminate low-quality and invalid training data.
[0059] After the training data is obtained, the initial image encoder can be trained using the training data. In practical applications, the specific method of training the initial image encoder using the training data can be set according to actual needs and is not limited here.
[0060] In S102, text prompt information of the target image is obtained, where the text prompt information is used to describe the features of the object in the target image.
[0061] In the embodiment of the present application, in addition to obtaining the image features corresponding to the target image to be labeled under multiple preset image parameters, the electronic device also needs to obtain text prompt information of the target image to be labeled.
[0062] It should be noted that the electronic device may first obtain the image features corresponding to the target image to be labeled under multiple preset image parameters, and then obtain the text prompt information of the target image to be labeled; the electronic device may first obtain the text prompt information of the target image to be labeled, and then obtain the image features corresponding to the target image to be labeled under multiple preset image parameters; the electronic device may also simultaneously obtain the image features corresponding to the target image to be labeled under multiple preset image parameters and the text prompt information of the target image to be labeled.
[0063] In a possible implementation, the electronic device may input the target image into a text prompt information output model, so that the text prompt information output model can output text prompt information of the target image.
[0064] The text prompt information can be used to describe the features of the object in the target image, and the text prompt information can be automatically generated by the electronic device according to the target image to be annotated. Exemplarily, the electronic device can generate text prompt information according to the target image to be annotated by using a pre-trained text prompt information output model. Exemplarily, the text prompt information of a target image can be "the target image includes two people, one of whom is a man and the other is a girl, and the target image also includes a car, a tree, a road and a sun."
[0065] In practical applications, the text prompt information output model may be trained in advance using training data pairs, so that the text prompt information output model outputs text prompt information corresponding to the target image after receiving the target image.
[0066] In S103, the target image is segmented and annotated according to the image features and the text prompt information.
[0067] In an embodiment of the present application, after obtaining the image features corresponding to the target image to be labeled under multiple preset image parameters and the text prompt information of the target image, the electronic device can segment and label the target image according to the image features corresponding to the target image to be labeled under multiple preset image parameters and the text prompt information of the target image.
[0068] In a possible implementation, after acquiring the image features corresponding to the respective preset image parameters, the electronic device may perform a first feature fusion process on the image features corresponding to the respective preset image parameters to obtain the overall image features.
[0069] By performing the first feature fusion processing on the image features corresponding to each preset image parameter, the electronic device can obtain the global information and local information of the target image under each preset image parameter, so that the image annotation result output by the electronic device can be more accurate.
[0070] In this implementation, the electronic device can perform a first feature fusion process on the image features corresponding to each preset image parameter through a pre-trained feature fuser to obtain an overall image feature.
[0071] After obtaining the overall image features, the electronic device can segment and annotate the target image according to the overall image features and text prompt information.
[0072] In an embodiment of the present application, after acquiring the text prompt information of the target image, the electronic device may also acquire the user demand information input by the user, and generate a semantic code according to the text prompt information and the user demand information.
[0073] The user demand information may be used to describe the target information and specific task information that the user wants to focus on. For example, the user demand information may be “marking the man and the girl in the target image with red and green, respectively”.
[0074] Exemplarily, the electronic device may obtain semantic coding according to text prompt information and user demand information through a pre-trained semantic encoder.
[0075] Based on this, after obtaining the overall image features and semantic coding, the electronic device can segment and annotate the target image according to the overall image features and semantic coding.
[0076] Specifically, the electronic device can perform a second feature fusion process on the total image features and the semantic coding through a cross-attention mechanism to obtain a multimodal feature. Specifically, the electronic device can strengthen the interaction between the total image features and the semantic coding through a cross-attention mechanism, thereby establishing a cross-modal relationship, so that the electronic device can not only understand the image content in the target image through the total image features, but also understand the contextual meaning of the text through the semantic coding, thereby improving the accuracy of the electronic device in image annotation.
[0077] In this implementation, the electronic device can perform the second feature fusion processing on the total image features and the semantic coding through a pre-trained feature fusion device. In practical applications, one feature fusion device can be used to perform the first feature fusion processing and the second feature fusion processing, or two feature fusion devices can be used to perform the first feature fusion processing and the second feature fusion processing respectively.
[0078] After obtaining the multimodal features, the electronic device can input the multimodal features into a trained segmentation decoder for processing, so that the segmentation decoder can output masks of several preset areas, and the electronic device can obtain the masks of each preset area output by the segmentation decoder.
[0079] After obtaining the masks of each preset area, the electronic device can segment and annotate the target image according to the masks of each preset area. In practical applications, the specific method of segmenting and annotating the target image according to the masks of each preset area can be set according to actual needs and is not specifically limited here.
[0080] The following provides an example to describe the specific steps of the image annotation method provided in the embodiment of the present application.
[0081] The electronic device may first obtain training data through various channels and methods. After obtaining the training data from various channels and methods, the electronic device may perform data cleaning on the training data to improve the quality of the training data. Afterwards, the electronic device may use the training data to train the initial image encoder, the initial text prompt information output model, the initial feature fusion device, the initial semantic encoder, and the initial segmentation decoder, thereby obtaining a trained image encoder, a trained text prompt information output model, a trained feature fusion device, a trained semantic encoder, and a trained segmentation decoder.
[0082] Afterwards, the electronic device can obtain the target image to be annotated, and can obtain the target images under multiple preset image parameters based on the target image to be annotated. For each preset image parameter, the target image under the preset image parameter can be input into the trained image encoder for processing, and the image encoder can output the corresponding image features under the preset image, so as to obtain the image features corresponding to the target image to be annotated under the multiple preset image parameters.
[0083] In addition to obtaining the image features corresponding to the target image to be annotated under multiple preset image parameters, the electronic device can also obtain the text prompt information of the target image. Specifically, the electronic device can input the target image into the trained text prompt information output model so that the text prompt information output model can output the text prompt information of the target image.
[0084] After obtaining the image features and text prompt information corresponding to the target image under multiple preset image parameters, the electronic device can segment and annotate the target image according to the image features and text prompt information. Specifically, the electronic device can input the image features corresponding to each preset image parameter into a trained feature fusion device, so that the feature fusion device outputs the overall image features. In addition, the electronic device can also obtain user demand information input by the user, and input the user demand information and text prompt information into a pre-trained semantic encoder, so that the semantic encoder outputs the semantic code.
[0085] After acquiring the total image features and semantic coding, the electronic device can segment and annotate the target image according to the total image features and semantic coding. Specifically, the electronic device can input the total image features and semantic coding into a trained feature fusion device so that the feature fusion device outputs multimodal features.
[0086] After acquiring the multimodal features, the electronic device can input the multimodal features into a trained segmentation decoder to enable the segmentation decoder to output masks of several preset areas. The electronic device can segment and annotate the target image according to the masks of the preset areas.
[0087] It can be seen from the above that in the image annotation method provided in the embodiment of the present application, the image features corresponding to the target image to be annotated under multiple preset image parameters are first obtained, and then the text prompt information of the target image is obtained, wherein the text prompt information is used to describe the object features in the target image, and finally the target image is segmented and annotated according to the image features and the text prompt information. The image annotation method of the present application enables the electronic device to segment and annotate the target image according to the extracted image features corresponding to the target image under multiple preset image parameters and the text prompt information of the target image, so the accuracy of the image annotation result output by the electronic device is high, and no manual proofreading is required, thereby improving the efficiency of image annotation.
[0088] Based on the image annotation method provided in the above embodiment, the present application further provides an image annotation device for implementing the above method embodiment. Figure 2 , Figure 2 This is a schematic diagram of the structure of an image annotation device provided in an embodiment of the present application. Figure 2 As shown, the image annotation device 20 may include: an image feature acquisition unit 21, a text prompt information acquisition unit 22, and an annotation unit 23. Among them:
[0089] The image feature acquisition unit 21 is used to acquire the image features corresponding to the target image to be labeled under a plurality of preset image parameters.
[0090] The text prompt information acquisition unit 22 is used to acquire text prompt information of the target image, where the text prompt information is used to describe the features of the object in the target image.
[0091] The labeling unit 23 is used to segment and label the target image according to the image features and text prompt information.
[0092] Optionally, the image feature acquisition unit 21 is specifically used for:
[0093] Performing downsampling and / or upsampling processing on the target image to obtain the target image under a plurality of preset image parameters;
[0094] According to the target image under the multiple preset image parameters, the image features corresponding to each of the multiple preset image parameters are determined.
[0095] Optionally, the image feature acquisition unit 21 is specifically used for:
[0096] Changing the resolution of the target image to obtain the target image under multiple preset image parameters;
[0097] According to the target image under the multiple preset image parameters, the image features corresponding to each of the multiple preset image parameters are determined.
[0098] Optionally, the image feature acquisition unit 21 is specifically used for:
[0099] For each preset image parameter, the target image under the preset image parameter is input into the trained image encoder for processing, and the image encoder outputs the corresponding image features under the preset image parameter.
[0100] Optionally, the marking unit 23 is specifically used for:
[0101] Performing a first feature fusion process on the image features corresponding to each of the plurality of preset image parameters to obtain an overall image feature;
[0102] According to the overall image features and text prompt information, the target image is segmented and labeled.
[0103] Optionally, the image annotation device 20 may further include a semantic encoding unit.
[0104] The semantic coding unit is specifically used for:
[0105] Obtain input user demand information;
[0106] Generate semantic coding based on text prompt information and user demand information.
[0107] Based on this, the marking unit 23 is specifically used for:
[0108] The image features and semantic encoding are fused into a second feature through a cross-attention mechanism to obtain multimodal features.
[0109] The multimodal features are input into the trained segmentation decoder for processing, and the segmentation decoder outputs masks of several preset areas;
[0110] According to each mask, the target image is segmented and labeled.
[0111] Optionally, the image features include spatial information and semantic information of the target image under various preset image parameters. Based on this, the labeling unit 23 is specifically used to:
[0112] The spatial information, semantic information and semantic encoding are fused into a second feature through the cross-attention mechanism to obtain multimodal features.
[0113] It should be noted that the information interaction, execution process and other contents between the above-mentioned units are based on the same concept as the method embodiment of the present application. Their specific functions and technical effects can be specifically referred to the method embodiment part and will not be repeated here.
[0114] See also Figure 3 , Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. Figure 3 As shown, the electronic device 3 provided in this embodiment may include: a processor 30, a memory 31, and a computer program 32 stored in the memory 31 and executable on the processor 30, such as a program corresponding to the image annotation method. When the processor 30 executes the computer program 32, the steps applied to the image annotation method embodiment are implemented, such as Figure 1 Alternatively, when the processor 30 executes the computer program 32, the functions of each module / unit in the above-mentioned image annotation device embodiment are realized, for example Figure 2 The functions of the units 21 to 23 are shown.
[0115] Exemplarily, the computer program 32 may be divided into one or more modules / units, one or more modules / units are stored in the memory 31 and executed by the processor 30 to complete the present application. One or more modules / units may be a series of computer program instruction segments that can complete specific functions, and the instruction segments are used to describe the execution process of the computer program 32 in the electronic device 3. For example, the computer program 32 may be divided into an image feature acquisition unit 21, a text prompt information acquisition unit 22, and a labeling unit 23. For the specific functions of each unit, please refer to Figure 2 The relevant descriptions in the corresponding embodiments are not repeated here.
[0116] Those skilled in the art will understand that Figure 3 This is only an example of the electronic device 3 and does not constitute a limitation on the electronic device 3 , which may include more or less components than those shown in the figure, or a combination of certain components, or different components.
[0117] The processor 30 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.
[0118] The memory 31 may be an internal storage unit of the electronic device 3, such as a hard disk or memory of the electronic device 3. The memory 31 may also be an external storage device of the electronic device 3, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, or a flash card, etc., equipped on the electronic device 3. Further, the memory 31 may also include both an internal storage unit and an external storage device of the electronic device 3. The memory 31 is used to store computer programs and other programs and data required by the electronic device. The memory 31 may also be used to temporarily store data that has been output or is to be output.
[0119] Those skilled in the art can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units is used as an example. In actual applications, the above-mentioned functions can be assigned to different functional units as needed, that is, the internal structure of the image annotation device can be divided into different functional units to complete all or part of the functions described above. The functional units in the embodiment can be integrated into a processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, which will not be repeated here.
[0120] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments can be implemented.
[0121] An embodiment of the present application provides a computer program product. When the computer program product is executed on a terminal device, the terminal device implements the steps in the above-mentioned various method embodiments.
[0122] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0123] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0124] The embodiments described above are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, a person skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. Such modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. An image annotation method, characterized in that: include: Obtaining image features corresponding to the target image to be annotated under multiple preset image parameters; Acquire text prompt information of the target image, where the text prompt information is used to describe the features of the object in the target image; The target image is segmented and annotated according to the image features and the text prompt information.
2. The method according to claim 1, characterized in that The step of obtaining the image features corresponding to the target image to be labeled under a plurality of preset image parameters includes: Performing downsampling processing and / or upsampling processing on the target image to obtain the target image under the plurality of preset image parameters; According to the target image under the multiple preset image parameters, the image features corresponding to each of the multiple preset image parameters are determined.
3. The method according to claim 1, characterized in that The step of obtaining the image features corresponding to the target image to be labeled under a plurality of preset image parameters includes: Changing the resolution of the target image to obtain the target image under the plurality of preset image parameters; According to the target image under the multiple preset image parameters, the image features corresponding to each of the multiple preset image parameters are determined.
4. The method according to claim 2 or 3, characterized in that: The determining, according to the target image under the plurality of preset image parameters, the image features corresponding to each of the plurality of preset image parameters comprises: For each of the preset image parameters, the target image under the preset image parameters is input into a trained image encoder for processing, and the image encoder outputs the corresponding image features under the preset image parameters.
5. The method according to claim 1, characterized in that The segmenting and labeling of the target image according to the image features and the text prompt information includes: Performing a first feature fusion process on the image features corresponding to each of the plurality of preset image parameters to obtain an overall image feature; The target image is segmented and labeled according to the overall image features and the text prompt information.
6. The method according to claim 5, characterized in that After acquiring the text prompt information of the target image, the method further includes: Obtain input user demand information; Generate semantic coding according to the text prompt information and the user demand information; The segmenting and labeling of the target image according to the image features and the text prompt information includes: Performing a second feature fusion process on the total image feature and the semantic code through a cross attention mechanism to obtain a multimodal feature; Inputting the multimodal features into a trained segmentation decoder for processing, and outputting masks of a plurality of preset areas through the segmentation decoder; The target image is segmented and labeled according to each of the masks.
7. The method according to claim 6, characterized in that The image features include spatial information and semantic information of the target image under each of the preset image parameters; the second feature fusion processing is performed on the image features and the semantic coding through the cross attention mechanism to obtain multimodal features, including: The spatial information, the semantic information and the semantic coding are subjected to a second feature fusion process through a cross-attention mechanism to obtain the multimodal feature.
8. An image annotation device, characterized in that: include: An image feature acquisition unit, used to acquire image features corresponding to the target image to be annotated under multiple preset image parameters; A text prompt information acquisition unit, used to acquire text prompt information of the target image, wherein the text prompt information is used to describe the features of the object in the target image; The labeling unit is used to segment and label the target image according to the image features and the text prompt information.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, each step of the image annotation method according to any one of claims 1 to 7 is implemented.
10. A computer program product, characterized in that When the computer program product is executed by a processor, each step of the image annotation method according to any one of claims 1 to 7 is implemented.