Image processing method and device, readable medium and electronic equipment
By integrating the tasks of generating and understanding referential expressions into a target image processing model, the problem of low model integration was solved, and the resource utilization and task accuracy of the computer system were improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING YOUZHUJU NETWORK TECH CO LTD
- Filing Date
- 2022-09-29
- Publication Date
- 2026-05-29
AI Technical Summary
In existing technologies, the integration of referential expression generation and referential expression understanding models is not high, resulting in low utilization of computer system resources.
By integrating the tasks of generating and understanding referential expressions through a target image processing model, and utilizing a visual encoder, a text encoder, a fusion encoder, a location detection network, and a text prediction network, a unified modeling of generating and understanding referential expressions is achieved.
It improved the integration of the model, increased the resource utilization of the computer system, and enhanced the accuracy of referential expression generation and referential expression comprehension.
Smart Images

Figure CN115578570B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of image processing technology, and more specifically, to an image processing method, apparatus, readable medium, and electronic device. Background Technology
[0002] Referential representation (also known as referential expression) generation refers to generating a natural language description for a specified object in a given image. This description accurately describes the object and clearly distinguishes it from other objects in the image. Referential representation understanding (also known as referential expression localization) is the identification of the referred object from a given image based on a natural language description. Referential representation generation and referential representation understanding are two highly related tasks. However, in related technologies, the models used for referential representation generation and referential representation understanding suffer from low integration, which is detrimental to improving the resource utilization of computer systems. Summary of the Invention
[0003] This summary section is provided to briefly introduce the concepts, which will be described in detail in the detailed description section below. This summary section is not intended to identify key or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.
[0004] This disclosure provides an image processing method, apparatus, readable medium, and electronic device.
[0005] In a first aspect, this disclosure provides an image processing method, the method comprising:
[0006] Acquire target detection data, which includes an image to be detected and image description information of the image to be detected. The image description information is either specified location information of the generated region in the image to be detected, or descriptive text of the object to be detected in the image to be detected.
[0007] The target detection data is input into the target image processing model to obtain the image processing result output by the target image processing model. The target image processing model is used to output the index expression text of the image corresponding to the index expression generation region when the image description information is the specified location information, and to output the target position of the index expression object in the image to be detected when the image description information is the description text.
[0008] Secondly, this disclosure provides an image processing apparatus, the apparatus comprising:
[0009] The acquisition module is configured to acquire target detection data, which includes an image to be detected and image description information of the image to be detected. The image description information is either specified location information of the generated area in the image to be detected, or descriptive text of the object to be detected in the image to be detected.
[0010] The determination module is configured to input the target detection data into a target image processing model to obtain the image processing result output by the target image processing model. The target image processing model is configured to output the indicator text of the image corresponding to the indicator generation region when the image description information is the specified location information, and output the target position of the indicator object in the image to be detected when the image description information is the description text.
[0011] Thirdly, this disclosure provides a computer-readable medium having a computer program stored thereon, which, when executed by a processing device, implements the steps of the method described in the first aspect above.
[0012] Fourthly, this disclosure provides an electronic device, comprising:
[0013] A storage device on which computer programs are stored;
[0014] A processing device for executing the computer program in the storage device to implement the steps of the method described in the first aspect above.
[0015] The above technical solution acquires target detection data, which includes an image to be detected and image description information of the image to be detected. The image description information is either the specified location information of the referential expression generation region in the image to be detected, or the descriptive text of the referential expression object in the image to be detected. The target detection data is input into a target image processing model to obtain the image processing result output by the target image processing model. The target image processing model is used to output the referential expression text corresponding to the image of the referential expression generation region when the image description information is the specified location information, and to output the target location of the referential expression object in the image to be detected when the image description information is the descriptive text. In this way, by integrating the two tasks of referential expression generation and referential expression understanding through the target image processing model, the integration degree of the model can be effectively improved in scenarios that require both referential expression generation and referential expression understanding, thereby effectively improving the resource utilization of the computer system.
[0016] Other features and advantages of this disclosure will be described in detail in the following detailed description section. Attached Figure Description
[0017] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale. In the drawings:
[0018] Figure 1 This is a flowchart illustrating an exemplary embodiment of the present disclosure of an image processing method;
[0019] Figure 2 This is a schematic diagram illustrating the structure of a target image processing model according to an exemplary embodiment of the present disclosure;
[0020] Figure 3 It is based on Figure 1 and Figure 2 The illustrated embodiment shows a flowchart of an image processing method;
[0021] Figure 4 It is based on Figure 1 and Figure 2 The illustrated embodiment shows a flowchart of another image processing method;
[0022] Figure 5 This is a flowchart illustrating a model training method according to an exemplary embodiment of the present disclosure;
[0023] Figure 6 This is a block diagram illustrating an image processing apparatus according to an exemplary embodiment of the present disclosure;
[0024] Figure 7 This is a block diagram illustrating an electronic device according to an exemplary embodiment of the present disclosure. Detailed Implementation
[0025] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0026] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.
[0027] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.
[0028] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.
[0029] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0030] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0031] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0032] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.
[0033] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0034] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.
[0035] Meanwhile, it is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.
[0036] Figure 1 This is a flowchart illustrating an exemplary embodiment of the present disclosure of an image processing method; as shown below. Figure 1 As shown, the method may include:
[0037] Step 101: Obtain target detection data. The target detection data includes the image to be detected and the image description information of the image to be detected. The image description information is either the specified location information of the generated region in the image to be detected, or the descriptive text of the object in the image to be detected.
[0038] The region to be generated by the pointer can be any designated region in the image to be detected, and the object to be represented by the pointer can be one or more of the multiple objects captured in the image to be detected.
[0039] It should be noted that in the application scenario of generating referential expressions, the image description information is the specified location information of the referential expression generation region in the image to be detected; in the scenario of interpreting referential expressions (i.e., locating or segmenting referential expressions), the image description information is the descriptive text of the referential expression object.
[0040] Step 102: Input the target detection data into the target image processing model to obtain the image processing result output by the target image processing model. The target image processing model is used to output the index expression text of the image corresponding to the index expression generation region when the image description information is the specified location information, and to output the target position of the index expression object in the image to be detected when the image description information is the description text.
[0041] The target image processing model may include a visual encoder, a text encoder, a fusion encoder, a location detection network, and a text prediction network. Figure 2 This is a schematic diagram illustrating the structure of a target image processing model according to an exemplary embodiment of this disclosure; as shown below. Figure 2 As shown, the output of the visual encoder is coupled to the input of the fusion encoder, the output of the text encoder is coupled to the input of the fusion encoder, the output of the fusion encoder is coupled to the input of the position detection network, and the output of the fusion encoder is also coupled to the input of the text prediction network.
[0042] It should be noted that the visual encoder can be an image extraction network based on ViT (Vision Transformer, a Transformer network applied to visual tasks) and initialized with the weights of CLIP-ViT[2] (Contrastive Language–Image Pre-training Vision Transformer, a cross-modal pre-trained model based on contrastive image-text learning). The visual encoder can uniformly divide the image to be detected into non-overlapping patches. Each patch can be regarded as a word in NLP (Natural Language Processing). Then, these patches are flattened into a sequence, and the embeddings corresponding to the segmented patches are input into the stacked Transformer encoder blocks. Through the self-attention interaction in the Transformer encoder, image features are generated. Where L I This refers to the number of image patches. The text encoder can be a text feature extraction network based on BERT (Bidirectional Encoder Representations from Transformer), used to segment the input descriptive text representing the object into multiple token embeddings, and then transform these multiple input embeddings into text features. Where L T It is the number of input token embeddings, z [cls]This refers to the text features corresponding to the special marker [cls]. The fusion encoder can include multiple decoding layers and multiple fusion modules. The multiple decoding layers are connected in series, and the multiple fusion modules are also connected in series. The input of the first decoding layer serves as the input of the fusion encoder. The output of the last decoding layer is coupled to the input of the first fusion module. The output of the last fusion module is the output of the fusion encoder. The decoding layer is a Dencoder in Transformer. The fusion module can include a self-attention layer, an image mutual attention layer, a region mutual attention layer, and a region prediction layer. The input of the self-attention layer serves as the input of the fusion module. The output of the image mutual attention layer is coupled to the input of the image mutual attention layer, and the output of the image mutual attention layer is coupled to the input of the region prediction layer. The output of the image mutual attention layer is also coupled to the input of the region mutual attention layer, and the output of the region prediction layer is coupled to the input of the region mutual attention layer. The output of the region mutual attention layer serves as the output of the fusion module. Both the image mutual attention layer and the region mutual attention layer are formed by self-attention networks. The input of the image mutual attention layer is the output of the self-attention layer and the output of the visual encoder. The input of the region mutual attention layer is the output of the image mutual attention layer, a portion of the visual encoder's output (the portion of image features belonging to the index expression generation region), and the output of the region prediction layer. The region prediction layer can be an MLP (Multilayer Perceptron) network. The position detection network can include fully connected layers to predict the target position of the index expression object in the image to be detected based on the output of the fusion encoder. The text prediction network can include a linear regression network to determine the index expression text of the image corresponding to the index expression generation region based on the output of the fusion encoder.
[0043] The above technical solution integrates two tasks—referential expression generation and referential expression understanding—into the target image processing model. This effectively improves the model's integration level in scenarios where both referential expression generation and understanding tasks are required. Consequently, it can effectively improve the resource utilization of the computer system and fully leverage the correlation between the two tasks, thereby effectively improving the accuracy of the results of referential expression generation and understanding.
[0044] Figure 3 It is based on Figure 1 and Figure 2 The illustrated embodiment presents a flowchart of an image processing method; as shown in the figure. Figure 3 As shown above, Figure 1Step 102, which involves inputting the target detection data into the target image processing model to obtain the image processing result output by the target image processing model, may include:
[0045] Step 1021: If the image description information is the specified location information, input the image to be detected and the specified location information into the visual encoder to obtain the image features output by the visual encoder and the regional location features of the region reached by the specified location information.
[0046] For example, if the visual encoder acquires the image features of the image to be detected, I, as follows: The regional location characteristics of the region where the delivery is generated are: Based on the location characteristics of this area Image features output from the visual encoder Determine the regional image features corresponding to the specified location information. Among them, L I v is the number of image patches in the image to be detected I. j Let p be the feature vector of the j-th image patch. i It is the image patch number that overlaps with the region, L R This represents the number of image blocks in the image to be detected that overlap with the region R corresponding to the specified location information.
[0047] Step 1022: Obtain a preset initial mask text and input the initial mask text into the text encoder to obtain the mask text features output by the text encoder.
[0048] The initial mask text can be a preset text that includes one or more MASKs.
[0049] Step 1023: The image features, region location features and mask text features are fused by the fusion encoder to obtain the first target features.
[0050] The fusion encoder includes multiple decoding layers and multiple fusion modules. The multiple decoding layers are connected in series, and the multiple fusion modules are also connected in series. The input of the first decoding layer serves as the input of the fusion encoder. The output of the last decoding layer is coupled to the input of the first fusion module. The output of the last fusion module serves as the output of the fusion encoder. Each fusion module includes a self-attention layer, an image mutual attention layer, a region mutual attention layer, and a region prediction layer. The input of the self-attention layer serves as the input of the fusion module. The output of the self-attention layer is coupled to the input of the image mutual attention layer. The output of the image mutual attention layer is coupled to the input of the region mutual attention layer. The output of the region prediction layer is coupled to the input of the region mutual attention layer. The output of the region mutual attention layer serves as the output of the fusion module.
[0051] This step 1023 can be implemented through the steps shown in S1 to S3:
[0052] S1, when the image description information is the specified location information, the region location features and the mask text features are fused through the multiple decoding layers to obtain the first fused feature.
[0053] For example, if the location characteristics of the area are The masked text feature is Where, p i It is an index that overlaps with the region, L T It is the number of input token embeddings, z [cls] This is the text feature corresponding to the special marker [cls]. By inputting the location feature P of this region and the masked text feature Z into the Dencoders in the Transformer, the first fused feature output by the last decoding layer in multiple decoding layers can be obtained.
[0054] S2, the first fusion module performs attention operations on the first fused feature and the regional location feature to obtain the first undetermined feature.
[0055] In S2, a first attention feature corresponding to the first fused feature can be obtained through the self-attention layer; a second attention feature can be obtained by performing a multi-head attention operation on the first attention feature and the image feature through the image mutual attention layer; a third attention feature can be obtained by performing a multi-head attention operation on the second attention feature and the region location feature through the region mutual attention layer; and the first undetermined feature can be determined based on the third attention feature.
[0056] For example, the processing of this self-attention layer can be represented as follows:
[0057] X Q =MHA(X,X,X)+X
[0058] Among them, MHA stands for multi-head attention, and X... Q The first attention feature is obtained by the self-attention layer based on the first fusion feature X;
[0059] Both the image mutual attention layer and the region mutual attention layer can include multi-head attention networks and gated linear networks. The execution process of the image mutual attention layer can be represented as follows:
[0060] Z I =MHA(X) Q V I V I )
[0061] X I =GLU([Z I ,X Q ])+X Q
[0062] Among them, Z I This is the intermediate representation of the output of the multi-head attention network in the image mutual attention layer, where [,] denotes the concatenation of vectors, and GLU is a gated linear network, GLU(X) = σ(XW). 1 )⊙XW 2 W 1 W 2 These are learnable parameters, σ is the sigmoid function, and ⊙ represents element-wise multiplication. I For gated linear networks to Z I ,X Q The result after processing is the second attention feature output by the mutual attention layer of the image.
[0063] The execution process of the mutual attention layer in this region can be represented as follows:
[0064] Z R =MHA(X) I V R VR )
[0065] X R =GLU([Z R ,X Q ])+X I
[0066] Among them, it can be based on the location characteristics of the region. Image features output from the visual encoder Determine the regional image features corresponding to the specified location information. L I v is the number of image patches in the image to be detected I. j Let p be the feature vector of the j-th image patch. i It is the image patch number that overlaps with the region, L R Z represents the number of image patches in the image to be detected that overlap with the region R corresponding to the specified location information. R This represents the intermediate representation of the multi-head attention network output in the region mutual attention layer, where [,] denotes the concatenation of vectors, GLU-gated linear network, and X. R For gated linear networks to Z R ,X Q X I The result after processing is the third attention feature output by the mutual attention layer of the image.
[0067] In addition, the fusion module may also include an FFN (feed-forward network) layer. The above-described method of determining the first undetermined feature based on the third attention feature can be implemented by inputting the third attention feature into the location feed-forward network layer to obtain the output feature of the location feed-forward network layer, and using the output feature of the location feed-forward network layer as the first undetermined feature.
[0068] S3, the first target feature is determined by other fusion modules among the plurality of fusion modules, excluding the first fusion module, based on the first undetermined feature and the regional location feature.
[0069] In this S3, attention operations can be performed on the specified undetermined features output by the previous fusion module and the region location features for each of the plurality of fusion modules except the first fusion module, so as to obtain the target undetermined features output by the current fusion module; and the target undetermined features output by the last fusion module are used as the first target feature.
[0070] Step 1024: Generate the index expression text of the image corresponding to the index expression generation region based on the first target features using the text prediction network.
[0071] In this step, the text prediction network can determine the first character of the indicated text based on the first target feature; update the initial mask text based on the first character; and determine the other characters in the indicated text based on the updated initial mask text, the image features, and the region location features to obtain the indicated text.
[0072] For example, when generating the pointer expression for the detection box corresponding to green grass in the image to be detected, the initial mask text is "mask mask mask". The first iteration determines the first character "green", and the updated initial mask text is "green mask mask". The second iteration determines the second character "color", and the updated initial mask text is "green mask". The third iteration determines the third character "grass", and the updated initial mask text is "green grass mask". This process is repeated until the pointer expression text "green grass" is obtained.
[0073] Figure 4 It is based on Figure 1 and Figure 2 The illustrated embodiment shows a flowchart of another image processing method; as shown Figure 4 As shown above, Figure 1 Step 102, which involves inputting the target detection data into the target image processing model to obtain the image processing result output by the target image processing model, may further include:
[0074] Step 1025: If the image description information is the description text, input the image to be detected into the visual encoder to obtain the image features output by the visual encoder, and input the description text into the text encoder to obtain the description text features output by the text encoder.
[0075] Step 1026: Input the image features and the descriptive text features into the fusion encoder to obtain the second target features output by the fusion encoder.
[0076] This step can be implemented through the steps shown in S21 to S22 below:
[0077] S21, when the image description information is the description text, the image features and the description text features are fused through the multiple decoding layers to obtain a second fused feature.
[0078] In this step, the decoding layer is the Dencoder in Transformer. The image features and the descriptive text features are input into the Dencoder to obtain the second fused feature.
[0079] S22, the second fusion feature and the image feature are fused by the multiple fusion modules to obtain the second target feature.
[0080] In this S22, attention operations can be performed on the second fused feature and the image feature by the first fusion module to obtain a second undetermined feature; the second undetermined feature can be determined by other fusion modules among the plurality of fusion modules other than the first fusion module based on the second undetermined feature and the image feature; the second target feature can be determined by other fusion modules among the plurality of fusion modules other than the first fusion module based on the second undetermined feature and the second image feature.
[0081] The above-described method of performing attention operations on the second fused feature and the image feature through the first fusion module to obtain a second undetermined feature may include: performing attention operations on the second fused feature through the self-attention layer in the first fusion module to obtain a first target attention feature; performing multi-head attention operations on the image feature and the first target attention feature through the image mutual attention layer in the first fusion module to obtain a second target attention feature; determining a predicted region image feature based on the second target attention feature through the region prediction layer in the first fusion module; performing multi-head attention operations on the image feature, the second target attention feature, and the predicted region image feature through the region mutual attention layer in the first fusion module to obtain a third target attention feature; and determining the second undetermined feature based on the third target attention feature.
[0082] The region prediction layer described above is a multilayer perceptron network. The step of determining the predicted region image features based on the second target attention features through the region prediction layer in the first fusion module includes:
[0083] The multilayer perceptron network determines the target probability that each unit region belongs to the region where the target expression object is located based on the full-text features in the second target attention features; the image features corresponding to the unit regions whose target probabilities are greater than or equal to a preset probability threshold are used as the image features of the predicted region.
[0084] For example, when performing indicative expression comprehension, since the input image description information is the descriptive text of the indicative object in the image to be detected, and there is no input specifying location information, it is necessary to predict the region based on the image features output by the visual encoder and the descriptive text features output by the text encoder. To ensure that the input for indicative expression comprehension is the same as that for indicative expression generation, a region prediction layer is set up to generate the predicted region position, which serves as the input to the region mutual attention layer. In implementation, for each image patch (i.e., unit region), based on... (Full text features) and the i-th image patch e i The position embedding is used to calculate the score α, and this score α is used to calculate the score α. i As the target probability that the unit region corresponding to the i-th image patch belongs to the region where the target object is located, image features of the regions where image patches with scores exceeding a preset probability threshold δ are selected to constitute the predicted region image features V. R This process can be represented as:
[0085]
[0086]
[0087] Where, α i The probability that the unit region corresponding to the i-th image patch belongs to the region where the indicated target object is located. Let be the image feature of the i-th image patch, and δ be the preset probability threshold.
[0088] The above-described method of determining the second target feature based on the second undetermined feature and the second image feature through other fusion modules among the plurality of fusion modules, excluding the first fusion module, may include:
[0089] When the image description information is the description text, for each of the plurality of fusion modules except the first fusion module, an attention operation is performed on the specified undetermined features output by the previous fusion module and the region location features to obtain the target undetermined features output by the current fusion module; the target undetermined features output by the last fusion module are used as the second target features.
[0090] Step 1027: Input the second target feature into the location detection network to obtain the target position of the target object in the image to be detected by the pointer output by the location detection network.
[0091] The location detection network may include a fully connected layer for predicting the target location of the pointer in the image to be detected based on the second target features.
[0092] The above technical solution extends the Transformer decoder by replacing the vanilla Transformer decoder layer with a fusion module, thereby bridging the gap between referential expression generation and referential expression understanding. By generating pseudo-input regions (regions output by the region prediction layer) in the referential expression generation task, the same representation space can be shared between referential expression generation and referential expression understanding in a unified manner. This fully utilizes the correlation between the two tasks, effectively achieving unified modeling of referential expression generation and referential expression understanding tasks, and improving the model integration while ensuring the accuracy of recognition results.
[0093] Figure 5 This is a flowchart illustrating a model training method according to an exemplary embodiment of this disclosure; as shown below. Figure 5 As shown, the target image processing model can be trained through the following steps:
[0094] Step 501: Obtain multiple pre-training datasets with different granularities. The pre-training datasets include multiple sets of training data. Each set of training data includes sample images, pointer-generated sample regions, and pointer-generated text samples corresponding to the pointer-generated sample regions. The pointer-generated text samples include a preset mask and text annotation data of the preset mask.
[0095] In this process, a portion of the text sample can be masked. For example, 25% of the text can be masked to obtain a text sample with a preset mask.
[0096] It should be noted that the pre-training datasets of different granularities may include the COCO (CommonObjects in Context, a commonly used dataset for object detection, segmentation, and keypoint detection) dataset in the prior art; the VisualGenome phrase dataset; the Visual Genome region description dataset, where the region description can be a phrase or sentence; and RefCOCO-MERGE (a dataset in the prior art). These datasets are commonly used in the prior art for training models for generating or locating referential expressions, and this disclosure does not limit their use.
[0097] Step 502: Using the multiple pre-training datasets of different granularities as training data, pre-train the preset initial model to obtain the image processing model to be determined.
[0098] The model structure of this preset initial model can be found in [reference needed]. Figure 2The model structure shown is not described in detail here. The model parameters in the preset initial model are the initial parameters before training.
[0099] This step can be obtained through training as shown in steps S31 to S35:
[0100] S31, sequentially input each set of training data from multiple pre-training datasets of different granularities into the preset initial model to obtain the predicted text data and prediction boundary data corresponding to the training data.
[0101] S32, calculate a first loss value using a first loss function based on the predicted text data and the text annotation data of the preset mask.
[0102] It should be noted that during training, the text content at the [mask] location can be predicted to obtain predicted text data. Then, based on this predicted text data and the text annotation data, the first loss value L of the first preset loss function can be calculated. VMLM :
[0103]
[0104] Wherein, sample image I refers to the generated sample region R and the corresponding generated sample text sample T form an image-region-text triple (I,R,T), E (I,R,T) θ represents the expected value of the data in a batch of data. G To generate network parameters for referential representation, Predict text data, Label the text with data.
[0105] S33, calculate the second loss value using the second loss function based on the predicted boundary data and the index-generated sample region, and determine the third loss value of the third loss function based on the labeled region mask and the predicted region mask.
[0106] For example, the second loss function could be:
[0107]
[0108] in, yes The generalized intersection union with b, It is the l1 norm. The actual labeled region represents the generated sample region, b represents the predicted bounding box location, and E represents the region of origin. (I,T) Let I be the expected value of the data in a batch of data formed by the sample image I and the text sample T.
[0109] The fourth loss function can be:
[0110]
[0111] in, denoted as the actual labeled mask regions, and m is the predicted mask region of the i-th token embedding (image-region-text fusion embedding) output by the region prediction layer during training.
[0112] S34, determine the fourth loss value using the fourth loss function based on the second loss value and the third loss value.
[0113] The fourth loss function can be L TRP :
[0114] L TRP =L bbox +L pred
[0115] L bbox For the second loss value, L pred Third loss value.
[0116] S35, the preset initial model is iteratively trained based on the first loss value and the fourth loss value to obtain the image processing model to be determined.
[0117] After one iteration of calculation is completed, the first loss value and the fourth loss value are calculated to determine whether the first loss value is greater than the first preset loss threshold and whether the fourth loss value is greater than the second preset loss threshold. If the first loss value is greater than the first preset loss threshold or the fourth loss value is greater than the second preset loss threshold, the model parameters of the preset initial model are updated, and the calculation of the first loss value and the fourth loss value is repeated until the step of determining whether the first loss value is greater than the first preset loss threshold and whether the fourth loss value is greater than the second preset loss threshold is completed. Until the first loss value is less than or equal to the first preset loss threshold and the fourth loss value is less than or equal to the second preset loss threshold, the current preset initial model is used as the image processing model to be determined.
[0118] Step 503: Obtain the target training dataset at the target granularity, and use the target training dataset as training data to refine the image processing model to be determined, so as to obtain the target image processing model.
[0119] In this step, any one of the following can be used as the target training dataset: COCO dataset, Visual Genome phrase dataset, Visual Genome region description dataset, and RefCOCO-MERGE. Other training datasets can also be constructed. The training method used in the pre-training process is still used for fine-tuning to obtain the target image processing model.
[0120] The above training methods can fully utilize the correlation between the two tasks of generating and understanding referential expressions. By unifying the modeling of referential expression generation and understanding, a more accurate model can be trained that can perform both tasks. In scenarios where both referential expression generation and understanding tasks are required, the model's integration can be effectively improved, thereby enhancing the resource utilization of the computer system.
[0121] Figure 6 This is a block diagram illustrating an image processing apparatus according to an exemplary embodiment of the present disclosure, such as... Figure 6 As shown, the device may include:
[0122] The acquisition module 601 is configured to acquire target detection data, which includes a target image to be detected and image description information of the target image. The image description information is either specified location information of the target image indicating the generated area, or descriptive text of the target image indicating the object.
[0123] The determination module 602 is configured to input the target detection data into a target image processing model to obtain the image processing result output by the target image processing model. The target image processing model is used to output the indicator text of the image corresponding to the indicator generation region when the image description information is the specified location information, and to output the target position of the indicator object in the image to be detected when the image description information is the description text.
[0124] The above technical solution integrates two tasks, namely, the generation and understanding of referential expressions, in the target image processing model. This effectively improves the integration of the model in scenarios where both referential expression generation and understanding tasks are required, thereby improving the resource utilization of the computer system.
[0125] Optionally, the target image processing model includes a visual encoder, a text encoder, a fusion encoder, a position detection network, and a text prediction network. The output of the visual encoder is coupled to the input of the fusion encoder, the output of the text encoder is coupled to the input of the fusion encoder, the output of the fusion encoder is coupled to the input of the position detection network, and the output of the fusion encoder is also coupled to the input of the text prediction network.
[0126] The determining module 602 is configured as follows:
[0127] When the image description information is the specified location information, the image to be detected and the specified location information are input into the visual encoder to obtain the image features output by the visual encoder and the regional location features of the region represented by the specified location information.
[0128] Obtain a preset initial mask text and input the initial mask text into the text encoder to obtain the mask text features output by the text encoder;
[0129] The image features, region location features, and masked text features are fused using the fusion encoder to obtain the first target features;
[0130] The text prediction network generates the index expression text of the image corresponding to the index expression generation region based on the first target features.
[0131] Optionally, the determining module 602 is configured to:
[0132] The text prediction network determines the first character of the target text based on the first target feature.
[0133] Update the initial mask text based on the first text;
[0134] Based on the updated initial mask text, the image features and the region location features determine other characters in the indicated text to obtain the indicated text.
[0135] Optionally, the determining module 602 is further configured to:
[0136] When the image description information is the description text, the image to be detected is input into the visual encoder to obtain the image features output by the visual encoder, and the description text is input into the text encoder to obtain the description text features output by the text encoder.
[0137] The image features and the descriptive text features are input into the fusion encoder to obtain the second target features output by the fusion encoder;
[0138] The second target feature is input into the location detection network to obtain the target location of the target object in the image to be detected, as indicated by the pointer output by the location detection network.
[0139] Optionally, the fusion encoder includes multiple decoding layers and multiple fusion modules, wherein the multiple decoding layers are connected in series and the multiple fusion modules are connected in series, the input terminal of the first decoding layer in the multiple decoding layers serves as the input terminal of the fusion encoder, the output terminal of the last decoding layer in the multiple decoding layers is coupled to the input terminal of the first fusion module, and the output terminal of the last fusion module in the multiple fusion modules serves as the output terminal of the fusion encoder;
[0140] The determining module 602 is configured as follows:
[0141] When the image description information is the specified location information, the region location features and the mask text features are fused through the multiple decoding layers to obtain a first fused feature;
[0142] The first fusion module performs an attention operation on the first fused feature and the region location feature to obtain the first undetermined feature;
[0143] The first target feature is determined by the first undetermined feature and the regional location feature through the other fusion modules among the plurality of fusion modules, excluding the first fusion module.
[0144] Optionally, the determining module 602 is configured to:
[0145] For each of the multiple fusion modules except the first fusion module, an attention operation is performed on the specified undetermined features output by the previous fusion module and the region location features to obtain the target undetermined features output by the current fusion module.
[0146] The target feature to be determined output by the last fusion module is used as the first target feature.
[0147] Optionally, the determining module 602 is configured to:
[0148] When the image description information is the description text, the image features and the description text features are fused through the multiple decoding layers to obtain a second fused feature;
[0149] The second fusion feature and the image feature are fused using the multiple fusion modules.
[0150] To obtain the second target feature.
[0151] Optionally, the fusion module includes a self-attention layer, an image mutual attention layer, a region mutual attention layer, and a region prediction layer. The input of the self-attention layer serves as the input of the fusion module. The output of the self-attention layer is coupled to the input of the image mutual attention layer. The output of the image mutual attention layer is coupled to the input of the region prediction layer. The output of the image mutual attention layer is also coupled to the input of the region mutual attention layer. The output of the region prediction layer is coupled to the input of the region mutual attention layer. The output of the region mutual attention layer serves as the output of the fusion module.
[0152] The determining module 602 is configured as follows:
[0153] The first attention feature corresponding to the first fusion feature is obtained through the self-attention layer;
[0154] The first attention feature and the image feature are subjected to multi-head attention operation through the image mutual attention layer to obtain the second attention feature;
[0155] A third attention feature is obtained by performing multi-head attention operations on the second attention feature and the regional location feature through the regional mutual attention layer.
[0156] The first undetermined feature is determined based on the third attention feature.
[0157] Optionally, the determining module 602 is configured to:
[0158] The first fusion module performs an attention operation on the second fused feature and the image feature to obtain the second undetermined feature;
[0159] The second undetermined feature is determined by the other fusion modules among the plurality of fusion modules, excluding the first fusion module, based on the second undetermined feature and the image features;
[0160] The second target feature is determined by the second undetermined feature and the second image feature through the other fusion modules among the plurality of fusion modules, excluding the first fusion module.
[0161] Optionally, the determining module 602 is configured to:
[0162] The second fused feature is subjected to attention operation through the self-attention layer in the first fusion module to obtain the first target attention feature;
[0163] The first target attention feature is obtained by performing a multi-head attention operation on the image features and the first target attention feature through the image mutual attention layer in the first fusion module;
[0164] The region prediction layer in the first fusion module determines the image features of the predicted region based on the second target attention features;
[0165] The image features, the second target attention features, and the predicted region image features are subjected to multi-head attention operations through the region mutual attention layer in the first fusion module to obtain the third target attention features;
[0166] The second undetermined feature is determined based on the third target attention feature.
[0167] Optionally, the region prediction layer is a multilayer perceptron network, and the determining module is configured to:
[0168] The multilayer perceptron network determines the target probability that each unit region belongs to the region where the target is located based on the full-text features in the second target attention features;
[0169] The image features corresponding to the unit regions whose target probability is greater than or equal to a preset probability threshold are used as the image features of the predicted region.
[0170] Optionally, the image processing apparatus further includes a model training module 603, configured to:
[0171] Obtain multiple pre-training datasets with different granularities. The pre-training datasets include multiple sets of training data. Each set of training data includes sample images, pointer-generated sample regions, and pointer-generated sample regions corresponding to pointer-generated sample regions. The pointer-generated text samples include a preset mask and text annotation data of the preset mask.
[0172] Using the multiple pre-training datasets of different granularities as training data, a preset initial model is pre-trained to obtain a pending image processing model.
[0173] Obtain the target training dataset at the target granularity, and use the target training dataset as training data to refine the image processing model to be determined, so as to obtain the target image processing model.
[0174] Optionally, the model training module 603 is configured as follows:
[0175] Each set of training data from multiple pre-training datasets of different granularities is sequentially input into the preset initial model to obtain the predicted text data and prediction boundary data corresponding to the training data;
[0176] A first loss value is calculated using a first loss function based on the predicted text data and the text annotation data of the preset mask.
[0177] The second loss value is calculated using the second loss function based on the predicted boundary data and the index-generated sample region, and the third loss value of the third loss function is determined based on the labeled region mask and the predicted region mask.
[0178] The fourth loss value is determined using the second loss value and the third loss value through the fourth loss function;
[0179] The preset initial model is iteratively trained based on the first loss value and the fourth loss value to obtain the image processing model to be determined.
[0180] The above technical solution extends the Transformer decoder by replacing the vanilla Transformer decoder layer with a fusion module, thereby bridging the gap between referential expression generation and referential expression understanding. By generating pseudo-input regions (regions output by the region prediction layer) in the referential expression generation task, the same representation space can be shared between referential expression generation and referential expression understanding in a unified manner. This effectively achieves unified modeling of referential expression generation and referential expression understanding tasks, ensuring the accuracy of recognition results while improving the model integration.
[0181] The following is for reference. Figure 7 The diagram illustrates a structural schematic of an electronic device 700 suitable for implementing embodiments of the present disclosure. Terminal devices in embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 7 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0182] like Figure 7As shown, the electronic device 700 may include a processing unit (e.g., a central processing unit, a graphics processor, etc.) 701, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage device 708 into a random access memory (RAM) 703. The RAM 703 also stores various programs and data required for the operation of the electronic device 700. The processing unit 701, ROM 702, and RAM 703 are interconnected via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.
[0183] Typically, the following devices can be connected to I / O interface 705: input devices 706 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 707 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 708 including, for example, magnetic tapes, hard disks, etc.; and communication devices 709. Communication device 709 allows electronic device 700 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 7 An electronic device 700 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0184] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 709, or installed from storage device 708, or installed from ROM 702. When the computer program is executed by processing device 701, it performs the functions defined in the methods of embodiments of this disclosure.
[0185] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0186] In some implementations, the client can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.
[0187] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0188] The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to: acquire target detection data, wherein the target detection data includes an image to be detected and image description information of the image to be detected, wherein the image description information is specified location information of the generated area in the image to be detected, or is descriptive text of the object to be detected in the image to be detected;
[0189] The target detection data is input into the target image processing model to obtain the image processing result output by the target image processing model. The target image processing model is used to output the index expression text of the image corresponding to the index expression generation region when the image description information is the specified location information, and to output the target position of the index expression object in the image to be detected when the image description information is the description text.
[0190] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including but not limited to object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0191] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0192] The modules described in the embodiments of this disclosure can be implemented in software or in hardware. The name of a module does not necessarily limit the module itself; for example, an acquisition module can also be described as a "module for acquiring target detection data".
[0193] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0194] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0195] According to one or more embodiments of this disclosure, Example 1 provides an image processing method, the method comprising:
[0196] Acquire target detection data, which includes an image to be detected and image description information of the image to be detected. The image description information is either specified location information of the generated region in the image to be detected, or descriptive text of the object to be detected in the image to be detected.
[0197] The target detection data is input into the target image processing model to obtain the image processing result output by the target image processing model. The target image processing model is used to output the index expression text of the image corresponding to the index expression generation region when the image description information is the specified location information, and to output the target position of the index expression object in the image to be detected when the image description information is the description text.
[0198] According to one or more embodiments of this disclosure, Example 2 provides the method of Example 1, wherein the target image processing model includes a visual encoder, a text encoder, a fusion encoder, a position detection network, and a text prediction network, wherein the output of the visual encoder is coupled to the input of the fusion encoder, the output of the text encoder is coupled to the input of the fusion encoder, the output of the fusion encoder is coupled to the input of the position detection network, and the output of the fusion encoder is also coupled to the input of the text prediction network;
[0199] The step of inputting the target detection data into the target image processing model to obtain the image processing result output by the target image processing model includes:
[0200] When the image description information is the specified location information, the image to be detected and the specified location information are input into the visual encoder to obtain the image features output by the visual encoder and the regional location features of the region represented by the specified location information.
[0201] Obtain a preset initial mask text and input the initial mask text into the text encoder to obtain the mask text features output by the text encoder;
[0202] The image features, region location features, and masked text features are fused using the fusion encoder to obtain the first target features;
[0203] The text prediction network generates the index expression text of the image corresponding to the index expression generation region based on the first target features.
[0204] According to one or more embodiments of this disclosure, Example 3 provides the method of Example 2, wherein determining the index expression text of the image corresponding to the index expression generation region based on the first target feature using the text prediction network includes:
[0205] The text prediction network determines the first character of the target text based on the first target feature.
[0206] Update the initial mask text based on the first text;
[0207] Based on the updated initial mask text, the image features and the region location features determine other characters in the indicated text to obtain the indicated text.
[0208] According to one or more embodiments of this disclosure, Example 4 provides the method of Example 2, wherein inputting the target detection data into a target image processing model to obtain the image processing result output by the target image processing model further includes:
[0209] When the image description information is the description text, the image to be detected is input into the visual encoder to obtain the image features output by the visual encoder, and the description text is input into the text encoder to obtain the description text features output by the text encoder.
[0210] The image features and the descriptive text features are input into the fusion encoder to obtain the second target features output by the fusion encoder;
[0211] The second target feature is input into the location detection network to obtain the target location of the target object in the image to be detected, as indicated by the pointer output by the location detection network.
[0212] According to one or more embodiments of this disclosure, Example 5 provides the method of Example 4, wherein the fusion encoder includes multiple decoding layers and multiple fusion modules, the multiple decoding layers are connected in series, and the multiple fusion modules are connected in series, the input terminal of the first decoding layer among the multiple decoding layers serves as the input terminal of the fusion encoder, the output terminal of the last decoding layer among the multiple decoding layers is coupled to the input terminal of the first fusion module, and the output terminal of the last fusion module among the multiple fusion modules serves as the output terminal of the fusion encoder;
[0213] The step of fusing the image features, region location features, and masked text features using the fusion encoder to obtain the first target feature includes:
[0214] When the image description information is the specified location information, the region location features and the mask text features are fused through the multiple decoding layers to obtain a first fused feature;
[0215] The first fusion module performs an attention operation on the first fused feature and the region location feature to obtain the first undetermined feature;
[0216] The first target feature is determined by the first undetermined feature and the regional location feature through the other fusion modules among the plurality of fusion modules, excluding the first fusion module.
[0217] According to one or more embodiments of this disclosure, Example 6 provides the method of Example 5, wherein determining the first target feature based on the first undetermined feature and the regional location feature through a fusion module other than the first fusion module among the plurality of fusion modules includes:
[0218] For each of the multiple fusion modules except the first fusion module, an attention operation is performed on the specified undetermined features output by the previous fusion module and the region location features to obtain the target undetermined features output by the current fusion module.
[0219] The target feature to be determined output by the last fusion module is used as the first target feature.
[0220] According to one or more embodiments of this disclosure, Example 7 provides the method of Example 5, wherein inputting the image features and the descriptive text features into the fusion encoder to obtain a second target feature output by the fusion encoder includes:
[0221] When the image description information is the description text, the image features and the description text features are fused through the multiple decoding layers to obtain a second fused feature;
[0222] The second target feature is obtained by fusing the second fusion feature and the image feature through the multiple fusion modules.
[0223] According to one or more embodiments of this disclosure, Example 8 provides the method of Example 7, wherein the fusion module includes a self-attention layer, an image mutual attention layer, a region mutual attention layer, and a region prediction layer. The input of the self-attention layer serves as the input of the fusion module, the output of the self-attention layer is coupled to the input of the image mutual attention layer, the output of the image mutual attention layer is coupled to the input of the region prediction layer, the output of the image mutual attention layer is also coupled to the input of the region mutual attention layer, the output of the region prediction layer is coupled to the input of the region mutual attention layer, and the output of the region mutual attention layer serves as the output of the fusion module.
[0224] The step of performing an attention operation on the first fused feature and the region location feature through the first fusion module to obtain the first undetermined feature includes:
[0225] The first attention feature corresponding to the first fusion feature is obtained through the self-attention layer;
[0226] The first attention feature and the image feature are subjected to multi-head attention operation through the image mutual attention layer to obtain the second attention feature;
[0227] A third attention feature is obtained by performing multi-head attention operations on the second attention feature and the regional location feature through the regional mutual attention layer.
[0228] The first undetermined feature is determined based on the third attention feature.
[0229] According to one or more embodiments of this disclosure, Example 9 provides the method of Example 8, wherein fusing the second fusion feature and the image feature through the plurality of fusion modules to obtain the second target feature includes:
[0230] The first fusion module performs an attention operation on the second fused feature and the image feature to obtain the second undetermined feature;
[0231] The second undetermined feature is determined by the other fusion modules among the plurality of fusion modules, excluding the first fusion module, based on the second undetermined feature and the image features;
[0232] The second target feature is determined by the second undetermined feature and the second image feature through the other fusion modules among the plurality of fusion modules, excluding the first fusion module.
[0233] According to one or more embodiments of this disclosure, Example 10 provides the method of Example 9, wherein the attention operation is performed on the second fused feature and the image feature by the first fusion module to obtain a second undetermined feature, including:
[0234] The second fused feature is subjected to attention operation through the self-attention layer in the first fusion module to obtain the first target attention feature;
[0235] The first target attention feature is obtained by performing a multi-head attention operation on the image features and the first target attention feature through the image mutual attention layer in the first fusion module;
[0236] The region prediction layer in the first fusion module determines the image features of the predicted region based on the second target attention features;
[0237] The image features, the second target attention features, and the predicted region image features are subjected to multi-head attention operations through the region mutual attention layer in the first fusion module to obtain the third target attention features;
[0238] The second undetermined feature is determined based on the third target attention feature.
[0239] According to one or more embodiments of this disclosure, Example 11 provides the method of Example 10, wherein the region prediction layer is a multilayer perceptron network, and the step of determining the predicted region image features based on the second target attention features through the region prediction layer in the first fusion module includes:
[0240] The multilayer perceptron network determines the target probability that each unit region belongs to the region where the target is located based on the full-text features in the second target attention features;
[0241] The image features corresponding to the unit regions whose target probability is greater than or equal to a preset probability threshold are used as the image features of the predicted region.
[0242] According to one or more embodiments of this disclosure, Example 12 provides a method as described in any one of Examples 1-11, wherein the target image processing model is trained in the following manner:
[0243] Obtain multiple pre-training datasets with different granularities. The pre-training datasets include multiple sets of training data. Each set of training data includes sample images, pointer-generated sample regions, and pointer-generated sample regions corresponding to pointer-generated sample regions. The pointer-generated text samples include a preset mask and text annotation data of the preset mask.
[0244] Using the multiple pre-training datasets of different granularities as training data, a preset initial model is pre-trained to obtain a pending image processing model.
[0245] Obtain the target training dataset at the target granularity, and use the target training dataset as training data to refine the image processing model to be determined, so as to obtain the target image processing model.
[0246] According to one or more embodiments of this disclosure, Example 13 provides the method of Example 12, wherein the method of pre-training a preset initial model using the plurality of pre-training datasets of different granularities as training data to obtain a pending image processing model includes:
[0247] Each set of training data from multiple pre-training datasets of different granularities is sequentially input into the preset initial model to obtain the predicted text data and prediction boundary data corresponding to the training data;
[0248] A first loss value is calculated using a first loss function based on the predicted text data and the text annotation data of the preset mask.
[0249] The second loss value is calculated using the second loss function based on the predicted boundary data and the index-generated sample region, and the third loss value of the third loss function is determined based on the labeled region mask and the predicted region mask.
[0250] The fourth loss value is determined using the second loss value and the third loss value through the fourth loss function;
[0251] The preset initial model is iteratively trained based on the first loss value and the fourth loss value to obtain the image processing model to be determined.
[0252] According to one or more embodiments of this disclosure, Example 14 provides an image processing apparatus, the apparatus comprising:
[0253] The acquisition module is configured to acquire target detection data, which includes an image to be detected and image description information of the image to be detected. The image description information is either specified location information of the generated area in the image to be detected, or descriptive text of the object to be detected in the image to be detected.
[0254] The determination module is configured to input the target detection data into a target image processing model to obtain the image processing result output by the target image processing model. The target image processing model is configured to output the indicator text of the image corresponding to the indicator generation region when the image description information is the specified location information, and output the target position of the indicator object in the image to be detected when the image description information is the description text.
[0255] According to one or more embodiments of the present disclosure, Example 15 provides a computer-readable medium having a computer program stored thereon that, when executed by a processing device, implements the steps of the method described in any one of Examples 1-13.
[0256] According to one or more embodiments of this disclosure, Example 16 provides an electronic device comprising:
[0257] A storage device on which computer programs are stored;
[0258] A processing device for executing the computer program in the storage device to implement the steps of any one of the methods in Examples 1-13.
[0259] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
[0260] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0261] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative forms of implementing the claims. Regarding the apparatus in the above embodiments, the specific manner in which the various modules perform their operations has been described in detail in the embodiments relating to the method, and will not be elaborated upon here.
Claims
1. An image processing method, characterized in that, The method includes: Acquire target detection data, which includes an image to be detected and image description information of the image to be detected. The image description information is either specified location information of the generated region in the image to be detected, or descriptive text of the object to be detected in the image to be detected. The target detection data is input into the target image processing model to obtain the image processing result output by the target image processing model. The target image processing model is used to output the index expression text of the image corresponding to the index expression generation area when the image description information is the specified location information, and to output the target position of the index expression object in the image to be detected when the image description information is the description text. The target image processing model includes a visual encoder, a text encoder, a fusion encoder, a location detection network, and a text prediction network. The output of the visual encoder is coupled to the input of the fusion encoder, the output of the text encoder is coupled to the input of the fusion encoder, the output of the fusion encoder is coupled to the input of the location detection network, and the output of the fusion encoder is also coupled to the input of the text prediction network. The step of inputting the target detection data into the target image processing model to obtain the image processing result output by the target image processing model includes: When the image description information is the specified location information, the image to be detected and the specified location information are input into the visual encoder to obtain the image features output by the visual encoder and the regional location features of the region represented by the specified location information. Obtain a preset initial mask text and input the initial mask text into the text encoder to obtain the mask text features output by the text encoder; The image features, region location features, and masked text features are fused using the fusion encoder to obtain the first target features; The text prediction network generates the index expression text of the image corresponding to the index expression generation region based on the first target features.
2. The method according to claim 1, characterized in that, The step of determining the index expression text of the image corresponding to the index expression generation region based on the first target feature through the text prediction network includes: The text prediction network determines the first character of the target text based on the first target feature. Update the initial mask text based on the first text; Based on the updated initial mask text, the image features and the region location features determine other characters in the indicated text to obtain the indicated text.
3. The method according to claim 1, characterized in that, The step of inputting the target detection data into the target image processing model to obtain the image processing result output by the target image processing model further includes: When the image description information is the description text, the image to be detected is input into the visual encoder to obtain the image features output by the visual encoder, and the description text is input into the text encoder to obtain the description text features output by the text encoder. The image features and the descriptive text features are input into the fusion encoder to obtain the second target features output by the fusion encoder; The second target feature is input into the location detection network to obtain the target location of the target object in the image to be detected, as indicated by the pointer output by the location detection network.
4. The method according to claim 3, characterized in that, The fusion encoder includes multiple decoding layers and multiple fusion modules. The multiple decoding layers are connected in series, and the multiple fusion modules are connected in series. The input terminal of the first decoding layer is used as the input terminal of the fusion encoder. The output terminal of the last decoding layer is coupled to the input terminal of the first fusion module. The output terminal of the last fusion module is used as the output terminal of the fusion encoder. The step of fusing the image features, region location features, and masked text features using the fusion encoder to obtain the first target feature includes: When the image description information is the specified location information, the region location features and the mask text features are fused through the multiple decoding layers to obtain a first fused feature; The first fusion module performs an attention operation on the first fused feature and the region location feature to obtain the first undetermined feature; The first target feature is determined by the first undetermined feature and the regional location feature through the other fusion modules among the plurality of fusion modules, excluding the first fusion module.
5. The method according to claim 4, characterized in that, The step of determining the first target feature based on the first undetermined feature and the regional location feature through a fusion module other than the first fusion module among the plurality of fusion modules includes: For each of the multiple fusion modules except the first fusion module, an attention operation is performed on the specified undetermined features output by the previous fusion module and the region location features to obtain the target undetermined features output by the current fusion module. The target feature to be determined output by the last fusion module is used as the first target feature.
6. The method according to claim 4, characterized in that, The step of inputting the image features and the descriptive text features into the fusion encoder to obtain the second target features output by the fusion encoder includes: When the image description information is the description text, the image features and the description text features are fused through the multiple decoding layers to obtain a second fused feature; The second target feature is obtained by fusing the second fusion feature and the image feature through the multiple fusion modules.
7. The method according to claim 6, characterized in that, The fusion module includes a self-attention layer, an image mutual attention layer, a region mutual attention layer, and a region prediction layer. The input of the self-attention layer serves as the input of the fusion module. The output of the self-attention layer is coupled to the input of the image mutual attention layer. The output of the image mutual attention layer is coupled to the input of the region prediction layer. The output of the image mutual attention layer is also coupled to the input of the region mutual attention layer. The output of the region prediction layer is coupled to the input of the region mutual attention layer. The output of the region mutual attention layer serves as the output of the fusion module. The step of performing an attention operation on the first fused feature and the region location feature through the first fusion module to obtain the first undetermined feature includes: The first attention feature corresponding to the first fusion feature is obtained through the self-attention layer; The first attention feature and the image feature are subjected to multi-head attention operation through the image mutual attention layer to obtain the second attention feature; A third attention feature is obtained by performing multi-head attention operations on the second attention feature and the regional location feature through the regional mutual attention layer. The first undetermined feature is determined based on the third attention feature.
8. The method according to claim 7, characterized in that, The step of fusing the second fusion feature and the image feature through the multiple fusion modules to obtain the second target feature includes: The first fusion module performs an attention operation on the second fused feature and the image feature to obtain the second undetermined feature; The second undetermined feature is determined by the other fusion modules among the plurality of fusion modules, excluding the first fusion module, based on the second undetermined feature and the image features; The second target feature is determined by the second undetermined feature and the second image feature through the other fusion modules among the plurality of fusion modules, excluding the first fusion module.
9. The method according to claim 8, characterized in that, The step of performing an attention operation on the second fused feature and the image feature through the first fusion module to obtain the second undetermined feature includes: The second fused feature is subjected to attention operation through the self-attention layer in the first fusion module to obtain the first target attention feature; The first target attention feature is obtained by performing a multi-head attention operation on the image features and the first target attention feature through the image mutual attention layer in the first fusion module; The region prediction layer in the first fusion module determines the image features of the predicted region based on the second target attention features; The image features, the second target attention features, and the predicted region image features are subjected to multi-head attention operations through the region mutual attention layer in the first fusion module to obtain the third target attention features; The second undetermined feature is determined based on the third target attention feature.
10. The method according to claim 9, characterized in that, The region prediction layer is a multilayer perceptron network. The process of determining the predicted region image features based on the second target attention features through the region prediction layer in the first fusion module includes: The multilayer perceptron network determines the target probability that each unit region belongs to the region where the target is located based on the full-text features in the second target attention features; The image features corresponding to the unit regions whose target probability is greater than or equal to a preset probability threshold are used as the image features of the predicted region.
11. The method according to any one of claims 1-10, characterized in that, The target image processing model is trained in the following way: Obtain multiple pre-training datasets with different granularities. The pre-training datasets include multiple sets of training data. Each set of training data includes sample images, pointer-generated sample regions, and pointer-generated sample regions corresponding to pointer-generated sample regions. The pointer-generated text samples include a preset mask and text annotation data of the preset mask. Using the multiple pre-training datasets of different granularities as training data, a preset initial model is pre-trained to obtain a pending image processing model. Obtain the target training dataset at the target granularity, and use the target training dataset as training data to refine the image processing model to be determined, so as to obtain the target image processing model.
12. The method according to claim 11, characterized in that, The step of using the multiple pre-training datasets of different granularities as training data to pre-train a preset initial model to obtain a pending image processing model includes: Each set of training data from multiple pre-training datasets of different granularities is sequentially input into the preset initial model to obtain the predicted text data and prediction boundary data corresponding to the training data; A first loss value is calculated using a first loss function based on the predicted text data and the text annotation data of the preset mask. The second loss value is calculated using the second loss function based on the predicted boundary data and the index-generated sample region, and the third loss value of the third loss function is determined based on the labeled region mask and the predicted region mask. The fourth loss value is determined using the second loss value and the third loss value through the fourth loss function; The preset initial model is iteratively trained based on the first loss value and the fourth loss value to obtain the image processing model to be determined.
13. An image processing apparatus, characterized in that, The device includes: The acquisition module is configured to acquire target detection data, which includes an image to be detected and image description information of the image to be detected. The image description information is either specified location information of the generated area in the image to be detected, or descriptive text of the object to be detected in the image to be detected. The determination module is configured to input the target detection data into the target image processing model to obtain the image processing result output by the target image processing model. The target image processing model is used to output the index expression text of the image corresponding to the index expression generation region when the image description information is the specified location information, and to output the target position of the index expression object in the image to be detected when the image description information is the description text. The target image processing model includes a visual encoder, a text encoder, a fusion encoder, a position detection network, and a text prediction network. The output of the visual encoder is coupled to the input of the fusion encoder, the output of the text encoder is coupled to the input of the fusion encoder, the output of the fusion encoder is coupled to the input of the position detection network, and the output of the fusion encoder is also coupled to the input of the text prediction network. The step of inputting the target detection data into the target image processing model to obtain the image processing result output by the target image processing model includes: When the image description information is the specified location information, the image to be detected and the specified location information are input into the visual encoder to obtain the image features output by the visual encoder and the regional location features of the region represented by the specified location information. Obtain a preset initial mask text and input the initial mask text into the text encoder to obtain the mask text features output by the text encoder; The image features, region location features, and masked text features are fused using the fusion encoder to obtain the first target features; The text prediction network generates the index expression text of the image corresponding to the index expression generation region based on the first target features.
14. A computer-readable medium having a computer program stored thereon, characterized in that, When the program is executed by the processing device, it implements the steps of the method described in any one of claims 1-12.
15. An electronic device, characterized in that, include: A storage device on which computer programs are stored; A processing device for executing the computer program in the storage device to implement the steps of the method according to any one of claims 1-12.