Method for Matching Processing of Image and Text and Electronic Device

By performing phase filling and multi-view feature extraction on the image data set, and combining the gating model for feature aggregation and alignment processing, the problem of image matching and text separation is solved, and accurate alignment and matching is achieved in complex scenarios.

CN119622368BActive Publication Date: 2025-08-01INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510164771.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-08-01
Estimated Expiration
2045-02-14

AI Technical Summary

Technical Problem

In the prior art, the matching of images and corresponding text is difficult to cope with complex scenarios, which easily leads to the separation of images and corresponding text matching.

Method used

By obtaining the image data set, performing phase filling processing, multiple viewing angle features are extracted, and a pre-generated gating model is input for multi-stage feature aggregation, text embedding features are extracted, and alignment is performed to determine the matching value between the image and the text, and optimize the matching degree.

Benefits of technology

It realizes accurate alignment between images and text in complex scenarios, solves the separation of images and corresponding text, and improves the accuracy and robustness of matching.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119622368B_ABST
    Figure CN119622368B_ABST
Patent Text Reader

Abstract

The present application provides a method and an electronic device for matching processing of images and texts, which relates to the field of computer technology. The method includes: obtaining an image data set; performing stage filling, multi-view feature extraction, and multi-stage feature aggregation on the images to obtain aggregated memory items; extracting text embedding features corresponding to the texts; performing alignment processing on the aggregated memory items and the embedding features to determine a matching value corresponding to the image set and the texts according to the aligned memory items and the aligned text embedding features; and optimizing the matching degree corresponding to the image set and the texts according to the matching value corresponding to the image set and the texts, so as to achieve precise alignment of the matching between the images and the corresponding text descriptions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular, to a method for matching and processing images and texts, and an electronic device. Background Art

[0002] An image can be a medical image, a traffic image, or other images. Taking a medical image as an example, it is to study the interaction between media such as X-rays and the human body, and present the internal tissue and organ structure, density, and other characteristics of the human body in the form of images. At the same time, the above-mentioned image and the corresponding text complement each other. The image is used to intuitively represent the image content, while the text is used to describe the corresponding image content.

[0003] In the related art, the matching of the image and the corresponding text is completed by integrating the image and the corresponding text. However, in the related art, it is difficult to handle complex scenarios by integrating the image and the corresponding text, and it is easy to cause the separation problem of the matching between the image and the corresponding text. Summary of the Invention

[0004] The embodiments of this application provide a method for matching and processing images and texts, and an electronic device, so as to achieve the effect of coping with complex scenarios and realizing accurate alignment of the matching between the image and the corresponding text description.

[0005] In a first aspect, the embodiments of this application provide a method for matching and processing images and texts, including:

[0006] Obtain an image data set, where the image data set includes one or more image samples, the image sample includes an image set and a text, and the image set includes images of one or more stages;

[0007] Perform stage filling processing on the image according to a preset threshold to obtain a filled image set;

[0008] Perform multi-view feature extraction processing on each stage image in the filled image set to obtain an extracted image set;

[0009] Input the extracted image set into a pre-generated gated model for multi-stage feature aggregation processing to obtain an aggregated memory item;

[0010] Extract the text embedding feature corresponding to the text;

[0011] Perform alignment processing on the aggregated memory item and the embedding feature to obtain an aligned memory item and an aligned text embedding feature;

[0012] Determine the matching value corresponding to the image set and the text according to the aligned memory item and the aligned text embedding feature;

[0013] Optimize the matching degree between the image set and the text according to the matching value corresponding to the image set and the text.

[0014] In a second aspect, an embodiment of the present application provides an electronic device, including: a memory and a processor;

[0015] The memory stores computer-executable instructions;

[0016] The processor executes the computer-executable instructions stored in the memory, so that the processor executes the above first aspect and / or various possible implementation manners of the first aspect.

[0017] In a third aspect, an embodiment of the present application provides a computer-readable storage medium, in which computer-executable instructions are stored, and when the computer-executable instructions are executed by a processor, they are used to implement the above first aspect and / or various possible implementation manners of the first aspect.

[0018] In a fourth aspect, a computer program product includes a computer program, characterized in that when the computer program is executed by a processor, it is used to implement the above first aspect and / or various possible implementation manners of the first aspect.

[0019] The image and text matching processing method and electronic device provided by the embodiments of the present application. The image and text matching processing method provided by this embodiment obtains an image data set, where the image data set includes one or more image samples, and the image samples include an image set and text, and the image set includes images of one or more stages; performs stage filling processing on the images according to a preset threshold to obtain a filled image set; performs multi-view feature extraction processing on the images of each stage in the filled image set to obtain an extracted image set; inputs the extracted image set into a pre-generated gating model for multi-stage feature aggregation processing to obtain an aggregated memory item; extracts the text embedding features corresponding to the text; performs alignment processing on the aggregated memory item and the embedding features to obtain an aligned memory item and an aligned text embedding feature; determines the matching value corresponding to the image set and the text according to the aligned memory item and the aligned text embedding feature; optimizes the matching degree between the image set and the text according to the matching value corresponding to the image set and the text, so as to handle complex scenarios and achieve precise alignment of the matching between the image and the corresponding text, and further solve the problem of separation between the image and the corresponding text. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The accompanying drawings here are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.

[0021] Figure 1Schematic diagram of the scenario of the image and text matching processing method provided by the embodiment of the present application;

[0022] Figure 2 Flow schematic of the image and text matching processing method provided by the embodiment of the present application Figure 1 ;

[0023] Figure 3 Flow schematic of the image and text matching processing method provided by the embodiment of the present application Figure 2 ;

[0024] Figure 4 Schematic diagram of the structure of the image and text matching processing device provided by the embodiment of the present application;

[0025] Figure 5 Schematic diagram of the structure of the electronic device provided by the embodiment of the present application.

[0026] Through the above-mentioned drawings, specific embodiments of the present application have been shown, and there will be more detailed descriptions hereinafter. These drawings and text descriptions are not intended to limit the scope of the concept of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. Detailed implementation manners

[0027] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0028] The image can be a medical image, a traffic image or other images. Taking a medical image as an example, it is to study the interaction between a certain medium (such as X-ray, electromagnetic field, ultrasonic wave, etc.) and the human body, and present the internal tissue and organ structure, density and other characteristics of the human body in the form of an image. At the same time, the image and the corresponding text complement each other. The image is used to intuitively represent the image content, while the text is used to describe the corresponding image content. In the related art, the matching of the image and the corresponding text is completed by integrating the image and the corresponding text. However, in the related art, it is difficult to handle complex scenarios by integrating the image and the corresponding text, and it is easy to cause the separation problem of the matching between the image and the corresponding text.

[0029] To solve the above technical problems, the embodiments of the present application propose the following technical concepts: The inventor considers the acquired image set, performs multi-view feature extraction processing on the images of one or more stages in the image set to obtain the extracted image set, and inputs the extracted image set into a gated model for multi-stage feature aggregation processing to obtain the aggregated memory items, extracts the text embedding features of the corresponding text, and uses the aligned memory items and the aligned text embedding features to determine the matching value between the image set and the text, so as to optimize the corresponding matching degree, enabling it to handle complex scenarios and achieve precise alignment of the matching between the image and the corresponding text description.

[0030] Figure 1 It is a schematic diagram of the application scenario of the image and text matching processing method provided by the embodiments of the present application.

[0031] Such as Figure 1 shown, this scenario includes: a display terminal 101 and a computer device 102.

[0032] Among them, the display terminal 101 can be a display or other display devices.

[0033] The computer device 102 can be an independent device or a cluster composed of multiple devices.

[0034] The computer device 102 acquires an image data set, performs stage filling processing on the images according to a preset threshold to obtain the filled image set; performs multi-view feature extraction processing on the images of each stage in the filled image set to obtain the extracted image set, and inputs it into a pre-generated gated model for multi-stage feature aggregation processing to obtain the aggregated memory items; extracts the text embedding features corresponding to the text; performs alignment processing on the aggregated memory items and the embedding features to obtain the aligned memory items and the aligned text embedding features; determines the matching value between the image set and the text according to the aligned memory items and the aligned text embedding features; optimizes the matching degree between the image set and the text according to the matching value between the image set and the text, and outputs it to the display terminal 101.

[0035] The following uses specific embodiments to elaborate in detail on the technical solutions of the present application and how the technical solutions of the present application solve the above technical problems. These several specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The following will describe the embodiments of the present application in conjunction with the accompanying drawings.

[0036] Figure 2 It is the flow schematic of the image and text matching processing method provided by the embodiments of the present application Figure 1 , the execution subject of this embodiment can be Figure 1The computer device in the illustrated embodiment is not particularly limited herein. For example, Figure 2 As shown, the method includes:

[0037] S201: Obtain an image data set, where the image data set includes one or more image samples, and each image sample includes an image set and text. The image set includes images of one or more stages.

[0038] In this embodiment, the image can be a medical image, a traffic image, or other images.

[0039] In this embodiment, the text is a corresponding medical record report, a traffic penalty statement, or other text.

[0040] In this embodiment, the image is a 3D image.

[0041] In this embodiment, the formula for obtaining the image data set includes:

[0042]

[0043] In the formula, is the image data set; is the image set in the i-th sample; is the description content of the text in the i-th sample; N is the number of image samples.

[0044] Correspondingly, the formula for the image set includes:

[0045]

[0046] In the formula, represents the medical image set of the i-th sample; is the medical image of the -th stage; ; is a real number; C is the number of channels; D is the depth of the image; H is the height of the image; W is the width of the image.

[0047] Exemplarily, for a grayscale image, C corresponds to 1 channel.

[0048] S202: Perform stage filling processing on the images according to a preset threshold to obtain a filled image set.

[0049] Specifically, determine whether the number of stages of the images in the image set is less than the preset threshold; if the number of stages of the images in the image set is less than the preset threshold, then perform stage filling processing on the images according to the preset threshold to obtain a filled image set.

[0050] Among them, the preset threshold is the preset value k.

[0051] In this embodiment, the image is subjected to stage filling processing according to a preset threshold to obtain a filled image set. The calculation formula includes:

[0052]

[0053] In the formula, is the filled image set corresponding to the i-th sample; is the number of images for the preset threshold pair; is a zero tensor, ; is a real number; C is the number of channels; D is the depth of the image; H is the height of the image; W is the width of the image.

[0054] S203: Perform multi-view feature extraction processing on each stage image in the filled image set to obtain an extracted image set.

[0055] Specifically, according to the visual coding network, multi-view feature extraction processing is performed on each stage image in the filled image set to obtain an extracted image set.

[0056] Exemplarily, the visual coding network is ViT (Vision Transformer) or 3D-CNN (3D Convolutional Neural Network).

[0057] In this embodiment, the calculation formula for performing multi-view feature extraction processing on each stage image set in the filled image set to obtain an extracted image set includes:

[0058]

[0059] In the formula, is the extracted image set; is each stage image set; ; is a real number; H is the dimension of the hidden feature.

[0060] S204: Input the extracted image set into a pre-generated gating model for multi-stage feature aggregation processing to obtain an aggregated memory item.

[0061] In this embodiment, the initial memory corresponding to the aggregated memory item is defined as: M0 = F1, where if the image set of the first stage is not available, it is initialized to zero.

[0062] S205: Extract the text embedding features corresponding to the text.

[0063] Specifically, the text embedding features corresponding to the text are extracted through a text encoder.

[0064] In this embodiment, the calculation formula for extracting the text embedding features corresponding to the text includes:

[0065]

[0066] Wherein, is the text embedding feature; is a real number; L is the length of the text sequence; H is the embedding dimension.

[0067] S206: Align the aggregated memory items and the embedding features to obtain the aligned memory items and the aligned text embedding features.

[0068] In this embodiment, the calculation formula for aligning the aggregated memory items and the text embedding features to obtain the aligned memory items and the aligned text embedding features includes:

[0069]

[0070] Wherein, is the aligned memory item, is determined according to the aggregated memory items; the aggregated memory items are obtained by performing multi-stage feature aggregation processing on through a gating model; is the aligned text embedding feature;

[0071] S207: Determine the matching value corresponding to the image set and the text according to the aligned memory items and the aligned text embedding features.

[0072] In this embodiment, the calculation formula for determining the matching value corresponding to the image set and the text according to the aligned memory items and the aligned text embedding features includes:

[0073]

[0074] Wherein, is the matching value corresponding to the image set and the text; is the number of image samples; is the similarity function; is the aligned memory item; is the i-th aligned text embedding feature; is the j-th aligned text embedding feature.

[0075] S208: Optimize the matching degree between the image set and the text according to the matching value corresponding to the image set and the text.

[0076] Specifically, according to the matching value corresponding to the image set and the text, use the contrastive learning loss method or the cross-entropy loss method to optimize the matching degree between the image set and the text.

[0077] In summary, the image-text matching processing method provided in this embodiment obtains an image dataset, where the image dataset includes one or more image samples, and each image sample includes an image set and text. The image set includes images of one or more stages. The images are subjected to stage filling processing according to a preset threshold to obtain a filled image set. Multiple perspective feature extraction processes are performed on the images of each stage in the filled image set to obtain an extracted image set. The extracted image set is input into a pre-generated gating model for multi-stage feature aggregation processing to obtain an aggregated memory item. The text embedding features corresponding to the text are extracted. Alignment processing is performed on the aggregated memory item and the embedding features to obtain an aligned memory item and aligned text embedding features. A matching value corresponding to the image set and the text is determined according to the aligned memory item and the aligned text embedding features. According to the matching value corresponding to the image set and the text, the matching degree between the image set and the text is optimized, so that complex scenarios can be handled, and accurate alignment of the matching between the image and the corresponding text can be achieved, further solving the problem of separation between the image and the corresponding text.

[0078] In addition, the image-text matching processing method provided in this embodiment performs alignment processing on the aggregated memory item and the embedding features to obtain an aligned memory item and aligned text embedding features, thereby realizing semantic mapping and precise association between the image and the text to ensure the maximization of matching similarity.

[0079] In addition, the image-text matching processing method provided in this embodiment performs stage filling processing on the images according to a preset threshold to obtain a filled image set. By filling in samples with insufficient stage numbers and at the same time adopting a dynamic selection strategy based on maximizing diversity and semantic relevance, the core stage information is ensured to be completely retained while reducing the interference of redundant stages.

[0080] In addition, the image-text matching processing method provided in this embodiment combines a dynamic gating mechanism with contrastive learning, enabling the gating model to learn compact and discriminative feature representations in the multi-modal embedding space, thereby improving the accuracy and robustness of image-text alignment.

[0081] In addition, the image-text matching processing method provided in this embodiment can capture the spatial association between different anatomical planes (such as axial, sagittal, and coronal) of the image and the temporal dynamics between image stages by using a visual encoder for multi-image feature extraction. The gating mechanism further enhances the dynamic fusion ability of key spatio-temporal semantic information, providing the gating model with comprehensive image understanding ability.

[0082] In addition, the image-text matching processing method provided in this embodiment can adapt to the diagnostic and auxiliary decision-making requirements in different clinical scenarios.

[0083] Figure 3 Schematic diagram of the process of the image and text matching processing method provided by the embodiment of the present application Figure 2 . In the embodiment of the present application, on the basis of the provided embodiment, a detailed description is given of the specific implementation method of the generation process of the gating model in step S205. As Figure 2 shown, the method includes: Figure 3 shown, the method includes:

[0084] S301: Extract images of multiple stages from any image sample in the image dataset, where the images of each stage include corresponding multiple perspective features.

[0085] S302: Perform extraction processing on each perspective feature to obtain corresponding first feature extraction results, second feature extraction results, and effective perspective masks.

[0086] In this embodiment, the calculation formulas for performing extraction processing on each perspective feature to obtain corresponding first feature extraction results, second feature extraction results, and effective perspective masks include:

[0087]

[0088] In the formula, is the first feature extraction result; is the visual feature tensor of the t-th perspective; is a real number; is the batch size; is the dimension of the hidden feature.

[0089]

[0090] In the formula, is the second feature extraction result; is the previous memory tensor;

[0091]

[0092] In the formula, is the effective perspective mask.

[0093] In this embodiment, the second feature extraction result is the old memory.

[0094] Among them, the first feature extraction result can be extracted from different scanning angles or modalities (such as CT, MRI).

[0095] Among them, CT is computed tomography; MRI is magnetic resonance imaging.

[0096] Among them, the second feature extraction result is the image information of the previous perspective extracted from the memory unit storage of the gating model.

[0097] Among them, the effective view mask is used to label some invalid or missing views, and the invalid or missing views are the un-imaged areas of the patient-specific region.

[0098] S303: Determine the final memory update item according to the first feature extraction result and the second feature extraction result.

[0099] Specifically, step S303 specifically includes:

[0100] S3031: Determine the reset gate vector and the update gate vector according to the first feature extraction result and the second feature extraction result.

[0101] In this embodiment, the calculation formulas for determining the reset gate vector and the update gate vector according to the first feature extraction result and the second feature extraction result include:

[0102]

[0103] In the formula, is the update gate vector; is the reset gate vector; is the first feature extraction result; is the second feature extraction result; is the activation function; is the learnable convolution kernel; * is the two-dimensional convolution operation.

[0104] Among them, the update gate vector and the reset gate vector dynamically filter information through weight adjustment, so as to ensure that the gated model can efficiently utilize important information when processing multi-view features.

[0105] Among them, the update gate vector and the reset gate vector are calculated through the Sigmoid activation function σ. The function of the Sigmoid function is to map the input value to the interval [0,1], so as to provide an appropriate proportional range for the weight.

[0106] Among them, is the learnable convolution kernel, and its shape is where K is the size of the convolution kernel.

[0107] Exemplarily, K = 3 represents a 3×3 convolution kernel.

[0108] S3032: Determine the corresponding candidate memory value according to the reset gate vector, the first feature extraction result and the second feature extraction result.

[0109] In this embodiment, the calculation formulas for determining the corresponding candidate memory value according to the reset gate vector, the first feature extraction result and the second feature extraction result include:

[0110]

[0111] Wherein, is the candidate memory value; is the hyperbolic tangent activation function; is the preset parameter; is the first feature extraction result; is the reset gate vector; is the element-wise multiplication; is the second feature extraction result; is the reset memory.

[0112] Among them, the candidate memory is generated by integrating the first feature extraction result and the reset memory. The purpose of this step is to effectively combine the current input information with the historical memory information, thereby constructing a candidate memory state. The reset operation here is controlled by the reset gate vector and implemented through the element-wise multiplication of the second feature extraction result. This operation can dynamically adjust the contribution degree of historical information.

[0113] In addition, in order to enhance the non-linear expression ability of the candidate memory, the hyperbolic tangent activation function is used, which can map the input value to range, endowing the model with stronger feature modeling ability to capture complex patterns, so as to better adapt to the changes and correlations of multi-perspective features.

[0114] S3033: Determine the corresponding memory update item according to the update gate vector, the second feature extraction result and the candidate memory value.

[0115] In this embodiment, the calculation formula for determining the corresponding memory update item according to the update gate vector, the second feature extraction result and the candidate memory value includes:

[0116]

[0117] Wherein, is the memory update item; is the update gate vector; is the element-wise multiplication; is the second feature extraction result; is the candidate memory value.

[0118] Among them, the memory update item is obtained by a weighted combination of the second feature extraction result and the candidate memory value. This weighting process is dynamically controlled by the update gate vector, where the value of the update gate vector determines the degree of memory update at the current moment. When the update gate vector approaches 1, it is more inclined to introduce the information of the candidate memory value; when the update gate vector approaches 0, it is more inclined to retain the old memory corresponding to the second feature extraction result. This weighting mechanism allows the gated model to dynamically adjust the information flow from different perspectives, thereby effectively balancing the contributions of historical information and new input information during the memory update process.

[0119] S3034: Normalize the memory update item to obtain a preliminarily normalized memory update item.

[0120] Specifically, step S3034 specifically includes:

[0121] S30341: Obtain the dimension of the hidden feature corresponding to the memory update item and each feature value.

[0122] S30342: Determine the mean value corresponding to the memory update item according to the dimension of the hidden feature and each feature value.

[0123] In this embodiment, the calculation formula for determining the mean value corresponding to the memory update item according to the dimension of the hidden feature and each feature value includes:

[0124]

[0125] In the formula, is the mean value; is the dimension of the hidden feature; is the value of the

[0126] Exemplarily, .

[0127] S30343: Determine the variance corresponding to the memory update item according to the dimension of the hidden feature, each feature value, and the mean value.

[0128] In this embodiment, the calculation formula for determining the variance corresponding to the memory update item according to the dimension of the hidden feature, each feature value, and the mean value includes:

[0129]

[0130] In the formula, is the variance; is the dimension of the hidden feature; is the value of the feature

[0131] S30344: Perform normalization based on the memory update item, mean, and variance to obtain a preliminarily normalized memory update item.

[0132] In this embodiment, the calculation formula for performing normalization based on the memory update item, mean, and variance to obtain a preliminarily normalized memory update item includes:

[0133]

[0134] In the formula, is the preliminarily normalized memory update item; is the memory update item; is the value of the th feature; is the mean; is the variance;

[0135]

[0136] where the variance is used to measure the range of change of the memory in the hidden dimension. S3035: Perform a linear transformation on the preliminarily normalized memory update item to obtain a finally normalized memory update item.

[0137] In this embodiment, the calculation formula for performing a linear transformation on the preliminarily normalized memory update item to obtain a finally normalized memory update item includes:

[0138]

[0139] In the formula, is the finally normalized memory update item; is the preliminarily normalized memory update item; and are learnable parameters.

[0140] where the shapes of γ and β are . Their role is to restore the feature distribution of the model before normalization, while retaining the numerical stability and optimization efficiency brought by normalization, so as to ensure that the updated memory not only has good numerical stability but also can flexibly adjust the feature distribution.

[0141] S3036: Perform a residual connection on the finally normalized memory update item according to the first feature extraction result to obtain the final memory update item.

[0142] In this embodiment, the calculation formula for performing a residual connection on the finally normalized memory update item according to the first feature extraction result to obtain the final memory update item includes:

[0143]

[0144] In the formula, is the final memory update item; is the finally normalized memory update item; is the parameter obtained by pre-training; is the first feature extraction result.

[0145] In this embodiment, the residual connection is used to directly add it to to form a skip connection. This operation can effectively prevent information from gradually being lost during multiple non-linear transformations, ensuring that the information of the original input features can be directly passed to subsequent layers. At the same time, the residual connection can also alleviate the problem of gradient disappearance in deep networks, so as to improve the training efficiency of the model and the stability of feature expression.

[0146] S304: Process the effective view mask according to the dimension of the hidden feature to obtain the processed effective view mask.

[0147] Specifically, step S304 specifically includes:

[0148] S3041: Expand the effective view mask according to the dimension of the hidden feature to obtain the expanded effective view mask;

[0149] In this embodiment, the formula for expanding the effective view mask according to the dimension of the hidden feature to obtain the expanded effective view mask includes:

[0150]

[0151] In the formula, is the expanded effective view mask; is the effective view mask; with a shape of ; is the batch size.

[0152] Among them, is a binary vector, where 1 indicates that the view is available and 0 indicates that the view is invalid. Through expansion and broadcasting, the mask is applied to all feature dimensions, thus ensuring that invalid views have no contribution to the result in subsequent calculations.

[0153] S3042: Perform broadcast processing on the expanded effective view mask to obtain the processed effective view mask.

[0154] In this embodiment, performing broadcast processing on the expanded effective view mask to obtain the processed effective view mask includes:

[0155]

[0156] In the formula, is the processed effective view mask; The expanded effective view mask; is a row vector of length consisting entirely of 1s; with a shape of ; is the batch size; is the dimension of the hidden feature.

[0157] Among them, serves to exclude the information of invalid views, thus avoiding interference with the model.

[0158] S305: Determine the reset gate vector of the mask and the update gate vector of the mask according to the first feature extraction result, the final memory update item, and the processed effective view mask.

[0159] In this embodiment, the calculation formula for determining the reset gate vector of the mask and the update gate vector of the mask according to the first feature extraction result, the final memory update item, and the processed effective view mask includes:

[0160]

[0161] In the formula, is the update gate vector of the mask; is the reset gate vector of the mask; is the first feature extraction result; is the second feature extraction result; is the activation function; is the learnable convolution kernel; is the processed effective view mask.

[0162] S306: Determine the memory update item of the corresponding mask according to the reset gate vector of the mask, the update gate vector of the mask, the final memory update item, the first feature extraction result, and the processed effective view mask.

[0163] Specifically, step S306 specifically includes:

[0164] S3061: Determine the candidate memory value of the mask according to the reset gate vector of the mask, the final memory update item, the first feature extraction result, and the processed effective view mask.

[0165] In this embodiment, the calculation formula for determining the candidate memory value of the mask according to the reset gate vector of the mask, the final memory update item, the first feature extraction result, and the processed effective view mask includes:

[0166]

[0167] In the formula, is the candidate memory value of the mask; is the hyperbolic tangent activation function; is the preset parameter; is the first feature extraction result; is the reset gate vector of the mask; is the element-wise multiplication; is the final memory update term.

[0168] Among them, this formula realizes the dynamic update of memory by weighted combination of the old memory and the candidate memory. The update gate vector ensures that the model can adjust the update degree of memory according to the importance of the current features, while ensures that the features of the invalid perspective do not participate in the memory update.

[0169] S3062: Determine the memory update term of the corresponding mask according to the update gate vector of the mask, the final memory update term, and the candidate memory value of the mask.

[0170] In this embodiment, the calculation formula for determining the memory update term of the corresponding mask according to the update gate vector of the mask, the final memory update term, and the candidate memory value of the mask includes:

[0171]

[0172] In the formula, is the memory update term of the mask; is the part of the old memory retained proportionally; is the part of the candidate memory introduced proportionally; is the update gate vector of the mask; is the final memory update term; is the element-wise multiplication; is the candidate memory value of the mask.

[0173] S307: Determine the memory update term of the mask as the gated model.

[0174] In summary, the image and text matching processing method provided in this embodiment extracts images at multiple stages from any image sample in the image data set, where the images at each stage include corresponding multiple perspective features; performs extraction processing on each perspective feature to obtain corresponding first feature extraction results, second feature extraction results, and effective perspective masks; determines the final memory update item according to the first feature extraction result and the second feature extraction result; processes the effective perspective mask according to the dimension of the hidden feature to obtain the processed effective perspective mask; determines the reset gate vector of the mask and the update gate vector of the mask according to the first feature extraction result, the final memory update item, and the processed effective perspective mask; determines the memory update item of the corresponding mask according to the reset gate vector of the mask, the update gate vector of the mask, the final memory update item, the first feature extraction result, and the processed effective perspective mask; determines the memory update item of the mask as the gated model, which can effectively capture the dynamic changes between stages, while suppressing the interference of noise and redundant features, and is particularly suitable for integrating multi-stage images reflecting temporal features such as contrast agent distribution and lesion dynamics.

[0175] In addition, the image and text matching processing method provided in this embodiment dynamically adjusts the update gate vector and the reset gate vector through a gating mechanism, and accurately screens important information and suppresses noise and invalid features when gradually integrating features from different perspectives and stages.

[0176] In addition, the image and text matching processing method provided in this embodiment controls the interference of invalid perspectives by using a validity mask, so as to achieve efficient fusion and dynamic update of multi-view image features. In addition, the validity mask can also ensure through element-wise operations that the features of invalid perspectives or stages do not affect memory update, thereby avoiding the interference of data heterogeneity and missing information on the model performance.

[0177] Figure 4 It is a schematic structural diagram of the image and text matching processing device provided in the embodiment of the present application. As Figure 4 shown, the image and text matching processing device provided in this embodiment includes: an acquisition module 401, a filling module 402, a first extraction module 403, an aggregation module 404, a second extraction module, an alignment module 406, a first determination module 407, and an optimization module 408.

[0178] The acquisition module 401 is used to acquire an image data set, where the image data set includes one or more image samples, the image samples include an image set and text, and the image set includes images at one or more stages;

[0179] The filling module 402 performs stage filling processing on the image according to a preset threshold to obtain a filled image set;

[0180] The first extraction module 403 is configured to perform multiple perspective feature extraction processes on each stage image in the filled image set to obtain an extracted image set;

[0181] The aggregation module 404 is configured to input the extracted image set into a pre-generated gating model for multi-stage feature aggregation processing to obtain an aggregated memory item;

[0182] The second extraction module 405 is configured to extract text embedding features corresponding to the text;

[0183] The alignment module 406 is configured to perform alignment processing on the aggregated memory item and the embedding features to obtain an aligned memory item and aligned text embedding features;

[0184] The first determination module 407 is configured to determine a matching value corresponding to the image set and the text according to the aligned memory item and the aligned text embedding features;

[0185] The optimization module 408 is configured to optimize the matching degree corresponding to the image set and the text according to the matching value corresponding to the image set and the text.

[0186] In a possible implementation manner, the formula for obtaining the image data set includes:

[0187]

[0188] Wherein, is the image data set; is the image set in the i-th sample; is the description content of the text in the i-th sample; N is the number of image samples;

[0189] Correspondingly, the formula for the image set includes:

[0190]

[0191] Wherein, represents the medical image set of the i-th sample; is the th stage of medical image; ; is a real number; C is the number of channels; D is the depth of the image; H is the height of the image; W is the width of the image;

[0192] Correspondingly, the calculation formula for performing stage filling processing on the image according to a preset threshold to obtain a filled image set includes:

[0193]

[0194] wherein is the image set after filling corresponding to the i-th sample; is the number of images corresponding to the preset threshold; is a zero tensor, ; is a real number; C is the number of channels; D is the depth of the image; H is the height of the image; W is the width of the image.

[0195] In a possible implementation manner, the calculation formula for performing multi-view feature extraction processing on each stage image set in the filled image set to obtain the extracted image set includes:

[0196]

[0197] wherein is the extracted image set; the is each stage image set; ; is a real number; H is the dimension of the hidden feature;

[0198] Correspondingly, the calculation formula for extracting the text embedding feature corresponding to the text includes:

[0199]

[0200] wherein is the text embedding feature; is a real number; L is the length of the text sequence; H is the embedding dimension;

[0201] Correspondingly, the calculation formula for aligning the aggregated memory item and the text embedding feature to obtain the aligned memory item and the aligned text embedding feature includes:

[0202]

[0203] wherein is the aligned memory item, is determined according to the aggregated memory item; the aggregated memory item is obtained by performing multi-stage feature aggregation processing on through the gating model; is the aligned text embedding feature;

[0204] Correspondingly, the calculation formula for determining the matching value between the image set and the text according to the aligned memory item and the aligned text embedding feature includes:

[0205]

[0206] wherein is the matching value corresponding to the image set and the text; is the number of image samples; is the similarity function; is the aligned memory item; is the i-th aligned text embedding feature; is the j-th aligned text embedding feature.

[0207] In a possible implementation, the device further includes:

[0208] A third extraction module 409, configured to extract images of multiple stages from any image sample in the image data set, where the images of each stage include corresponding multiple perspective features;

[0209] A fourth extraction module 410, configured to perform extraction processing on each perspective feature to obtain corresponding first feature extraction results, second feature extraction results, and an effective perspective mask;

[0210] A second determination module 411, configured to determine a final memory update item according to the first feature extraction result and the second feature extraction result;

[0211] A processing module 412, configured to process the effective perspective mask according to the dimension of the hidden feature to obtain a processed effective perspective mask;

[0212] A third determination module 413, configured to determine a reset gate vector of the mask and an update gate vector of the mask according to the first feature extraction result, the final memory update item, and the processed effective perspective mask;

[0213] A fourth determination module 414, configured to determine a memory update item of the corresponding mask according to the reset gate vector of the mask, the update gate vector of the mask, the final memory update item, the first feature extraction result, and the processed effective perspective mask;

[0214] A fifth determination module 415, configured to determine the memory update item of the mask as a gated model.

[0215] In a possible implementation, the calculation formula for performing extraction processing on each perspective feature to obtain corresponding first feature extraction results, second feature extraction results, and an effective perspective mask includes:

[0216]

[0217] In the formula, is the first feature extraction result; is the visual feature tensor of the t-th perspective; is a real number; is the batch size; is the dimension of the hidden feature;

[0218]

[0219] In the formula, is the second feature extraction result; is the previous memory tensor;

[0220]

[0221] In the formula, is the effective view mask.

[0222] In a possible implementation manner, the calculation formula for determining the mask reset gate vector and the mask update gate vector according to the first feature extraction result, the final memory update item, and the processed effective view mask includes:

[0223]

[0224] In the formula, is the mask update gate vector; is the mask reset gate vector; is the first feature extraction result; is the second feature extraction result; is the activation function; is the learnable convolution kernel; is the processed effective view mask.

[0225] In a possible implementation manner, the second determination module 411 specifically includes:

[0226] The first determination unit 4111 is configured to determine the reset gate vector and the update gate vector according to the first feature extraction result and the second feature extraction result;

[0227] The second determination unit 4112 is configured to determine the corresponding candidate memory value according to the reset gate vector, the first feature extraction result, and the second feature extraction result;

[0228] The third determination unit 4113 is configured to determine the corresponding memory update item according to the update gate vector, the second feature extraction result, and the candidate memory value;

[0229] The normalization unit 4114 is configured to perform normalization processing on the memory update item to obtain a preliminarily normalized memory update item;

[0230] The transformation unit 4115 is configured to perform a linear transformation process on the preliminarily normalized memory update item to obtain a finally normalized memory update item;

[0231] The connection unit 4116 is configured to perform a residual connection process on the finally normalized memory update item according to the first feature extraction result to obtain a final memory update item.

[0232] In a possible implementation manner, the calculation formula for determining the reset gate vector and the update gate vector according to the first feature extraction result and the second feature extraction result includes:

[0233]

[0234] In the formula, is the update gate vector; is the reset gate vector; is the first feature extraction result; is the second feature extraction result; is the activation function; is the learnable convolution kernel;

[0235] Correspondingly, the calculation formula for determining the corresponding candidate memory value according to the reset gate vector, the first feature extraction result, and the second feature extraction result includes:

[0236]

[0237] In the formula, is the candidate memory value; is the hyperbolic tangent activation function; is the preset parameter; is the first feature extraction result; is the reset gate vector; is the element-wise multiplication; is the second feature extraction result;

[0238] Correspondingly, the calculation formula for determining the corresponding memory update item according to the update gate vector, the second feature extraction result, and the candidate memory value includes:

[0239]

[0240] In the formula, is the memory update item; is the update gate vector; is the element-wise multiplication; is the second feature extraction result; is the candidate memory value;

[0241] Correspondingly, the calculation formula for linearly transforming the preliminarily normalized memory update item to obtain the finally normalized memory update item includes:

[0242]

[0243] In the formula, is the finally normalized memory update item; is the preliminarily normalized memory update item; and are learnable parameters;

[0244] Correspondingly, the calculation formula for performing residual connection processing on the finally normalized memory update item according to the first feature extraction result to obtain the final memory update item includes:

[0245]

[0246] In the formula, is the final memory update item; is the finally normalized memory update item; is the parameter obtained by pre-training; is the first feature extraction result.

[0247] In a possible implementation manner, the normalization unit 4114 specifically includes:

[0248] An acquisition unit 41141, configured to acquire the dimension of the hidden feature corresponding to the memory update item and each feature value;

[0249] A first determination unit 41142, configured to determine the mean value corresponding to the memory update item according to the dimension of the hidden feature and each feature value;

[0250] A second determination unit 41143, configured to determine the variance corresponding to the memory update item according to the dimension of the hidden feature, each feature value, and the mean value;

[0251] A normalization unit 41144, configured to perform normalization processing according to the memory update item, the mean value, and the variance to obtain a preliminarily normalized memory update item.

[0252] In a possible implementation manner, the calculation formula for determining the mean value corresponding to the memory update item according to the dimension of the hidden feature and each feature value includes:

[0253]

[0254] In the formula, is the mean value; is the dimension of the hidden feature; is the value of the th feature;

[0255] Accordingly, the calculation formula for determining the variance corresponding to the memory update item according to the dimension of the hidden feature, each feature value, and the mean value includes:

[0256]

[0257] In the formula, is the variance; is the dimension of the hidden feature; is the th feature value [[ID=2②]]is the mean value;

[0258] Accordingly, the calculation formula for performing normalization processing according to the memory update item, the mean value, and the variance to obtain a preliminarily normalized memory update item includes:

[0259]

[0260] In the formula, is the preliminarily normalized memory update item; is the memory update item; is the th feature value is the mean value; is the variance; is a constant used for numerical stability to avoid a zero denominator.

[0261] In a possible implementation manner, the processing module 412 specifically includes:

[0262] An expansion unit 4121, configured to expand the effective view mask according to the dimension of the hidden feature to obtain an expanded effective view mask;

[0263] A processing unit 4122, configured to perform broadcast processing on the expanded effective view mask to obtain a processed effective view mask.

[0264] In a possible implementation manner, the calculation formula for expanding the effective view mask according to the dimension of the hidden feature to obtain an expanded effective view mask includes:

[0265]

[0266] In the formula, is the expanded effective view mask; is the effective view mask; with a shape of ; is the batch size;

[0267] Correspondingly, the broadcast processing of the expanded valid view mask is performed to obtain the processed valid view mask, including:

[0268]

[0269] wherein is the processed valid view mask; the expanded valid view mask; is a row vector of length consisting entirely of 1s; with a shape of ; is the batch size; is the dimension of the hidden feature.

[0270] The image and text matching processing device provided in this embodiment can execute the method provided in the above method embodiment, and its implementation principle and technical effect are similar, which will not be elaborated here in this embodiment.

[0271] Figure 5 is a schematic structural diagram of the electronic device provided in this application. As Figure 5 shown, the electronic device 50 provided in this embodiment includes: at least one processor 501 and a memory 502. Optionally, the device 50 further includes a communication component 503. Among them, the processor 501, the memory 502, and the communication component 503 are connected through a bus.

[0272] In a specific implementation process, at least one processor 501 executes the computer execution instructions stored in the memory 502, so that at least one processor 501 executes the above method.

[0273] For the specific implementation process of the processor 501, reference can be made to the above method embodiment, and its implementation principle and technical effect are similar, which will not be elaborated here in this embodiment.

[0274] In the above embodiment, it should be understood that the processor may be a central processing unit (Central Processing Unit, abbreviated as: CPU), or other general-purpose processors, digital signal processors (Digital Signal Processor, abbreviated as: DSP), application specific integrated circuits (Application Specific Integrated Circuit, abbreviated as: ASIC), etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the application can be directly implemented by the execution of the hardware processor, or can be implemented by the combination of the hardware and software modules in the processor.

[0275] The memory may include a Random Access Memory (RAM), and may also include a Non-volatile Memory (NVM), such as at least one disk memory.

[0276] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, the buses in the drawings of this application are not limited to only one bus or one type of bus.

[0277] The embodiments of this application also provide a computer program product, including a computer program, which implements the above method when executed by a processor.

[0278] The embodiments of this application also provide a computer-readable storage medium, in which computer-executable instructions are stored, and when the processor executes the computer-executable instructions, the above method is implemented.

[0279] The above-mentioned readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic memory, flash memory, a magnetic disk or an optical disc. The readable storage medium can be any available medium accessible by a general-purpose or special-purpose computer.

[0280] An exemplary readable storage medium is coupled to the processor, enabling the processor to read information from the readable storage medium and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can be located in an Application Specific Integrated Circuit (ASIC). Of course, the processor and the readable storage medium can also exist as discrete components in a device.

[0281] The division of units is merely a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the couplings or direct couplings or communication connections shown or discussed among each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical or other forms.

[0282] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0283] In addition, in each embodiment of this application, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0284] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art or part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of this application. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks or optical discs and other various media that can store program codes.

[0285] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed through hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When this program is executed, it executes the steps including the above method embodiments; and the aforementioned storage medium includes: ROM, RAM, magnetic disks or optical discs and other various media that can store program codes.

[0286] Finally, it should be noted that those skilled in the art will readily conceive of other embodiments of the present application after considering the specification and practicing the application disclosed herein. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include common general knowledge or conventional technical means in the technical field not disclosed in the present application. It is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.

Claims

1. A method for matching and processing images and texts, characterized in that, Including: Obtain an image dataset, where the image dataset includes one or more image samples, the image samples include an image set and text, and the image set includes images of one or more stages; Perform stage filling processing on the images according to a preset threshold to obtain a filled image set; Extract multi-view features from the images of each stage in the filled image set to obtain an extracted image set; Input the extracted image set into a pre-generated gated model for multi-stage feature aggregation to obtain an aggregated memory item; Extract the text embedding features corresponding to the text; Perform alignment processing on the aggregated memory item and the embedding features to obtain an aligned memory item and aligned text embedding features; Determine a matching value corresponding to the image set and the text according to the aligned memory item and the aligned text embedding features; Optimize the corresponding matching degree according to the matching value; The generation process of the gated model includes: extracting first feature extraction results, second feature extraction results, and valid view masks for each view feature; the first feature extraction result is the current image information, and the second feature extraction result is the historical image information; Determine the final memory update item according to the first feature extraction result and the second feature extraction result; Process the valid view mask according to the dimension of the hidden feature to obtain a processed valid view mask; Determine a reset gate vector and an update gate vector of the mask according to the first feature extraction result, the final memory update item, and the processed valid view mask; Determine a memory update item of the corresponding mask according to the reset gate vector and update gate vector of the mask, the final memory update item, the first feature extraction result, and the processed valid view mask; Determine the memory update item of the mask as the gated model.

2. The method according to claim 1, characterized in that, The formula for obtaining the image dataset includes: In the formula, is the image data set; is the image set in the i-th sample; is the description content of the text in the i-th sample; N is the number of image samples; Correspondingly, the formula for the image set includes: In the formula, represents the medical image set of the i-th sample; is the medical image of the -th stage; ; is a real number; C is the number of channels; D is the depth of the image; H is the height of the image; W is the width of the image; Correspondingly, the calculation formula for performing stage filling processing on the images according to a preset threshold to obtain a filled image set includes: Wherein, is the filled image set corresponding to the i-th sample; is the image corresponding to the preset threshold pair quantity; is a zero tensor, ; is a real number; C is the number of channels; D is the depth of the image; H is the height of the image; W is the width of the image.

3. The method according to claim 1, wherein The calculation formula for extracting multi-view features from the images of each stage in the filled image set to obtain an extracted image set includes: In the formula, is the extracted image set; the is the image set at each stage; ; is a real number; H is the dimension of the hidden feature; Correspondingly, the calculation formula for extracting the text embedding features corresponding to the text includes: In the formula, is the text embedding feature; is a real number; L is the length of the text sequence; H is the embedding dimension; Correspondingly, the calculation formula for performing alignment processing on the aggregated memory item and the text embedding features to obtain an aligned memory item and aligned text embedding features includes: In the formula, is the aligned memory item, is determined according to the aggregated memory items; the aggregated memory items are obtained by performing multi-stage feature aggregation processing on; is the aligned text embedding feature; Correspondingly, the calculation formula for determining the matching value corresponding to the image set and the text according to the aligned memory item and the aligned text embedding features includes: Wherein, is the matching value corresponding to the image set and the text; is the number of image samples; is the similarity function; is the aligned memory item; is the i-th aligned text embedding feature; is the j-th aligned text embedding feature.

4. The method according to claim 1, characterized in that, The calculation formula for extracting each view feature to obtain the corresponding first feature extraction result, second feature extraction result, and valid view mask includes: Wherein, is the first feature extraction result; is the visual feature tensor of the t-th perspective; is a real number; is the batch size; is the dimension of the hidden feature; In the formula, is the second feature extraction result; is the previous memory tensor; In the formula, is the effective viewing angle mask.

5. The method according to claim 1, characterized in that The calculation formula for determining the reset gate vector of the mask and the update gate vector of the mask according to the first feature extraction result, the final memory update item, and the processed valid view mask includes: In the formula, is the update gate vector of the mask; is the reset gate vector of the mask; is the first feature extraction result; is the second feature extraction result; is the activation function; is the learnable convolution kernel; is the processed effective view mask.

6. The method according to claim 1, wherein Determining the final memory update item according to the first feature extraction result and the second feature extraction result includes: Determining a reset gate vector and an update gate vector according to the first feature extraction result and the second feature extraction result; Determining a corresponding candidate memory value according to the reset gate vector, the first feature extraction result and the second feature extraction result; Determining a corresponding memory update item according to the update gate vector, the second feature extraction result and the candidate memory value; Normalizing the memory update item to obtain a preliminarily normalized memory update item; Performing a linear transformation on the preliminarily normalized memory update item to obtain a finally normalized memory update item; Performing a residual connection on the finally normalized memory update item according to the first feature extraction result to obtain the final memory update item.

7. The method according to claim 6, wherein The calculation formula for determining the reset gate vector and the update gate vector according to the first feature extraction result and the second feature extraction result includes: In the formula, is the update gate vector; is the reset gate vector; is the first feature extraction result; is the second feature extraction result; is the activation function; is the learnable convolution kernel; Correspondingly, the calculation formula for determining a corresponding candidate memory value according to the reset gate vector, the first feature extraction result and the second feature extraction result includes: Wherein, is the candidate memory value; is the hyperbolic tangent activation function; is the preset parameter; is the first feature extraction result; is the reset gate vector; is the element-wise multiplication; is the second feature extraction result; Correspondingly, the calculation formula for determining a corresponding memory update item according to the update gate vector, the second feature extraction result and the candidate memory value includes: In the formula, is the memory update item; is the update gate vector; is the element-wise multiplication; is the second feature extraction result; is the candidate memory value; Correspondingly, the calculation formula for performing a linear transformation on the preliminarily normalized memory update item to obtain a finally normalized memory update item includes: wherein, is the finally normalized memory update term; is the preliminarily normalized memory update term; and are learnable parameters; Correspondingly, the calculation formula for performing a residual connection on the finally normalized memory update item according to the first feature extraction result to obtain the final memory update item includes: Wherein, is the final memory update item; is the finally normalized memory update item; is the parameter obtained by pre-training; is the first feature extraction result.

8. The method according to claim 6, wherein Normalizing the memory update item to obtain a preliminarily normalized memory update item includes: Obtaining the dimension of the hidden feature corresponding to the memory update item and each eigenvalue; Determining the mean value corresponding to the memory update item according to the dimension of the hidden feature and each eigenvalue; Determining the variance corresponding to the memory update item according to the dimension of the hidden feature, each eigenvalue and the mean value; Performing a normalization process on the memory update item, the mean value and the variance to obtain a preliminarily normalized memory update item.

9. The method according to claim 8, characterized in that, The calculation formula for determining the mean value corresponding to the memory update item according to the dimension of the hidden feature and each eigenvalue includes: In the formula, is the mean value; is the dimension of the hidden feature; is the th value of the feature; Correspondingly, the calculation formula for determining the variance corresponding to the memory update item according to the dimension of the hidden feature, each eigenvalue and the mean value includes: In the formula, is the variance; is the dimension of the hidden feature; is the -th value of the feature is the mean; Correspondingly, the calculation formula for performing a normalization process on the memory update item, the mean value and the variance to obtain a preliminarily normalized memory update item includes: In the formula, is the preliminary normalized memory update term; is the memory update term; is the value of the th feature; is the mean; is a constant used for numerical stability to avoid a zero denominator.

10. The method according to claim 1, characterized in that, Processing the effective view mask according to the dimension of the hidden feature to obtain a processed effective view mask includes: Expanding the effective view mask according to the dimension of the hidden feature to obtain an expanded effective view mask; Performing a broadcast process on the expanded effective view mask to obtain a processed effective view mask.

11. The method according to claim 10, wherein The calculation formula for expanding the effective view mask according to the dimension of the hidden feature to obtain the expanded effective view mask includes: In the formula, is the expanded effective view mask; is the effective view mask; the shape is ; is the batch size; Correspondingly, the broadcast processing of the expanded effective view mask to obtain the processed effective view mask includes: wherein, is the processed effective view mask; is the expanded effective view mask; is a row vector of length consisting entirely of 1s; with a shape of ; is the batch size; is the dimension of the hidden feature.

12. An electronic device, characterized in that, Including: A memory and a processor; The memory stores computer execution instructions; The processor executes the computer execution instructions stored in the memory to implement the image and text matching processing method according to any one of claims 1 to 11.

13. A computer-readable storage medium, characterized in that, Computer execution instructions are stored in the computer-readable storage medium, and when the computer execution instructions are executed by a processor, they are used to implement the image and text matching processing method according to any one of claims 1 to 11.

14. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the image and text matching processing method according to any one of claims 1 to 11.

Citation Information

Patent Citations

  • Image report generation method, apparatus and device, and computer readable storage medium

    CN118136197A