Image retrieval enhancement method and device based on bimodal matching, equipment and medium

By adopting the dual-modal matching method in the image retrieval technology, and using the language image model to perform shard coding and contrast training of images and text, the problem of insufficient image retrieval accuracy and efficiency in the prior art is solved, and more efficient and accurate image retrieval is achieved.

CN120067355AActive Publication Date: 2025-05-30SHENZHEN PINGAN COMM TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510510206.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-05-30
Estimated Expiration
2045-04-23

AI Technical Summary

Technical Problem

The existing image retrieval technology in the fields of healthcare and finance has limitations in terms of accuracy and efficiency, and it is necessary to improve the accuracy of image retrieval.

Method used

Using an image retrieval enhancement method based on dual-modal matching, images and text are fragmented encoding through the preset language image model, sharded image feature sets and sharded text feature sets are generated, and the language image model is compared and trained to obtain a bimodal matching model. This model can realize cross-modal contrast learning between text features and image features, improving the flexibility and accuracy of image retrieval.

Benefits of technology

Through the application of the dual-modal matching model, efficient matching between images and text is achieved, and the efficiency and accuracy of medical image analysis and financial image retrieval are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067355A_ABST
    Figure CN120067355A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image matching, can be applied to business system platforms such as financial science and technology and medical health, and discloses an image retrieval enhancement method, device and equipment based on bimodal matching and a medium. Performing fragment text coding on the matched text set according to the fragment image feature set to obtain a fragment text feature set; performing comparison training on the language image model by using the fragment image feature set and the fragment text feature set to obtain a bimodal matching model; performing preliminary image matching on the matching image set according to the user historical image and the user historical text by using a bimodal matching model to obtain a primary recall image sequence; and performing cross coding sorting on the primary recalled image sequence according to the user historical image and the user historical text to obtain a standard recalled image sequence. According to the invention, the image retrieval accuracy can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image matching, and in particular, to an enhanced method, device, equipment and medium for image retrieval based on bimodal matching. Background Art

[0002] Image retrieval refers to retrieving images similar to a query input from a large number of image databases, which can be matched based on text descriptions or the images themselves. Among them, text-description-based matching means matching the corresponding images according to the text describing the image content, and image-based matching means matching images with similar content according to the input image. With the development of artificial intelligence technology, image retrieval has been more and more widely applied.

[0003] For example, in the field of medical and health, doctors need to find similar case images in a vast medical image database to assist in diagnosis and treatment. For example, the corresponding pathological ultrasound images are matched through the text description of the case for analysis.

[0004] Again, for example, in the field of fintech business, in order to conduct risk analysis on users, it is necessary to find suspicious transaction pattern diagrams in the database, so as to detect abnormal transaction patterns and give transaction warnings to users.

[0005] Currently, the image retrieval technologies in the medical and health and financial fields mainly rely on deep learning feature extraction and vector retrieval technologies, but there are certain limitations in terms of retrieval accuracy and efficiency, etc., and it is necessary to further improve the accuracy of image retrieval. Summary of the Invention

[0006] Image retrieval refers to retrieving images similar to a query input from a large number of image databases, which can be matched based on text descriptions or the images themselves. Among them, text-description-based matching means matching the corresponding images according to the text describing the image content, and image-based matching means matching images with similar content according to the input image. With the development of artificial intelligence technology, image retrieval has been more and more widely applied.

[0007] For example, in the field of medical and health, doctors need to find similar case images in a vast medical image database to assist in diagnosis and treatment. For example, the corresponding pathological ultrasound images are matched through the text description of the case for analysis.

[0008] Again, for example, in the field of fintech business, in order to conduct risk analysis on users, it is necessary to find suspicious transaction pattern diagrams in the database, so as to detect abnormal transaction patterns and give transaction warnings to users.

[0009] At present, image retrieval technologies in the fields of medical health and finance mainly rely on deep learning feature extraction and vector retrieval technologies, but there are certain limitations in aspects such as retrieval accuracy and efficiency, and it is necessary to further improve the accuracy of image retrieval. Brief Description of the Drawings

[0010] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0011] Figure 1 is a schematic diagram of an application environment of an image retrieval enhancement method based on bimodal matching in an embodiment of the present invention; Figure 2 is a schematic flowchart of an image retrieval enhancement method based on bimodal matching in an embodiment of the present invention; Figure 3 is Figure 1 a schematic flowchart of a specific implementation manner of step S20 in Figure 4 is Figure 1 a schematic flowchart of a specific implementation manner of step S30 in Figure 5 is Figure 1 a schematic flowchart of a specific implementation manner of step S50 in Figure 6 is a schematic structural diagram of an image retrieval enhancement device based on bimodal matching in an embodiment of the present invention; Figure 7 is a schematic structural diagram of a computer device in an embodiment of the present invention; Figure 8 is another schematic structural diagram of a computer device in an embodiment of the present invention. Detailed Description of the Embodiments

[0012] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, rather than all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0013] The image retrieval enhancement method based on bimodal matching provided by the embodiments of the present invention can be applied in, for example, Figure 1In the application environment, the client communicates with the server through the network. The server can communicate with the server through the client's network. The server can obtain the matching image set and the matching text set corresponding to the matching image set through the client; use a preset language-image model to perform sliced image encoding on the matching image set to obtain a set of sliced image feature groups, and perform sliced text encoding on the matching text set according to the set of sliced image feature groups to obtain a set of sliced text feature groups; use the set of sliced image feature groups and the set of sliced text feature groups to perform contrastive training on the language-image model to obtain a bimodal matching model; obtain the user's historical images and historical texts of the user to be queried; use the bimodal matching model to perform preliminary image matching on the matching image set according to the user's historical images and historical texts to obtain a primary recall image sequence; perform cross-encoding sorting on the primary recall image sequence according to the user's historical images and historical texts to obtain a standard recall image sequence. In the present invention, for the contrastive training of the language-image model using the set of sliced image feature groups and the set of sliced text feature groups to obtain a bimodal matching model, it includes: selecting one by one the sliced image feature groups in the set of sliced image feature groups as the target image feature groups, and using the sliced text feature groups corresponding to the target image feature groups in the set of sliced text feature groups as the target text feature groups; generating a feature matching matrix according to the target image feature groups and the target text feature groups; respectively extracting a set of normal matching feature pairs and a set of abnormal matching feature pairs from the feature matching matrix; respectively determining the normal matching similarity of the set of normal matching feature pairs and the abnormal matching similarity of the set of abnormal matching feature pairs; calculating a matching loss value according to the normal matching similarity and the abnormal matching similarity; calculating a contrastive loss value according to the matching loss values of all the target image feature groups in the set of sliced image feature groups; performing iterative training on the language-image model according to the contrastive loss value to obtain a bimodal matching model, which can realize cross-modal contrastive learning between text features and image features, thereby improving the flexibility of image retrieval and the accuracy of image retrieval. Among them, the client can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, and portable wearable devices. The server can be implemented by an independent server or a server cluster composed of multiple servers. The present invention will be described in detail through specific embodiments below.

[0014] Please refer to Figure 2 as shown in Figure 2 which is a schematic flowchart of an image retrieval enhancement method based on bimodal matching provided by an embodiment of the present invention, including the following steps: S10: Obtain a matching image set and a matching text set corresponding to the matching image set.

[0015] Specifically, the matching image set is a set composed of a large number of customer image data with privacy data removed, and the image types of each matching image in the matching image set are the same as the image type of the image to be retrieved and enhanced. The matching text set is a set composed of a large number of customer text data with privacy data removed. Each matching text in the matching text set is descriptive text for describing each matching image in the matching image set, and the matching text corresponds to the matching image one by one.

[0016] Example: In the field of medical and health, doctors often need to improve the efficiency of medical diagnosis and treatment through patient medical record text and patients' medical imaging images to help doctors more accurately identify diseases and provide treatment plans. At this time, each matching image in the matching image set refers to the patient's x-ray image, microscopic image, ultrasound image, magnetic resonance image, and radionuclide image. Each matching text in the matching text set is text such as the diagnostic report, patient medical record, progress record, and test report corresponding to each matching image in the matching image set.

[0017] In the financial field, financial companies or financial organizations need to analyze the transaction behavior images of users in order to accurately evaluate the credit status of customers, monitor suspicious transaction patterns, and prevent fraud. At this time, each matching image in the matching image set refers to the time series curve, K-line chart, capital flow path chart, etc. of the user's transaction behavior. Each matching text in the matching text set is text such as financial news, market analysis reports, company financial reports, and social media financial sentiment corresponding to each matching image in the matching image set.

[0018] In the embodiment of the present invention, by obtaining the matching image set and the matching text set corresponding to the matching image set, mutually matching text data and image data can be obtained, providing a basis for subsequent contrast learning between modal features.

[0019] S20: Use a preset language-image model to perform sliced image encoding on the matching image set to obtain a set of sliced image feature groups, and perform sliced text encoding on the matching text set according to the set of sliced image feature groups to obtain a set of sliced text feature groups.

[0020] Specifically, each sliced image feature in the set of sliced image feature groups refers to the feature data after each matching image in the matching image set is sliced and extracted. Each sliced text feature in the set of sliced text feature groups refers to the feature after each matching text in the matching text set is sliced and encoded.

[0021] In the embodiment of the present invention, with reference to Figure 3As shown, using a preset language-image model to perform piecewise image encoding on the set of matching images to obtain a set of piecewise image feature groups, including: S21. Select each matching image in the set of matching images as a target matching image, perform image segmentation on the target matching image to obtain a set of segmented matching tiles; S22. Perform linear projection on the set of segmented matching tiles to obtain a set of segmented image features; S23. Perform local attention calculation on the set of segmented image features to obtain a set of image local features; S24. Perform global attention calculation on the set of image local features to obtain a set of image global features; S25. Use a preset language-image model to perform multi-scale feature extraction and fully connected operation on the set of image global features to obtain a set of standard image features; S26. Aggregate the sets of standard image features of all target matching images in the set of matching images into a set of standard image features; S27. Perform clustering storage on each of the standard image feature groups in the set of standard image features to obtain a set of piecewise image feature groups.

[0022] Specifically, the image segmentation refers to dividing the target matching image into multiple tiles of uniform size according to a preset window size, that is, each segmented matching tile in the set of segmented matching tiles. The linear projection refers to mapping a vector in a vector space to a subspace of the vector space, that is, mapping two-dimensional image features into one-dimensional linear features.

[0023] Specifically, the local attention calculation refers to using a multi-head self-attention mechanism to calculate the attention feature representation in each of the segmented image features in the set of segmented image features. The global attention calculation refers to using a preset sliding window to calculate the attention features of the interaction between each adjacent segmented image feature, thereby enhancing the global feature representation of the set of segmented image features.

[0024] Specifically, the multi-scale feature extraction refers to using a method of gradually downsampling to perform multi-size feature extraction on each of the image global features in the set of image global features step by step, obtaining multiple global feature layers in a pyramid structure, and performing a feature fully connected operation on each global feature layer to obtain piecewise image features.

[0025] Specifically, the performing clustering storage on each of the standard image feature groups in the set of standard image features to obtain a set of piecewise image feature groups includes: Performing a feature splicing operation on each of the standard image feature groups in the set of standard image features to obtain a set of spliced image features; Randomly split the spliced image feature set into multiple spliced image feature groups; Randomly select spliced image features as the spliced image center features from each spliced image feature group to obtain a spliced image center feature group; Respectively determine the feature distances between each spliced image feature in the spliced image feature set and each spliced image center feature in the spliced image center feature group; According to the principle of proximity, allocate each spliced image feature in the spliced image feature set to the spliced image feature group where the corresponding spliced image center feature is located based on the feature distance to obtain an updated image feature group set; Perform feature center calculation on each updated image feature group in the updated image feature group set to obtain an updated center feature group; Obtain the center feature distances between each updated center feature in the updated center feature group and the corresponding spliced image center feature in the spliced image center feature group, and take the mean of all the center feature distances as the standard center feature distance; Update the updated image feature group set according to the standard center feature distance to obtain a clustered image feature group set; Perform index storage on each clustered image feature group in the clustered image feature group set to obtain an indexed image feature set; Perform feature splitting on each indexed image feature in the indexed image feature set to obtain a sliced image feature group set.

[0026] Specifically, the operation of performing feature splicing on each standard image feature group in the standard image feature group set to obtain a spliced image feature set means splicing the standard image features in each standard image feature group into spliced image features, and aggregating all the spliced image features into a spliced image feature set.

[0027] In detail, the following mathematical formula can be used to respectively determine the feature distances between each spliced image feature in the spliced image feature set and each spliced image center feature in the spliced image center feature group:

[0028] Among them, refers to the spliced image feature and the spliced image center feature the feature distance between them, is the exponential function symbol, refers to the spliced image feature, refers to the spliced image center feature, is the transpose symbol, refers to the spliced image feature and the covariance matrix of the central features of the stitched image , is a preset regularization coefficient, and is the identity matrix.

[0029] Specifically, the cosine distance algorithm or the Euclidean distance algorithm can be used to calculate the central feature distance. Updating the updated image feature set according to the standard central feature distance to obtain the clustered image feature set means determining whether the standard central feature distance is greater than a preset central distance threshold; if so, using the updated central feature set to update the central feature set of the stitched image, and returning to the step of respectively determining the feature distances between each stitched image feature in the stitched image feature set and each stitched image central feature in the central feature set of the stitched image; if not, using the updated image feature set as the clustered image feature set.

[0030] Specifically, the index storage means splitting and storing each clustered image feature set in the clustered image feature set, and storing the central feature of each clustered image feature set as the feature index of the corresponding clustered image feature set, and storing all the clustered image features in a queue to obtain the indexed image feature set. The method of feature splitting is the inverse step of the method of feature stitching, which will not be elaborated here.

[0031] Example: In the field of medical and health, when the matching images in the matching image set are nuclear magnetic resonance images, the image segmentation means slicing the matching image set and dividing it into small tiles of 32×32 size. Each tile is a segmented matching tile in the segmented matching tile group. Each segmented matching tile group may correspond to different brain region structures in the nuclear magnetic resonance image, such as the cortex, white matter, and ventricles, etc. The linear projection means projecting each segmented matching tile into a fixed feature dimension using a convolutional network or a multi-layer perceptron model. The local attention calculation means calculating the attention features corresponding to each segmented image feature one by one using a window of the same size as the segmented image feature, and aggregating the attention features into an image local feature group as the image local features, so that the model can learn local structural information, such as the details of a certain lesion area. The global attention calculation means sliding a sliding window on the feature sequence composed of the image local feature group and calculating the attention features in each corresponding window, so that the local features can be propagated globally to construct a complete nuclear magnetic resonance image feature.

[0032] In the financial field, when the matching image in the set of matching images is a fund flow network image, the image segmentation refers to slicing the set of matching images according to the transaction area or time window of a fixed window, and dividing it into multiple small tiles. Each tile is a segmented matching tile in the segmented matching tile group. Each segmented matching tile group may correspond to transaction information in different transaction time windows in the fund flow network image, etc. The linear projection refers to using a convolutional network or a multi-layer perception model to project each segmented matching tile into a fixed feature dimension. For example, each transaction node can be represented by features such as transaction amount, time interval, transaction frequency, etc., and then mapped to a 768-dimensional feature vector using a linear layer. The local attention calculation refers to calculating the attention features corresponding to each segmented image feature one by one using a window with the same size as the segmented image feature, and aggregating the attention features into an image local feature group as the local feature of the image, so that the model can learn local structural information, such as the details of multiple suspicious transactions of a certain enterprise. The global attention calculation refers to using a sliding window to slide on the feature sequence composed of the image local feature group, and calculating the attention features within each corresponding window, so that the local features can spread globally and larger fund flow patterns can be discovered, such as whether there are suspicious behaviors of layer-by-layer fund transfer.

[0033] Specifically, the process of performing sharded text encoding on the set of matching texts according to the set of sharded image feature groups to obtain a set of sharded text feature groups includes: Select each matching text in the set of matching texts as the target matching text one by one, and use the sharded image feature group corresponding to the target matching text in the set of sharded image feature groups as the target sharded image feature group; Extract the number of sharded features from the target sharded image feature group; Perform text segmentation on the target matching text according to the number of sharded features to obtain a set of segmented matching texts; Perform text tokenization and feature embedding on the set of segmented matching texts to obtain a set of segmented text features; Perform local attention calculation on the set of segmented text features to obtain a set of text local features; Perform global attention calculation on the set of text local features to obtain a set of text global features; Use the language image model to perform multi-scale feature extraction and fully connected operations on the set of text global features to obtain a set of standard text features; Aggregate the sets of standard text features of all target matching texts in the set of matching texts into a set of standard text features; Perform clustering storage on each standard text feature group in the set of standard text features to obtain a set of sharded text feature groups.

[0034] Specifically, the number of piecewise features refers to the total number of image features in the target piecewise image feature group. Segmenting the target matching text according to the number of piecewise features to obtain a segmented matching text group means segmenting the target matching text into the number of text segments equal to the number of piecewise features, and the text can be segmented according to symbols or paragraphs of the text.

[0035] Specifically, text tokenization refers to dividing a continuous text into individual words or sub-word units. Tokenization methods based on word segmentation such as bidirectional maximum matching, regular expression tokenization, or Byte Pair Encoding, WordPiece, and SentencePiece can be used for text tokenization. Feature embedding refers to the process of converting the tokenized words into word features, that is, representing discrete text data as continuous numerical vectors. Feature embedding methods based on statistics such as Bag of Words (BoW) and TF IDF, feature embedding methods based on prediction such as Word2Vec and GloVe, and feature embedding methods based on context-dependent dynamic word features such as ELMo and Transformer based models can be used for feature embedding.

[0036] In the embodiment of the present invention, by using a preset language-image model to perform piecewise image encoding on the matching image set to obtain a set of piecewise image feature groups, and performing piecewise text encoding on the matching text set according to the set of piecewise image feature groups to obtain a set of piecewise text feature groups, local features and global features of the matching image set and the matching text set can be effectively extracted, high-quality feature representations can be obtained, and subsequent contrast training between text features and image features can be facilitated.

[0037] S30: Use the set of piecewise image feature groups and the set of piecewise text feature groups to perform contrast training on the language-image model to obtain a bimodal matching model.

[0038] Specifically, the language-image model refers to the Contrastive Language Image Pretraining (CLIP) model. The Contrastive Language Image Pretraining model is a multimodal model proposed by OpenAI, which can understand the relationship between images and texts. The Contrastive Language Image Pretraining model uses contrastive learning methods to train on large-scale image and text pairing data, enabling the model to directly match, retrieve, and classify images and texts.

[0039] In the embodiment of the present invention, with reference to Figure 4As shown, the contrast training of the language image model using the set of segmented image feature groups and the set of segmented text feature groups to obtain a bimodal matching model includes: S31. Select each segmented image feature group in the set of segmented image feature groups as the target image feature group one by one, and use the corresponding segmented text feature group of the target image feature group in the set of segmented text feature groups as the target text feature group; S32. Generate a feature matching matrix according to the target image feature group and the target text feature group; S33. Extract the normal matching feature pair set and the abnormal matching feature pair set from the feature matching matrix respectively; S34. Determine the normal matching similarity set of the normal matching feature pair set and the abnormal matching similarity set of the abnormal matching feature pair set respectively; S35. Calculate the matching loss value according to the normal matching similarity set and the abnormal matching similarity set; S36. Calculate the contrast loss value according to the matching loss values of all target image feature groups in the set of segmented image feature groups; S37. Iteratively train the language image model according to the contrast loss value to obtain a bimodal matching model.

[0040] Specifically, the feature matching matrix refers to a matrix established by enumerating and pairwise matching each target image feature in the target image feature group as a row vector and the target text feature group as a column vector.

[0041] Specifically, the step of respectively extracting the normal matching feature pair set and the abnormal matching feature pair set from the feature matching matrix means taking the feature pairs in the main diagonal position of the feature matching matrix as the normal matching feature pairs to form a normal matching feature pair set, and taking the feature pairs not in the main diagonal position of the feature matching matrix as the abnormal matching feature pairs to form an abnormal matching feature pair set.

[0042] In detail, the cosine similarity algorithm can be used to determine the normal matching similarity set of the normal matching feature pair set and the abnormal matching similarity set of the abnormal matching feature pair set respectively.

[0043] Specifically, the following matching algorithm is used to calculate the matching loss value according to the normal matching similarity set and the abnormal matching similarity set:

[0044] where, refers to the matching loss value, refers to the total number of features in the target image feature group, and the total number of features in the target image feature group is equal to the total number of features in the target text feature group. , is the feature index. is the logarithmic function symbol. is the exponential function symbol. is the cosine similarity calculation symbol. is the th image feature in the target image feature group. is the th text feature in the target text feature group. is the th image feature in the target image feature group. is the th text feature in the target text feature group. is the preset matching weight. is the dot product symbol. is the modulo symbol.

[0045] Specifically, the normal matching feature pair set refers to the set of feature pairs located on the main diagonal, that is, in the formula, and the abnormal matching feature pair set refers to the set of feature pairs in the remaining positions, that is, and in the formula. , and the contrast loss value is the average of the matching loss values of all target image feature groups in the sliced image feature group set.

[0046] Specifically, the iterative training of the language-image model according to the contrast loss value to obtain the bimodal matching model means determining whether the contrast loss value is greater than a preset loss value threshold. If so, the model parameters of the language-image model are updated according to the contrast loss value using the backpropagation algorithm, and the step of obtaining the sliced image feature group set by performing sliced image encoding on the matching image set using the preset language-image model is returned. If not, the language-image model is used as the bimodal matching model.

[0047] In the embodiments of the present invention, by performing contrast training on the language-image model using the sliced image feature group set and the sliced text feature group set to obtain the bimodal matching model, a retrieval model capable of realizing the matching between images and texts can be obtained, achieving efficient cross-modal retrieval, thereby improving the efficiency of medical image analysis.

[0048] S40: Obtain the user's historical images and historical texts to be queried.

[0049] Specifically, the user to be queried refers to the user who needs to perform corresponding image retrieval operations. The user's historical images refer to the images stored by the user who needs to perform image retrieval in the database in the past. The user's historical texts refer to the texts stored by the user who needs to perform image retrieval in the database in the past.

[0050] Example: In the field of medical and health, the user to be queried refers to the user who needs to perform a health condition diagnosis. The user's historical images refer to the X-ray images, microscopic images, ultrasound images, nuclear magnetic resonance images, and radionuclide images retained during the past diagnoses of the user to be queried. The user's historical texts refer to the texts such as diagnosis reports, patient medical records, course records, and test reports retained during the past diagnoses of the user to be queried.

[0051] In the financial field, the user to be queried refers to the user who needs to perform financial risk analysis. The user's historical images refer to the time series curves, K-line charts, fund flow path charts, etc. that record the trading behaviors uploaded by the user to be queried in the past. The user's historical texts refer to the texts such as financial news, market analysis reports, company financial reports, and financial sentiment on social media of the user to be queried in the past.

[0052] In the embodiments of the present invention, obtaining the user's historical images and user's historical texts of the user to be queried includes: obtaining the user identifier of the user to be queried; querying the historical image data of the user to be queried according to the user identifier to obtain the user's historical images; querying the historical text data of the user to be queried according to the user identifier to obtain the user's historical texts.

[0053] Specifically, the user identifier refers to the unique identifier used to confirm the user identity of the user to be queried. For example, the ID card number of the user to be queried. Querying the historical image data of the user to be queried according to the user identifier to obtain the user's historical images means retrieving the associated image data according to the user identifier in the preset database to obtain the user's historical images. Querying the historical text data of the user to be queried according to the user identifier to obtain the user's historical texts means retrieving the associated text data according to the user identifier in the preset database to obtain the user's historical texts.

[0054] In the embodiments of the present invention, by obtaining the user's historical images and user's historical texts of the user to be queried, the user to be queried can be obtained, and the past information of the user to be queried can be quickly extracted, thereby improving the efficiency of medical image analysis.

[0055] S50: Using the dual-modal matching model, perform preliminary image matching on the matching image set according to the user's historical images and the user's historical texts to obtain a primary recall image sequence.

[0056] Specifically, each primary recall image included in the primary recall image sequence is a matching image in the matching image set related to the user historical image and the user historical text, arranged according to the recall rate.

[0057] In the embodiment of the present invention, with reference to Figure 5 as shown, the method of using the bimodal matching model to perform preliminary image matching on the matching image set according to the user historical image and the user historical text to obtain a primary recall image sequence includes: S51. Use the bimodal matching model to perform sliced image encoding on the user historical image to obtain a historical image feature group; S52. Use the bimodal matching model to perform sliced text encoding on the user historical text to obtain a historical text feature group; S53. Obtain the set of sliced image feature groups corresponding to the bimodal matching model; S54. Use the historical text feature group to perform cross-modal image matching on the set of sliced image feature groups to obtain a cross-modal recall image sequence; S55. Use the historical image feature group to perform similarity image matching on the set of sliced image feature groups to obtain a similarity recall image sequence; S56. Generate a primary recall image sequence according to the cross-modal recall image sequence and the similarity recall image sequence.

[0058] Specifically, the method of sliced image encoding is the same as the method of sliced image encoding in step S20 above, which will not be elaborated here. The method of sliced text encoding is the same as the method of sliced image encoding in step S20 above, which will not be elaborated here. The set of sliced image feature groups corresponding to the bimodal matching model refers to the set of sliced image feature groups corresponding when the language image model is updated to the bimodal matching model.

[0059] Specifically, the method of using the historical text feature group to perform cross-modal image matching on the set of sliced image feature groups to obtain a cross-modal recall image sequence means using the historical text feature group to perform similarity matching with each sliced image feature group in the set of sliced image feature groups respectively, screening out the sliced image feature groups with a similarity greater than a preset similarity threshold as the recall image feature groups, and aggregating all the recall image feature groups into a recall image feature group sequence in the order of similarity from large to small, and taking the respective matching images corresponding to the recall image feature group sequence as cross-modal recall images to aggregate into a cross-modal recall image sequence.

[0060] Specifically, the step of performing similarity image matching on the set of sliced image feature groups using the historical image feature group to obtain a similarity recall image sequence means performing similarity matching on the historical image feature group with each sliced image feature group in the set of sliced image feature groups, screening out the sliced image feature groups with a similarity greater than a preset similarity threshold as similarity recall image feature groups, and aggregating all the similarity recall image feature groups into a similarity recall image feature group sequence in descending order of similarity. Then, using the matching images corresponding to the similarity recall image feature group sequence as similarity recall images to aggregate into a similarity recall image sequence.

[0061] Specifically, the step of generating a primary recall image sequence based on the cross-modal recall image sequence and the similarity recall image sequence means performing a global sorting according to the similarity of each image in the cross-modal recall image sequence and the similarity recall image sequence to obtain a primary recall image sequence.

[0062] In the embodiment of the present invention, by using the dual-modal matching model to perform preliminary image matching on the set of matching images according to the user's historical image and the user's historical text to obtain a primary recall image sequence, cross-modal retrieval of text and image and image retrieval between the same modalities can be realized, thereby improving the flexibility of image retrieval.

[0063] S60: Perform cross-coding sorting on the primary recall image sequence according to the user's historical image and the user's historical text to obtain a standard recall image sequence.

[0064] Specifically, each standard recall image in the standard recall image sequence refers to the final relevant image result obtained after performing image retrieval for the user's historical image and the user's historical text.

[0065] In the embodiment of the present invention, the step of performing cross-coding sorting on the primary recall image sequence according to the user's historical image and the user's historical text to obtain a standard recall image sequence includes: Obtain the historical image feature group corresponding to the user's historical image and the historical text feature group corresponding to the user's historical text; Perform cross-coding on the historical image feature group and the historical text feature group to obtain an encoded user feature; Extract the primary recall text sequence corresponding to the primary recall image sequence from the set of matching texts; Extract a recall image feature group sequence from the primary recall image sequence and a recall text feature group sequence from the primary recall text sequence respectively; Generate a sequence of recall feature group pairs based on the sequence of recalled image feature groups and the sequence of recalled text feature groups; Perform cross-encoding on each recall feature group pair in the sequence of recall feature group pairs to obtain an encoded recall feature sequence; Use the encoded user features to perform feature matching on each encoded recall feature in the encoded recall feature sequence to obtain a sequence of matching degrees; Re-rank the primary recall image sequence according to the sequence of matching degrees to obtain a standard recall image sequence.

[0066] Specifically, the cross-encoding refers to calculating the attention feature encoding between an image feature group and the corresponding text feature group using a cross-attention mechanism. Generating a sequence of recall feature group pairs based on the sequence of recalled image feature groups and the sequence of recalled text feature groups means combining each recalled image feature group in the sequence of recalled image feature groups with the corresponding recalled text feature group in the sequence of recalled text feature groups to obtain recall feature group pairs, and aggregating all the recall feature group pairs into a sequence of recall feature group pairs.

[0067] Specifically, the feature matching refers to calculating the matching degree between features using a similarity calculation formula, and aggregating all the matching degrees into a sequence of matching degrees. Re-ranking the primary recall image sequence according to the sequence of matching degrees to obtain a standard recall image sequence means performing a weighted summation operation based on the sequence of matching degrees and the similarity of each primary recall image in the primary recall image sequence to obtain the final recall degree, and re-ranking the primary recall image sequence according to the recall degree to obtain a standard recall image sequence.

[0068] In the embodiments of the present invention, by performing cross-encoding and sorting on the primary recall image sequence according to the user's historical images and the user's historical texts to obtain a standard recall image sequence, it is possible to further accurately sort the indexing results according to the context relationship between the text and the image, improving the accuracy of image retrieval.

[0069] It can be seen that in the above solution, by obtaining the matching image set and the corresponding matching text set of the matching image set, it is possible to obtain the mutually matching text data and image data, providing a basis for subsequent contrastive learning between modal features. By using a preset language-image model to perform sliced image encoding on the matching image set to obtain a set of sliced image feature groups, and performing sliced text encoding on the matching text set according to the set of sliced image feature groups to obtain a set of sliced text feature groups, it is possible to effectively extract the local and global features of the matching image set and the matching text set, obtain an efficient feature representation, and facilitate subsequent contrastive training between text features and image features. By using the set of sliced image feature groups and the set of sliced text feature groups to perform contrastive training on the language-image model to obtain a bimodal matching model, a retrieval model that can achieve the matching between images and texts can be obtained, realizing efficient cross-modal retrieval, thereby improving the efficiency of medical image analysis.

[0070] By obtaining the user's historical images and historical texts of the user to be queried, the user to be queried can be obtained, and the past information of the user to be queried can be quickly extracted, thereby improving the efficiency of medical image analysis. By using the bimodal matching model to perform preliminary image matching on the matching image set according to the user's historical images and historical texts of the user to be queried to obtain a primary recall image sequence, cross-modal retrieval between text and image and image retrieval between the same modalities can be realized, thereby improving the flexibility of image retrieval. By performing cross-coding sorting on the primary recall image sequence according to the user's historical images and historical texts of the user to be queried to obtain a standard recall image sequence, the index results can be further accurately sorted according to the context relationship between text and image, improving the accuracy of image retrieval.

[0071] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The order of execution of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.

[0072] In one embodiment, an image retrieval enhancement device based on bimodal matching is provided. The image retrieval enhancement device based on bimodal matching corresponds one-to-one to the image retrieval enhancement method based on bimodal matching in the above embodiment. As Figure 6 shown, the image retrieval enhancement device based on bimodal matching includes a data acquisition module 101, a sliced encoding module 102, a contrastive training module 103, a retrieval input module 104, a preliminary matching module 105, and an accurate sorting module 106. The detailed description of each functional module is as follows: The data acquisition module 101 is used to obtain a matching image set and the corresponding matching text set of the matching image set; The sharding encoding module 102 is used to perform sharding image encoding on the matching image set by using a preset language-image model to obtain a set of sharding image feature groups, and perform sharding text encoding on the matching text set according to the set of sharding image feature groups to obtain a set of sharding text feature groups; The contrast training module 103 is used to perform contrast training on the language-image model by using the set of sharding image feature groups and the set of sharding text feature groups to obtain a bimodal matching model; The retrieval input module 104 is used to obtain the user's historical image and historical text of the user to be queried; The preliminary matching module 105 is used to perform preliminary image matching on the matching image set according to the user's historical image and historical text by using the bimodal matching model to obtain a primary recall image sequence; The precise sorting module 106 is used to perform cross-encoding sorting on the primary recall image sequence according to the user's historical image and historical text to obtain a standard recall image sequence.

[0073] In one embodiment, when the sharding encoding module 102 performs sharding image encoding on the matching image set by using a preset language-image model to obtain a set of sharding image feature groups, it is used for: Select each matching image in the matching image set as a target matching image, perform image segmentation on the target matching image to obtain a set of segmented matching tiles; Perform linear projection on the set of segmented matching tiles to obtain a set of segmented image features; Perform local attention calculation on the set of segmented image features to obtain a set of image local features; Perform global attention calculation on the set of image local features to obtain a set of image global features; Use a preset language-image model to perform multi-scale feature extraction and fully connected operation on the set of image global features to obtain a set of standard image features; Aggregate the sets of standard image features of all target matching images in the matching image set into a set of standard image features; Perform clustering storage on each of the standard image feature groups in the set of standard image features to obtain a set of sharding image feature groups.

[0074] In one embodiment, when the sharding encoding module 102 performs clustering storage on each of the standard image feature groups in the set of standard image features to obtain a set of sharding image feature groups, it is used for: Perform feature splicing operation on each of the standard image feature groups in the set of standard image features to obtain a set of spliced image features; Randomly split the spliced image feature set into multiple spliced image feature groups; Randomly select spliced image features as the center features of the spliced images from each spliced image feature group to obtain a spliced image center feature group; Respectively determine the feature distances between each spliced image feature in the spliced image feature set and each spliced image center feature in the spliced image center feature group; According to the principle of proximity, allocate each spliced image feature in the spliced image feature set to the spliced image feature group where the corresponding spliced image center feature is located to obtain an updated image feature group set; Perform feature center calculation on each updated image feature group in the updated image feature group set to obtain an updated center feature group; Obtain the center feature distances between each updated center feature in the updated center feature group and the corresponding spliced image center feature in the spliced image center feature group, and take the mean of all the center feature distances as the standard center feature distance; Update the updated image feature group set according to the standard center feature distance to obtain a clustered image feature group set; Perform index storage on each clustered image feature group in the clustered image feature group set to obtain an indexed image feature set; Perform feature splitting on each indexed image feature in the indexed image feature set to obtain a sharded image feature group set.

[0075] In one embodiment, when the sharded encoding module 102 performs sharded text encoding on the matching text set according to the sharded image feature group set to obtain a sharded text feature group set, it is used for: Select each matching text in the matching text set one by one as the target matching text, and use the sharded image feature group corresponding to the target matching text in the sharded image feature group set as the target sharded image feature group; Extract the number of sharded features from the target sharded image feature group; Perform text segmentation on the target matching text according to the number of sharded features to obtain a segmented matching text group; Perform text tokenization and feature embedding on the segmented matching text group to obtain a segmented text feature group; Perform local attention calculation on the segmented text feature group to obtain a text local feature group; Perform global attention calculation on the text local feature group to obtain a text global feature group; Use the language image model to perform multi-scale feature extraction and fully connected operation on the text global feature group to obtain a standard text feature group; Assemble the standard text feature groups of all target matching texts in the matching text set into a standard text feature group set; Cluster and store each of the standard text feature groups in the standard text feature group set to obtain a sharded text feature group set.

[0076] In one embodiment, when the contrast training module 103 performs contrast training on the language image model using the sharded image feature group set and the sharded text feature group set to obtain a bimodal matching model, it is used for: Select each sharded image feature group in the sharded image feature group set as a target image feature group one by one, and use the sharded text feature group corresponding to the target image feature group in the sharded text feature group set as a target text feature group; Generate a feature matching matrix according to the target image feature group and the target text feature group; Extract a normal matching feature pair set and an abnormal matching feature pair set from the feature matching matrix respectively; Determine the normal matching similarity of the normal matching feature pair set and the abnormal matching similarity of the abnormal matching feature pair set respectively; Calculate a matching loss value according to the normal matching similarity and the abnormal matching similarity; Calculate a contrast loss value according to the matching loss values of all target image feature groups in the sharded image feature group set; Iteratively train the language image model according to the contrast loss value to obtain a bimodal matching model.

[0077] In one embodiment, when the preliminary matching module 105 performs preliminary image matching on the matching image set using the bimodal matching model according to the user's historical image and the user's historical text to obtain a primary recall image sequence, it is used for: Perform sharded image encoding on the user's historical image using the bimodal matching model to obtain a historical image feature group; Perform sharded text encoding on the user's historical text using the bimodal matching model to obtain a historical text feature group; Obtain the sharded image feature group set corresponding to the bimodal matching model; Perform cross-modal image matching on the sharded image feature group set using the historical text feature group to obtain a cross-modal recall image sequence; Perform similarity image matching on the sharded image feature group set using the historical image feature group to obtain a similarity recall image sequence; Generate a primary recall image sequence according to the cross-modal recall image sequence and the similarity recall image sequence.

[0078] In one embodiment, when the precise sorting module 106 performs cross-coding and sorting on the primary recall image sequence according to the user's historical images and the user's historical texts to obtain a standard recall image sequence, it is used for: Obtain the historical image feature group corresponding to the user's historical images and the historical text feature group corresponding to the user's historical texts; Perform cross-coding on the historical image feature group and the historical text feature group to obtain encoded user features; Extract the primary recall text sequence corresponding to the primary recall image sequence from the matching text set; Extract a recall image feature group sequence from the primary recall image sequence respectively and extract a recall text feature group sequence from the primary recall text sequence; Generate a recall feature group pair sequence according to the recall image feature group sequence and the recall text feature group sequence; Perform cross-coding on each recall feature group pair in the recall feature group pair sequence to obtain an encoded recall feature sequence; Use the encoded user features to perform feature matching on each encoded recall feature in the encoded recall feature sequence to obtain a matching degree sequence; Re-sort the primary recall image sequence according to the matching degree sequence to obtain a standard recall image sequence.

[0079] The present invention provides an image retrieval enhancement device based on bimodal matching. By first obtaining a matching image set and the matching text set corresponding to the matching image set, it is possible to obtain mutually matching text data and image data, providing a basis for subsequent contrast learning between modal features. By using a preset language-image model to perform sliced image coding on the matching image set to obtain a sliced image feature group set, and performing sliced text coding on the matching text set according to the sliced image feature group set to obtain a sliced text feature group set, it is possible to effectively extract the local features and global features of the matching image set and the matching text set, obtain an efficient feature representation, and facilitate subsequent contrast training between text features and image features. By using the sliced image feature group set and the sliced text feature group set to perform contrast training on the language-image model to obtain a bimodal matching model, it is possible to obtain a retrieval model that realizes the matching between images and texts, realizing efficient cross-modal retrieval, thereby improving the efficiency of medical image analysis.

[0080] By obtaining the user's historical images and historical texts of the user to be queried, the user to be queried can be obtained, and the past information of the user to be queried can be quickly extracted, thereby improving the efficiency of medical image analysis. By using the bimodal matching model, based on the user's historical images and the user's historical texts, a preliminary image matching is performed on the matching image set to obtain a primary recall image sequence, which can achieve cross-modal retrieval of text and images and image retrieval between the same modalities, thereby improving the flexibility of image retrieval. By performing cross-coding sorting on the primary recall image sequence according to the user's historical images and the user's historical texts to obtain a standard recall image sequence, the index results can be further accurately sorted according to the context relationship between the text and the image, improving the accuracy of image retrieval.

[0081] For the specific limitations of the image retrieval enhancement device based on bimodal matching, reference can be made to the limitations of the method for intelligent question answering in the above text, which will not be elaborated here. Each module in the above image retrieval enhancement device based on bimodal matching can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form, so as to facilitate the processor to call and execute the operations corresponding to the above respective modules.

[0082] In one embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 7 shown. The computer device includes a processor, a memory, a network interface, and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage media. The network interface of the computer device is used to communicate with an external client through a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the server side of an image retrieval enhancement method based on bimodal matching.

[0083] In one embodiment, a computer device is provided. The computer device can be a client, and its internal structure diagram can be as Figure 8As shown. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected by a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server through a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the client side of an image retrieval enhancement method based on dual-modal matching.

[0084] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the following steps are implemented: Obtain a set of matching images and a set of matching texts corresponding to the set of matching images; Use a preset language-image model to perform slice image encoding on the set of matching images to obtain a set of slice image feature groups, and perform slice text encoding on the set of matching texts according to the set of slice image feature groups to obtain a set of slice text feature groups; Use the set of slice image feature groups and the set of slice text feature groups to perform contrast training on the language-image model to obtain a dual-modal matching model; Obtain the user's historical images and historical texts of the user to be queried; Use the dual-modal matching model to perform preliminary image matching on the set of matching images according to the user's historical images and historical texts to obtain a primary recall image sequence; Perform cross-encoding sorting on the primary recall image sequence according to the user's historical images and historical texts to obtain a standard recall image sequence.

[0085] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by the processor, the following steps are implemented: Obtain a set of matching images and a set of matching texts corresponding to the set of matching images; Use a preset language-image model to perform slice image encoding on the set of matching images to obtain a set of slice image feature groups, and perform slice text encoding on the set of matching texts according to the set of slice image feature groups to obtain a set of slice text feature groups; Use the set of slice image feature groups and the set of slice text feature groups to perform contrast training on the language-image model to obtain a dual-modal matching model; Obtain the user's historical images and historical texts of the user to be queried; Using the dual-modal matching model, perform preliminary image matching on the matching image set according to the user's historical images and the user's historical texts, to obtain a primary recall image sequence; Perform cross-coding sorting on the primary recall image sequence according to the user's historical images and the user's historical texts, to obtain a standard recall image sequence.

[0086] It should be noted that for the functions or steps that the above computer-readable storage medium or computer device can achieve, reference can be made to the relevant descriptions on the server side and the client side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.

[0087] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above method embodiments. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0088] Those skilled in the art can clearly understand that for the convenience and brevity of description, only the above division of each functional unit and module is used for illustration. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above.

[0089] The above-described embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention. It should be noted that if non-company software tools or components appear in the embodiments of the present application, they are only used for example introduction and do not represent actual use.

Claims

1. An image retrieval enhancement method based on bimodal matching, characterized in that: include: Acquire a matching image set and a matching text set corresponding to the matching image set; Performing slice image encoding on the matching image set using a preset language image model to obtain a slice image feature set, and performing slice text encoding on the matching text set according to the slice image feature set to obtain a slice text feature set; Using the segmented image feature set and the segmented text feature set to perform comparative training on the language image model to obtain a bimodal matching model; Obtain user history images and user history texts of the user to be queried; Using the bimodal matching model, performing preliminary image matching on the matching image set according to the user historical images and the user historical texts to obtain a primary recall image sequence; The primary recall image sequence is cross-coded and sorted according to the user history images and the user history texts to obtain a standard recall image sequence.

2. The image retrieval enhancement method based on bimodal matching according to claim 1, characterized in that: The step of performing segmented image encoding on the matching image set by using a preset language image model to obtain a segmented image feature set includes: Selecting matching images in the matching image set one by one as target matching images, performing image segmentation on the target matching images, and obtaining a segmented matching image block group; Performing linear projection on the segmentation matching block group to obtain a segmentation image feature group; Performing local attention calculation on the segmented image feature group to obtain an image local feature group; Performing global attention calculation on the local feature group of the image to obtain a global feature group of the image; Using a preset language image model, multi-scale feature extraction and full connection operation are performed on the global feature group of the image to obtain a standard image feature group; Aggregating the standard image feature groups of all target matching images in the matching image set into a standard image feature group set; Each of the standard image feature groups in the standard image feature group set is clustered and stored to obtain a fragmented image feature group set.

3. The image retrieval enhancement method based on bimodal matching according to claim 2, characterized in that: The step of clustering and storing each of the standard image feature groups in the standard image feature group set to obtain a fragmented image feature group set includes: Performing a feature splicing operation on each standard image feature group in the standard image feature group set to obtain a spliced ​​image feature set; Randomly splitting the stitched image feature set into a plurality of stitched image feature groups; Randomly select stitched image features from each stitched image feature group as stitched image center features to obtain a stitched image center feature group; respectively determining feature distances between each stitched image feature in the stitched image feature set and each stitched image center feature in the stitched image center feature group; Allocating each stitched image feature in the stitched image feature set to a stitched image feature group where a corresponding stitched image center feature is located according to the feature distance according to the proximity principle, to obtain an updated image feature group set; Performing feature center calculation on each updated image feature group in the updated image feature group set to obtain an updated center feature group; Obtaining the center feature distance between each update center feature in the update center feature group and the corresponding stitched image center feature in the stitched image center feature group, and taking the average of all center feature distances as the standard center feature distance; updating the updated image feature set according to the standard center feature distance to obtain a clustered image feature set; Indexing and storing each clustering image feature group in the clustering image feature group set to obtain an indexed image feature set; Feature splitting is performed on each index image feature in the index image feature set to obtain a fragmented image feature group set.

4. The image retrieval enhancement method based on bimodal matching according to claim 1, characterized in that: The step of performing segmented text encoding on the matching text set according to the segmented image feature set to obtain the segmented text feature set includes: Selecting matching texts in the matching text set one by one as target matching texts, and taking the segment image feature group corresponding to the target matching text in the segment image feature group set as the target segment image feature group; Extracting a fragment feature number from the target fragment image feature group, and performing text segmentation on the target matching text according to the fragment feature number to obtain a segmented matching text group; Performing text segmentation and feature embedding on the segmented matching text group to obtain a segmented text feature group; Performing local attention calculation on the segmented text feature group to obtain a text local feature group; Performing global attention calculation on the local feature group of the text to obtain a global feature group of the text; Using the language image model, performing multi-scale feature extraction and full connection operation on the global text feature group to obtain a standard text feature group; Gathering the standard text feature groups of all target matching texts in the matching text set into a standard text feature group set; Each of the standard text feature groups in the standard text feature group set is clustered and stored to obtain a fragmented text feature group set.

5. The image retrieval enhancement method based on bimodal matching according to claim 1, characterized in that: The step of performing comparative training on the language image model by using the segmented image feature set and the segmented text feature set to obtain a bimodal matching model includes: Selecting the slice image feature groups in the slice image feature group set one by one as the target image feature group, and selecting the slice text feature group corresponding to the target image feature group in the slice text feature group set as the target text feature group; Generate a feature matching matrix according to the target image feature group and the target text feature group; Extracting a normal matching feature pair set and an abnormal matching feature pair set from the feature matching matrix respectively; respectively determining a normal matching similarity of the normal matching feature pair set and an abnormal matching similarity of the abnormal matching feature pair set; Determine a matching loss value according to the normal matching similarity and the abnormal matching similarity; Calculating a contrast loss value based on the matching loss values ​​of all target image feature groups in the segmented image feature group set; The language image model is iteratively trained according to the contrast loss value to obtain a bimodal matching model.

6. The image retrieval enhancement method based on bimodal matching according to claim 1, characterized in that: The using the bimodal matching model to perform preliminary image matching on the matching image set according to the user historical images and the user historical texts to obtain a primary recall image sequence includes: Using the bimodal matching model to perform image segmentation encoding on the user's historical images to obtain a historical image feature group; Using the bimodal matching model to perform text segmentation encoding on the user's historical text to obtain a historical text feature group; Obtaining a set of slice image feature groups corresponding to the bimodal matching model; Using the historical text feature group to perform cross-modal image matching on the segmented image feature group set to obtain a cross-modal recall image sequence; Using the historical image feature group to perform similarity image matching on the segmented image feature group set, to obtain a similarity recall image sequence; A primary recall image sequence is generated according to the cross-modal recall image sequence and the similarity recall image sequence.

7. The image retrieval enhancement method based on bimodal matching according to claim 1, characterized in that: The step of cross-coding and sorting the primary recall image sequence according to the user historical images and the user historical texts to obtain a standard recall image sequence includes: Acquire a historical image feature group corresponding to the user's historical image and a historical text feature group corresponding to the user's historical text; Cross-coding the historical image feature group and the historical text feature group to obtain coded user features; Extracting a primary recall text sequence corresponding to the primary recall image sequence from the matching text set; extracting a recalled image feature group sequence from the primary recalled image sequence, and extracting a recalled text feature group sequence from the primary recalled text sequence; generating a recalled feature group pair sequence according to the recalled image feature group sequence and the recalled text feature group sequence; Cross-coding each recall feature group pair in the recall feature group pair sequence to obtain a coded recall feature sequence; Using the coded user features, feature matching is performed on each coded recall feature in the coded recall feature sequence to obtain a matching degree sequence; The primary recalled image sequence is reordered according to the matching degree sequence to obtain a standard recalled image sequence.

8. An image retrieval enhancement device based on bimodal matching, characterized in that: include: A data acquisition module, used to acquire a matching image set and a matching text set corresponding to the matching image set; A fragment encoding module, used for performing fragment image encoding on the matching image set by using a preset language image model to obtain a fragment image feature set, and performing fragment text encoding on the matching text set according to the fragment image feature set to obtain a fragment text feature set; A comparative training module, used for performing comparative training on the language image model using the segmented image feature set and the segmented text feature set to obtain a bimodal matching model; A retrieval input module is used to obtain the user history image and user history text of the user to be queried; A preliminary matching module, configured to perform preliminary image matching on the matching image set according to the user historical images and the user historical texts using the bimodal matching model to obtain a primary recall image sequence; The precise sorting module is used to perform cross-coding sorting on the primary recall image sequence according to the user history images and the user history texts to obtain a standard recall image sequence.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the image retrieval enhancement method based on bimodal matching as described in any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the image retrieval enhancement method based on bimodal matching as claimed in any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Block chain fragmentation method and system based on multi-modal behavior information

    CN117082077A

  • Remote sensing image retrieval method and device, electronic equipment and computer storage medium

    CN117972126A

  • Heterogeneous XML document support in a sharded database

    US12158870B1

  • Transformer with multi-scale multi-context attentions

    WO2024263633A1