Enhanced Method, Device, Equipment and Medium for Image Retrieval Based on Bimodal Matching

Through shard coding and contrast training of image and text features, a dual-modal matching model is established, which solves the problem of low image retrieval efficiency and accuracy in the existing technology, and achieves efficient cross-modal retrieval and precise matching, improving the analysis capabilities in the medical and financial fields.

CN120067355BActive Publication Date: 2025-07-22SHENZHEN PINGAN COMM TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510510206.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-07-22
Estimated Expiration
2045-04-23

AI Technical Summary

Technical Problem

The existing image retrieval technology has problems with low retrieval efficiency and low accuracy in the fields of medical health and financial technology, mainly due to the lack of cross-modal feature analysis.

Method used

By obtaining matching image sets and text sets, using the preset language image model for shard encoding, generating sharded images and text feature sets, and performing comparison training, establishing a dual-modal matching model, combining user historical images and text for preliminary matching and cross-coding sorting, improving retrieval accuracy and flexibility.

Benefits of technology

It realizes efficient cross-modal retrieval, improves the efficiency of medical image analysis and the accuracy of financial risk analysis, and enhances the flexibility and accuracy of image retrieval.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067355B_ABST
    Figure CN120067355B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of image matching technology and can be applied to business system platforms such as fintech and medical and health. It discloses an enhanced method, device, equipment and medium for image retrieval based on bimodal matching, including: performing sharded image encoding on a set of matching images to obtain a set of sharded image feature groups, and performing sharded text encoding on a set of matching texts according to the set of sharded image feature groups to obtain a set of sharded text feature groups; using the set of sharded image feature groups and the set of sharded text feature groups to perform contrastive training on a language-image model to obtain a bimodal matching model; using the bimodal matching model to perform preliminary image matching on the set of matching images according to the user's historical images and historical texts to obtain a primary recall image sequence; performing cross-encoding sorting on the primary recall image sequence according to the user's historical images and historical texts to obtain a standard recall image sequence. The present invention can improve the accuracy of image retrieval.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image matching, and in particular, to an enhanced method, device, equipment and medium for image retrieval based on bimodal matching. Background Art

[0002] Image retrieval refers to retrieving images similar to a query input from a large image database, which can be matched based on text descriptions or the images themselves. Among them, text-description-based matching means matching the corresponding images according to the text describing the image content, and image-based matching means matching the images with similar content according to the input images. With the development of artificial intelligence technology, image retrieval has been more and more widely applied.

[0003] For example, in the field of medical health, doctors need to find similar case images in a vast medical image database to assist in diagnosis and treatment. For example, the corresponding pathological ultrasound images are matched through the text description of the case for analysis.

[0004] For another example, in the field of fintech business, in order to conduct risk analysis on users, it is necessary to find suspicious transaction pattern images in the database, so as to detect abnormal transaction patterns and give transaction warnings to users.

[0005] At present, the image retrieval technologies in the medical health and financial fields mainly rely on deep learning feature extraction and vector retrieval technologies, but there are certain limitations in terms of retrieval accuracy and efficiency, etc., and it is necessary to further improve the accuracy of image retrieval. Summary of the Invention

[0006] The present invention provides an enhanced method, device, equipment and medium for image retrieval based on bimodal matching to solve the technical problems of low efficiency and low accuracy in image retrieval in the existing image retrieval technology due to the lack of cross-modal feature analysis.

[0007] In a first aspect, an enhanced method for image retrieval based on bimodal matching is provided, including:

[0008] Obtaining a set of matching images and a set of matching texts corresponding to the set of matching images;

[0009] Performing sliced image encoding on the set of matching images by using a preset language-image model to obtain a set of sliced image feature groups, and performing sliced text encoding on the set of matching texts according to the set of sliced image feature groups to obtain a set of sliced text feature groups;

[0010] Comparatively training the language-image model by using the set of sliced image feature groups and the set of sliced text feature groups to obtain a bimodal matching model;

[0011] Obtain the user's historical images and historical texts of the user to be queried;

[0012] Use the bimodal matching model to perform preliminary image matching on the matching image set according to the user's historical images and the user's historical texts, and obtain a primary recall image sequence;

[0013] Perform cross-coding sorting on the primary recall image sequence according to the user's historical images and the user's historical texts to obtain a standard recall image sequence.

[0014] In a second aspect, an image retrieval enhancement device based on bimodal matching is provided, including:

[0015] A data acquisition module, configured to acquire a matching image set and a matching text set corresponding to the matching image set;

[0016] A sharding coding module, configured to perform sharding image coding on the matching image set by using a preset language-image model to obtain a set of sharded image feature groups, and perform sharding text coding on the matching text set according to the set of sharded image feature groups to obtain a set of sharded text feature groups;

[0017] A contrast training module, configured to perform contrast training on the language-image model by using the set of sharded image feature groups and the set of sharded text feature groups to obtain a bimodal matching model;

[0018] A retrieval input module, configured to acquire the user's historical images and the user's historical texts of the user to be queried;

[0019] A preliminary matching module, configured to perform preliminary image matching on the matching image set by using the bimodal matching model according to the user's historical images and the user's historical texts to obtain a primary recall image sequence;

[0020] An accurate sorting module, configured to perform cross-coding sorting on the primary recall image sequence according to the user's historical images and the user's historical texts to obtain a standard recall image sequence.

[0021] In a third aspect, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the above-mentioned image retrieval enhancement method based on bimodal matching is implemented.

[0022] In a fourth aspect, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned image retrieval enhancement method based on bimodal matching is implemented.

[0023] In the solution implemented by the above-described image retrieval enhancement method, device, computer device, and storage medium based on bimodal matching, communication can be carried out between the client and the server through the network of the client. The server can obtain the matching image set and the corresponding matching text set through the client, and can obtain mutually matching text data and image data, providing a basis for subsequent contrast learning between modal features; using a preset language-image model to perform sliced image encoding on the matching image set to obtain a set of sliced image feature groups, and performing sliced text encoding on the matching text set according to the set of sliced image feature groups to obtain a set of sliced text feature groups, which can effectively extract the local and global features of the matching image set and the matching text set, obtain efficient feature representations, and facilitate subsequent contrast training between text features and image features; using the set of sliced image feature groups and the set of sliced text feature groups to perform contrast training on the language-image model to obtain a bimodal matching model, which can achieve efficient cross-modal retrieval, thereby improving the efficiency of medical image analysis; obtaining the user's historical images and historical texts of the user to be queried; using the bimodal matching model to perform preliminary image matching on the matching image set according to the user's historical images and historical texts to obtain a primary recall image sequence; performing cross-encoding sorting on the primary recall image sequence according to the user's historical images and historical texts to obtain a standard recall image sequence. By using the set of sliced image feature groups and the set of sliced text feature groups to perform contrast training on the language-image model, the obtained bimodal matching model can achieve cross-modal contrast learning between text features and image features, thereby enhancing the flexibility of image retrieval and improving the accuracy of image retrieval. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings without creative efforts.

[0025] Figure 1 is an application environment diagram of an image retrieval enhancement method based on bimodal matching in an embodiment of the present invention;

[0026] Figure 2 is a flowchart of an image retrieval enhancement method based on bimodal matching in an embodiment of the present invention;

[0027] Figure 3 is Figure 1 a specific implementation flowchart of step S20 in

[0028] Figure 4 is Figure 1 a schematic flowchart of a specific implementation manner of step S30 in

[0029] Figure 5 is Figure 1 a schematic flowchart of a specific implementation manner of step S50 in

[0030] Figure 6 a schematic structural diagram of an image retrieval enhancement device based on bimodal matching in an embodiment of the present invention;

[0031] Figure 7 a schematic structural diagram of a computer device in an embodiment of the present invention;

[0032] Figure 8 is another schematic structural diagram of a computer device in an embodiment of the present invention. Specific implementation manner

[0033] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0034] The image retrieval enhancement method based on bimodal matching provided by the embodiments of the present invention can be applied in, for example, Figure 1In the application environment, the client communicates with the server through the network. The server can communicate with the server through the client's network. The server can obtain the matching image set and the matching text set corresponding to the matching image set through the client; perform sliced image encoding on the matching image set by using a preset language-image model to obtain a sliced image feature group set, and perform sliced text encoding on the matching text set according to the sliced image feature group set to obtain a sliced text feature group set; perform contrast training on the language-image model by using the sliced image feature group set and the sliced text feature group set to obtain a bimodal matching model; obtain the user's historical images and historical texts of the user to be queried; perform preliminary image matching on the matching image set by using the bimodal matching model according to the user's historical images and historical texts to obtain a primary recall image sequence; perform cross-encoding sorting on the primary recall image sequence according to the user's historical images and historical texts to obtain a standard recall image sequence. In the present invention, for the contrast training of the language-image model by using the sliced image feature group set and the sliced text feature group set to obtain a bimodal matching model, it includes: sequentially selecting the sliced image feature groups in the sliced image feature group set as target image feature groups, and using the sliced text feature groups corresponding to the target image feature groups in the sliced text feature group set as target text feature groups; generating a feature matching matrix according to the target image feature groups and the target text feature groups; respectively extracting a normal matching feature pair set and an abnormal matching feature pair set from the feature matching matrix; respectively determining the normal matching similarity of the normal matching feature pair set and the abnormal matching similarity of the abnormal matching feature pair set; calculating a matching loss value according to the normal matching similarity and the abnormal matching similarity; calculating a contrast loss value according to the matching loss values of all the target image feature groups in the sliced image feature group set; performing iterative training on the language-image model according to the contrast loss value to obtain a bimodal matching model, which can realize cross-modal contrast learning between text features and image features, thereby improving the flexibility of image retrieval and the accuracy of image retrieval. Among them, the client can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, and portable wearable devices. The server can be implemented by an independent server or a server cluster composed of multiple servers. The present invention will be described in detail through specific embodiments below.

[0035] Please refer to Figure 2 as shown in Figure 2 FIG. 1 is a schematic flowchart of an enhanced method for image retrieval based on bimodal matching provided by an embodiment of the present invention, including the following steps:

[0036] S10: Obtain a matching image set and a matching text set corresponding to the matching image set.

[0037] Specifically, the matching image set is a set composed of a large amount of customer image data with privacy data removed, and the image types of the respective matching images in the matching image set are the same as the image type of the image for which retrieval enhancement is required. The matching text set is a set composed of a large amount of customer text data with privacy data removed. Each matching text in the matching text set is descriptive text for describing each matching image in the matching image set, and the matching text and the matching image correspond one by one.

[0038] Example: In the field of medical and health, doctors often need to use patient medical record texts and patients' medical imaging images to improve the efficiency of medical diagnosis and treatment, helping doctors more accurately identify diseases and provide treatment plans. At this time, each matching image in the matching image set refers to the patient's x-ray image, microscopic image, ultrasound image, magnetic resonance image, and radionuclide image. Each matching text in the matching text set is text such as the diagnostic report, patient medical record, progress record, and test report corresponding to each matching image in the matching image set.

[0039] In the financial field, financial companies or financial organizations need to analyze the transaction behavior images of users in order to accurately evaluate the credit status of customers, monitor suspicious transaction patterns, and prevent fraud. At this time, each matching image in the matching image set refers to the time series curve of the user's transaction behavior, K-line chart, fund flow path chart, etc. Each matching text in the matching text set is text such as financial news, market analysis reports, company financial reports, and social media financial sentiment corresponding to each matching image in the matching image set.

[0040] In the embodiment of the present invention, by obtaining the matching image set and the matching text set corresponding to the matching image set, mutually matching text data and image data can be obtained, providing a basis for subsequent contrast learning between modal features.

[0041] S20: Use a preset language-image model to perform sliced image encoding on the matching image set to obtain a set of sliced image feature groups, and perform sliced text encoding on the matching text set according to the set of sliced image feature groups to obtain a set of sliced text feature groups.

[0042] Specifically, each sliced image feature in the set of sliced image feature groups refers to the feature data after each matching image in the matching image set is sliced and extracted. Each sliced text feature in the set of sliced text feature groups refers to the feature after each matching text in the matching text set is sliced and encoded.

[0043] In the embodiment of the present invention, with reference to Figure 3As shown, using a preset language-image model to perform piecewise image encoding on the matching image set to obtain a set of piecewise image feature groups, including:

[0044] S21. Select each matching image in the matching image set as a target matching image one by one, perform image segmentation on the target matching image to obtain a set of segmented matching tiles;

[0045] S22. Perform linear projection on the set of segmented matching tiles to obtain a set of segmented image features;

[0046] S23. Perform local attention calculation on the set of segmented image features to obtain a set of image local features;

[0047] S24. Perform global attention calculation on the set of image local features to obtain a set of image global features;

[0048] S25. Use the preset language-image model to perform multi-scale feature extraction and fully connected operation on the set of image global features to obtain a set of standard image features;

[0049] S26. Aggregate the sets of standard image features of all target matching images in the matching image set into a set of standard image features;

[0050] S27. Perform clustering storage on each of the standard image features in the set of standard image features to obtain a set of piecewise image feature groups.

[0051] Specifically, the image segmentation refers to segmenting the target matching image into multiple tiles of uniform size according to a preset window size, that is, each segmented matching tile in the set of segmented matching tiles. The linear projection refers to mapping a vector in a vector space to a subspace of the vector space, that is, mapping two-dimensional image features into one-dimensional linear features.

[0052] Specifically, the local attention calculation refers to using the multi-head self-attention mechanism to calculate the attention feature representation in each segmented image feature in the set of segmented image features. The global attention calculation refers to using a preset sliding window to calculate the attention features of the interaction between adjacent segmented image features, so as to enhance the global feature representation of the set of segmented image features.

[0053] Specifically, the multi-scale feature extraction refers to using a method of gradually downsampling to perform multi-size feature extraction on each image global feature in the set of image global features step by step, obtaining multiple global feature layers with a pyramid structure, and performing a fully connected operation on each global feature layer to obtain piecewise image features.

[0054] Specifically, clustering and storing each of the standard image feature groups in the standard image feature group set to obtain a fragmented image feature group set includes:

[0055] Performing a feature splicing operation on each standard image feature group in the standard image feature group set to obtain a spliced image feature set;

[0056] Randomly splitting the spliced image feature set into multiple spliced image feature groups;

[0057] Randomly screening out a spliced image feature as a spliced image center feature in each spliced image feature group to obtain a spliced image center feature group;

[0058] Respectively determining the feature distances between each spliced image feature in the spliced image feature set and each spliced image center feature in the spliced image center feature group;

[0059] According to the principle of proximity, allocating each spliced image feature in the spliced image feature set to the spliced image feature group where the corresponding spliced image center feature is located based on the feature distance to obtain an updated image feature group set;

[0060] Performing feature center calculation on each updated image feature group in the updated image feature group set to obtain an updated center feature group;

[0061] Obtaining the center feature distances between each updated center feature in the updated center feature group and the corresponding spliced image center feature in the spliced image center feature group, and taking the mean of all the center feature distances as the standard center feature distance;

[0062] Updating the updated image feature group set according to the standard center feature distance to obtain a clustered image feature group set;

[0063] Performing index storage on each clustered image feature group in the clustered image feature group set to obtain an indexed image feature set;

[0064] Performing feature splitting on each indexed image feature in the indexed image feature set to obtain a fragmented image feature group set.

[0065] Specifically, the performing a feature splicing operation on each standard image feature group in the standard image feature group set to obtain a spliced image feature set means splicing the standard image features in each standard image feature group into spliced image features, and aggregating all the spliced image features into a spliced image feature set.

[0066] Specifically, the following mathematical formula can be used to respectively determine the feature distances between each spliced image feature in the spliced image feature set and each spliced image center feature in the spliced image center feature group:

[0067]

[0068] Among them, refers to the feature distance between the spliced image feature and the center feature of the spliced image ; is the exponential function symbol, refers to the spliced image feature, refers to the center feature of the spliced image, is the transpose symbol, refers to the spliced image feature and the covariance matrix of the center feature of the spliced image ; is a preset regularization coefficient, is the identity matrix.

[0069] Specifically, the central feature distance can be calculated by using the cosine distance algorithm or the Euclidean distance algorithm. Updating the set of updated image feature groups according to the standard central feature distance to obtain a set of clustered image feature groups means determining whether the standard central feature distance is greater than a preset central distance threshold; if so, using the updated central feature group to update the central feature group of the spliced image, and returning to the step of respectively determining the feature distances between each spliced image feature in the spliced image feature set and each central feature of the spliced image central feature group; if not, using the set of updated image feature groups as the set of clustered image feature groups.

[0070] Specifically, the index storage means splitting and storing each clustered image feature group in the set of clustered image feature groups, and storing the central feature of each clustered image feature group as the feature index of the corresponding clustered image feature group, and storing all the clustered image features in a queue to obtain an indexed image feature set. The method of feature splitting is the inverse step of the method of feature splicing, which will not be elaborated here.

[0071] Example illustration: In the field of healthcare, when the matching images in the matching image set are magnetic resonance images, the image segmentation refers to slicing the matching image set and dividing it into small tiles of size 32×32. Each tile is a segmented matching tile in the segmented matching tile group. Each segmented matching tile group may correspond to different brain region structures in the magnetic resonance image, such as the cortex, white matter, and ventricles. The linear projection refers to projecting each segmented matching tile into a fixed feature dimension using a convolutional network or a multi-layer perceptron model. The local attention calculation refers to calculating the attention features corresponding to each segmented image feature one by one using a window of the same size as the segmented image feature, and aggregating the attention features as the local image features into a local image feature group, enabling the model to learn local structural information, such as the details of a certain lesion area. The global attention calculation refers to sliding a window over the feature sequence composed of the local image feature group and calculating the attention features within each corresponding window, so that the local features can spread to the global and construct a complete magnetic resonance imaging feature.

[0072] In the financial field, when the matching images in the matching image set are fund flow network images, the image segmentation refers to slicing the matching image set according to the trading areas or time windows of a fixed window, and dividing it into multiple small tiles. Each tile is a segmented matching tile in the segmented matching tile group. Each segmented matching tile group may correspond to the trading information of different trading time windows in the fund flow network image. The linear projection refers to projecting each segmented matching tile into a fixed feature dimension using a convolutional network or a multi-layer perceptron model. For example, each trading node can be represented by features such as trading amount, time interval, and trading frequency, and then mapped to a 768-dimensional feature vector using a linear layer. The local attention calculation refers to calculating the attention features corresponding to each segmented image feature one by one using a window of the same size as the segmented image feature, and aggregating the attention features as the local image features into a local image feature group, enabling the model to learn local structural information, such as the details of multiple suspicious transactions of a certain enterprise. The global attention calculation refers to sliding a window over the feature sequence composed of the local image feature group and calculating the attention features within each corresponding window, so that the local features can spread to the global and detect larger fund flow patterns, such as whether there are suspicious behaviors of layer-by-layer fund transfer.

[0073] Specifically, the step of performing segmented text encoding on the matching text set according to the set of segmented image feature groups to obtain a set of segmented text feature groups includes:

[0074] Select the matching texts in the matching text set one by one as the target matching texts, and use the shard image feature groups corresponding to the target matching texts in the shard image feature group set as the target shard image feature groups;

[0075] Extract the number of shard features from the target shard image feature groups;

[0076] Segment the target matching texts according to the number of shard features to obtain a segmented matching text set;

[0077] Perform text tokenization and feature embedding on the segmented matching text set to obtain a segmented text feature set;

[0078] Perform local attention calculation on the segmented text feature set to obtain a text local feature set;

[0079] Perform global attention calculation on the text local feature set to obtain a text global feature set;

[0080] Use the language-image model to perform multi-scale feature extraction and fully connected operations on the text global feature set to obtain a standard text feature set;

[0081] Aggregate the standard text feature sets of all target matching texts in the matching text set into a standard text feature set;

[0082] Perform clustering storage on each standard text feature set in the standard text feature set to obtain a shard text feature set.

[0083] Specifically, the number of shard features refers to the total number of image features in the target shard image feature group. The segmenting the target matching texts according to the number of shard features to obtain a segmented matching text set means splitting the target matching texts into the number of text segments equal to the number of shard features, and the text can be segmented according to symbols or paragraphs of the text.

[0084] Specifically, text tokenization means dividing a continuous text into individual words or sub-word units. Text tokenization can be performed using bidirectional maximum matching, regular expression tokenization, or tokenization methods based on character segmentation such as Byte Pair Encoding, WordPiece, and SentencePiece. Feature embedding refers to the process of converting the tokenized words into word features, that is, representing discrete text data as continuous numerical vectors. Feature embedding can be performed using statistical-based feature embedding methods such as Bag of Words (BoW) and TF IDF, prediction-based feature embedding methods such as Word2Vec and GloVe, and context-based dynamic word feature-based feature embedding methods such as ELMo and Transformer based models.

[0085] In the embodiments of the present invention, by using a preset language-image model to perform sliced image encoding on the set of matching images to obtain a set of sliced image feature groups, and performing sliced text encoding on the set of matching texts according to the set of sliced image feature groups to obtain a set of sliced text feature groups, local features and global features of the set of matching images and the set of matching texts can be effectively extracted, high-efficiency feature representations can be obtained, and subsequent contrast training between text features and image features can be facilitated.

[0086] S30: Use the set of sliced image feature groups and the set of sliced text feature groups to perform contrast training on the language-image model to obtain a bimodal matching model.

[0087] Specifically, the language-image model refers to a Contrastive Language Image Pretraining (CLIP for short) model. The Contrastive Language Image Pretraining model is a multimodal model proposed by OpenAI, which can understand the relationship between images and texts. The Contrastive Language Image Pretraining model uses a contrastive learning method to train on a large-scale image and text pairing data, enabling the model to directly perform matching, retrieval, and classification on images and texts.

[0088] In the embodiments of the present invention, as shown in Figure 4 The step of using the set of sliced image feature groups and the set of sliced text feature groups to perform contrast training on the language-image model to obtain a bimodal matching model includes:

[0089] S31: Select each sliced image feature group in the set of sliced image feature groups as a target image feature group one by one, and use the corresponding sliced text feature group of the target image feature group in the set of sliced text feature groups as a target text feature group;

[0090] S32: Generate a feature matching matrix according to the target image feature group and the target text feature group;

[0091] S33: Extract a set of normal matching feature pairs and a set of abnormal matching feature pairs from the feature matching matrix respectively;

[0092] S34: Determine a set of normal matching similarities of the set of normal matching feature pairs and a set of abnormal matching similarities of the set of abnormal matching feature pairs respectively;

[0093] S35: Calculate a matching loss value according to the set of normal matching similarities and the set of abnormal matching similarities;

[0094] S36. Calculate a contrast loss value based on the matching loss values of all target image feature groups in the sliced image feature group set;

[0095] S37. Iteratively train the language image model according to the contrast loss value to obtain a bimodal matching model.

[0096] Specifically, the feature matching matrix refers to a matrix established by enumerating and pairwise matching each target image feature in the target image feature group as a row vector and the target text feature group as a column vector.

[0097] Specifically, the step of respectively extracting a normal matching feature pair set and an abnormal matching feature pair set from the feature matching matrix means taking the feature pairs in the main diagonal position of the feature matching matrix as normal matching feature pairs to form a normal matching feature pair set, and taking the feature pairs not in the main diagonal position of the feature matching matrix as abnormal matching feature pairs to form an abnormal matching feature pair set.

[0098] In detail, the cosine similarity algorithm can be used to respectively determine the normal matching similarity set of the normal matching feature pair set and the abnormal matching similarity set of the abnormal matching feature pair set.

[0099] Specifically, use the following matching algorithm to calculate the matching loss value according to the normal matching similarity set and the abnormal matching similarity set:

[0100]

[0101] where, refers to the matching loss value, refers to the total number of features in the target image feature group, and the total number of features in the target image feature group is equal to the total number of features in the target text feature group, 、 is the feature index, is the logarithmic function symbol, is the exponential function symbol, is the cosine similarity calculation symbol, is the th image feature in the target image feature group, is the th text feature in the target text feature group, is the th image feature in the target image feature group, is the th text feature in the target text feature group, is a preset matching weight, is the dot product symbol, is the modulo symbol.

[0102] Specifically, the normal matching feature pair set refers to the set of feature pairs located at the main diagonal position, that is, , the abnormal matching feature pair set refers to the set of feature pairs at other positions, that is, and , , the contrast loss value is the average of the matching loss values of all target image feature groups in the segmented image feature group set.

[0103] In detail, the iterative training of the language image model according to the contrast loss value to obtain the bimodal matching model refers to determining whether the contrast loss value is greater than a preset loss value threshold. If so, the contrast loss value is used to update the model parameters of the language image model using a back propagation algorithm, and the step of performing piecewise image encoding on the matching image set using the preset language image model to obtain a piecewise image feature set is returned. If not, the language image model is used as a bimodal matching model.

[0104] In an embodiment of the present invention, by using the segmented image feature set and the segmented text feature set to perform comparative training on the language image model, a bimodal matching model is obtained, and a retrieval model that achieves matching between image and text can be obtained, thereby achieving efficient cross-modal retrieval and improving the efficiency of medical image analysis.

[0105] S40: Acquire the user history image and user history text of the user to be queried.

[0106] In detail, the user to be queried refers to the user who needs to perform the corresponding image retrieval operation, the user historical image refers to the image stored in the database in the past by the user who needs to perform image retrieval, and the user historical text refers to the text stored in the database in the past by the user who needs to perform image retrieval.

[0107] Example description: In the field of medical health, the user to be queried refers to a user who needs to be diagnosed with a health condition, the user historical images refer to the X-ray images, microscopic images, ultrasound images, magnetic resonance images, and radionuclide images retained when the user to be queried was diagnosed in the past, and the user historical texts refer to the diagnostic reports, patient medical records, medical records, test reports and other texts retained when the user to be queried was diagnosed in the past.

[0108] In the financial field, the user to be queried refers to a user who needs to conduct financial risk analysis. The user historical images refer to time series curves, K-line charts, fund flow path diagrams, etc. that record the trading behaviors uploaded by the user to be queried in the past. The user historical texts refer to texts such as financial news, market analysis reports, company financial reports, and financial sentiment on social media of the user to be queried in the past.

[0109] In an embodiment of the present invention, the obtaining of the user historical images and user historical texts of the user to be queried includes: obtaining the user identifier of the user to be queried; querying historical image data of the user to be queried according to the user identifier to obtain user historical images; querying historical text data of the user to be queried according to the user identifier to obtain user historical texts.

[0110] Specifically, the user identifier refers to a unique identifier used to confirm the user identity of the user to be queried. For example, the ID card number of the user to be queried. Querying historical image data of the user to be queried according to the user identifier to obtain user historical images means retrieving associated image data according to the user identifier in a preset database to obtain user historical images. Querying historical text data of the user to be queried according to the user identifier to obtain user historical texts means retrieving associated text data according to the user identifier in a preset database to obtain user historical texts.

[0111] In an embodiment of the present invention, by obtaining the user historical images and user historical texts of the user to be queried, the user to be queried can be obtained, and the past information of the user to be queried can be quickly extracted, thereby improving the efficiency of medical image analysis.

[0112] S50: Using the bimodal matching model, based on the user historical images and the user historical texts, perform preliminary image matching on the matching image set to obtain a primary recall image sequence.

[0113] Specifically, each primary recall image included in the primary recall image sequence is a matching image in the matching image set related to the user historical images and the user historical texts arranged according to the recall rate.

[0114] In an embodiment of the present invention, as shown in Figure 5 Using the bimodal matching model, based on the user historical images and the user historical texts, perform preliminary image matching on the matching image set to obtain a primary recall image sequence includes:

[0115] S51: Using the bimodal matching model to perform sliced image encoding on the user historical images to obtain a historical image feature group;

[0116] S52. Piecewise encode the user's historical text using the bimodal matching model to obtain a historical text feature group;

[0117] S53. Obtain a set of piecewise image feature groups corresponding to the bimodal matching model;

[0118] S54. Perform cross-modal image matching on the set of piecewise image feature groups using the historical text feature group to obtain a cross-modal recalled image sequence;

[0119] S55. Perform similarity image matching on the set of piecewise image feature groups using the historical image feature group to obtain a similarity recalled image sequence;

[0120] S56. Generate a primary recalled image sequence based on the cross-modal recalled image sequence and the similarity recalled image sequence.

[0121] Specifically, the method of piecewise image encoding is the same as the method of piecewise image encoding in step S20 above, which will not be elaborated here. The method of piecewise text encoding is the same as the method of piecewise image encoding in step S20 above, which will not be elaborated here. The set of piecewise image feature groups corresponding to the bimodal matching model refers to the set of piecewise image feature groups corresponding when the language-image model is updated to the bimodal matching model.

[0122] Specifically, the process of performing cross-modal image matching on the set of piecewise image feature groups using the historical text feature group to obtain a cross-modal recalled image sequence means using the historical text feature group to perform similarity matching with each piecewise image feature group in the set of piecewise image feature groups, screening out the piecewise image feature groups with similarity greater than a preset similarity threshold as recalled image feature groups, and aggregating all the recalled image feature groups into a recalled image feature group sequence in descending order of similarity. The respective matching images corresponding to the recalled image feature group sequence are aggregated as cross-modal recalled images into a cross-modal recalled image sequence.

[0123] Specifically, the process of performing similarity image matching on the set of piecewise image feature groups using the historical image feature group to obtain a similarity recalled image sequence means using the historical image feature group to perform similarity matching with each piecewise image feature group in the set of piecewise image feature groups, screening out the piecewise image feature groups with similarity greater than a preset similarity threshold as similarity recalled image feature groups, and aggregating all the similarity recalled image feature groups into a similarity recalled image feature group sequence in descending order of similarity. The respective matching images corresponding to the similarity recalled image feature group sequence are aggregated as similarity recalled images into a similarity recalled image sequence.

[0124] Specifically, generating the primary recall image sequence based on the cross-modal recall image sequence and the similarity recall image sequence means globally sorting the images in the cross-modal recall image sequence and the similarity recall image sequence according to their similarities to obtain the primary recall image sequence.

[0125] In the embodiment of the present invention, by using the bimodal matching model to perform preliminary image matching on the matching image set according to the user's historical images and the user's historical texts, a primary recall image sequence is obtained, which can realize cross-modal retrieval of text and images and image retrieval between the same modalities, thereby improving the flexibility of image retrieval.

[0126] S60: Perform cross-coding sorting on the primary recall image sequence according to the user's historical images and the user's historical texts to obtain a standard recall image sequence.

[0127] Specifically, each standard recall image in the standard recall image sequence refers to the final relevant image result obtained after performing image retrieval on the user's historical images and the user's historical texts.

[0128] In the embodiment of the present invention, performing cross-coding sorting on the primary recall image sequence according to the user's historical images and the user's historical texts to obtain a standard recall image sequence includes:

[0129] Obtain the historical image feature group corresponding to the user's historical images and the historical text feature group corresponding to the user's historical texts;

[0130] Perform cross-coding on the historical image feature group and the historical text feature group to obtain encoded user features;

[0131] Extract the primary recall text sequence corresponding to the primary recall image sequence from the matching text set;

[0132] Extract the recall image feature group sequence from the primary recall image sequence and the recall text feature group sequence from the primary recall text sequence respectively;

[0133] Generate a recall feature group pair sequence according to the recall image feature group sequence and the recall text feature group sequence;

[0134] Perform cross-coding on each recall feature group pair in the recall feature group pair sequence to obtain an encoded recall feature sequence;

[0135] Use the encoded user features to perform feature matching on each encoded recall feature in the encoded recall feature sequence to obtain a matching degree sequence;

[0136] Reorder the primary recall image sequence according to the matching degree sequence to obtain a standard recall image sequence.

[0137] Specifically, the cross-coding refers to calculating the attention feature coding between the image feature group and the corresponding text feature group by using the cross-attention mechanism. Generating the recall feature group pair sequence according to the recall image feature group sequence and the recall text feature group sequence means combining each recall image feature group in the recall image feature group sequence with the corresponding recall text feature group in the recall text feature group sequence to obtain a recall feature group pair, and aggregating all the recall feature group pairs into a recall feature group pair sequence.

[0138] Specifically, the feature matching refers to calculating the matching degree between features by using a similarity calculation formula, and aggregating all the matching degrees into a matching degree sequence. Reordering the primary recall image sequence according to the matching degree sequence to obtain a standard recall image sequence means performing a weighted summation operation according to the matching degree sequence and the similarity of each primary recall image in the primary recall image sequence to obtain the final recall degree, and reordering the primary recall image sequence according to the recall degree to obtain a standard recall image sequence.

[0139] In the embodiment of the present invention, by cross-coding and sorting the primary recall image sequence according to the user's historical image and the user's historical text to obtain a standard recall image sequence, it is possible to further accurately sort the indexing results according to the context relationship between the text and the image, thereby improving the accuracy of image retrieval.

[0140] It can be seen that in the above solution, by obtaining the matching image set and the matching text set corresponding to the matching image set, it is possible to obtain mutually matching text data and image data, providing a basis for subsequent contrast learning between modal features. By using a preset language-image model to perform slice image coding on the matching image set to obtain a slice image feature group set, and performing slice text coding on the matching text set according to the slice image feature group set to obtain a slice text feature group set, it is possible to effectively extract the local features and global features of the matching image set and the matching text set, obtain an efficient feature representation, and facilitate subsequent contrast training between the text features and the image features. By using the slice image feature group set and the slice text feature group set to perform contrast training on the language-image model to obtain a bimodal matching model, it is possible to obtain a retrieval model that realizes the matching between images and texts, realizing efficient cross-modal retrieval, thereby improving the efficiency of medical image analysis.

[0141] By obtaining the user's historical images and historical texts of the user to be queried, the user to be queried can be obtained, and the past information of the user to be queried can be quickly extracted, thereby improving the efficiency of medical image analysis. By using the dual-modal matching model to perform preliminary image matching on the matching image set according to the user's historical images and the user's historical texts, a primary recall image sequence is obtained, which can realize cross-modal retrieval of text and images and image retrieval between the same modalities, thereby improving the flexibility of image retrieval. By cross-coding and sorting the primary recall image sequence according to the user's historical images and the user's historical texts, a standard recall image sequence is obtained, which can further accurately sort the indexing results according to the context relationship between the text and the image, and improve the accuracy of image retrieval.

[0142] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.

[0143] In one embodiment, an image retrieval enhancement device based on dual-modal matching is provided. The image retrieval enhancement device based on dual-modal matching corresponds one-to-one with the image retrieval enhancement method based on dual-modal matching in the above embodiment. As Figure 6 shown, the image retrieval enhancement device based on dual-modal matching includes a data acquisition module 101, a sharding encoding module 102, a contrast training module 103, a retrieval input module 104, a preliminary matching module 105, and an accurate sorting module 106. The detailed description of each functional module is as follows:

[0144] The data acquisition module 101 is used to acquire a matching image set and a matching text set corresponding to the matching image set;

[0145] The sharding encoding module 102 is used to perform sharding image encoding on the matching image set by using a preset language image model to obtain a sharding image feature group set, and perform sharding text encoding on the matching text set according to the sharding image feature group set to obtain a sharding text feature group set;

[0146] The contrast training module 103 is used to perform contrast training on the language image model by using the sharding image feature group set and the sharding text feature group set to obtain a dual-modal matching model;

[0147] The retrieval input module 104 is used to acquire the user's historical images and historical texts of the user to be queried;

[0148] The preliminary matching module 105 is configured to perform preliminary image matching on the matching image set according to the user's historical image and the user's historical text by using the bimodal matching model, so as to obtain a primary recall image sequence;

[0149] The precise sorting module 106 is configured to perform cross-coding sorting on the primary recall image sequence according to the user's historical image and the user's historical text, so as to obtain a standard recall image sequence.

[0150] In one embodiment, when the sharding encoding module 102 executes sharding image encoding on the matching image set by using a preset language-image model to obtain a set of sharded image feature groups, it is configured to:

[0151] Select each matching image in the matching image set as a target matching image one by one, perform image segmentation on the target matching image to obtain a set of segmented matching tiles;

[0152] Perform linear projection on the set of segmented matching tiles to obtain a set of segmented image features;

[0153] Perform local attention calculation on the set of segmented image features to obtain a set of image local features;

[0154] Perform global attention calculation on the set of image local features to obtain a set of image global features;

[0155] Use a preset language-image model to perform multi-scale feature extraction and fully connected operations on the set of image global features to obtain a set of standard image features;

[0156] Aggregate the sets of standard image features of all target matching images in the matching image set into a set of standard image features;

[0157] Perform clustering storage on each of the sets of standard image features in the set of standard image features to obtain a set of sharded image feature groups.

[0158] In one embodiment, when the sharding encoding module 102 executes clustering storage on each of the sets of standard image features in the set of standard image features to obtain a set of sharded image feature groups, it is configured to:

[0159] Perform feature splicing operations on each of the sets of standard image features in the set of standard image features to obtain a set of spliced image features;

[0160] Randomly split the set of spliced image features into multiple sets of spliced image features;

[0161] Randomly select a spliced image feature as a spliced image central feature in each set of spliced image features to obtain a set of spliced image central features;

[0162] Determine the feature distances between each stitching image feature in the stitching image feature set and each stitching image center feature in the stitching image center feature group respectively;

[0163] According to the principle of proximity, allocate each stitching image feature in the stitching image feature set to the stitching image feature group where the corresponding stitching image center feature is located based on the feature distances, to obtain an updated image feature group set;

[0164] Perform feature center calculation on each updated image feature group in the updated image feature group set to obtain an updated center feature group;

[0165] Obtain the center feature distances between each updated center feature in the updated center feature group and the corresponding stitching image center feature in the stitching image center feature group, and take the mean value of all the center feature distances as the standard center feature distance;

[0166] Update the updated image feature group set according to the standard center feature distance to obtain a clustered image feature group set;

[0167] Perform index storage on each clustered image feature group in the clustered image feature group set to obtain an indexed image feature set;

[0168] Perform feature splitting on each indexed image feature in the indexed image feature set to obtain a sliced image feature group set.

[0169] In one embodiment, when the slicing encoding module 102 performs slicing text encoding on the matching text set according to the sliced image feature group set to obtain a sliced text feature group set, it is used for:

[0170] Select each matching text in the matching text set one by one as the target matching text, and use the sliced image feature group corresponding to the target matching text in the sliced image feature group set as the target sliced image feature group;

[0171] Extract the number of sliced features from the target sliced image feature group;

[0172] Perform text segmentation on the target matching text according to the number of sliced features to obtain a segmented matching text group;

[0173] Perform text word segmentation and feature embedding on the segmented matching text group to obtain a segmented text feature group;

[0174] Perform local attention calculation on the segmented text feature group to obtain a text local feature group;

[0175] Perform global attention calculation on the text local feature group to obtain a text global feature group;

[0176] Performing multi-scale feature extraction and fully-connected operation on the text global feature group by using the language image model to obtain a standard text feature group;

[0177] Aggregating the standard text feature groups of all target matching texts in the matching text set into a standard text feature group set;

[0178] Performing clustering storage on each standard text feature group in the standard text feature group set to obtain a sharded text feature group set.

[0179] In one embodiment, when the contrast training module 103 performs contrast training on the language image model by using the sharded image feature group set and the sharded text feature group set to obtain a bimodal matching model, it is used for:

[0180] Sequentially selecting the sharded image feature groups in the sharded image feature group set as target image feature groups, and using the sharded text feature groups corresponding to the target image feature groups in the sharded text feature group set as target text feature groups;

[0181] Generating a feature matching matrix according to the target image feature group and the target text feature group;

[0182] Respectively extracting a normal matching feature pair set and an abnormal matching feature pair set from the feature matching matrix;

[0183] Respectively determining the normal matching similarity of the normal matching feature pair set and the abnormal matching similarity of the abnormal matching feature pair set;

[0184] Calculating a matching loss value according to the normal matching similarity and the abnormal matching similarity;

[0185] Calculating a contrast loss value according to the matching loss values of all target image feature groups in the sharded image feature group set;

[0186] Performing iterative training on the language image model according to the contrast loss value to obtain a bimodal matching model.

[0187] In one embodiment, when the preliminary matching module 105 performs preliminary image matching on the matching image set according to the user historical image and the user historical text by using the bimodal matching model to obtain a primary recall image sequence, it is used for:

[0188] Performing sharded image encoding on the user historical image by using the bimodal matching model to obtain a historical image feature group;

[0189] Performing sharded text encoding on the user historical text by using the bimodal matching model to obtain a historical text feature group;

[0190] Obtain the set of shard image feature groups corresponding to the bimodal matching model;

[0191] Perform cross-modal image matching on the set of shard image feature groups using the historical text feature groups to obtain a cross-modal recalled image sequence;

[0192] Perform similarity image matching on the set of shard image feature groups using the historical image feature groups to obtain a similarity recalled image sequence;

[0193] Generate a primary recalled image sequence according to the cross-modal recalled image sequence and the similarity recalled image sequence.

[0194] In one embodiment, when the precise sorting module 106 performs cross-encoding and sorting on the primary recalled image sequence according to the user's historical image and the user's historical text to obtain a standard recalled image sequence, it is used for:

[0195] Obtain the historical image feature group corresponding to the user's historical image and the historical text feature group corresponding to the user's historical text;

[0196] Perform cross-encoding on the historical image feature group and the historical text feature group to obtain encoded user features;

[0197] Extract the primary recalled text sequence corresponding to the primary recalled image sequence from the matching text set;

[0198] Extract a recalled image feature group sequence from the primary recalled image sequence and a recalled text feature group sequence from the primary recalled text sequence respectively;

[0199] Generate a sequence of recalled feature group pairs according to the recalled image feature group sequence and the recalled text feature group sequence;

[0200] Perform cross-encoding on each recalled feature group pair in the sequence of recalled feature group pairs to obtain an encoded recalled feature sequence;

[0201] Perform feature matching on each encoded recalled feature in the encoded recalled feature sequence using the encoded user features to obtain a matching degree sequence;

[0202] Re-sort the primary recalled image sequence according to the matching degree sequence to obtain a standard recalled image sequence.

[0203] The present invention provides an enhanced image retrieval device based on bimodal matching. By obtaining a matching image set and a corresponding matching text set, it is possible to obtain mutually matching text data and image data, providing a basis for subsequent contrast learning between modal features. By using a preset language-image model to perform sliced image encoding on the matching image set to obtain a sliced image feature group set, and performing sliced text encoding on the matching text set according to the sliced image feature group set to obtain a sliced text feature group set, it is possible to effectively extract the local and global features of the matching image set and the matching text set, obtain an efficient feature representation, and facilitate subsequent contrast training between text features and image features. By using the sliced image feature group set and the sliced text feature group set to perform contrast training on the language-image model to obtain a bimodal matching model, a retrieval model capable of realizing the matching between images and texts can be obtained, achieving efficient cross-modal retrieval, thereby improving the efficiency of medical image analysis.

[0204] By obtaining the user's historical images and historical texts of the user to be queried, the user to be queried can be obtained, and the past information of the user to be queried can be quickly extracted, thereby improving the efficiency of medical image analysis. By using the bimodal matching model to perform preliminary image matching on the matching image set according to the user's historical images and historical texts to obtain a primary recall image sequence, cross-modal retrieval between text and image and image retrieval between the same modalities can be realized, thereby improving the flexibility of image retrieval. By performing cross-encoding sorting on the primary recall image sequence according to the user's historical images and historical texts to obtain a standard recall image sequence, the index results can be further accurately sorted according to the context relationship between text and image, improving the accuracy of image retrieval.

[0205] For the specific limitations of the enhanced image retrieval device based on bimodal matching, reference can be made to the limitations of the method for intelligent question answering in the above text, which will not be elaborated here. Each module in the above enhanced image retrieval device based on bimodal matching can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form, so as to facilitate the processor to call and execute the operations corresponding to the above respective modules.

[0206] In one embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 7As shown in the figure. The computer device includes a processor, a memory, a network interface, and a database connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external client via a network connection. When the computer program is executed by the processor, it realizes the functions or steps of the server side of an image retrieval enhancement method based on dual-modal matching.

[0207] In one embodiment, a computer device is provided. The computer device can be a client, and its internal structure diagram can be as Figure 8 As shown in the figure. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it realizes the functions or steps of the client side of an image retrieval enhancement method based on dual-modal matching.

[0208] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the following steps are implemented:

[0209] Obtain a set of matching images and a set of matching texts corresponding to the set of matching images;

[0210] Use a preset language-image model to perform slice image encoding on the set of matching images to obtain a set of slice image feature groups, and perform slice text encoding on the set of matching texts according to the set of slice image feature groups to obtain a set of slice text feature groups;

[0211] Use the set of slice image feature groups and the set of slice text feature groups to perform contrast training on the language-image model to obtain a dual-modal matching model;

[0212] Obtain the user's historical images and historical texts of the user to be queried;

[0213] Use the dual-modal matching model to perform preliminary image matching on the set of matching images according to the user's historical images and the user's historical texts to obtain a primary recall image sequence;

[0214] Cross - encode and sort the primary recall image sequence according to the user historical image and the user historical text to obtain a standard recall image sequence.

[0215] In one embodiment, a computer - readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0216] Obtain a set of matching images and a set of matching texts corresponding to the set of matching images;

[0217] Use a preset language - image model to perform sliced - image encoding on the set of matching images to obtain a set of sliced - image feature groups, and perform sliced - text encoding on the set of matching texts according to the set of sliced - image feature groups to obtain a set of sliced - text feature groups;

[0218] Use the set of sliced - image feature groups and the set of sliced - text feature groups to perform contrast training on the language - image model to obtain a bimodal matching model;

[0219] Obtain the user historical image and the user historical text of the user to be queried;

[0220] Use the bimodal matching model to perform preliminary image matching on the set of matching images according to the user historical image and the user historical text to obtain a primary recall image sequence;

[0221] Cross - encode and sort the primary recall image sequence according to the user historical image and the user historical text to obtain a standard recall image sequence.

[0222] It should be noted that for the functions or steps that the above - mentioned computer - readable storage medium or computer device can achieve, reference can be made to the relevant descriptions on the server side and the client side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.

[0223] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0224] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above.

[0225] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention. It should be noted that if non-company software tools or components appear in the embodiments of the present application, they are only used for illustrative introduction and do not represent actual use.

Claims

1. An enhanced method for image retrieval based on bimodal matching, characterized in that Including: Obtain a set of matching images and a set of matching texts corresponding to the set of matching images; Use a preset language-image model to perform piecewise image encoding on the set of matching images to obtain a set of piecewise image feature groups, and perform piecewise text encoding on the set of matching texts according to the set of piecewise image feature groups to obtain a set of piecewise text feature groups; wherein, the using the preset language-image model to perform piecewise image encoding on the set of matching images to obtain a set of piecewise image feature groups includes: successively select the matching images in the set of matching images as target matching images, perform image segmentation on the target matching images to obtain a set of segmented matching tiles; perform linear projection on the set of segmented matching tiles to obtain a set of segmented image features; perform local attention calculation on the set of segmented image features to obtain a set of image local features; perform global attention calculation on the set of image local features to obtain a set of image global features; use the preset language-image model to perform multi-scale feature extraction and fully connected operation on the set of image global features to obtain a set of standard image features; aggregate the sets of standard image features of all target matching images in the set of matching images into a set of standard image feature sets; perform clustering storage on each of the standard image feature sets in the set of standard image feature sets to obtain a set of piecewise image feature groups; Use the set of piecewise image feature groups and the set of piecewise text feature groups to perform contrastive training on the language-image model to obtain a bimodal matching model; Obtain the user's historical images and historical texts of the user to be queried; Use the bimodal matching model to perform preliminary image matching on the set of matching images according to the user's historical images and the user's historical texts to obtain a primary recall image sequence; Perform cross-encoding sorting on the primary recall image sequence according to the user's historical images and the user's historical texts to obtain a standard recall image sequence.

2. The image retrieval enhancement method based on dual-modal matching according to claim 1, wherein The performing clustering storage on each of the standard image feature sets in the set of standard image feature sets to obtain a set of piecewise image feature groups includes: Perform feature splicing operation on each of the standard image feature sets in the set of standard image feature sets to obtain a set of spliced image features; Randomly split the set of spliced image features into multiple groups of spliced image features; Randomly select a spliced image feature as a spliced image central feature in each group of spliced image features to obtain a set of spliced image central features; Respectively determine the feature distances between each of the spliced image features in the set of spliced image features and each of the spliced image central features in the set of spliced image central features; According to the principle of proximity, allocate each of the spliced image features in the set of spliced image features to the group of spliced image features where the corresponding spliced image central feature is located to obtain an updated set of image feature groups; Perform feature center calculation on each of the updated image feature groups in the updated set of image feature groups to obtain an updated set of central features; Obtain the central feature distances between each update center feature in the update center feature group and the corresponding stitching image center feature in the stitching image center feature group, and take the average of all the central feature distances as the standard central feature distance; Update the update image feature group set according to the standard central feature distance to obtain a clustered image feature group set; Index and store each clustered image feature group in the clustered image feature group set to obtain an indexed image feature set; Perform feature splitting on each indexed image feature in the indexed image feature set to obtain a sliced image feature group set.

3. The enhanced method for image retrieval based on bimodal matching according to claim 1, wherein The step of performing sliced text encoding on the matching text set according to the sliced image feature group set to obtain a sliced text feature group set includes: Select each matching text in the matching text set as a target matching text one by one, and use the sliced image feature group corresponding to the target matching text in the sliced image feature group set as the target sliced image feature group; Extract the number of sliced features from the target sliced image feature group, and perform text segmentation on the target matching text according to the number of sliced features to obtain a segmented matching text group; Perform text tokenization and feature embedding on the segmented matching text group to obtain a segmented text feature group; Perform local attention calculation on the segmented text feature group to obtain a text local feature group; Perform global attention calculation on the text local feature group to obtain a text global feature group; Use the language-image model to perform multi-scale feature extraction and fully connected operations on the text global feature group to obtain a standard text feature group; Aggregate the standard text feature groups of all target matching texts in the matching text set into a standard text feature group set; Perform clustering storage on each standard text feature group in the standard text feature group set to obtain a sliced text feature group set.

4. The enhanced method for image retrieval based on bimodal matching according to claim 1, characterized in that The step of using the sliced image feature group set and the sliced text feature group set to perform contrastive training on the language-image model to obtain a bimodal matching model includes: Select each sliced image feature group in the sliced image feature group set as a target image feature group one by one, and use the sliced text feature group corresponding to the target image feature group in the sliced text feature group set as the target text feature group; Generate a feature matching matrix according to the target image feature group and the target text feature group; Extract a normal matching feature pair set and an abnormal matching feature pair set from the feature matching matrix respectively; Determine the normal matching similarity of the normal matching feature pair set and the abnormal matching similarity of the abnormal matching feature pair set respectively; Determine a matching loss value according to the normal matching similarity and the abnormal matching similarity; Calculate a contrastive loss value according to the matching loss values of all target image feature groups in the sliced image feature group set; Perform iterative training on the language-image model according to the contrastive loss value to obtain a bimodal matching model.

5. The enhanced method for image retrieval based on bimodal matching according to claim 1, characterized in that, The step of using the bimodal matching model to perform preliminary image matching on the matching image set according to the user's historical image and the user's historical text to obtain a primary recall image sequence includes: Encode the user historical image into shard image features by using the dual-modal matching model to obtain a historical image feature group; Encode the user historical text into shard text features by using the dual-modal matching model to obtain a historical text feature group; Obtain the shard image feature group set corresponding to the dual-modal matching model; Perform cross-modal image matching on the shard image feature group set by using the historical text feature group to obtain a cross-modal recalled image sequence; Perform similarity image matching on the shard image feature group set by using the historical image feature group to obtain a similarity recalled image sequence; Generate a primary recalled image sequence according to the cross-modal recalled image sequence and the similarity recalled image sequence.

6. The enhanced image retrieval method based on bimodal matching according to claim 1, characterized in that The step of performing cross-coding sorting on the primary recalled image sequence according to the user historical image and the user historical text to obtain a standard recalled image sequence includes: Obtain the historical image feature group corresponding to the user historical image and the historical text feature group corresponding to the user historical text; Perform cross-coding on the historical image feature group and the historical text feature group to obtain an encoded user feature; Extract the primary recalled text sequence corresponding to the primary recalled image sequence from the matching text set; Extract a recalled image feature group sequence from the primary recalled image sequence, and extract a recalled text feature group sequence from the primary recalled text sequence; Generate a recalled feature group pair sequence according to the recalled image feature group sequence and the recalled text feature group sequence; Perform cross-coding on each recalled feature group pair in the recalled feature group pair sequence to obtain an encoded recalled feature sequence; Perform feature matching on each encoded recalled feature in the encoded recalled feature sequence by using the encoded user feature to obtain a matching degree sequence; Reorder the primary recalled image sequence according to the matching degree sequence to obtain a standard recalled image sequence.

7. An enhanced device for image retrieval based on bimodal matching, characterized in that, It includes: A data acquisition module for acquiring a matching image set and a matching text set corresponding to the matching image set; The sharding encoding module is used to perform sharding image encoding on the set of matching images by using a preset language-image model to obtain a set of sharded image feature groups, and perform sharding text encoding on the set of matching texts according to the set of sharded image feature groups to obtain a set of sharded text feature groups; wherein, the performing sharding image encoding on the set of matching images by using a preset language-image model to obtain a set of sharded image feature groups includes: sequentially selecting the matching images in the set of matching images as target matching images, performing image segmentation on the target matching images to obtain a set of segmented matching patches; performing linear projection on the set of segmented matching patches to obtain a set of segmented image features; performing local attention calculation on the set of segmented image features to obtain a set of image local features; performing global attention calculation on the set of image local features to obtain a set of image global features; performing multi-scale feature extraction and fully connected operations on the set of image global features by using a preset language-image model to obtain a set of standard image features; aggregating the sets of standard image features of all target matching images in the set of matching images into a set of standard image feature sets; performing clustering storage on each of the standard image feature sets in the set of standard image feature sets to obtain a set of sharded image feature groups; The contrast training module is used to perform contrast training on the language-image model by using the set of sharded image feature groups and the set of sharded text feature groups to obtain a bimodal matching model; The retrieval input module is used to obtain the user's historical images and historical texts to be queried; The preliminary matching module is used to perform preliminary image matching on the set of matching images by using the bimodal matching model according to the user's historical images and historical texts to obtain a primary recalled image sequence; The precise sorting module is used to perform cross-coding sorting on the primary recalled image sequence according to the user's historical images and historical texts to obtain a standard recalled image sequence.

8. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the image retrieval enhancement method based on bimodal matching according to any one of claims 1 to 6.

9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the image retrieval enhancement method based on bimodal matching according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Remote sensing image retrieval method and device, electronic equipment and computer storage medium

    CN117972126A

  • Transformer with multi-scale multi-context attentions

    WO2024263633A1