Image-text retrieval method and system based on multi-modal fusion and depth spectral clustering
By employing multimodal fusion and deep spectral clustering, the problems of multimodal feature misalignment and insufficient generalization ability in image-text retrieval are solved, achieving efficient cross-modal retrieval and rapid adaptation to new sample image-text matching.
Patent Information
- Application Number
- CN202511637451.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-10
- Publication Date
- 2026-01-09
AI Technical Summary
Existing image and text retrieval methods struggle to capture fine-grained common semantics between modalities in multimodal data, exhibiting issues such as multimodal feature misalignment and poor generalization ability, especially lacking rapid matching capabilities in unknown scenarios.
This paper employs a multimodal fusion and deep spectral clustering approach. By extracting image and text features through an encoder, constructing global and local features, aligning them using a multimodal contrastive learning model, and performing cluster analysis using a deep spectral clustering algorithm, a cross-modal semantic center prototype set is constructed to achieve cross-modal retrieval of images and text.
It improves the representation ability and robustness of multimodal data, enhances the semantic alignment of image and text modalities, enables cross-modal retrieval that can quickly adapt to new sample data, and improves retrieval accuracy and efficiency.
Smart Images

Figure CN121301595A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data retrieval technology, and in particular to a text and image retrieval method and system based on multimodal fusion and deep spectral clustering. Background Technology
[0002] With the rapid development of the internet and multimedia technologies, multimodal data, including images and text, is experiencing explosive growth. On social media platforms, users upload hundreds of millions of images and texts daily; enterprise databases have also accumulated massive amounts of product images and documentation. How to quickly and accurately retrieve the required image and text information from such a vast amount of data has become a key technical challenge in the field of information processing. Image and text retrieval technology aims to break down the modal barriers between images and text, achieving cross-modal information retrieval and matching, and providing users with a more efficient and intelligent way to obtain information. This technology is widely used in various fields such as intelligent security, e-commerce, and digital libraries, and is of great significance in improving information processing efficiency and user experience.
[0003] Currently, the field of image and text retrieval has accumulated rich research results and technical methods. Traditional image and text retrieval methods are mainly based on manually designed features, such as using the bag-of-words model to process text, extracting image features using scale-invariant feature transformation and accelerated robust features, and then performing retrieval matching by calculating the similarity between features. However, the extraction of manually designed features depends on specific design rules and prior knowledge, making it difficult to accurately capture the essential features of the data when faced with complex and ever-changing image and text data, resulting in limited retrieval accuracy and generalization ability. With the rise of deep learning technology, deep learning-based image and text retrieval methods have gradually become mainstream. These methods automatically learn feature representations from large-scale image and text data by constructing deep neural networks, effectively improving the performance of image and text retrieval. According to different network structures and implementation methods, deep learning image and text retrieval methods can be divided into methods based on single-modal feature extraction and fusion, and methods based on cross-modal mapping. The former extracts deep features from images and text separately, and then fuses features through concatenation, weighted summation, etc.; the latter aims to construct a unified cross-modal space, mapping images and text into this space, and achieving cross-modal retrieval by calculating the distance in the space.
[0004] However, existing image and text retrieval methods based on deep learning technology have the following drawbacks: 1. Technical challenges in learning multi-granularity features for multimodal data. Existing multimodal data feature extraction methods generally only consider global features, ignoring the complementary relationships between different modal features and local fine-grained features, making it difficult for the model to capture the fine-grained common semantics between modalities.
[0005] 2. Technical challenges in feature comparison and alignment to address the semantic gap in multimodal data. Different modalities express the same semantic meaning significantly differently, and the same semantic meaning is expressed differently in images and text. The ambiguity of natural language and the polysemy of image content significantly increase the difficulty of accurate multimodal semantic alignment, making it difficult for models to accurately capture the semantic correspondences between different modalities.
[0006] 3. Technical challenges in rapidly matching multimodal data in unknown scenarios. In practical applications, obtaining sufficient and effective samples is extremely difficult, making it hard to cover the diverse distributions of samples outside the training set. Traditional matching models have limited generalization ability, making it difficult to quickly adapt to new situations, and retraining the model is time-consuming. Summary of the Invention
[0007] This application provides a text and image retrieval method and system based on multimodal fusion and deep spectral clustering to solve the problems of insufficient consideration of local features, misalignment of multimodal features, and poor generalization ability in the prior art.
[0008] On the one hand, embodiments of this application provide a text and image retrieval method based on multimodal fusion and deep spectral clustering, including: Obtain training samples, which include image training samples and text training samples; The encoder is used to extract image training features and text training features from image training samples and text training samples, respectively. The global features are obtained by concatenating the image training features and the text training features. Construct a global similarity matrix based on global features; Local image training features and local text training features are constructed using image training features, text training features, and a global similarity matrix. A multimodal contrastive learning model is used to align local image training features and local text training features. In the alignment process, the multimodal contrastive learning model is first trained using positive and negative sample pairs, and then alignment is performed based on the cosine similarity between the local image training features and the local text training features. The global features, aligned local image training features, and local text training features are fused to obtain the fused features; The fused features are normalized to map them into a unified shared semantic space, resulting in normalized features. A deep spectral clustering algorithm was used to perform cluster analysis on the normalized features, resulting in a cross-modal semantic center prototype set composed of multiple cluster centers; Obtain query samples, which can be either image or text query samples; The encoder is used to extract query features from the query samples, and the query features are mapped to a shared semantic space. The cosine similarity between the query features and each cluster center in the cross-modal semantic center prototype set is calculated, and the retrieval results matching the query features are determined based on the magnitude of the cosine similarity.
[0009] On the other hand, embodiments of this application also provide a text and image retrieval system based on multimodal fusion and deep spectral clustering, including: The training sample acquisition module is used to acquire training samples, which include image training samples and text training samples. The feature extraction module is used to extract image training features and text training features from image training samples and text training samples respectively using the encoder; The feature concatenation module is used to concatenate image training features and text training features to obtain global features; The matrix construction module is used to construct a global similarity matrix based on global features; The local feature construction module is used to construct local image training features and local text training features using image training features, text training features, and the global similarity matrix. The feature alignment module is used to align local image training features and local text training features using a multimodal contrastive learning model. In the alignment process, the multimodal contrastive learning model is first trained using positive and negative sample pairs, and then alignment is performed based on the cosine similarity between the local image training features and the local text training features. The feature fusion module is used to fuse global features, aligned local image training features, and local text training features to obtain fused features; The feature mapping module is used to normalize the fused features so as to map the fused features to a unified shared semantic space to obtain normalized features. The clustering analysis module is used to perform clustering analysis on normalized features using a deep spectral clustering algorithm to obtain a cross-modal semantic center prototype set composed of multiple cluster centers; The query sample acquisition module is used to acquire query samples, which can be image query samples or text query samples; The image and text retrieval module is used to extract query features from query samples using an encoder, map the query features to a shared semantic space, calculate the cosine similarity between the query features and each cluster center in the cross-modal semantic center prototype set, and determine the retrieval results that match the query features based on the magnitude of the cosine similarity.
[0010] On the other hand, embodiments of this application also provide a computer storage medium storing a plurality of computer instructions for causing a computer to execute the above-described method.
[0011] The image and text retrieval method and system based on multimodal fusion and deep spectral clustering in this application have the following advantages: 1. By employing data-level fusion and local structural constraints, global features from multimodal data and local features from unimodal data are extracted, enabling the model to autonomously learn the optimal local-global fusion features and capture sample association features at different granularities. Simultaneously, it can better uncover semantic relationships within and between modalities, improving the model's representation ability and robustness for multimodal data.
[0012] 2. By constructing positive and negative sample pairs between modal data, different modal data are mapped to a shared semantic space, achieving self-supervised alignment of image and text modal data. Simultaneously, it can establish semantic correspondences between image regions and text words, further enhancing the alignment of image and text features and enabling cross-modal retrieval.
[0013] 3. By learning the semantic center prototype of image and text data through multimodal deep spectral clustering, given a query for an image or text, the nearest cluster center is found in the embedded semantic space, and images or text belonging to the cluster center are further retrieved. This application accelerates the image and text retrieval process by reducing the low-dimensional sample space to a high-dimensional cluster space through cluster centers. At the same time, the cluster centers can learn high-level semantic relationships and quickly adapt to new sample data. Attached Figure Description
[0014] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0015] Figure 1 A flowchart illustrating a text and image retrieval method based on multimodal fusion and deep spectral clustering, provided for embodiments of this application. Detailed Implementation
[0016] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0017] Figure 1A flowchart illustrating an image and text retrieval method based on multimodal fusion and deep spectral clustering, provided in an embodiment of this application. This embodiment of the application provides an image and text retrieval method based on multimodal fusion and deep spectral clustering, including: S100, Obtain training samples, which include image training samples and text training samples.
[0018] S101, the encoder is used to extract image training features and text training features from image training samples and text training samples respectively.
[0019] For example, spatial transformation parameters are learned through a pre-constructed lightweight structured autoencoder, which then extracts image training features. and text training features :
[0020]
[0021] in, These are image training samples. These are text training samples. and These are encoder parameters. To improve the accuracy of the embedded features, this application further constructs a feature decoder. and Used to reconstruct multimodal data, accordingly, and These are the decoder parameters.
[0022] S102, concatenate the image training features and text training features to obtain the global features.
[0023] For example, global features are represented as .
[0024] S103, construct a global similarity matrix based on global features.
[0025] For example, the global similarity matrix is represented as:
[0026] in, k This is the number of neighboring samples selected, typically set to 3. For the first The image region and the first Semantic similarity between words in a text a ij The global similarity matrix is formed. Indicates the first i Global features Z iand the j Global features Z j The distance between them.
[0027] The calculation expression is:
[0028] in, It represents the 2-norm.
[0029] S104 utilizes image training features, text training features, and the global similarity matrix to construct local image training features and local text training features.
[0030] For example, local image training features are constructed based on image training features and a global similarity matrix using a shared nonlinear network, and local text training features are constructed based on text training features and a global similarity matrix using a nonlinear network.
[0031] Specifically, nonlinear networks are used This indicates that the network can employ a three-layer MLP (Multilayer Perceptron) network, thus training features from local images. It can be represented as Local text training features It can be represented as Furthermore, a loss function of the following form is used for the nonlinear network. Train it to learn good feature representations from both image and text data simultaneously. and :
[0032]
[0033] in, N For the sample size, and These are the image features and text features in the positive sample pair, respectively. This is a hyperparameter used to balance the loss term.
[0034] S105. A multimodal contrastive learning model is used to align local image training features and local text training features. In the alignment process, the multimodal contrastive learning model is first trained using positive and negative sample pairs, and then the alignment is performed based on the cosine similarity between the local image training features and the local text training features.
[0035] For example, a positive sample pair representing different modalities from the same instance (such as image features and their corresponding textual description features) can be represented as follows: Negative sample pairs, on the other hand, come from modal data from different instances and can be represented as... By constructing positive and negative sample pairs in this way, positive and negative samples can be automatically generated using the transformation of the data itself, thereby improving the model's generalization ability under unsupervised or weakly supervised conditions.
[0036] The cosine similarity between local image training features and local text training features is expressed as:
[0037] in, Training features for local images and local text training features Cosine similarity between them This is a temperature coefficient used to adjust the smoothness of the similarity distribution.
[0038] in, Training features for local images and local text training features Cosine similarity between them This is a temperature coefficient used to adjust the smoothness of the similarity distribution.
[0039] Furthermore, to enable the model to effectively distinguish between positive and negative samples and learn local feature representations with discriminative ability regarding modal semantic information, this embodiment employs the InfoNCE loss function during the training of the multimodal contrastive learning model to maximize the mutual information between image and text modalities. This loss function achieves effective alignment of cross-modal features by bringing positive sample pairs closer together and distancing negative sample pairs further apart in the semantic space. It is illustrated below:
[0040] in, For the InfoNCE loss function, and These are the image features and text features in the positive sample pair, respectively. Represents cosine similarity. This refers to all text features within a batch, including both positive and negative samples. This is a temperature parameter used to control the degree of sharpening of classification boundaries.
[0041] In S104, through a nonlinear network Image and text data are mapped to a unified embedding space. Furthermore, the multimodal contrastive learning model brings cross-modal samples with the same semantics (i.e., positive sample pairs) closer together in the unified embedding space (i.e., cosine similarity), while pushing away irrelevant samples (i.e., negative sample pairs), thereby achieving semantic alignment of cross-modal data.
[0042] The multimodal contrastive learning model described above can learn more discriminative cross-modal semantic representations in the absence of manual annotation, thereby significantly improving the semantic matching and retrieval performance in multimodal tasks.
[0043] S106, the global features, aligned local image training features, and local text training features are fused to obtain the fused features.
[0044] For example, in cross-modal global features Z With local features y Building upon this foundation, an adaptive weighting mechanism-based feature fusion method further enhances the image-text retrieval model's ability to represent multimodal data. This application's embodiments introduce a shared nonlinear mapping function implemented by a single-layer MLP network. This method maps local features of images and text into a shared semantic space.
[0045] Furthermore, the fusion feature is represented as:
[0046] in, As a feature of fusion, These are learnable adaptive weight coefficients used to control the fusion ratio of global and local features. Z As a global feature, This represents a pooling operation to obtain a vector of fixed dimensions. This represents local image training features or local text training features. ,when Pick I hour Training features for local images ,when Pick T hour Training features for local text , It is a non-linear mapping function.
[0047] Furthermore, to ensure that the fusion process can capture fine-grained semantic relationships within and between modalities, the fused features are optimized using a metric consistency constraint loss function after acquisition, so that the fused features maintain structural consistency in the embedding space. The metric consistency constraint loss function is expressed as:
[0048] in, To measure the consistency constraint loss function, This represents the kernel similarity matrix calculated from global features. This represents the kernel similarity matrix calculated from local features, including local image training features and local text training features. Let Frobenius norm be represented. If... and Unified Indicate, then , For hyperparameters, and They represent the first i and j A fusion feature.
[0049] After obtaining optimized fusion features Then, it can be used Construct a shared semantic space across modalities.
[0050] S107, normalize the fused features to map them into a unified shared semantic space to obtain normalized features.
[0051] For example, the first i Normalized features Represented as:
[0052] Note that before performing normalization, the similarity matrix between modes is used first. As the initial input, this matrix reflects the global semantic correlation between the image and text. Then, local features of the image and text are learned through S104 and fused. Based on this, the fused features of the image and text are analyzed. Normalization is performed to map it to a unified shared semantic space.
[0053] S108 uses a deep spectral clustering algorithm to perform cluster analysis on the normalized features, resulting in a cross-modal semantic center prototype set composed of multiple cluster centers.
[0054] For example, after completing the extraction of local features and semantic alignment of images and text, in order to further improve the efficiency and generalization ability of image and text retrieval, this application introduces a multimodal semantic center learning mechanism based on deep spectral clustering to achieve fast matching and efficient retrieval of cross-modal data.
[0055] After completing the fusion features After normalization, a deep spectral clustering algorithm is used to perform cluster analysis on the normalized fusion features in order to construct a cross-modal semantic center prototype set. ,in M The preset number of clusters, Then it is the first m Cluster centers. This process involves constructing a graph model and calculating the graph Laplacian matrix. L This allows for the acquisition of more discriminative clustering structures within a low-dimensional embedding space. Specifically, embodiments of this application utilize a global similarity matrix... A ij In the construction process of the graph weight matrix, cross-modal consistency of clustering results is enhanced.
[0056] S109, Obtain the query sample, which can be an image query sample or a text query sample.
[0057] S110: The encoder is used to extract the query features of the query samples, the query features are mapped to the shared semantic space, the cosine similarity between the query features and each cluster center in the cross-modal semantic center prototype set is calculated, and the retrieval result matching the query features is determined based on the magnitude of the cosine similarity.
[0058] For example, when mapping query features to a shared semantic space, the query features need to be used as input, and steps S102-107 need to be repeated.
[0059] Furthermore, in order to improve the relevance and robustness of the retrieval results, this application introduces a hierarchical filtering retrieval strategy: the first layer of filtering selects candidate cluster centers based on the similarity threshold between the cluster centers and the query features; the second layer of filtering performs fine-grained matching of the target modality data within the candidate cluster centers to further identify the target object that best matches the query features semantically.
[0060] Specifically, the retrieval results matching the query features are determined based on the magnitude of cosine similarity, including: Clusters whose cosine similarity to each cluster center in the query features and cross-modal semantic center prototype set is greater than a similarity threshold are selected as candidate cluster centers. The candidate cluster centers are matched with the query features in a fine-grained manner, and the candidate cluster center with the best modality match is taken as the category to which the search result belongs.
[0061] In the first step, candidate cluster centers It can be represented as:
[0062] in, Represents cosine similarity. To map query features to results in a shared semantic space, This is the similarity threshold.
[0063] This application also provides an image and text retrieval system based on multimodal fusion and deep spectral clustering, including: The training sample acquisition module is used to acquire training samples, which include image training samples and text training samples. The feature extraction module is used to extract image training features and text training features from image training samples and text training samples respectively using the encoder; The feature concatenation module is used to concatenate image training features and text training features to obtain global features; The matrix construction module is used to construct a global similarity matrix based on global features; The local feature construction module is used to construct local image training features and local text training features using image training features, text training features, and the global similarity matrix. The feature alignment module is used to align local image training features and local text training features using a multimodal contrastive learning model. In the alignment process, the multimodal contrastive learning model is first trained using positive and negative sample pairs, and then alignment is performed based on the cosine similarity between the local image training features and the local text training features. The feature fusion module is used to fuse global features, aligned local image training features, and local text training features to obtain fused features; The feature mapping module is used to normalize the fused features so as to map the fused features to a unified shared semantic space to obtain normalized features. The clustering analysis module is used to perform clustering analysis on normalized features using a deep spectral clustering algorithm to obtain a cross-modal semantic center prototype set composed of multiple cluster centers; The query sample acquisition module is used to acquire query samples, which can be image query samples or text query samples; The image and text retrieval module is used to extract query features from query samples using an encoder, map the query features to a shared semantic space, calculate the cosine similarity between the query features and each cluster center in the cross-modal semantic center prototype set, and determine the retrieval results that match the query features based on the magnitude of the cosine similarity.
[0064] This application also provides a computer storage medium storing a plurality of computer instructions for causing a computer to execute the above-described method.
[0065] Although preferred embodiments of this application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this application.
[0066] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.
Claims
1. A text and image retrieval method based on multimodal fusion and deep spectral clustering, characterized in that, include: Obtain training samples, which include image training samples and text training samples; The image training features and text training features in the image training samples and the text training samples are extracted respectively using the encoder; The global features are obtained by concatenating the image training features and the text training features. Construct a global similarity matrix based on the global features; Local image training features and local text training features are constructed using the image training features, the text training features, and the global similarity matrix; A multimodal contrastive learning model is used to align the local image training features and the local text training features. In the alignment process, the multimodal contrastive learning model is first trained using positive and negative sample pairs, and then the alignment is performed based on the cosine similarity between the local image training features and the local text training features. The global features, the aligned local image training features, and the local text training features are fused to obtain the fused features; The fused features are normalized to map them into a unified shared semantic space, resulting in normalized features. The normalized features are clustered using a deep spectral clustering algorithm to obtain a cross-modal semantic center prototype set composed of multiple cluster centers; Obtain a query sample, which may be an image query sample or a text query sample; The encoder is used to extract the query features of the query sample, the query features are mapped to the shared semantic space, the cosine similarity between the query features and each cluster center in the cross-modal semantic center prototype set is calculated, and the retrieval result matching the query features is determined based on the magnitude of the cosine similarity.
2. The image and text retrieval method based on multimodal fusion and deep spectral clustering according to claim 1, characterized in that, The global similarity matrix is represented as follows: in, k The number of neighboring samples selected. For the first The image region and the first Semantic similarity between words in a text a ij The global similarity matrix is formed. Indicates the first i The global features Z i and the j The global features Z j The distance between them.
3. The image and text retrieval method based on multimodal fusion and deep spectral clustering according to claim 1, characterized in that, The local image training features are constructed based on the image training features and the global similarity matrix using a shared nonlinear network, and the local text training features are constructed based on the text training features and the global similarity matrix using the same nonlinear network.
4. The image and text retrieval method based on multimodal fusion and deep spectral clustering according to claim 1, characterized in that, The cosine similarity between the local image training features and the local text training features is expressed as follows: in, Training features for the local image and the local text training features Cosine similarity between them It is a 2-norm. This is the temperature coefficient.
5. The image and text retrieval method based on multimodal fusion and deep spectral clustering according to claim 4, characterized in that, The InfoNCE loss function is used when training the multimodal contrastive learning model, as follows: in, For the InfoNCE loss function, and These are the image features and text features in the positive sample pair, respectively. Represents cosine similarity. N For the sample size, This refers to all text features within a batch, including both positive and negative samples. This is a temperature parameter used to control the degree of sharpening of classification boundaries.
6. The image and text retrieval method based on multimodal fusion and deep spectral clustering according to claim 1, characterized in that, The fusion feature is represented as follows: in, For the fusion feature, For learnable adaptive weight coefficients, Z For the global feature, This indicates a pooling operation. This refers to the local image training features or the local text training features. ,when Pick I hour Training features for the local image ,when Pick T hour Training features for the local text , It is a non-linear mapping function.
7. The image and text retrieval method based on multimodal fusion and deep spectral clustering according to claim 1, characterized in that, After obtaining the fused features, the fused features are optimized using a metric consistency constraint loss function, which is expressed as follows: in, Let the metric consistency constraint loss function be... This represents the kernel similarity matrix calculated from the global features. This represents the kernel similarity matrix calculated from local features, including the local image training features and the local text training features. This represents the Frobenius norm.
8. The image and text retrieval method based on multimodal fusion and deep spectral clustering according to claim 1, characterized in that, Determining the retrieval results that match the query features based on the magnitude of cosine similarity includes: Clusters whose cosine similarity to each cluster center in the cross-modal semantic center prototype set is greater than a similarity threshold are selected as candidate cluster centers. The candidate cluster centers are matched with the query features in a fine-grained manner, and the candidate cluster center with the best modality match is taken as the category to which the search result belongs.
9. A system applying the image and text retrieval method based on multimodal fusion and deep spectral clustering as described in any one of claims 1-8, characterized in that, include: The training sample acquisition module is used to acquire training samples, which include image training samples and text training samples; The feature extraction module is used to extract image training features and text training features from the image training samples and text training samples respectively using the encoder; The feature concatenation module is used to concatenate the image training features and the text training features to obtain global features; The matrix construction module is used to construct a global similarity matrix based on the global features; The local feature construction module is used to construct local image training features and local text training features using the image training features, the text training features, and the global similarity matrix. The feature alignment module is used to align the local image training features and the local text training features using a multimodal contrastive learning model. In the alignment process, the multimodal contrastive learning model is first trained using positive sample pairs and negative sample pairs, and then the alignment is performed based on the cosine similarity between the local image training features and the local text training features. The feature fusion module is used to fuse the global features, the aligned local image training features, and the local text training features to obtain fused features; The feature mapping module is used to normalize the fused features so as to map the fused features to a unified shared semantic space to obtain normalized features. The clustering analysis module is used to perform clustering analysis on the normalized features using a deep spectral clustering algorithm to obtain a cross-modal semantic center prototype set composed of multiple cluster centers; The query sample acquisition module is used to acquire query samples, which are either image query samples or text query samples; The image and text retrieval module is used to extract query features of the query sample using the encoder, map the query features to the shared semantic space, calculate the cosine similarity between the query features and each cluster center in the cross-modal semantic center prototype set, and determine the retrieval result matching the query features based on the magnitude of the cosine similarity.
10. A computer storage medium, characterized in that, The computer storage medium stores a plurality of computer instructions, which are used to cause the computer to perform the method described in any one of claims 1-8.
Citation Information
Cited By
Training method of retrieval model, retrieval method and electronic equipment
CN122198022A