Cross-modal-based medicine logistics retrieval method and system, terminal and storage medium
By employing a cross-modal drug logistics retrieval method, which utilizes image and text feature fusion, graph neural networks, and hash loss optimization, the problems of semantic mismatch and modal imbalance in drug logistics scenarios are solved, achieving high-precision retrieval results.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GUANGDONG UNIV OF TECH
- Filing Date
- 2026-02-27
- Publication Date
- 2026-06-02
AI Technical Summary
Existing technologies for cross-modal data retrieval in pharmaceutical and food logistics scenarios suffer from semantic mismatch, modal imbalance, and difficulty in feature fusion, resulting in low retrieval accuracy.
A cross-modal drug logistics retrieval method is adopted. By fusing image modal feature representation and text modal feature representation, a label semantic graph is constructed and encoded using a graph neural network. A cross-modal attention mechanism is introduced to construct a hash code and optimize the feature extraction network. The method combines semantic neighborhood perception contrastive hashing and label distribution perception semantic alignment loss to improve retrieval accuracy.
It significantly improved the semantic alignment of multimodal features and the generalization ability of the model, thereby enhancing the accuracy of drug logistics retrieval.
Smart Images

Figure CN122132505A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of logistics management technology, and in particular to a cross-modal pharmaceutical logistics retrieval method, system, terminal, and computer-readable storage medium. Background Technology
[0002] With the rapid development of the e-commerce logistics industry, a large amount of multimodal data (such as images and text) has accumulated. Cross-modal hash retrieval significantly reduces storage and computing overhead and improves retrieval efficiency by mapping high-dimensional features to low-dimensional binary hash codes, and its importance is becoming increasingly prominent.
[0003] However, in the pharmaceutical and food logistics scenario, the data sources are extensive, involving warehouse management systems, warehouse control systems and various supplier systems, and are characterized by cross-industry, cross-database and cross-data domain, which makes it difficult for models trained in a single data domain to maintain stable performance in other domains.
[0004] Therefore, existing technologies still need to be improved and developed. Summary of the Invention
[0005] The main objective of this invention is to provide a cross-modal pharmaceutical logistics retrieval method, system, terminal, and computer-readable storage medium, aiming to solve the problems in the prior art, such as the lack of cross-modal logistics retrieval methods, semantic mismatch, modal imbalance, and difficulty in feature fusion leading to low logistics retrieval accuracy.
[0006] To achieve the above objectives, the present invention provides a cross-modal drug logistics retrieval method, which includes the following steps: Acquire multiple drug data sets, construct a logistics dataset using all the drug data sets, extract features from the logistics dataset to obtain multiple modal feature representations; A label semantic graph is constructed based on all semantic labels in the logistics dataset, and the label semantic graph is encoded using a graph neural network to obtain multiple semantic embedding representations. Attention dynamic weights are added to each semantic embedding to semantically enhance each modal feature representation and obtain the corresponding modal features. Construct a hash code for each modal feature, construct a cross-modal semantic adjacency matrix based on each modal feature, and construct multiple semantic neighborhood-aware comparative hash losses for each modal feature across different modalities based on the cross-modal semantic adjacency matrix and each hash code; A weighted label representation is constructed based on all the semantic labels, multiple similarity matrices are constructed based on the weighted label representations, and a label distribution-aware semantic alignment loss function is constructed based on each similarity matrix and each hash code; A cosine triplet loss is constructed, and the constructed perceptual visual feature extraction network is optimized based on all the contrastive hash loss, the semantic alignment loss function, and the cosine triplet loss. Multiple test drug data are input into the optimized target feature extraction network, and the corresponding retrieval results are output.
[0007] Optionally, in the cross-modal drug logistics retrieval method, the modal feature representation includes: image modal feature representation and text modal feature representation; The process of acquiring multiple drug data sets, constructing a logistics dataset using all the drug data sets, and extracting features from the logistics dataset to obtain multiple modal feature representations specifically includes: A feature fusion mechanism and a regression classifier are introduced into the encoder to obtain a feature extraction network, and a cross-modal attention mechanism is added to the feature extraction network to obtain a perceptual visual feature extraction network. Multiple drug data points are acquired, and each drug data point is input into the perceptual visual feature extraction network to obtain image data and text data corresponding to each drug data point. Semantic labels are constructed based on each image data point and each text data point, and a logistics dataset is constructed based on all the image data points, all the text data points, and all the semantic labels. ; ; ; in, D Represents a logistics dataset. N Indicates the quantity of logistics data. Represents the first in the logistics dataset i One logistics data point, Indicates the first i Image data, Indicates the first i A text data, Indicates the first i Semantic labels for image-text pairs, C Indicates the number of semantic tag categories. , and They represent the first i The first, second, and third image-text pairs C Semantic labels for each category; For each of the drug data sets, the perceptual visual feature extraction network extracts text information from the image data, and performs feature extraction on the text information and the image data respectively to obtain corresponding in-image text feature representations and image feature representations. The in-image text feature representations and image feature representations are then fused to obtain the image modal feature representation. : ; ; in, P Indicates the fusion weight. This represents the sigmoid activation function. and All of these represent learnable parameters. Representing image features, This represents the text feature representation within the image. Indicates to and Perform feature splicing. Indicates the use of fusion weights Modulation; For each of the drug data, the perceptual visual feature extraction network divides the text data into multiple sub-word units, maps each sub-word unit to a text embedding representation, and encodes the text embedding representation to obtain the text modal feature representation of the drug data. .
[0008] Optionally, the cross-modal drug logistics retrieval method, wherein constructing a tag semantic graph based on all semantic tags in the logistics dataset, encoding the tag semantic graph using a graph neural network to obtain multiple semantic embedding representations, and adding dynamic attention weights to each semantic embedding to semantically enhance each modal feature representation to obtain the corresponding modal features, specifically includes: Based on the prior semantic relationships between each semantic label, a label semantic graph is constructed using each semantic label as a node; The semantic graph of the labels is input into a graph neural network for encoding to obtain the semantic embedding representation of each semantic label: ; in, Indicates the first c Semantic embedding representation of a semantic tag Indicates the first c The initial feature matrix of semantic labels, The adjacency matrix representing semantic tags. GNN Represents a graph neural network; Construct corresponding key vectors and value vectors based on each semantic embedding representation, and construct query vectors using the modal feature representations: ; in, Q Represents the query vector. Indicates the first c A key vector of semantic labels, Indicates the first c A vector of semantic label values, Indicates the first i Image modal feature representation or text modal feature representation, , and All represent learnable parameters; Assign dynamic attention weights to the current semantic labels: ; in, Indicates the first c Dynamic attention weights for each semantic label. Indicates transpose. Indicates the first i The data for the first drug includes the first c A semantic tag, C Indicates the number of semantic tag categories. Indicates the first i The data for the first drug includes the first j A semantic tag, Indicates the dimension of the key vector; Based on each dynamic attention weight, each value vector is weighted and aggregated to obtain the semantic context representation of each semantic label: ; in, Indicates the first i Each label represents a semantic context; Each of the label semantic context representations is added to the corresponding modality feature representation to obtain the semantically enhanced modality features: ; in, Indicates the first i Modal features.
[0009] Optionally, the cross-modal drug logistics retrieval method, wherein constructing a hash code for each modal feature, constructing a cross-modal semantic adjacency matrix based on each modal feature, and constructing multiple semantic neighborhood-aware contrastive hash losses for each modal feature across different modalities based on the cross-modal semantic adjacency matrix and each hash code, specifically includes: After performing a linear transformation on each modal feature, a nonlinear activation function is used to further transform it, yielding the corresponding hash code: ; in, This represents the hash code, and tanh represents the non-linear activation function. Represents the projection matrix. Representing modal features, Represents the bias vector; The current drug data is defined as anchor samples. For each modality of the anchor samples, a semantic domain of another modality is constructed. Based on the two semantic domains, a cross-modal semantic adjacency matrix is constructed. ; Based on each of the cross-modal semantic adjacency matrices, a semantic neighborhood-aware contrastive hash loss is constructed for each image modal feature representation to its corresponding text modal feature representation and for each text modal feature representation to its corresponding image modal feature representation. ; ; in, Indicates the first i Drug data from images x To text y Semantic neighborhood perception in a given direction contrasts with hash loss. Indicates the first j Drug data from text y To image x Semantic neighborhood perception in a given direction contrasts with hash loss. Indicates the first j The text modality feature representation and the first i Each image modal feature represents label information with a semantic similarity higher than a threshold in the drug data. n Represents a collection of tag information. It is the sigmoid activation function. Indicates the first i A hash code representing the modal features of an image. Indicates the first j The hash code representing each text modal feature, where T is the transpose. This represents the temperature parameter.
[0010] Optionally, the cross-modal drug logistics retrieval method, wherein constructing a weighted label representation based on all the semantic labels, constructing multiple similarity matrices based on the weighted label representations, and constructing a label distribution-aware semantic alignment loss function based on each similarity matrix and each hash code, specifically includes: A label matrix is constructed based on all the semantic labels, and the label matrix is added to a weighted strategy for modeling to obtain a weighted label representation. A similarity matrix is then constructed based on the weighted label representation. ; ; in, This indicates a weighted label representation. L Represents the label matrix, diag Represents the category-level weight matrix. Indicates the first i The data for the first drug includes the first c A semantic tag, N Indicates the quantity of logistics data. C Indicates the number of semantic tag categories. Represents the semantic similarity matrix of tags. Indicates transpose; Based on each similarity matrix and each hash code, construct inter-modal semantic alignment loss function and intra-modal semantic alignment loss function for each drug data, and construct a semantic alignment loss function based on the inter-modal semantic alignment loss function and the intra-modal semantic alignment loss function.
[0011] Optionally, the cross-modal drug logistics retrieval method, wherein constructing an inter-modal semantic alignment loss function and an intra-modal semantic alignment loss function for each drug data based on each similarity matrix and each hash code, and constructing a semantic alignment loss function based on the inter-modal semantic alignment loss function and the intra-modal semantic alignment loss function, specifically includes: For each of the drug data sets, based on each hash code, an intra-modal hash similarity matrix is constructed between the modal features of the drug data set and the modal features of other drug data sets. The difference between the intra-modal hash similarity matrix and the corresponding similarity matrix is calculated to obtain the intra-modal semantic alignment loss function. ; in, This represents the intramodal semantic alignment loss function. and These represent the first balance coefficient and the second balance coefficient, respectively. The intra-modal hash similarity matrix represents the dimension of the image modal features. The intra-modal hash similarity matrix represents the dimension of the text modal features; For each of the drug data, based on each hash code, a modal hash similarity matrix is constructed between the modal features of the drug data and the modal features of different types of other drug data. The difference between the modal hash similarity matrix and the corresponding similarity matrix is calculated to obtain the inter-modal semantic alignment loss function. ; ; in, This represents the inter-modal semantic alignment loss function. This represents the third balance coefficient. This represents the inter-modal hash similarity matrix between image modal feature representations and text modal feature representations. Indicates the first i The hash code representing the image modality feature and the first image modality feature. j The similarity matrix between hash codes represented by text modal features Indicates the first i The data for the first drug includes the first j A semantic tag, e The base of the natural logarithm; A semantic alignment loss function is constructed using the inter-modal semantic alignment loss function and the intra-modal semantic alignment loss function. : .
[0012] Optionally, the cross-modal drug logistics retrieval method, wherein constructing a cosine triplet loss, optimizing the constructed perceptual visual feature extraction network based on all the contrastive hash losses, the semantic alignment loss function, and the cosine triplet loss, and inputting multiple test drug data into the optimized target feature extraction network to output the corresponding retrieval results, specifically includes: For each type of modal feature of each of the drug data, determine the hash code of another modal feature that has a positive and negative correlation with the hash code of the modal feature, so as to construct a triplet of the modal feature: ; in, Indicates the first i A hash code representing the modal features of an image. Indicates and There is a positive correlation. j A hash code representing a text modality feature. Indicates and There is a negative correlation. k A hash code representing a text modality feature; Based on each of the triples, construct the cosine triplet loss for the image modality dimension and the text modality dimension respectively: ; ; in, The cosine triplet loss represents the modality dimension of the image. Indicates the first i A hash code representing the modal features of an image. Indicates and There is a positive correlation. j A hash code representing a text modality feature. Indicates and There is a negative correlation. k A hash code representing a text modality feature. The cosine triplet loss represents the text modality dimension. Indicates the first j A hash code representing a text modality feature. Indicates and There is a positive correlation. i A hash code representing a text modality feature. Indicates and There is a negative correlation. k A hash code representing a text modality feature. m Indicates the margin parameter; We construct the overall cosine triplet loss by utilizing the cosine triplet losses of the image modality dimension and the text modality dimension: ; in, This represents the overall cosine triplet loss; The perceptual visual feature extraction network is optimized using the overall cosine triplet loss, semantic alignment loss function, and all semantic neighborhood perceptual contrastive hash loss to obtain the target feature extraction network. Acquire test drug data, input all the test drug data into the target feature extraction network, and output the retrieval results corresponding to each test drug data.
[0013] Furthermore, to achieve the above objectives, the present invention also provides a cross-modal pharmaceutical logistics retrieval system, wherein the cross-modal pharmaceutical logistics retrieval system includes: The feature extraction module is used to acquire multiple drug data, construct a logistics dataset using all the drug data, and extract features from the logistics dataset to obtain multiple modal feature representations. The graph encoding module is used to construct a label semantic graph based on all semantic labels in the logistics dataset, and to encode the label semantic graph using a graph neural network to obtain multiple semantic embedding representations. Attention dynamic weights are added to each semantic embedding to semantically enhance each modal feature representation and obtain the corresponding modal features. The first loss calculation module is used to construct a hash code for each modal feature, construct a cross-modal semantic adjacency matrix based on each modal feature, and construct multiple semantic neighborhood-aware comparative hash losses for each modal feature across different modalities based on the cross-modal semantic adjacency matrix and each hash code. The second loss calculation module is used to construct a weighted label representation based on all the semantic labels, construct multiple similarity matrices based on the weighted label representations, and construct a label distribution-aware semantic alignment loss function based on each similarity matrix and each hash code; The model optimization module is used to construct a cosine triplet loss, optimize the constructed perceptual visual feature extraction network based on all the contrastive hash loss, the semantic alignment loss function and the cosine triplet loss, and input multiple test drug data into the optimized target feature extraction network to output the corresponding retrieval results.
[0014] Furthermore, to achieve the above objectives, the present invention also provides a terminal, wherein the terminal includes: a memory, a processor, and a cross-modal drug logistics retrieval program stored in the memory and executable on the processor, wherein when the cross-modal drug logistics retrieval program is executed by the processor, it implements the steps of the cross-modal drug logistics retrieval method described above.
[0015] Furthermore, to achieve the above objectives, the present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a cross-modal pharmaceutical logistics retrieval program, which, when executed by a processor, implements the steps of the cross-modal pharmaceutical logistics retrieval method described above.
[0016] In this invention, multiple drug data are acquired, and a logistics dataset is constructed using all the drug data. Feature extraction is performed on the logistics dataset to obtain multiple modal feature representations. A tag semantic graph is constructed based on all semantic labels in the logistics dataset, and a graph neural network is used to encode the tag semantic graph to obtain multiple semantic embedding representations. Dynamic attention weights are added to each semantic embedding to semantically enhance each modal feature representation, resulting in corresponding modal features. A hash code for each modal feature is constructed, and a cross-modal semantic adjacency matrix is constructed based on each modal feature. The present invention constructs multiple semantic neighborhood-aware contrastive hash losses for each modal feature across different modalities using the array and each hash code; constructs weighted label representations based on all semantic labels, constructs multiple similarity matrices based on the weighted label representations, and constructs a label distribution-aware semantic alignment loss function based on each similarity matrix and each hash code; constructs a cosine triplet loss, and optimizes the constructed perceptual visual feature extraction network based on all the contrastive hash losses, the semantic alignment loss function, and the cosine triplet loss; inputs multiple test drug data into the optimized target feature extraction network, and outputs the corresponding retrieval results. This invention can improve the semantic alignment of multimodal features, enhance model generalization ability, and significantly improve the retrieval accuracy of drug logistics. Attached Figure Description
[0017] Figure 1 This is a flowchart of a preferred embodiment of the cross-modal drug logistics retrieval method of the present invention; Figure 2 This is a flowchart of a preferred embodiment of the cross-modal drug logistics retrieval method of the present invention; Figure 3 This is a schematic diagram of the fusion mechanism of a preferred embodiment of the cross-modal drug logistics retrieval method of the present invention; Figure 4 This is a schematic diagram of the tag map of a preferred embodiment of the cross-modal drug logistics retrieval method of the present invention; Figure 5 This is a retrieval result diagram of a preferred embodiment of the cross-modal drug logistics retrieval method of the present invention; Figure 6 This is a structural diagram of a preferred embodiment of the cross-modal drug logistics retrieval system of the present invention; Figure 7 This is a structural diagram of a preferred embodiment of the terminal of the present invention. Detailed Implementation
[0018] To make the objectives, technical solutions, and advantages of this invention clearer and more explicit, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0019] In pharmaceutical and food logistics scenarios, data sources are wide-ranging, involving warehouse management systems, warehouse control systems, and various supplier systems. This data is characterized by its cross-industry, cross-database, and cross-data domain nature. This makes it difficult for models trained in a single data domain to maintain stable performance in other domains. Existing methods are also susceptible to problems such as differences in feature distribution between modalities, semantic misalignment, and imbalance between positive and negative samples, resulting in limited retrieval accuracy.
[0020] Therefore, this invention, based on the Transformer model (encoder), introduces OCR (Optical Character Recognition) and feature fusion mechanisms to extract visual representations of textual information within fused images. Furthermore, a cross-modal attention mechanism guided by label graphs is used to achieve dynamic alignment and semantic enhancement of multimodal features, improving the model's generalization ability. To address the issues of modality imbalance and semantic inconsistency, a semantic neighborhood-aware contrastive hashing and label distribution-aware semantic alignment loss are designed to mitigate the impact of sample skewness and the semantic gap. Simultaneously, a cosine triplet loss is introduced to enhance the discriminative power of the hash code, indirectly promoting cross-modal feature alignment.
[0021] The preferred embodiment of the drug logistics retrieval method based on cross-modal processing described in this invention, such as... Figure 1 As shown, the cross-modal drug logistics retrieval method includes the following steps: Step S10: Obtain multiple drug data, construct a logistics dataset using all the drug data, extract features from the logistics dataset, and obtain multiple modal feature representations.
[0022] The modal feature representation includes image modal feature representation and text modal feature representation. Based on the encoder disclosed in this invention, each collected drug data (test sample) is first decomposed into image and text, and features are extracted separately to obtain image modal feature representation and text modal feature representation. These two modal data are then used to construct image-text pairs. For each image-text pair, its semantic label is constructed, and a label matrix can be constructed based on all semantic labels.
[0023] Specifically, a feature fusion mechanism and a regression classifier are introduced into the encoder to obtain a feature extraction network, and a cross-modal attention mechanism is added to the feature extraction network to obtain a perceptual visual feature extraction network. Multiple drug data points are acquired, and each drug data point is input into the perceptual visual feature extraction network to obtain image data and text data corresponding to each drug data point. Semantic labels are constructed based on each image data point and each text data point, and a logistics dataset is constructed based on all the image data points, all the text data points, and all the semantic labels. ; ; ; in, D Represents a logistics dataset. N Indicates the quantity of logistics data. Represents the first in the logistics dataset i One logistics data point, Indicates the first i Image data, Indicates the first i A text data, Indicates the first i Semantic labels for image-text pairs, C Indicates the number of semantic tag categories. , and They represent the first i The first, second, and third image-text pairs C Semantic labels for each category; For each of the drug data sets, the perceptual visual feature extraction network extracts text information from the image data, and performs feature extraction on the text information and the image data respectively to obtain corresponding in-image text feature representations and image feature representations. The in-image text feature representations and image feature representations are then fused to obtain the image modal feature representation. : ; ; in, P Indicates the fusion weight. This represents the sigmoid activation function. and All of these represent learnable parameters. Representing image features, This represents the text feature representation within the image. Indicates to and Perform feature splicing. Indicates the use of fusion weights Modulation; For each of the drug data, the perceptual visual feature extraction network divides the text data into multiple sub-word units, maps each sub-word unit to a text embedding representation, and encodes the text embedding representation to obtain the text modal feature representation of the drug data. .
[0024] In the image feature learning network, ORC technology is first used to extract image data from drug data ( Extract text information from the image (which can be denoted as) Subsequently, the image data and the text information within the image are input into the encoder for feature encoding, which yields the image feature representation and the text feature representation within the image. Then, the two feature representations are jointly modeled using a gated feature fusion mechanism to obtain the fused image modal feature representation.
[0025] Furthermore, in the text feature learning network, the text data is first segmented into multiple sub-word units using a byte-encoding-based word segmentation strategy. Each sub-word unit is then mapped to a corresponding low-dimensional continuous vector representation through a learnable word embedding matrix, thus obtaining the text embedding representation. After encoding by an encoder, the corresponding text modal feature representation is obtained. At this point, a drug data set yields both image modal feature representation and text modal feature representation. In subsequent descriptions, the image modal feature representation and text modal feature representation will be uniformly referred to as... .
[0026] Specifically, such as Figure 2 As shown, the overall framework of the model for implementing the cross-modal drug logistics retrieval method disclosed in this invention is illustrated. The image feature extraction network of this framework extracts key information from the image data of drug data and encodes the text information into text semantic features. Subsequently, a gated fusion mechanism is used to adaptively fuse the visual features and text semantic features. The fusion process requires first performing splicing processing, and then using fusion weights to fuse the features. The fusion weights modulate the text semantic information and inject it into the visual features in a residual manner to obtain the fused image feature representation.
[0027] Among them, such as Figure 3 As shown, the gated fusion mechanism can dynamically adjust the injection level of OCR text information according to the text quality and semantic relevance of different drug images. While making full use of the semantics of drug packaging text, it can effectively suppress noise interference, thereby obtaining a more robust and discriminative image feature representation.
[0028] Step S20: Construct a label semantic graph based on all semantic labels in the logistics dataset, and encode the label semantic graph using a graph neural network to obtain multiple semantic embedding representations. Add dynamic attention weights to each semantic embedding to semantically enhance each modal feature representation and obtain the corresponding modal features.
[0029] In pharmaceutical logistics scenarios, data typically originates from different suppliers' systems. These data sources exhibit variations in distribution and labeling standards, limiting the model's cross-domain generalization ability. Therefore, this invention introduces a label graph-guided cross-modal attention mechanism. A structured semantic graph is constructed at the label level, and semantic relationships between labels are modeled using a Graph Neural Network (GNN) to obtain label embeddings with global structural information. Based on this, multi-label information from samples guides cross-modal attention computation, enabling the model to focus on the label context relevant to the current sample's semantics, thereby effectively mitigating the impact of cross-domain distribution differences and semantic noise.
[0030] Specifically, based on the prior semantic relationships between each semantic tag, a tag semantic graph is constructed using each semantic tag as a node; The semantic graph of the labels is input into a graph neural network for encoding to obtain the semantic embedding representation of each semantic label: ; in, Indicates the first c Semantic embedding representation of a semantic tag Indicates the first c The initial feature matrix of semantic labels, The adjacency matrix representing semantic tags. GNN Represents a graph neural network; Construct corresponding key vectors and value vectors based on each semantic embedding representation, and construct query vectors using the modal feature representations: ; in, Q Represents the query vector. Indicates the first c A key vector of semantic labels, Indicates the first c A vector of semantic label values, Indicates the first i Image modal feature representation or text modal feature representation, , and All represent learnable parameters; Assign dynamic attention weights to the current semantic labels: ; in, Indicates the first c Dynamic attention weights for each semantic label. Indicates transpose. Indicates the first i The data for the first drug includes the first c A semantic tag, C Indicates the number of semantic tag categories. Indicates the first i The data for the first drug includes the first j A semantic tag, Indicates the dimension of the key vector; Based on each dynamic attention weight, each value vector is weighted and aggregated to obtain the semantic context representation of each semantic label: ; in, Indicates the first i Each label represents a semantic context; Each of the label semantic context representations is added to the corresponding modality feature representation to obtain the semantically enhanced modality features: ; in, Indicates the first i Modal features.
[0031] Among them, such as Figure 4 As shown, the nodes of the label semantic graph are composed of semantic tags. Semantic tags can be represented as semantic entities such as drug name, shape, brand and use, while the prior semantic relationships between semantic tags (such as category affiliation, component association, functional similarity, etc.) constitute the edges of the label semantic graph.
[0032] Specifically, the number of nodes in the label semantic graph corresponds to the types of semantic labels. The feature dimensions of the initially generated graph's label nodes are consistent with the features of the generated samples, and its adjacency matrix is used to characterize the prior semantic relationships between label nodes. Based on node features and adjacency structure, a GNN is used to encode the label semantic graph, obtaining a structured semantic embedding representation for each label node.
[0033] The core of the graph-guided cross-modal attention mechanism lies in evaluating the relationship between the input modality (such as image and text) and the structured semantic nodes in the graph. By assigning dynamic attention weights to each modality instance, this mechanism can measure the importance of graph nodes to the multimodal joint representation, thereby efficiently capturing cross-modal semantic associations and domain knowledge. In the graph-guided cross-modal attention computation, modal features are used as queries, and the embedded labels of graph nodes are used as keys and values. To explicitly introduce label constraints, attention weights are assigned only to label nodes relevant to the current sample. This effectively suppresses the interference of irrelevant label nodes on attention computation, guiding the model to focus on the semantically consistent label subspace.
[0034] Furthermore, based on the aforementioned attention weights, the value vectors of the label nodes are weighted and aggregated to obtain the label semantic context representation most relevant to the current modal input. Subsequently, this label semantic context is injected into the original modal features to achieve semantic enhancement. Through the cross-modal attention mechanism guided by the label graph, the model can dynamically align and enhance multimodal features under the constraints of structured label semantics, making the image and text representations have stronger consistency and discriminability in the semantic space, providing reliable semantic support for subsequent hash learning and high-precision cross-modal retrieval.
[0035] Step S30: Construct a hash code for each modal feature, construct a cross-modal semantic adjacency matrix based on each modal feature, and construct multiple semantic neighborhood-aware comparative hash losses for each modal feature across different modalities based on the cross-modal semantic adjacency matrix and each hash code.
[0036] Since the pharmaceutical and food datasets suffer from modal imbalance and semantic misalignment, the method disclosed in this invention proposes semantic neighborhood-aware contrastive hashing and label distribution-aware semantic alignment loss to mitigate the impact of modal imbalance and semantic misalignment.
[0037] Specifically, after performing a linear transformation on each modal feature, a nonlinear activation function is used to further transform it to obtain the corresponding hash code: ; in, This represents the hash code, and tanh represents the non-linear activation function. Represents the projection matrix. Representing modal features, Represents the bias vector; The current drug data is defined as anchor samples. For each modality of the anchor samples, a semantic domain of another modality is constructed. Based on the two semantic domains, a cross-modal semantic adjacency matrix is constructed. ; Based on each of the cross-modal semantic adjacency matrices, a semantic neighborhood-aware contrastive hash loss is constructed for each image modal feature representation to its corresponding text modal feature representation and for each text modal feature representation to its corresponding image modal feature representation. ; ; in, Indicates the first i Drug data from images x To text y Semantic neighborhood perception in a given direction contrasts with hash loss. Indicates the first j Drug data from text y To image x Semantic neighborhood perception in a given direction contrasts with hash loss. Indicates the first j The text modality feature representation and the first i Each image modal feature represents label information with a semantic similarity higher than a threshold in the drug data. n Represents a collection of tag information. It is the sigmoid activation function. Indicates the first i A hash code representing the modal features of an image. Indicates the first j The hash code representing each text modal feature, where T is the transpose. This represents the temperature parameter.
[0038] In order to address the problem that significant differences exist in the image appearance, text description and annotation specifications of the same drug across different systems, leading to a high imbalance in the distribution of samples and semantic correspondences between modalities, the semantic neighborhood-aware contrastive hash loss proposed in this invention explicitly models the semantic adjacency relationship between cross-modal samples, extending cross-modal alignment from a single positive pair to a structured many-to-many semantic contrastive relationship.
[0039] Specifically, a drug data point is randomly defined from all drug data as an anchor sample. Within this anchor sample, a semantic domain of another modality is constructed in a random modality (image or text). This allows the model to simultaneously learn by comparing multiple visual samples and text descriptions under the same drug category, dosage form, or function. Based on this, the present invention introduces a cross-modal semantic adjacency matrix. This matrix transforms the binary determination of "whether it is a positive match" into a weighted comparative modeling process of multiple adjacent positive samples, thereby obtaining the semantic neighborhood-aware contrastive hash loss in the text-to-image direction and the semantic neighborhood-aware contrastive hash loss in the image-to-text direction.
[0040] Among them, the semantic neighborhood-aware contrastive hash loss uses a semantic adjacency matrix to weight multiple semantically related samples, enabling each text anchor point to establish a contrast relationship with multiple semantically consistent drug image samples at the same time, thereby significantly improving the positive sample coverage and semantic density, which is more in line with the real semantic distribution characteristics of "one-to-many and many-to-many" in the drug dataset.
[0041] Furthermore, for the total loss of the two semantic neighborhood-aware comparative hash losses, the semantic adjacency-guided comparative relationship reconstruction explicitly preserves the semantic consistency of drug samples in different modalities and data domains during the cross-modal hash learning process. This effectively alleviates the alignment degradation problem caused by modal imbalance, alignment sparsity, and description differences, thereby improving the robustness and discriminative ability of cross-modal retrieval in the pharmaceutical e-commerce logistics scenario.
[0042] Step S40: Construct a weighted label representation based on all the semantic labels, construct multiple similarity matrices based on the weighted label representations, and construct a label distribution-aware semantic alignment loss function based on each similarity matrix and each hash code.
[0043] Although semantic neighborhood-aware contrastive hash loss alleviates the local contrastive bias problem to some extent, cross-modal hash learning still needs to further characterize the global semantic relationship between samples; therefore, this invention introduces a weighting strategy based on global label distribution into the existing label matrix.
[0044] Specifically, a label matrix is constructed based on all the semantic labels, and the label matrix is added to a weighted strategy for modeling to obtain a weighted label representation. A similarity matrix is then constructed based on the weighted label representation. ; ; in, This indicates a weighted label representation. L Represents the label matrix, diag Represents the category-level weight matrix. Indicates the first i The data for the first drug includes the first c A semantic tag, N Indicates the quantity of logistics data. C Indicates the number of semantic tag categories. Represents the semantic similarity matrix of tags. Indicates transpose; Based on each similarity matrix and each hash code, construct inter-modal semantic alignment loss function and intra-modal semantic alignment loss function for each drug data, and construct a semantic alignment loss function based on the inter-modal semantic alignment loss function and the intra-modal semantic alignment loss function.
[0045] In this process, the label semantics are normalized and modeled to obtain a weighted label representation, and a label semantic similarity matrix is constructed based on this representation to characterize the similarity relationship of samples in the global label semantic space, effectively alleviating the bias caused by label imbalance in semantic modeling.
[0046] Furthermore, for each of the drug data, based on each hash code, an intra-modal hash similarity matrix is constructed between the modal features of the drug data and the modal features of other drug data. The difference between the intra-modal hash similarity matrix and the corresponding similarity matrix is calculated to obtain the intra-modal semantic alignment loss function. ; in, This represents the intramodal semantic alignment loss function. and These represent the first balance coefficient and the second balance coefficient, respectively. The intra-modal hash similarity matrix represents the dimension of the image modal features. The intra-modal hash similarity matrix represents the dimension of the text modal features; For each of the drug data, based on each hash code, a modal hash similarity matrix is constructed between the modal features of the drug data and the modal features of different types of other drug data. The difference between the modal hash similarity matrix and the corresponding similarity matrix is calculated to obtain the inter-modal semantic alignment loss function. ; ; in, This represents the inter-modal semantic alignment loss function. This represents the third balance coefficient. This represents the inter-modal hash similarity matrix between image modal feature representations and text modal feature representations. Indicates the first i The hash code representing the image modality feature and the first image modality feature. j The similarity matrix between hash codes represented by text modal features Indicates the first i The data for the first drug includes the first j A semantic tag, e The base of the natural logarithm; A semantic alignment loss function is constructed using the inter-modal semantic alignment loss function and the intra-modal semantic alignment loss function. : .
[0047] In the embodiments disclosed in this invention, joint alignment of hash similarity structure and label semantic structure at both intramodal and intermodal levels can effectively improve the semantic consistency of cross-modal hash codes. At the intramodal level, by minimizing the difference between the same modal hash similarity matrix and the weighted label similarity matrix, the semantic structure of samples in a single modal space is constrained to remain consistent. At the intermodal level, to further enhance the consistency of image and text representations in the semantic space, a cross-modal label distribution-aware semantic alignment loss is introduced to constrain cross-modal hash similarity and label semantic similarity to remain consistent.
[0048] Step S50: Construct a cosine triplet loss, optimize the constructed perceptual visual feature extraction network based on all the contrastive hash loss, the semantic alignment loss function and the cosine triplet loss, and input multiple test drug data into the optimized target feature extraction network to output the corresponding retrieval results.
[0049] To further optimize the relative semantic ranking relationship between samples, especially for the diverse image and text representations of the same drug in e-commerce pharmaceutical logistics scenarios due to differences in packaging form, specifications, or description granularity, this invention further introduces cosine triplet loss.
[0050] Specifically, for each type of modal feature of each of the drug data, the hash code of another modal feature that has a positive and negative correlation with the hash code of the modal feature is determined to construct a triple of the modal feature: ; in, Indicates the first i A hash code representing the modal features of an image. Indicates and There is a positive correlation. j A hash code representing a text modality feature. Indicates and There is a negative correlation. k A hash code representing a text modality feature; Based on each of the triples, construct the cosine triplet loss for the image modality dimension and the text modality dimension respectively: ; ; in, The cosine triplet loss represents the modality dimension of the image. Indicates the first i A hash code representing the modal features of an image. Indicates and There is a positive correlation. jA hash code representing a text modality feature. Indicates and There is a negative correlation. k A hash code representing a text modality feature. The cosine triplet loss represents the text modality dimension. Indicates the first j A hash code representing a text modality feature. Indicates and There is a positive correlation. i A hash code representing a text modality feature. Indicates and There is a negative correlation. k A hash code representing a text modality feature. m Indicates the margin parameter; We construct the overall cosine triplet loss by utilizing the cosine triplet losses of the image modality dimension and the text modality dimension: ; in, This represents the overall cosine triplet loss; The perceptual visual feature extraction network is optimized using the overall cosine triplet loss, semantic alignment loss function, and all semantic neighborhood perceptual contrastive hash loss to obtain the target feature extraction network. Acquire test drug data, input all the test drug data into the target feature extraction network, and output the retrieval results corresponding to each test drug data.
[0051] In the cross-modal retrieval task, for each image data hash code, a triplet is constructed based on the positive and negative correlation hash codes of the corresponding text data. Then, the triplet loss of the image data can be calculated based on the cosine similarity metric. The same similarity method is used for the text data to calculate the triplet loss of the text data. Based on the triplet losses of the two modalities, the overall cosine triplet loss of the drug data can be calculated. The overall cosine triplet loss enhances the aggregation of semantically consistent samples in the embedding space by constraining the relative similarity of cross-modal samples, and explicitly widens the distance between semantically irrelevant drug samples, thereby improving the discriminativeness and ranking stability of cross-modal retrieval.
[0052] Furthermore, after optimizing the perceptual visual feature extraction network based on the aforementioned multiple loss functions, the optimized target feature extraction network is used to perform logistics retrieval on the test drug data, resulting in the following: Figure 5The search results shown indicate that the semantic matching degree of the images and text is relatively high among the first four search results. This shows that the model of the present invention performs well in cross-modal feature alignment and semantic consistency, effectively bridging the feature gap between modalities, thereby improving the accuracy and reliability of the search.
[0053] The multimodal retrieval model proposed in this invention demonstrates significant advantages in drug image feature modeling, cross-data domain generalization ability, and mitigation of modality imbalance and semantic mismatch. By introducing a label graph-guided cross-modal attention mechanism into the Transformer framework for drug image features, the model effectively utilizes structured semantic information to enhance multimodal feature representation, thereby improving generalization performance in complex cross-domain scenarios. Simultaneously, by combining semantic neighborhood-aware contrastive hashing and label distribution-aware semantic alignment loss, the model explicitly maintains semantic consistency during cross-modal hashing learning, effectively addressing the common problems of uneven modality distribution and semantic alignment difficulties in drug datasets. This makes the target feature extraction network highly adaptable and valuable for intelligent picking and retrieval tasks of cross-domain multimodal e-commerce drug logistics objects.
[0054] Furthermore, such as Figure 6 As shown, based on the above-described cross-modal drug logistics retrieval method, this invention also provides a cross-modal drug logistics retrieval system, wherein the cross-modal drug logistics retrieval system includes: Feature extraction module 51 is used to acquire multiple drug data, construct a logistics dataset using all the drug data, extract features from the logistics dataset, and obtain multiple modal feature representations. The graph encoding module 52 is used to construct a label semantic graph based on all semantic labels in the logistics dataset, and to encode the label semantic graph using a graph neural network to obtain multiple semantic embedding representations. Attention dynamic weights are added to each semantic embedding to semantically enhance each modal feature representation and obtain the corresponding modal features. The first loss calculation module 53 is used to construct a hash code for each modal feature, construct a cross-modal semantic adjacency matrix based on each modal feature, and construct multiple semantic neighborhood-aware comparative hash losses for each modal feature in different modalities based on the cross-modal semantic adjacency matrix and each hash code. The second loss calculation module 54 is used to construct a weighted label representation based on all the semantic labels, construct multiple similarity matrices based on the weighted label representations, and construct a label distribution-aware semantic alignment loss function based on each similarity matrix and each hash code; The model optimization module 55 is used to construct a cosine triplet loss, optimize the constructed perceptual visual feature extraction network based on all the contrastive hash loss, the semantic alignment loss function and the cosine triplet loss, and input multiple test drug data into the optimized target feature extraction network to output the corresponding retrieval results.
[0055] Furthermore, such as Figure 7 As shown, based on the above-mentioned cross-modal drug logistics retrieval method and system, the present invention also provides a terminal, which includes a processor 10, a memory 20 and a display 30. Figure 7 Only some of the terminal components are shown; however, it should be understood that it is not required to implement all of the components shown, and more or fewer components may be implemented instead.
[0056] In some embodiments, the memory 20 may be an internal storage unit of the terminal, such as a hard disk or memory. In other embodiments, the memory 20 may be an external storage device of the terminal, such as a plug-in hard disk, smart media card (SMC), secure digital card (SD), flash card, etc., equipped on the terminal. Further, the memory 20 may include both internal and external storage devices. The memory 20 is used to store application software and various types of data installed on the terminal, such as the program code installed on the terminal. The memory 20 may also be used to temporarily store data that has been output or will be output. In one embodiment, the memory 20 stores a cross-modal drug logistics retrieval program 40, which can be executed by the processor 10 to implement the cross-modal drug logistics retrieval method of this application.
[0057] In some embodiments, the processor 10 may be a central processing unit (CPU), a microprocessor, or other data processing chip, used to run program code stored in the memory 20 or process data, such as executing the cross-modal drug logistics retrieval method.
[0058] In some embodiments, the display 30 may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, or an OLED (Organic Light-Emitting Diode) touchscreen. The display 30 is used to display information on the terminal and to display a visual user interface. The components of the terminal communicate with each other via a system bus.
[0059] In one embodiment, when the processor 10 executes the cross-modal drug logistics retrieval program 40 in the memory 20, it implements the steps of the cross-modal drug logistics retrieval method as described above.
[0060] The present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a cross-modal pharmaceutical logistics retrieval program, which, when executed by a processor, implements the steps of the cross-modal pharmaceutical logistics retrieval method described above.
[0061] In summary, this invention provides a cross-modal pharmaceutical logistics retrieval method and related equipment. The method includes: acquiring multiple pharmaceutical data sets; constructing a logistics dataset using all the pharmaceutical data sets; extracting features from the logistics dataset to obtain multiple modal feature representations; constructing a tag semantic graph based on all semantic tags in the logistics dataset; encoding the tag semantic graph using a graph neural network to obtain multiple semantic embedding representations; adding dynamic attention weights to each semantic embedding to semantically enhance each modal feature representation to obtain corresponding modal features; constructing a hash code for each modal feature; and constructing cross-modal semantics based on each modal feature. The network employs an adjacency matrix, constructs multiple semantic neighborhood-aware contrastive hash losses for each modal feature across different modalities based on the cross-modal semantic adjacency matrix and each hash code, constructs weighted label representations based on all semantic labels, constructs multiple similarity matrices based on the weighted label representations, and constructs a label distribution-aware semantic alignment loss function based on each similarity matrix and each hash code. A cosine triplet loss is constructed, and the constructed perceptual visual feature extraction network is optimized based on the semantic alignment loss function and the cosine triplet loss. Multiple test drug data are input into the optimized target feature extraction network, and the corresponding retrieval results are output. This invention can improve the semantic alignment of multimodal features, enhance model generalization ability, and significantly improve the retrieval accuracy of drug logistics.
[0062] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal that includes that element.
[0063] Of course, those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware (such as a processor, controller, etc.). The program can be stored in a computer-readable storage medium, and when executed, it can include the processes described in the above method embodiments. The computer-readable storage medium can be a memory, magnetic disk, optical disk, etc.
[0064] It should be understood that the application of the present invention is not limited to the examples above. Those skilled in the art can make improvements or modifications based on the above description, and all such improvements and modifications should fall within the protection scope of the appended claims.
Claims
1. A cross-modal drug logistics retrieval method, characterized in that, The cross-modal drug logistics retrieval method includes: Acquire multiple drug data sets, construct a logistics dataset using all the drug data sets, extract features from the logistics dataset to obtain multiple modal feature representations; A label semantic graph is constructed based on all semantic labels in the logistics dataset, and the label semantic graph is encoded using a graph neural network to obtain multiple semantic embedding representations. Attention dynamic weights are added to each semantic embedding to semantically enhance each modal feature representation and obtain the corresponding modal features. Construct a hash code for each modal feature, construct a cross-modal semantic adjacency matrix based on each modal feature, and construct multiple semantic neighborhood-aware comparative hash losses for each modal feature across different modalities based on the cross-modal semantic adjacency matrix and each hash code; A weighted label representation is constructed based on all the semantic labels, multiple similarity matrices are constructed based on the weighted label representations, and a label distribution-aware semantic alignment loss function is constructed based on each similarity matrix and each hash code; A cosine triplet loss is constructed, and the constructed perceptual visual feature extraction network is optimized based on all the contrastive hash loss, the semantic alignment loss function, and the cosine triplet loss. Multiple test drug data are input into the optimized target feature extraction network, and the corresponding retrieval results are output.
2. The cross-modal drug logistics retrieval method according to claim 1, characterized in that, The modal feature representation includes: image modal feature representation and text modal feature representation; The process of acquiring multiple drug data sets, constructing a logistics dataset using all the drug data sets, and extracting features from the logistics dataset to obtain multiple modal feature representations specifically includes: A feature fusion mechanism and a regression classifier are introduced into the encoder to obtain a feature extraction network, and a cross-modal attention mechanism is added to the feature extraction network to obtain a perceptual visual feature extraction network. Multiple drug data points are acquired, and each drug data point is input into the perceptual visual feature extraction network to obtain image data and text data corresponding to each drug data point. Semantic labels are constructed based on each image data point and each text data point, and a logistics dataset is constructed based on all the image data points, all the text data points, and all the semantic labels. ; ; ; in, D Represents a logistics dataset. N Indicates the quantity of logistics data. Represents the first in the logistics dataset i One logistics data point, Indicates the first i Image data, Indicates the first i A text data, Indicates the first i Semantic labels for image-text pairs, C Indicates the number of semantic tag categories. , and They represent the first i The first, second, and third image-text pairs C Semantic labels for each category; For each of the drug data sets, the perceptual visual feature extraction network extracts text information from the image data, and performs feature extraction on the text information and the image data respectively to obtain corresponding in-image text feature representations and image feature representations. The in-image text feature representations and image feature representations are then fused to obtain the image modal feature representation. : ; ; in, P Indicates the fusion weight. This represents the sigmoid activation function. and All of these represent learnable parameters. Representing image features, This represents the text feature representation within the image. Indicates to and Perform feature splicing. Indicates the use of fusion weights Modulation; For each of the drug data, the perceptual visual feature extraction network divides the text data into multiple sub-word units, maps each sub-word unit to a text embedding representation, and encodes the text embedding representation to obtain the text modal feature representation of the drug data. .
3. The cross-modal drug logistics retrieval method according to claim 1, characterized in that, The step of constructing a label semantic graph based on all semantic labels in the logistics dataset, encoding the label semantic graph using a graph neural network to obtain multiple semantic embedding representations, and adding dynamic attention weights to each semantic embedding to semantically enhance each modal feature representation to obtain the corresponding modal features specifically includes: Based on the prior semantic relationships between each semantic label, a label semantic graph is constructed using each semantic label as a node; The semantic graph of the labels is input into a graph neural network for encoding to obtain the semantic embedding representation of each semantic label: ; in, Indicates the first c Semantic embedding representation of a semantic tag Indicates the first c The initial feature matrix of semantic labels, The adjacency matrix representing semantic tags. GNN Represents a graph neural network; Construct corresponding key vectors and value vectors based on each semantic embedding representation, and construct query vectors using the modal feature representations: ; in, Q Represents the query vector. Indicates the first c A key vector of semantic labels, Indicates the first c A vector of semantic label values, Indicates the first i Image modal feature representation or text modal feature representation, , and All represent learnable parameters; Assign dynamic attention weights to the current semantic labels: ; in, Indicates the first c Dynamic attention weights for each semantic label. Indicates transpose. Indicates the first i The data for the first drug includes the first c A semantic tag, C Indicates the number of semantic tag categories. Indicates the first i The data for the first drug includes the first j A semantic tag, Indicates the dimension of the key vector; Based on each dynamic attention weight, each value vector is weighted and aggregated to obtain the semantic context representation of each semantic label: ; in, Indicates the first i Each label represents a semantic context; Each of the label semantic context representations is added to the corresponding modality feature representation to obtain the semantically enhanced modality features: ; in, Indicates the first i Modal features.
4. The cross-modal drug logistics retrieval method according to claim 1, characterized in that, The process of constructing a hash code for each modal feature, constructing a cross-modal semantic adjacency matrix based on each modal feature, and constructing multiple semantic neighborhood-aware contrastive hash losses for each modal feature across different modalities based on the cross-modal semantic adjacency matrix and each hash code specifically includes: After performing a linear transformation on each modal feature, a nonlinear activation function is used to further transform it, yielding the corresponding hash code: ; in, This represents the hash code, and tanh represents the non-linear activation function. Represents the projection matrix. Representing modal features, Represents the bias vector; The current drug data is defined as anchor samples. For each modality of the anchor samples, a semantic domain of another modality is constructed. Based on the two semantic domains, a cross-modal semantic adjacency matrix is constructed. ; Based on each of the cross-modal semantic adjacency matrices, a semantic neighborhood-aware contrastive hash loss is constructed for each image modal feature representation to its corresponding text modal feature representation and for each text modal feature representation to its corresponding image modal feature representation. ; ; in, Indicates the first i Drug data from images x To text y Semantic neighborhood perception in a given direction contrasts with hash loss. Indicates the first j Drug data from text y To image x Semantic neighborhood perception in a given direction contrasts with hash loss. Indicates the first j The text modality feature representation and the first i Each image modal feature represents label information with a semantic similarity higher than a threshold in the drug data. n Represents a collection of tag information. It is the sigmoid activation function. Indicates the first i A hash code representing the modal features of an image. Indicates the first j The hash code representing each text modal feature, where T is the transpose. This represents the temperature parameter.
5. The cross-modal drug logistics retrieval method according to claim 1, characterized in that, The step of constructing a weighted label representation based on all the semantic labels, constructing multiple similarity matrices based on the weighted label representations, and constructing a label distribution-aware semantic alignment loss function based on each similarity matrix and each hash code specifically includes: A label matrix is constructed based on all the semantic labels, and the label matrix is added to a weighted strategy for modeling to obtain a weighted label representation. A similarity matrix is then constructed based on the weighted label representation. ; ; in, This indicates a weighted label representation. L Represents the label matrix, diag Represents the category-level weight matrix. Indicates the first i The data for the first drug includes the first c A semantic tag, N Indicates the quantity of logistics data. C Indicates the number of semantic tag categories. Represents the semantic similarity matrix of tags. Indicates transpose; Based on each similarity matrix and each hash code, construct inter-modal semantic alignment loss function and intra-modal semantic alignment loss function for each drug data, and construct a semantic alignment loss function based on the inter-modal semantic alignment loss function and the intra-modal semantic alignment loss function.
6. The cross-modal drug logistics retrieval method according to claim 5, characterized in that, The step of constructing an inter-modal semantic alignment loss function and an intra-modal semantic alignment loss function for each drug data based on each similarity matrix and each hash code, and constructing a semantic alignment loss function based on the inter-modal semantic alignment loss function and the intra-modal semantic alignment loss function, specifically includes: For each of the drug data sets, based on each hash code, an intra-modal hash similarity matrix is constructed between the modal features of the drug data set and the modal features of other drug data sets. The difference between the intra-modal hash similarity matrix and the corresponding similarity matrix is calculated to obtain the intra-modal semantic alignment loss function. ; in, This represents the intramodal semantic alignment loss function. and These represent the first balance coefficient and the second balance coefficient, respectively. The intra-modal hash similarity matrix represents the dimension of the image modal features. The intra-modal hash similarity matrix represents the dimension of the text modal features; For each of the drug data, based on each hash code, a modal hash similarity matrix is constructed between the modal features of the drug data and the modal features of different types of other drug data. The difference between the modal hash similarity matrix and the corresponding similarity matrix is calculated to obtain the inter-modal semantic alignment loss function. ; ; in, This represents the inter-modal semantic alignment loss function. This represents the third balance coefficient. This represents the inter-modal hash similarity matrix between image modal feature representations and text modal feature representations. Indicates the first i The hash code representing the image modality feature and the first image modality feature. j The similarity matrix between hash codes represented by text modal features Indicates the first i The data for the first drug includes the first j A semantic tag, e The base of the natural logarithm; A semantic alignment loss function is constructed using the inter-modal semantic alignment loss function and the intra-modal semantic alignment loss function. : 。 7. The cross-modal drug logistics retrieval method according to claim 1, characterized in that, The construction of the cosine triplet loss optimizes the constructed perceptual visual feature extraction network based on all the contrastive hash losses, the semantic alignment loss function, and the cosine triplet loss. Multiple test drug data are then input into the optimized target feature extraction network, outputting the corresponding retrieval results. Specifically, this includes: For each type of modal feature of each of the drug data, determine the hash code of another modal feature that has a positive and negative correlation with the hash code of the modal feature, so as to construct a triplet of the modal feature: ; in, Indicates the first i A hash code representing the modal features of an image. Indicates and There is a positive correlation. j A hash code representing a text modality feature. Indicates and There is a negative correlation. k A hash code representing a text modality feature; Based on each of the triples, construct the cosine triplet loss for the image modality dimension and the text modality dimension respectively: ; ; in, The cosine triplet loss represents the modality dimension of the image. Indicates the first i A hash code representing the modal features of an image. Indicates and There is a positive correlation. j A hash code representing a text modality feature. Indicates and There is a negative correlation. k A hash code representing a text modality feature. The cosine triplet loss represents the text modality dimension. Indicates the first j A hash code representing a text modality feature. Indicates and There is a positive correlation. i A hash code representing a text modality feature. Indicates and There is a negative correlation. k A hash code representing a text modality feature. m Indicates the margin parameter; We construct the overall cosine triplet loss by utilizing the cosine triplet losses of the image modality dimension and the text modality dimension: ; in, This represents the overall cosine triplet loss; The perceptual visual feature extraction network is optimized using the overall cosine triplet loss, semantic alignment loss function, and all semantic neighborhood perceptual contrastive hash loss to obtain the target feature extraction network. Acquire test drug data, input all the test drug data into the target feature extraction network, and output the retrieval results corresponding to each test drug data.
8. A cross-modal pharmaceutical logistics retrieval system, characterized in that, The cross-modal drug logistics retrieval system is used to implement the cross-modal drug logistics retrieval method as described in any one of claims 1-7, wherein the cross-modal drug logistics retrieval system comprises: The feature extraction module is used to acquire multiple drug data, construct a logistics dataset using all the drug data, and extract features from the logistics dataset to obtain multiple modal feature representations. The graph encoding module is used to construct a label semantic graph based on all semantic labels in the logistics dataset, and to encode the label semantic graph using a graph neural network to obtain multiple semantic embedding representations. Attention dynamic weights are added to each semantic embedding to semantically enhance each modal feature representation and obtain the corresponding modal features. The first loss calculation module is used to construct a hash code for each modal feature, construct a cross-modal semantic adjacency matrix based on each modal feature, and construct multiple semantic neighborhood-aware comparative hash losses for each modal feature across different modalities based on the cross-modal semantic adjacency matrix and each hash code. The second loss calculation module is used to construct a weighted label representation based on all the semantic labels, construct multiple similarity matrices based on the weighted label representations, and construct a label distribution-aware semantic alignment loss function based on each similarity matrix and each hash code; The model optimization module is used to construct a cosine triplet loss, optimize the constructed perceptual visual feature extraction network based on all the contrastive hash loss, the semantic alignment loss function and the cosine triplet loss, and input multiple test drug data into the optimized target feature extraction network to output the corresponding retrieval results.
9. A terminal, characterized in that, The terminal includes: a memory, a processor, and a cross-modal drug logistics retrieval program stored in the memory and executable on the processor. When executed by the processor, the cross-modal drug logistics retrieval program implements the steps of the cross-modal drug logistics retrieval method as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a cross-modal pharmaceutical logistics retrieval program, which, when executed by a processor, implements the steps of the cross-modal pharmaceutical logistics retrieval method as described in any one of claims 1-7.