Document retrieval method, system and equipment and medium

By cross-modally fusing the features of user interaction logs, text and image data to generate fused features, the problem of insufficient multimodal data integration in traditional document retrieval systems is solved, and more accurate document retrieval result sorting and query optimization are achieved.

CN120632057APending Publication Date: 2025-09-12PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510722340.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Traditional document retrieval systems are unable to effectively integrate multimodal data, resulting in a disconnect between image features and text semantics, retrieval results being out of touch with business scenarios, and an inability to accurately locate relevant documents.

Method used

By obtaining user interaction logs and input text data and image data, semantic features, visual features and temporal features are extracted, cross-modal fusion is performed, fusion features are generated, and query rewriting and document retrieval are performed based on the fusion features. The multi-head attention mechanism and similarity matrix training are used to calculate the relevance score for sorting.

Benefits of technology

It improves the accuracy and precision of document retrieval, can better combine multimodal data for retrieval, generate more accurate query text, and improve the relevance and matching of retrieval results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120632057A_ABST
    Figure CN120632057A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of natural language processing, can be applied to business system platforms of financial science and technology, medical health and the like, and discloses a document retrieval method, system, equipment and medium. Extracting semantic features of the text data and visual features of the image data, extracting time sequence features of the interaction log, carrying out cross-modal fusion on the semantic features, the visual features and the time sequence features to obtain fusion features, carrying out query rewriting on the text data according to the fusion features to obtain a query text, and sending the query text to the server. The method comprises the steps of obtaining a query text, performing document retrieval according to the query text to obtain a retrieval result set, determining the correlation between each retrieval result in the retrieval result set and the query text to obtain a correlation score set, and sorting the retrieval result set according to the correlation score set to obtain a final retrieval result. And the document retrieval accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of natural language processing technology, and in particular to a document retrieval method, system, device and medium. Background Art

[0002] In the healthcare sector, efficient and accurate document retrieval is a core requirement for clinical decision-making, medical research, and patient management. With the development of medical informatization, data modalities have become significantly diversified. Electronic medical records contain structured text (diagnostic conclusions, medication regimens), unstructured text (medical records), medical images (CT / MRI images, pathological sections), biological signals (ECG waveforms), and patient interaction data (interview records, follow-up feedback). Traditional document retrieval systems rely on text retrieval technology based on keyword matching, which makes it difficult to meet the semantic fusion requirements of multimodal data. For example, when doctors search for literature related to "differential diagnosis of lung nodules", they need to combine the characteristics of lung CT images (such as nodule size, edge morphology), pathology report text descriptions, and patient historical medical records. However, traditional methods can only process a single text modality, resulting in a separation between image features and text semantics. The retrieval results often show "text related but image features do not match".

[0003] In the FinTech sector, particularly in the insurance industry, document retrieval permeates core business processes such as claims processing, policy management, and customer service. Insurance data is typically multimodal. Policy terms contain both legal text and tabular data, while claims applications often involve accident scene photos (such as vehicle damage images in auto insurance), images of medical receipts (such as treatment reports in liability insurance), scanned copies of handwritten signatures, and user interaction feedback (such as claims adjusters' highlighting of key terms and historical query session logs). Traditional retrieval systems rely solely on text keyword matching and fail to effectively integrate image evidence with user behavior data, resulting in retrieval results that are disconnected from the business scenario. For example, in an auto insurance claim, claims adjusters must combine the damage location in accident photos (such as a cracked bumper), the report text description ("low-speed collision"), and historical interaction records from similar cases (such as reviewing compensation standards for similar injuries) to retrieve relevant policy terms. However, traditional systems only return general terms based on text keyword matching, making it difficult to accurately locate relevant documents that contain specific damage image features and historical processing logic. Summary of the Invention

[0004] The present invention provides an artificial intelligence document retrieval method, system, computer device and medium to solve the problem of low accuracy of existing document retrieval methods.

[0005] In a first aspect, a document retrieval method is provided, comprising:

[0006] Obtain user interaction logs and text data and image data input by users;

[0007] Extracting semantic features of the text data and visual features of the image data, and performing time series encoding on the interaction log to obtain time series features;

[0008] Performing cross-modal fusion on the semantic features, the visual features, and the temporal features to obtain fused features;

[0009] Performing query rewriting on the text data according to the fusion feature to obtain a query text, and performing document retrieval according to the query text to obtain a retrieval result set;

[0010] Determining the relevance of each search result in the search result set to the query text to obtain a relevance score set;

[0011] The search result set is sorted according to the relevance score set to obtain a final search result.

[0012] In a second aspect, a document retrieval system is provided, comprising:

[0013] A data acquisition module is used to obtain the user's interaction log and the text data and image data input by the user;

[0014] A feature extraction module, configured to extract semantic features of the text data and visual features of the image data, and perform time series encoding on the interaction log to obtain time series features;

[0015] A feature fusion module, configured to perform cross-modal fusion of the semantic features, the visual features, and the temporal features to obtain fused features;

[0016] A query rewriting module, configured to perform query rewriting on the text data according to the fusion features to obtain a query text, perform document retrieval according to the query text, and obtain a retrieval result set;

[0017] The retrieval output module is used to determine the relevance between each retrieval result in the retrieval result set and the query text, obtain a relevance score set, and sort the retrieval result set according to the relevance score set to obtain a final retrieval result.

[0018] In a third aspect, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the document retrieval method when executing the computer program.

[0019] In a fourth aspect, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned document retrieval method are implemented.

[0020] In the solution implemented by the above-mentioned document retrieval method, system, computer device and storage medium, the semantic features of the text data and the visual features of the image data can be extracted by obtaining the user's interaction log and the text data and image data input by the user, and the interaction log can be temporally encoded to obtain temporal features. The semantic features, the visual features and the temporal features can be cross-modally fused to obtain fused features. The text data can be query-rewritten based on the fused features to obtain a query text. Document retrieval can be performed based on the query text to obtain a retrieval result set. The relevance of each retrieval result in the retrieval result set to the query text can be determined to obtain a relevance score set. The retrieval result set can be sorted based on the relevance score set to obtain a final retrieval result. This improves the accuracy of document retrieval. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments of the present invention. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0022] Figure 1 is a schematic diagram of an application environment of a document retrieval method according to an embodiment of the present invention;

[0023] Figure 2 is a flowchart of a document retrieval method according to an embodiment of the present invention;

[0024] Figure 3 is a structural diagram of a document retrieval system according to an embodiment of the present invention;

[0025] Figure 4 is a structural diagram of a computer device in one embodiment of the present invention;

[0026] Figure 5 FIG. 2 is another structural diagram of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0027] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0028] The document retrieval method provided by the embodiment of the present invention can be applied in Figure 1In an application environment, the client communicates with the server through a network. The server can obtain text data, image data and interaction logs input by the user through the client, extract the semantic features of the text data and the visual features of the image data, extract the temporal features of the interaction log, perform cross-modal fusion of the semantic features, the visual features and the temporal features to obtain fusion features, perform query rewriting on the text data according to the fusion features to obtain a query text, perform document retrieval according to the query text to obtain a retrieval result set, determine the relevance of each retrieval result in the retrieval result set with the query text, obtain a relevance score set, sort the retrieval result set according to the relevance score set to obtain a final retrieval result. The accuracy of document retrieval is improved, and the client can be, but is not limited to, various personal computers, laptops, smart phones, tablet computers and portable wearable devices. The server can be implemented with an independent server or a server cluster consisting of multiple servers. The present invention is described in detail below through specific embodiments.

[0029] See also Figure 2 As shown, Figure 2 A flowchart of a document retrieval method provided by an embodiment of the present invention includes the following steps:

[0030] S1. Obtain the user's interaction log and the text data and image data input by the user.

[0031] In the embodiment of the present invention, the text data and image data input by the user refer to the query text and the corresponding image input by the user for document retrieval.

[0032] In the field of financial technology, when a user uses an insurance claim system, the text data may be the "car insurance glass damage claim process" and the image data may be a photo of a cracked windshield uploaded by the user.

[0033] In the healthcare field, when a user uses a medical query system, the text data can be a description of the user's symptoms, and the image data can be an image corresponding to the symptoms. For example, if a user develops an abnormal red spot on their skin, the text data can be "What is the cause of the red spot on the skin?", and the image data can be an image of the red spot on the skin.

[0034] In the embodiment of the present invention, the interaction log may refer to a user's viewing record of related documents or a historical query record.

[0035] In detail, in the fintech field, when a user uses an insurance claims system, the interaction log may be a record of the user's retrieval of relevant insurance documents. For example, when a user inquires about the claim process for broken glass in automobile insurance, the interaction log may be a record of viewing documents related to the glass insurance terms.

[0036] In detail, in the medical and health field, when a user uses a medical query system, the interaction log may be a record of the user's viewing of relevant disease documents. For example, when a user queries "What is the cause of erythema on the skin", the interaction record may be a record of the user viewing popular science documents on skin diseases.

[0037] In the embodiment of the present invention, by obtaining the user's interaction log and the text data and image data input by the user, the efficiency of subsequent feature extraction of the text data, image data and interaction log can be improved.

[0038] S2. Extracting semantic features of the text data and visual features of the image data, and performing time series encoding on the interaction log to obtain time series features.

[0039] In an embodiment of the present invention, the semantic features of the text data are extracted by standardizing the text data and then sending the data to a text encoder for feature extraction.

[0040] Specifically, the text encoder uses an improved DeBERTa-v3 architecture. This architecture, based on the original DeBERTa model, has been optimized for domain adaptation. It targets the characteristics of documents in specialized fields like finance and healthcare, enhancing the semantic modeling and domain terminology representation capabilities of long texts.

[0041] In an embodiment of the present invention, extracting the semantic features of the text data includes:

[0042] Performing standardization processing on the text data to obtain standardized text;

[0043] Performing word segmentation processing on the standardized text to obtain a word segmentation sequence;

[0044] Mapping the word segmentation sequence into a word frequency vector according to the occurrence position and number of occurrences of each word in the word segmentation sequence in a preset dictionary;

[0045] Using a preset text encoder to perform professional term recognition on the word segmentation sequence to obtain a recognition result;

[0046] The word frequency vector is marked with professional terms according to the recognition result to obtain semantic features.

[0047] In this embodiment of the present invention, the extraction of visual features from the image data is performed automatically using a pre-defined image encoder. This image encoder is based on the ViT-Large model, a large-scale version of the Vision Transformer (ViT) technology system, designed to efficiently handle image classification and complex vision tasks.

[0048] The technical architecture of the ViT-Large model is significantly innovative and targeted. As a visual domain adaptation of the Transformer encoder, it overcomes the local perception limitations of traditional convolutional neural networks and constructs a global feature modeling system suitable for image input. In its implementation, the input image is first segmented into a sequence of fixed-size image patches. Each patch is converted into an embedding vector via a linear mapping, forming a sequence data format suitable for Transformer processing. To preserve the spatial position information of the image, the model introduces a learnable position encoding mechanism that incorporates spatial coordinate information into the sequence of embedding vectors. The model's core processing unit consists of a stack of multiple layers of Transformer encoders, each of which incorporates a multi-head self-attention mechanism and a feed-forward neural network. The multi-head self-attention mechanism concurrently captures image patch dependencies within different subspaces, enabling cross-region semantic association modeling. The feed-forward neural network performs nonlinear transformations on the attention mechanism outputs to enhance feature representation. Through this hierarchical feature processing pipeline, the model is able to extract multi-granular visual features from the original image layer by layer, from edge texture to semantic concepts.

[0049] Furthermore, in its implementation, the image encoder's workflow can be broken down into the following steps: First, the input image undergoes standardized preprocessing and resizing to fit the model's input specifications. Subsequently, the image is converted into a two-dimensional grid of image blocks through a block-by-block operation, with each block mapped to an embedding vector of fixed dimension. After positional encoding is injected into the embedding vector sequence, the vector is fed into the ViT-Large model for multi-layer feature transformation, ultimately extracting a highly abstract and semantically representative visual feature vector from the model's deep output. This feature vector integrates the image's global structural information with local detail features, providing high-quality feature input for subsequent visual tasks such as image classification and object detection.

[0050] In the embodiment of the present invention, performing time series encoding on the interaction log to obtain time series features includes:

[0051] Extracting user operation data from the interaction log;

[0052] Performing feature encoding on the user operation data based on a preset operation type to obtain an operation coding feature;

[0053] Dividing the operation coding feature into a time series based on the timestamp of the interaction log to obtain a sequence operation feature;

[0054] Based on the preset time series neural network, the time vector of each time series is spliced ​​with the corresponding sequence operation feature to obtain the time series feature.

[0055] In detail, the user operation data may include user query, click, confirm, cancel and other operation data.

[0056] In the embodiment of the present invention, the time series encoding of the interaction log to obtain the time series features is performed by utilizing a temporal series neural network (TGNN) to time series encode the interaction log.

[0057] In detail, the temporal neural network includes three layers of temporal convolution layers and an attention pooling layer, which can capture the long-term dependencies of user behaviors.

[0058] In detail, the temporal neural network adopts a hierarchical temporal feature modeling architecture, the core components of which are 3 layers of cascaded temporal convolutional layers (Temporal Convolutional Layers) and an innovatively designed attention pooling layer (Attention Pooling Layer), forming a multi-scale temporal dependency modeling system for user behavior sequences. In the temporal feature extraction module, the 3 layers of temporal convolutional layers adopt a design paradigm that combines causal convolution (Causal Convolution) and dilated convolution (Dilated Convolution). The first temporal convolution layer is configured with a convolution kernel of size 5, and causal padding (Causal Padding) is used to ensure that the output at the current moment depends only on the input at the historical moment, maintaining the temporal causality of sequence processing; a conventional convolution operation with a dilation rate of 1 is used to realize basic temporal local feature extraction. The second convolution layer increases the dilation rate to 2, and while maintaining the same number of parameters, the receptive field is expanded to 9 time steps, which can capture temporal dependencies with a larger span. The third layer further sets the expansion rate to 4, so that the effective receptive field covers 17 time steps, and with the 256-dimensional feature channel number, it realizes the hierarchical abstraction of long-distance temporal patterns in user behavior sequences. After each layer of convolution, layer normalization (LayerNormalization) and GELU activation function are connected to construct a nonlinear transformation unit to enhance feature expression capabilities. The attention pooling operation realizes the selective aggregation of historical information through weighted summation, especially strengthening the feature weights of long-term dependent related time steps, effectively solving the problem of temporal information loss caused by fixed windows in traditional pooling operations. The network architecture expands the receptive field layer by layer through three layers of time convolution, gradually integrating local temporal features into contextual representations containing long-distance dependencies, and then through the adaptive weight allocation of the attention pooling layer, it can accurately capture the cross-time scale dependencies in user behavior sequences.

[0059] In the embodiment of the present invention, dividing the operation coding features into time series based on the timestamp of the interaction log is to divide the operation coding features according to preset time intervals.

[0060] In an embodiment of the present invention, by extracting the semantic features of the text data and the visual features of the image data, and performing time series encoding on the interaction log to obtain time series features, the efficiency of subsequent cross-modal fusion of the semantic features, the visual features, and the time series features can be improved.

[0061] S3. Perform cross-modal fusion on the semantic features, the visual features, and the temporal features to obtain fused features.

[0062] In an embodiment of the present invention, the cross-modal fusion of the semantic features, the visual features and the temporal features to obtain the fused features is to first calculate the respective weights of the semantic features, the visual features and the temporal features, and then perform weighted fusion based on the calculated weights.

[0063] In the embodiment of the present invention, cross-modal fusion of the semantic features, the visual features, and the temporal features to obtain fused features includes:

[0064] Determining attention weights of the semantic features, the visual features, and the temporal features to obtain a weight set;

[0065] The semantic features, the visual features and the temporal features are weightedly fused according to the weight set to obtain a fused feature.

[0066] In an embodiment of the present invention, by cross-modally fusing the semantic features, the visual features, and the temporal features to obtain fused features, the accuracy of query rewriting of the text data based on the fused features can be improved.

[0067] In detail, the calculation formula of the weight set is as follows:

[0068]

[0069] Among them, α t is the semantic feature weight contained in the weight set, α i is the visual feature weight contained in the weight set, α u is the time series feature weight contained in the weight set, h t is the semantic feature, h i is the visual feature, h u is the timing feature, T represents the transposition operation, W t is the learnable matrix of the semantic features, W i is the learnable matrix of the visual features, W u is the learnable matrix of the time series features, and exp() is the natural exponential function.

[0070] S4. Perform query rewriting on the text data according to the fusion feature to obtain a query text, perform document retrieval according to the query text, and obtain a retrieval result set.

[0071] In an embodiment of the present invention, the query rewriting of the text data according to the fusion feature is to obtain user feedback data through a multimodal interactive interface. The feedback data may include explicit feedback (such as revised query terms, related document annotations) and implicit feedback (such as click behavior, browsing time). A multi-head attention mechanism is used to extract features from the feedback data, construct a feature similarity measurement function, and calculate the semantic distance between the feedback feature and the fusion feature. Based on the calculated loss value, a dynamic alignment module is designed to realize feature space calibration. In the aligned embedding space, an encoder-decoder architecture is used to realize the generation of query text.

[0072] In the embodiment of the present invention, the query rewriting of the text data according to the fusion feature to obtain the query text includes:

[0073] Obtaining user feedback data, and extracting feedback features of the feedback data based on an attention mechanism;

[0074] Calculating a loss value of the feedback feature and the fusion feature;

[0075] Performing spatial alignment on the feedback feature and the fusion feature based on the loss value to obtain an aligned embedding space;

[0076] Query enhancement is performed according to the aligned embedding space and the text data to obtain a query text.

[0077] In the embodiment of the present invention, query enhancement based on the aligned embedding space and the text data refers to performing keyword splicing and insertion on the text data based on the aligned embedding space to obtain a more accurate query text.

[0078] In the embodiment of the present invention, the user feedback data refers to the user's feedback data on the query result obtained from the text data.

[0079] In the embodiment of the present invention, the calculating the loss value of the feedback feature and the fusion feature includes:

[0080] Respectively calculating the product of the fusion feature and the feedback feature, the product of the fusion feature and each negative sample in the preset negative sample set, and the product of each negative sample in the negative sample set and the feedback feature to obtain a first product term, a second product term, and a third product term;

[0081] Respectively calculating the product of the norm of the fusion feature and the norm of the feedback feature, the product of the norm of the fusion feature and the norm corresponding to each negative sample in the negative sample set, and the product of the norm corresponding to each negative sample in the negative sample set and the norm of the feedback feature to obtain a first norm product term, a second norm product term, and a third norm product term;

[0082] After respectively calculating the ratio of the first product term to the first norm product term, the ratio of the second product term to the second norm product term, and the ratio of the third product term to the third norm product term, the calculated ratios are divided by a preset adjustment parameter to obtain a first ratio term, a second ratio term, and a third ratio term;

[0083] Sum the second ratio items corresponding to all negative samples in the negative sample set to obtain a first sum item;

[0084] Summing the third ratio terms corresponding to all negative samples in the negative sample set to obtain a second summation term, and adding the first summation term to the second summation term to obtain a denominator term;

[0085] The ratio of the first ratio term to the denominator term is calculated to obtain a comprehensive ratio, and the reciprocal of the natural logarithm of the comprehensive ratio is calculated to obtain the loss value.

[0086] In the embodiment of the present invention, the calculating the loss value of the feedback feature and the fusion feature includes:

[0087] The loss value is calculated using the following formula:

[0088]

[0089] in, is the loss value, exp() is the natural exponential function, h a is the fusion feature, h b is the feedback feature, ||h a || represents the norm of the fusion feature, ||h b || represents the norm of the feedback feature, M is the number of samples in the preset negative sample set, h k represents the kth negative sample in the negative sample set, ||h k || represents the norm of the kth negative sample in the negative sample set, and τ is a preset adjustment parameter.

[0090] In detail, the adjustment parameter may be 0.07.

[0091] Specifically, in machine learning and deep learning, negative samples are the opposite of target samples or positive samples. They typically refer to samples that do not belong to the target category or meet specific criteria. Their core function is to help the model learn to "distinguish differences," preventing the model from only remembering the characteristics of positive samples, thereby improving generalization capabilities.

[0092] In an embodiment of the present invention, by performing query rewriting on the text data according to the fusion feature to obtain a query text, and performing document retrieval according to the query text to obtain a retrieval result set, the accuracy of document retrieval can be improved.

[0093] S5. Determine the relevance between each search result in the search result set and the query text to obtain a relevance score set.

[0094] In an embodiment of the present invention, determining the relevance between each search result in the search result set and the query text to obtain a relevance score set includes:

[0095] Constructing positive and negative sample pairs according to the query text;

[0096] Generating a feature vector of the query text using a text encoder to obtain query features;

[0097] Iteratively optimizing a preset trainable similarity matrix based on the positive and negative sample pairs and the query features to obtain a trained similarity matrix;

[0098] Performing feature extraction on each search result in the search result set to obtain a search feature set;

[0099] The trained similarity matrix is ​​used to determine the relevance between each retrieval result in the retrieval result set and the query text according to the query feature and the retrieval feature set, thereby obtaining a relevance score set.

[0100] In an embodiment of the present invention, the constructing of positive and negative sample pairs according to the query text is to construct retrieval results related to the query text and retrieval results unrelated to the query text based on the query text. For example, in a medical query system, if the query text is "common skin diseases", the positive sample in the positive and negative sample pair can be a text related to skin diseases, and the negative sample in the positive and negative sample pair can be a text related to orthopedics.

[0101] In an embodiment of the present invention, the iterative optimization of a preset trainable similarity matrix based on the positive and negative sample pairs and the query features to obtain a trained similarity matrix is ​​performed by initializing the trainable similarity matrix as a model parameter, encoding the query feature and its corresponding positive sample and negative sample as feature vectors respectively; then, using the current similarity matrix to calculate the similarity between the query feature and the positive and negative sample features, and measuring the difference between the positive sample similarity and the negative sample similarity by designing a loss function, requiring that the positive sample similarity may be higher than the negative sample similarity; then, calculating the loss gradient based on the sample batch, updating the parameters of the similarity matrix through the back propagation algorithm, and adjusting the weights in the matrix to shorten the distance between the positive sample pairs and increase the distance between the negative sample pairs; finally, repeating the above process, iteratively optimizing on the training data until the loss function converges or reaches a preset training round, thereby obtaining a trained similarity matrix that can effectively distinguish the similarity between positive and negative samples.

[0102] S6. Sort the search result set according to the relevance score set to obtain a final search result.

[0103] In an embodiment of the present invention, sorting the retrieval result set according to the relevance score set to obtain the final retrieval result means sorting the corresponding retrieval results in the retrieval result set in descending order of all relevance scores in the relevance score set, and the result obtained by sorting is the final retrieval result.

[0104] As can be seen, in the above solution, user input text data, image data, and interaction logs are first collected. Text data represents search queries (e.g., "Auto insurance glass damage claim process" in the financial sector, or "What causes skin erythema" in the medical sector). Image data corresponds to contextual images (e.g., photos of cracked windshields, images of skin erythema). Interaction logs record user behavior history, including document viewing (e.g., insurance policy review, disease science documentation browsing), providing behavioral temporal information for subsequent feature analysis. An improved DeBERTa-v3 architecture encoder is used to generate feature vectors containing domain semantics through normalization, word segmentation, word frequency vector mapping, and terminology recognition. An image encoder based on the ViT-Large model extracts local and global visual features for image classification and vision tasks. A temporal sequence neural network (TGNN, consisting of three temporal convolutional layers and an attention pooling layer) is used to segment interaction logs into sequences based on time intervals, encoding dependencies between operation types and timestamps to capture long-term patterns in user behavior. The attention mechanism is used to calculate the weights of semantic, visual, and temporal features, and multimodal information is integrated through a weighted fusion formula to form fusion features that include cross-domain associations, providing a comprehensive representation for query optimization. Combined with user feedback data, feedback features are extracted through the attention mechanism, and the loss value of the fusion feature and the feedback feature is calculated (MMCL loss function). After spatial alignment, the original query text is keyword enhanced (splicing / inserting) to generate a more accurate query text and improve the matching degree of the retrieval intent. Positive and negative sample pairs are constructed (for example, in medical scenarios, "common skin diseases" correspond to dermatology text as positive and orthopedics text as negative), and the similarity matrix is ​​trained to distinguish sample correlations. Based on the trained matrix, the relevance score of the retrieval results and the query text is calculated, sorted from high to low by score, and the final retrieval results are output. This improves the accuracy of document retrieval.

[0105] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0106] In one embodiment, a document retrieval system is provided, which corresponds one-to-one to the document retrieval method in the above embodiment. Figure 3 As shown, the document retrieval system includes a data acquisition module 101, a feature extraction module 102, a feature fusion module 103, a query rewriting module 104, and a retrieval output module 105. The functional modules are described in detail as follows:

[0107] The data acquisition module 101 is used to acquire the user's interaction log and the text data and image data input by the user;

[0108] A feature extraction module 102 is configured to extract semantic features of the text data and visual features of the image data, and perform time series coding on the interaction log to obtain time series features;

[0109] A feature fusion module 103 is configured to perform cross-modal fusion on the semantic features, the visual features, and the temporal features to obtain fused features;

[0110] A query rewriting module 104 is configured to perform query rewriting on the text data according to the fusion feature to obtain a query text, perform document retrieval according to the query text, and obtain a retrieval result set;

[0111] The retrieval output module 105 is configured to determine the relevance between each retrieval result in the retrieval result set and the query text, obtain a relevance score set, and sort the retrieval result set according to the relevance score set to obtain a final retrieval result.

[0112] In one embodiment, the feature extraction module 102, when performing the extraction of semantic features of the text data, is specifically configured to:

[0113] Performing standardization processing on the text data to obtain standardized text;

[0114] Performing word segmentation processing on the standardized text to obtain a word segmentation sequence;

[0115] Mapping the word segmentation sequence into a word frequency vector according to the occurrence position and number of occurrences of each word in the word segmentation sequence in a preset dictionary;

[0116] Using a preset text encoder to perform professional term recognition on the word segmentation sequence to obtain a recognition result;

[0117] The word frequency vector is marked with professional terms according to the recognition result to obtain semantic features.

[0118] In one embodiment, the feature extraction module 102, when executing and performing time series coding on the interaction log to obtain time series features, is specifically configured to:

[0119] Extracting user operation data from the interaction log;

[0120] Performing feature encoding on the user operation data based on a preset operation type to obtain an operation coding feature;

[0121] Dividing the operation coding feature into a time series based on the timestamp of the interaction log to obtain a sequence operation feature;

[0122] Based on the preset time series neural network, the time vector of each time series is spliced ​​with the corresponding sequence operation feature to obtain the time series feature.

[0123] In one embodiment, the feature fusion module 103, when performing the cross-modal fusion of the semantic features, the visual features, and the temporal features to obtain the fused features, is specifically configured to:

[0124] Determining attention weights of the semantic features, the visual features, and the temporal features to obtain a weight set;

[0125] The semantic features, the visual features and the temporal features are weightedly fused according to the weight set to obtain a fused feature.

[0126] In one embodiment, the query rewriting module 104, when performing query rewriting on the text data according to the fusion feature to obtain the query text, is specifically configured to:

[0127] Obtaining user feedback data, and extracting feedback features of the feedback data based on an attention mechanism;

[0128] Calculating a loss value of the feedback feature and the fusion feature;

[0129] Performing spatial alignment on the feedback feature and the fusion feature based on the loss value to obtain an aligned embedding space;

[0130] Query enhancement is performed according to the aligned embedding space and the text data to obtain a query text.

[0131] In one embodiment, the query rewriting module 104, when calculating the loss value of the feedback feature and the fusion feature, is specifically configured to:

[0132] Respectively calculating the product of the fusion feature and the feedback feature, the product of the fusion feature and each negative sample in the preset negative sample set, and the product of each negative sample in the negative sample set and the feedback feature to obtain a first product term, a second product term, and a third product term;

[0133] Respectively calculating the product of the norm of the fusion feature and the norm of the feedback feature, the product of the norm of the fusion feature and the norm corresponding to each negative sample in the negative sample set, and the product of the norm corresponding to each negative sample in the negative sample set and the norm of the feedback feature to obtain a first norm product term, a second norm product term, and a third norm product term;

[0134] After respectively calculating the ratio of the first product term to the first norm product term, the ratio of the second product term to the second norm product term, and the ratio of the third product term to the third norm product term, the calculated ratios are divided by a preset adjustment parameter to obtain a first ratio term, a second ratio term, and a third ratio term;

[0135] Sum the second ratio items corresponding to all negative samples in the negative sample set to obtain a first sum item;

[0136] Summing the third ratio terms corresponding to all negative samples in the negative sample set to obtain a second summation term, and adding the first summation term to the second summation term to obtain a denominator term;

[0137] The ratio of the first ratio term to the denominator term is calculated to obtain a comprehensive ratio, and the reciprocal of the natural logarithm of the comprehensive ratio is calculated to obtain the loss value.

[0138] In one embodiment, the search output module 105, when determining the relevance of each search result in the search result set with the query text to obtain a relevance score set, is specifically configured to:

[0139] Constructing positive and negative sample pairs according to the query text;

[0140] Generating a feature vector of the query text using a text encoder to obtain query features;

[0141] Iteratively optimizing a preset trainable similarity matrix based on the positive and negative sample pairs and the query features to obtain a trained similarity matrix;

[0142] Performing feature extraction on each search result in the search result set to obtain a search feature set;

[0143] The trained similarity matrix is ​​used to determine the relevance between each retrieval result in the retrieval result set and the query text according to the query feature and the retrieval feature set, thereby obtaining a relevance score set.

[0144] The present invention provides a document retrieval system, which first collects text data, image data and interaction logs input by users. The text data is the retrieval query text (such as "car insurance glass damage claim process" in the financial field and "what is the cause of skin erythema" in the medical field), and the image data corresponds to scene-based pictures (such as photos of windshield cracks and images of skin erythema). The interaction log records the user's historical behavior, including document viewing records (such as insurance clause review and disease science document browsing), providing behavioral time series information for subsequent feature analysis. The improved DeBERTa-v3 architecture encoder is used to generate feature vectors containing domain semantics through standardization, word segmentation, word frequency vector mapping and professional terminology recognition. The image encoder based on the ViT-Large model extracts local and global visual features of the image and adapts to image classification and visual tasks. Through the temporal neural network (TGNN, containing 3 layers of temporal convolutional layers and attention pooling layers), the interaction log is divided into sequences according to time intervals, the operation type and timestamp dependency are encoded, and the long-term pattern of user behavior is captured. The attention mechanism is used to calculate the weights of semantic, visual, and temporal features, and multimodal information is integrated through a weighted fusion formula to form fusion features that include cross-domain associations, providing a comprehensive representation for query optimization. Combined with user feedback data, feedback features are extracted through the attention mechanism, and the loss value of the fusion feature and the feedback feature is calculated (MMCL loss function). After spatial alignment, the original query text is keyword enhanced (splicing / inserting) to generate a more accurate query text and improve the matching degree of the retrieval intent. Positive and negative sample pairs are constructed (for example, in medical scenarios, "common skin diseases" correspond to dermatology text as positive and orthopedics text as negative), and the similarity matrix is ​​trained to distinguish sample correlations. Based on the trained matrix, the relevance score of the retrieval results and the query text is calculated, sorted from high to low by score, and the final retrieval results are output. This improves the accuracy of document retrieval.

[0145] The specific definition of the document retrieval system can be found in the definition of the document retrieval method above and will not be repeated here. Each module in the above-mentioned document retrieval system can be implemented in whole or in part through software, hardware, or a combination thereof. Each of the above-mentioned modules can be embedded in or independent of the processor of the computer device in hardware form, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the corresponding operations of each of the above modules.

[0146] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 4As shown. The computer device includes a processor, a memory, a network interface and a database connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external client via a network connection. When the computer program is executed by the processor, it implements the functions or steps on the server side of a document retrieval method.

[0147] In one embodiment, a computer device is provided. The computer device may be a client, and its internal structure diagram may be as follows: Figure 5 As shown. The computer device includes a processor, memory, network interface, display screen and input system connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it implements the functions or steps of the client side of a document retrieval method.

[0148] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are performed:

[0149] Obtain user interaction logs and text data and image data input by users;

[0150] Extracting semantic features of the text data and visual features of the image data, and performing time series encoding on the interaction log to obtain time series features;

[0151] Performing cross-modal fusion on the semantic features, the visual features, and the temporal features to obtain fused features;

[0152] Performing query rewriting on the text data according to the fusion feature to obtain a query text, and performing document retrieval according to the query text to obtain a retrieval result set;

[0153] Determining the relevance of each search result in the search result set to the query text to obtain a relevance score set;

[0154] The search result set is sorted according to the relevance score set to obtain a final search result.

[0155] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0156] Obtain user interaction logs and text data and image data input by users;

[0157] Extracting semantic features of the text data and visual features of the image data, and performing time series encoding on the interaction log to obtain time series features;

[0158] Performing cross-modal fusion on the semantic features, the visual features, and the temporal features to obtain fused features;

[0159] Performing query rewriting on the text data according to the fusion feature to obtain a query text, and performing document retrieval according to the query text to obtain a retrieval result set;

[0160] Determining the relevance of each search result in the search result set to the query text to obtain a relevance score set;

[0161] The search result set is sorted according to the relevance score set to obtain a final search result.

[0162] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or computer device can be found in the relevant descriptions of the server side and the client side in the aforementioned method embodiment. To avoid repetition, they will not be described one by one here.

[0163] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0164] Those skilled in the art will clearly understand that for the sake of convenience and brevity in description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the system can be divided into different functional units or modules to complete all or part of the functions described above.

[0165] Finally, it should be noted that if software tools or components other than those of our company appear in the application examples, they are merely for illustration and do not represent actual use. The above-described embodiments are intended only to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the above-described embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above-described embodiments or replace some of the technical features therein with equivalents. Such modifications or replacements do not deviate from the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention and should be included within the scope of protection of the present invention.

Claims

1. A document retrieval method, characterized in that: include: Obtain user interaction logs and text data and image data input by users; Extracting semantic features of the text data and visual features of the image data, and performing time series encoding on the interaction log to obtain time series features; Performing cross-modal fusion on the semantic features, the visual features, and the temporal features to obtain fused features; Performing query rewriting on the text data according to the fusion feature to obtain a query text, and performing document retrieval according to the query text to obtain a retrieval result set; Determining the relevance of each search result in the search result set to the query text to obtain a relevance score set; The search result set is sorted according to the relevance score set to obtain a final search result.

2. The document retrieval method according to claim 1, wherein: In the step of extracting the semantic features of the text data and the visual features of the image data, extracting the semantic features of the text data includes: Performing standardization processing on the text data to obtain standardized text; Performing word segmentation processing on the standardized text to obtain a word segmentation sequence; Mapping the word segmentation sequence into a word frequency vector according to the occurrence position and number of occurrences of each word in the word segmentation sequence in a preset dictionary; Using a preset text encoder to perform professional term recognition on the word segmentation sequence to obtain a recognition result; The word frequency vector is marked with professional terms according to the recognition result to obtain semantic features.

3. The document retrieval method according to claim 1, wherein: The interaction log is time-series encoded to obtain time-series features, including: Extracting user operation data from the interaction log; Performing feature encoding on the user operation data based on a preset operation type to obtain an operation coding feature; Dividing the operation coding feature into a time series based on the timestamp of the interaction log to obtain a sequence operation feature; Based on the preset time series neural network, the time vector of each time series is spliced ​​with the corresponding sequence operation feature to obtain the time series feature.

4. The document retrieval method according to claim 1, wherein: The cross-modal fusion of the semantic features, the visual features, and the temporal features to obtain fused features includes: Determining attention weights of the semantic features, the visual features, and the temporal features to obtain a weight set; The semantic features, the visual features and the temporal features are weightedly fused according to the weight set to obtain a fused feature.

5. The document retrieval method according to claim 1, wherein: The step of performing query rewriting on the text data according to the fusion feature to obtain a query text includes: Obtaining user feedback data, and extracting feedback features of the feedback data based on an attention mechanism; Calculating a loss value of the feedback feature and the fusion feature; Performing spatial alignment on the feedback feature and the fusion feature based on the loss value to obtain an aligned embedding space; Query enhancement is performed according to the aligned embedding space and the text data to obtain a query text.

6. The document retrieval method according to claim 5, wherein: In the embodiment of the present invention, the calculating the loss value of the feedback feature and the fusion feature includes: Respectively calculating the product of the fusion feature and the feedback feature, the product of the fusion feature and each negative sample in the preset negative sample set, and the product of each negative sample in the negative sample set and the feedback feature to obtain a first product term, a second product term, and a third product term; Respectively calculating the product of the norm of the fusion feature and the norm of the feedback feature, the product of the norm of the fusion feature and the norm corresponding to each negative sample in the negative sample set, and the product of the norm corresponding to each negative sample in the negative sample set and the norm of the feedback feature to obtain a first norm product term, a second norm product term, and a third norm product term; After respectively calculating the ratio of the first product term to the first norm product term, the ratio of the second product term to the second norm product term, and the ratio of the third product term to the third norm product term, the calculated ratios are divided by a preset adjustment parameter to obtain a first ratio term, a second ratio term, and a third ratio term; Sum the second ratio items corresponding to all negative samples in the negative sample set to obtain a first sum item; Summing the third ratio terms corresponding to all negative samples in the negative sample set to obtain a second summation term, and adding the first summation term to the second summation term to obtain a denominator term; The ratio of the first ratio term to the denominator term is calculated to obtain a comprehensive ratio, and the reciprocal of the natural logarithm of the comprehensive ratio is calculated to obtain the loss value.

7. The document retrieval method according to claim 1, wherein: Determining the relevance between each search result in the search result set and the query text to obtain a relevance score set includes: Constructing positive and negative sample pairs according to the query text; Generating a feature vector of the query text using a text encoder to obtain query features; Iteratively optimizing a preset trainable similarity matrix based on the positive and negative sample pairs and the query features to obtain a trained similarity matrix; Performing feature extraction on each search result in the search result set to obtain a search feature set; The trained similarity matrix is ​​used to determine the relevance between each retrieval result in the retrieval result set and the query text according to the query feature and the retrieval feature set, thereby obtaining a relevance score set.

8. A document retrieval system, characterized in that: include: A data acquisition module is used to obtain the user's interaction log and the text data and image data input by the user; A feature extraction module, configured to extract semantic features of the text data and visual features of the image data, and perform time series encoding on the interaction log to obtain time series features; A feature fusion module, configured to perform cross-modal fusion of the semantic features, the visual features, and the temporal features to obtain fused features; A query rewriting module, configured to perform query rewriting on the text data according to the fusion features to obtain a query text, perform document retrieval according to the query text, and obtain a retrieval result set; The retrieval output module is used to determine the relevance between each retrieval result in the retrieval result set and the query text, obtain a relevance score set, and sort the retrieval result set according to the relevance score set to obtain a final retrieval result.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the document retrieval method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the document retrieval method according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Power document generation method, system and equipment based on sentence drawing retrieval and medium

    CN120996014A

  • Skin cancer risk assessment method and system combining vision and language model

    CN121639704A