Document retrieval method and device, electronic equipment and storage medium

By training the document embedding and document coding model of financial text, the problem of low accuracy in financial text retrieval is solved, and more efficient document retrieval is achieved.

CN120067055APending Publication Date: 2025-05-30SHENZHEN SECURITIES INFORMATION
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510102848.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

In financial text retrieval, due to the long text announcements, complex directory structure and similar semantics in a large number of texts, the search accuracy of only semantic matching is not high.

Method used

By obtaining sample document data, document embedding is performed to obtain document embedding features, and training the preset document encoding model based on these features to obtain the target document encoding model. Then, the preset document database is encoded based on the target document encoding model, and document search is performed in combination with the target search method.

Benefits of technology

Through the use of the target document coding model, the target search term and candidate document coding can be more accurately matched, and the search accuracy of financial text can be improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067055A_ABST
    Figure CN120067055A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a document retrieval method and device, electronic equipment and a storage medium, and belongs to the technical field of text processing. The method comprises the following steps: acquiring sample document data; performing document embedding on the sample document data to obtain document embedding features; training a preset document coding model according to the document embedding features and the sample document data to obtain a target document coding model; performing document coding on a preset document database based on the target document coding model to obtain candidate document codes; and obtaining a target retrieval formula, and performing document retrieval on a preset document database based on the target document coding model, the target retrieval formula and the candidate document codes. According to the embodiment of the invention, the retrieval accuracy of the financial text can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of text processing, and in particular, to a document retrieval method and apparatus, an electronic device, and a storage medium. Background Art

[0002] Document retrieval can be to retrieve a preset database or a preset external data source based on a given retrieval formula to obtain a target document. For example, in a financial scenario, document retrieval is performed on a financial document database based on a given title. Generally, semantic retrieval is performed based on the retrieval formula during document retrieval. However, due to the characteristics of long financial text announcements, complex directory structures, and similar semantics in a large amount of text, the accuracy of semantic matching retrieval between the retrieval formula and the target text is not high. Therefore, how to improve the retrieval accuracy of financial texts has become an urgent problem to be solved. Summary of the Invention

[0003] The main purpose of the embodiments of this application is to propose a document retrieval method and apparatus, an electronic device, and a storage medium, aiming to improve the retrieval accuracy of financial texts.

[0004] To achieve the above object, a first aspect of the embodiments of this application proposes a document retrieval method, and the method includes:

[0005] Obtain sample document data;

[0006] Perform document embedding on the sample document data to obtain document embedding features;

[0007] Train a preset document encoding model according to the document embedding features and the sample document data to obtain a target document encoding model;

[0008] Perform document encoding on a preset document database based on the target document encoding model to obtain candidate document encodings;

[0009] Obtain a target retrieval formula, and perform document retrieval on the preset document database based on the target document encoding model, the target retrieval formula, and the candidate document encodings.

[0010] In some embodiments, the performing document embedding on the sample document data to obtain document embedding features includes:

[0011] Perform document word segmentation on the sample document data to obtain a document word sequence;

[0012] Perform paragraph segmentation on the sample document data to obtain a document paragraph sequence;

[0013] Perform directory structure extraction on the sample document data to obtain document structure information;

[0014] Perform document embedding based on the document word sequence, the document paragraph sequence, and the document structure information to obtain the document embedding features.

[0015] In some embodiments, performing document embedding based on the document word sequence, the document paragraph sequence, and the document structure information to obtain the document embedding features includes:

[0016] Perform text embedding on the document word sequence to obtain text vector embeddings;

[0017] Perform directory tree position embedding on the document paragraph sequence according to the document structure information to obtain directory tree vector embeddings;

[0018] Perform index embedding on the document structure information based on a preset directory index encoder to obtain directory index embeddings;

[0019] Perform feature splicing according to the text vector embeddings, the directory tree vector embeddings, and the directory index embeddings to obtain the document embedding features.

[0020] In some embodiments, the document word sequence includes target words and the target positions of the target words in the document word sequence. Performing text embedding on the document word sequence to obtain text vector embeddings includes:

[0021] Perform word vector embedding on the target words to obtain target word vectors;

[0022] Perform position vector embedding on the target positions to obtain target position vectors;

[0023] Perform feature summation on the target word vectors and the target position vectors to obtain the text vector embeddings.

[0024] In some embodiments, training a preset document encoding model according to the document embedding features and the sample document data to obtain a target document encoding model includes:

[0025] Perform document masking processing on the document embedding features to obtain document mask features;

[0026] Perform document prediction restoration on the document mask features according to the preset document encoding model to obtain predicted document data;

[0027] Optimize the parameters of the preset document encoding model according to the predicted document data and the sample document data to obtain the target document encoding model.

[0028] In some embodiments, the preset document encoding model includes an encoding layer and a restoration prediction layer. The process of performing document prediction and restoration on the document mask feature according to the preset document encoding model to obtain predicted document data includes:

[0029] Performing document encoding on the document mask feature according to the encoding layer to obtain a document encoding feature;

[0030] Performing document restoration prediction on the document encoding feature according to the restoration prediction layer to obtain the predicted document data.

[0031] In some embodiments, the process of performing document retrieval on the preset document database based on the target document encoding model, the target retrieval formula, and the candidate document encoding includes:

[0032] Performing document embedding on the target retrieval formula to obtain a target embedding feature;

[0033] Performing document encoding on the retrieval formula embedding feature based on the target document encoding model to obtain a target encoding feature;

[0034] Performing similarity retrieval based on the target encoding feature and the candidate document encoding to obtain a target document encoding;

[0035] Filtering the preset document database according to the target document encoding to obtain a target retrieved document.

[0036] To achieve the above object, a second aspect of the embodiments of the present application proposes a document retrieval device, which includes:

[0037] A data acquisition module, configured to acquire sample document data;

[0038] A document embedding module, configured to perform document embedding on the sample document data to obtain a document embedding feature;

[0039] A model training module, configured to train a preset document encoding model according to the document embedding feature and the sample document data to obtain a target document encoding model;

[0040] A document encoding module, configured to perform document encoding on a preset document database based on the target document encoding model to obtain candidate document encodings;

[0041] A document retrieval module, configured to obtain a target retrieval formula and perform document retrieval on the preset document database based on the target document encoding model, the target retrieval formula, and the candidate document encoding.

[0042] To achieve the above object, a third aspect of the embodiments of the present application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the method described in the first aspect above is implemented.

[0043] To achieve the above object, a fourth aspect of the embodiments of the present application provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the method described in the first aspect above is implemented.

[0044] A document retrieval method, device, electronic device, and storage medium provided by the present application obtain sample document data, then perform document embedding on the sample document data to obtain document embedding features, and then train a preset document encoding model according to the document embedding features and the sample document data to obtain a target document encoding model that can well perceive the potential connection between the document and the encoding, so that the target document encoding model can mine the semantic and structural information in the document data, thereby transforming the semantic and structural information of the document into the encoding of the document; further, perform document encoding on a preset document database based on the target document encoding model to obtain candidate document encodings, then obtain a target retrieval formula, and perform document retrieval on the preset document database based on the target document encoding model, the target retrieval formula, and the candidate document encodings, so as to realize matching the target retrieval formula and the candidate document encodings based on the semantic and structural information of the document by the target document encoding model, and screening the matching information from the preset document database, so as to improve the retrieval accuracy of financial texts. Description of the Drawings

[0045] Figure 1 is a flowchart of the document retrieval method provided by the embodiments of the present application;

[0046] Figure 2 is Figure 1 a flowchart of step S102 in

[0047] Figure 3 is Figure 2 a flowchart of step S204 in

[0048] Figure 4 is Figure 3 a flowchart of step S301 in

[0049] Figure 5 is Figure 1 a flowchart of step S103 in

[0050] Figure 6 is Figure 5 a flowchart of step S502 in

[0051] Figure 7 Yes Figure 1 is the flowchart of step S105 in

[0052] Figure 8 is the structural schematic diagram of the document retrieval device provided by the embodiment of the present application;

[0053] Figure 9 is the hardware structural schematic diagram of the electronic device provided by the embodiment of the present application. Detailed implementation manners

[0054] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0055] It should be noted that although functional module division is performed in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order from the module division in the device or the order in the flowchart. Terms such as "first" and "second" in the specification, claims and the above-mentioned drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence.

[0056] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.

[0057] First, several nouns involved in the present application are analyzed:

[0058] Text embedding: It is a technology in natural language processing that converts text into numerical vectors, enabling computers to understand and process language data. Text embedding belongs to a branch of artificial intelligence, especially the cross-field of computer science and linguistics, and is widely regarded as one of the core technologies of computational linguistics. This technology can capture semantic and syntactic information in text by learning the dense vector representations of words, phrases, or documents in language, making text data available for various machine learning models. The applications of text embedding technology include but are not limited to sentiment analysis, text classification, machine translation, and information retrieval, etc. Through text embedding, machines can not only recognize the surface form of words but also understand the similarities and relationships between different words, thus greatly improving the performance of natural language processing systems. This technology also shows its importance and utility in many technical fields such as automatic summarization, chatbots, and recommendation systems.

[0059] Text masking: It is a technique in natural language processing used to hide or replace specific information in text data to support different processing tasks, such as language model training or privacy protection. Text masking belongs to a branch of artificial intelligence, especially in the intersection of computer science and linguistics, and is widely applied in various aspects of computational linguistics. This technique helps enhance the model's language understanding and generation capabilities by temporarily hiding words or phrases in the text and requiring the model to predict or reconstruct these masked parts in subsequent processing. When training deep learning models such as BERT, text masking is a core step that can improve the depth and accuracy of the model's language processing. In addition, text masking is also used to protect sensitive information and ensure compliance with privacy standards during text analysis and data sharing. In this way, text masking not only supports the development of machine learning models but also promotes ethical data processing.

[0060] Masked language model (MLM): It is a model in natural language processing used to train deep learning networks to predict randomly masked words in text through a masking mechanism. Masked language models belong to a branch of artificial intelligence, especially in the intersection of computer science and linguistics, and are widely applied in computational linguistics. The core of this model is that it does not require traditional sequence-to-sequence prediction. Instead, by randomly masking a certain proportion of words (usually 15%) in the input sequence and then training the model to predict these masked words, it promotes the ability to understand the context. Masked language models are a key component of pre-trained language processing models such as BERT (Bidirectional Encoder Representations from Transformers) and are widely used to improve the performance of applications such as machine translation, text summarization, question answering systems, and text generation. Through this method, the model can learn more profound language rules and complex semantic relationships, thus effectively improving the accuracy and efficiency of natural language processing tasks.

[0061] Multi-head self-attention mechanism: It is an advanced natural language processing technology used to improve the performance of models in processing and understanding text data. This mechanism belongs to an important branch in the field of artificial intelligence, especially in the interdisciplinary field of computer science and linguistics. The core of the multi-head self-attention mechanism lies in simultaneously processing different representation subspaces of data in parallel to capture various relationships and features in the text. By dispersing attention across multiple "heads", each "head" focuses on different parts of the input data, enabling a more detailed understanding of the semantic and syntactic structures of the text. Multi-head self-attention is a key component of the Transformer architecture and is widely used in advanced language models such as BERT and GPT, significantly improving the accuracy and efficiency of tasks such as machine translation, text summarization, question answering systems, and text generation. Through this mechanism, the model can better understand and process complex language phenomena, promoting the development of natural language processing technology.

[0062] Document retrieval can be performed by retrieving a preset database or a preset external data source based on a given retrieval formula to obtain the target document. For example, in a financial scenario, document retrieval is performed on a financial document database based on a given title. Usually, semantic retrieval is performed based on the retrieval formula during document retrieval. However, due to the characteristics of long financial text announcements, complex directory structures, and similar semantics in a large amount of text, the accuracy of semantic matching retrieval between the retrieval formula and the target text is not high. Therefore, how to improve the retrieval accuracy of financial texts has become an urgent problem to be solved.

[0063] Based on this, the embodiments of the present application provide a document retrieval method, device, electronic device, and storage medium, aiming to improve the retrieval accuracy of financial texts.

[0064] A document retrieval method, device, electronic device, and storage medium provided by the embodiments of the present application are specifically described through the following embodiments. First, the document retrieval method in the embodiments of the present application is described.

[0065] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results.

[0066] The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technologies, operation / interaction systems, mechatronics, etc. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0067] The document retrieval method provided by the embodiments of this application relates to the field of text processing technologies. The document retrieval method provided by the embodiments of this application can be applied to a terminal, can also be applied to a server side, or can be software running on a terminal or a server side. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, etc.; the server side can be configured as an independent physical server, can also be configured as a server cluster or a distributed system composed of multiple physical servers, or can be configured as a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application that implements the document retrieval method, etc., but is not limited to the above forms.

[0068] This application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet-type devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This application can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0069] It should be noted that in each specific embodiment of the present application, when it comes to relevant processing based on data related to the user's identity or characteristics, such as user information, user behavior data, user historical data, and user location information, the user's permission or consent will be obtained first. Moreover, the collection, use, and processing of these data will comply with relevant laws, regulations, and standards. In addition, when the embodiments of the present application need to obtain the user's sensitive personal information, the user's separate permission or separate consent will be obtained through methods such as pop-up windows or redirecting to a confirmation page. After clearly obtaining the user's separate permission or separate consent, the necessary user-related data for the normal operation of the embodiments of the present application will be obtained.

[0070] Figure 1 is an optional flowchart of the document retrieval method provided by the embodiments of the present application. Figure 1 The method in may include but is not limited to steps S101 to S105.

[0071] Step S101, obtain sample document data;

[0072] Step S102, perform document embedding on the sample document data to obtain document embedding features;

[0073] Step S103, train a preset document encoding model according to the document embedding features and the sample document data to obtain a target document encoding model;

[0074] Step S104, perform document encoding on a preset document database based on the target document encoding model to obtain candidate document encodings;

[0075] Step S105, obtain a target retrieval formula, and perform document retrieval on a preset document database based on the target document encoding model, the target retrieval formula, and the candidate document encodings.

[0076] Steps S101 to S105 illustrated in the embodiments of the present application, by obtaining sample document data, then performing document embedding on the sample document data to obtain document embedding features, and then training a preset document encoding model based on the document embedding features and the sample document data to obtain a target document encoding model that can well perceive the potential connection between the document and the encoding, so that the target document encoding model can mine the semantic and structural information in the document data, thereby transforming the semantic and structural information of the document into the encoding of the document; further, performing document encoding on a preset document database based on the target document encoding model to obtain candidate document encodings, then obtaining a target retrieval formula, and performing document retrieval on the preset document database based on the target document encoding model, the target retrieval formula, and the candidate document encodings, so as to realize matching the target retrieval formula and the candidate document encodings based on the semantic and structural information of the document by the target document encoding model, and screening the matched information from the preset document database, so as to improve the retrieval accuracy of financial texts.

[0077] In step S101 of some embodiments, the sample document data is a structure including multiple texts and corresponding directory trees. The document includes multiple chapters, sub-chapters, paragraphs, and items, and the directory tree represents the hierarchical relationship between these texts. For example, in the financial field, the sample document data is a report, and the directory tree includes chapters "Market Overview", "Risk Assessment", "Strategy", and "Data Analysis". There may be multiple sub-chapters or items under each chapter. For example, there may be two sub-chapters, "Global Trends" and "Industry Development Trends", under "Market Overview". Each sub-chapter contains detailed analysis texts below.

[0078] Please refer to Figure 2 , in some embodiments, step S102 may include but is not limited to steps S201 to S204:

[0079] Step S201, performing document word segmentation on the sample document data to obtain a document word sequence;

[0080] Step S202, performing paragraph segmentation on the sample document data to obtain a document paragraph sequence;

[0081] Step S203, extracting the directory structure from the sample document data to obtain document structure information;

[0082] Step S204, performing document embedding based on the document word sequence, the document paragraph sequence, and the document structure information to obtain document embedding features.

[0083] In the steps S201 to S204 illustrated in the embodiments of the present application, by performing document word segmentation on the sample document data, a document word sequence is obtained. At the same time, the sample document data is segmented into paragraphs to obtain a document paragraph sequence. At the same time, the table of contents structure of the sample document data is extracted to obtain document structure information. Finally, document embedding is performed based on the document word sequence, the document paragraph sequence, and the document structure information to obtain document embedding features, thereby realizing multi-dimensional mining of document content. By combining word, paragraph, and structure information, the document embedding features can more accurately express the semantics and hierarchical relationships of the document, improving the effect and efficiency in subsequent processing of complex documents.

[0084] In step S201 of some embodiments, document word segmentation refers to dividing the continuous text in the sample document into words or phrases. The document word sequence refers to the sequence of words obtained through the document word segmentation process, which reflects the order and structure of the words in the document. Each document word sequence consists of a series of segmented words arranged in the original order of the document.

[0085] In step S202 of some embodiments, paragraph segmentation refers to dividing the sample document into several paragraphs according to the logical structure or punctuation marks. Each paragraph usually contains one or more sentences. The document paragraph sequence refers to the sequence of paragraphs obtained through the paragraph segmentation process, and each paragraph is a sub-part of the document. The document paragraph sequence retains the order information of the paragraphs in the document, reflecting the structural hierarchy of the document content.

[0086] In step S203 of some embodiments, table of contents structure extraction refers to identifying and extracting the table of contents tree or the hierarchical structure of the table of contents from the sample document. The document structure information refers to the hierarchical structure data extracted from the document, including the order and relationship of the chapter titles, sub-chapter titles, paragraphs, etc. in the document. The document structure information reflects the organizational structure of the document presented in a tree structure, that is, the form of the table of contents tree, where each node represents a chapter or a sub-chapter, and the edge represents the relationship between different levels. For example, if the sample document is "1. Project Overview 1.1 Background Introduction 1.2 Objectives and Scope 2. Methodology 2.1 Research Design 2.2 Data Collection", the table of contents tree structure is "[“Project Overview”]→[“Background Introduction”, “Objectives and Scope”], [“Methodology”]→[“Research Design”, “Data Collection”]".

[0087] Please refer to Figure 3 , in some embodiments, step S204 may include but is not limited to steps S301 to S304:

[0088] Step S301, perform text embedding on the document word sequence to obtain text vector embedding;

[0089] Step S302: Embed the document paragraph sequence into the directory tree position according to the document structure information to obtain the directory tree vector embedding.

[0090] Step S303: Perform index embedding on the document structure information based on a preset directory index encoder to obtain the directory index embedding.

[0091] Step S304: Concatenate the features according to the text vector embedding, the directory tree vector embedding, and the directory index embedding to obtain the document embedding feature.

[0092] Steps S301 to S304 shown in the embodiments of the present application, by performing text embedding on the document word sequence to obtain the text vector embedding, at the same time embedding the document paragraph sequence into the directory tree position according to the document structure information to obtain the directory tree vector embedding, and performing index embedding on the document structure information based on a preset directory index encoder to obtain the directory index embedding, and finally concatenating the features according to the text vector embedding, the directory tree vector embedding, and the directory index embedding to obtain the document embedding feature, thereby realizing the comprehensive representation of the sample document. Through multi-dimensional embedding features, the semantic content and structural relationship of the document can be captured more accurately, improving the effect of the subsequent model in encoding the document.

[0093] Please refer to Figure 4 , in some embodiments, the document word sequence includes a target word and the target position of the target word in the document word sequence. Step S301 may include but is not limited to steps S401 to S403:

[0094] Step S401: Perform word vector embedding on the target word to obtain the target word vector.

[0095] Step S402: Perform position vector embedding on the target position to obtain the target position vector.

[0096] Step S403: Sum the features of the target word vector and the target position vector to obtain the text vector embedding.

[0097] Steps S401 to S403 shown in the embodiments of the present application, by performing word vector embedding on the target word to obtain the target word vector, at the same time performing position vector embedding on the target position to obtain the target position vector, and finally summing the features of the target word vector and the target position vector to obtain the text vector embedding, thereby realizing the efficient representation of the document word sequence. By combining word semantic information and position information, the text vector embedding can better express the context relationship and grammatical structure in the text.

[0098] In step S401 of some embodiments, the target word refers to the specific word to be embedded in the document word sequence. Word vector embedding is to map a word into a real number vector of a fixed dimension. The target word vector refers to the vector representation corresponding to the target word obtained through the word vector embedding process. In one embodiment, the preset document encoding model includes a word vector embedding model. The target word vector is input into the word vector embedding model to obtain the target word vector. And during the model training process of the preset document encoding model, the word vector embedding model will be trained together to improve the word embedding ability for the target word.

[0099] In step S402 of some embodiments, the target position refers to the specific position or index of the target word in the document word sequence, indicating the position order of the target word in the document and used to identify the relative position relationship of the word in the context. Position vector embedding is to map the target position into a vector of a fixed dimension. The target position vector is the vector representation corresponding to the target position obtained through the position vector embedding process, reflecting the position relationship of the target word relative to other words in the document. In one embodiment, the preset document encoding model includes a position vector embedding model. The target position is input into the position vector embedding model to obtain the target position vector. And during the model training process, the position vector embedding model will be trained together to improve the position embedding ability for the target position.

[0100] It should be noted that the data dimension of the target position vector is the same as that of the target word vector.

[0101] In step S403 of some embodiments, feature summation means adding multiple feature vectors element by element, that is, adding the target word vector and the target position vector element by element to obtain a new set of numerical element representations, that is, text vector embedding.

[0102] In step S302 of some embodiments, the table of contents tree position embedding means converting the table of contents tree of the text in the document into a vector representation. This vector describes the position of the text in the document, including hierarchical information such as chapters, sub-chapters, and paragraphs. The table of contents tree vector embedding means the vector representation obtained by aggregating multiple position vectors that have undergone the table of contents tree position embedding in the document structure information.

[0103] In one embodiment, generating a table of contents tree for a document paragraph sequence based on document structure information specifically includes obtaining the specific hierarchical relationship of a paragraph in the document structure information, and obtaining the specific positions of the beginning and end of the paragraph at their respective hierarchical levels. Finally, the hierarchical relationship and specific positions are aggregated to form a table of contents tree vector embedding. For example, the document structure information includes "1. General Introduction 1.1 Background Information 1.2 Project Objectives 2. Methodology 2.1 Data Analysis 2.2 Model Design 3. Result Analysis 3.1 Experimental Results 3.2 Data Interpretation 4. Discussion 5. Conclusion". For the first chapter in the document paragraph sequence, the fifth subsection under the first chapter, and the third paragraph under the sixth title in the fifth subsection, the beginning of the third paragraph is at 55% of the fifth subsection, and the end of the third paragraph is at 60% of the fifth subsection.

[0104] For the beginning of the paragraph, there are the following position values: chapter position: the first chapter out of a total of five chapters → 1 / 5 = 0.2, subsection position: the fifth subsection out of a total of ten subsections → 5 / 10 = 0.5, title position: the sixth title out of a total of six titles → 6 / 6 = 1, the beginning of the paragraph is at 55% of the fifth subsection → 0.55. The table of contents tree position embedding for the beginning of the paragraph is: [0.2, 0.5, 1, 0.55].

[0105] For the end of the paragraph, there are the following position values: chapter position: the first chapter out of a total of five chapters → 1 / 5 = 0.2, subsection position: the fifth subsection out of a total of ten subsections → 5 / 10 = 0.5, title position: the sixth title out of a total of six titles → 6 / 6 = 1, the beginning of the paragraph is at 60% of the fifth subsection → 0.60. The table of contents tree position embedding for the beginning of the paragraph is: [0.2, 0.5, 1, 0.60].

[0106] In step S303 of some embodiments, the preset table of contents index encoder is a pre-trained neural network model, which is composed of four layers of encoders and one layer of linear layer. The index embedding specifically means inputting the document structure information into the table of contents index encoder to obtain the numerical vector output by the linear layer in the table of contents index encoder, that is, the table of contents index embedding.

[0107] In step S304 of some embodiments, feature concatenation refers to concatenating multiple different feature vectors in sequence into a longer vector. By concatenating multiple feature vectors, different types of information are fused into a unified vector numerical representation, that is, the document embedding feature is obtained.

[0108] Please refer to Figure 5 , in some embodiments, step S103 includes but is not limited to steps S501 to S503:

[0109] Step S501, performing document masking processing on the document embedding feature to obtain a document mask feature;

[0110] Step S502: Perform document prediction restoration on the document mask feature according to the preset document encoding model to obtain predicted document data;

[0111] Step S503: Optimize the parameters of the preset document encoding model according to the predicted document data and the sample document data to obtain the target document encoding model.

[0112] Steps S501 to S503 shown in the embodiments of the present application, by performing document masking processing on the document embedding feature to obtain the document mask feature, then performing document prediction restoration on the document mask feature according to the preset document encoding model to obtain the predicted document data, and finally optimizing the parameters of the preset document encoding model according to the predicted document data and the sample document data to obtain the target document encoding model, successively enable the target document encoding model to have good encoding ability for the document embedding feature, and further enable the target encoding model to have the ability of document prediction restoration according to the encoding. Driven by the document restoration prediction, the target document encoding model discovers the potential associations between the document embedding feature and the encoding, as well as between the encoding and the document prediction restoration, that is, the target document encoding model can have good perception of the context relationship and structural information in the document embedding feature, so as to convert the document embedding feature into the encoding feature.

[0113] In step S501 of some embodiments, the document masking processing is to perform matrix dot multiplication on the preset mask matrix or a randomly generated mask matrix and the document embedding feature, and use the dot-multiplied numerical matrix as the document mask feature, thereby hiding some information in the document embedding feature, providing a data basis for the subsequent target document encoding model to encode based on incomplete information, enabling the target document encoding model to perform context reasoning and document structure detection based on incomplete information, and performing encoding to provide encoding data for subsequent document restoration prediction.

[0114] It should be noted that during the inference process of the target document encoding model, that is, during the application process, the target retrieval formula is embedded into a document to obtain the embedding representation of the target retrieval formula, and then the embedding representation of the target retrieval formula is encoded based on the target encoding model, thereby utilizing the ability of the target encoding model to perform context reasoning and document structure detection based on incomplete information, so that the embedding representation of the target retrieval formula, that is, the encoded representation of the target retrieval formula, has context information and document structure information after encoding. Thus, it solves the problem that in the financial scenario, due to the characteristics of long financial text announcements, complex directory structures, and similar semantics in a large amount of text, the accuracy of semantic matching retrieval between the retrieval formula and the target text is not high.

[0115] Please refer to Figure 6, in some embodiments, the preset document encoding model includes an encoding layer and a restoration prediction layer, and step S502 includes but is not limited to steps S601 to S602:

[0116] Step S601, perform document encoding on the document mask feature according to the encoding layer to obtain a document encoding feature;

[0117] Step S602, perform document restoration prediction on the document encoding feature according to the restoration prediction layer to obtain predicted document data.

[0118] Steps S601 to S602 illustrated in the embodiments of the present application perform document encoding on the document mask feature according to the encoding layer to obtain a document encoding feature, and then perform document restoration prediction on the document encoding feature according to the restoration prediction layer to obtain predicted document data, so that the target encoding model can reason and encode according to context information and structural information under incomplete information, and improve the encoding ability of the target encoding model driven by the task of document restoration prediction.

[0119] In step S601 of some embodiments, the encoding layer is stacked by 12 encoders, and each encoder includes two sub-layers. Among them, the first sub-layer is a multi-head attention sub-layer, and the second sub-layer is a feed-forward neural network layer. The hidden information in the document mask feature is inferred through the multi-head attention sub-layer, and feature mining is performed through the feed-forward neural network layer. Document encoding is to input the document mask feature into the encoding layer, and the output numerical matrix is the document encoding feature.

[0120] In step S602 of some embodiments, the restoration prediction layer is formed by combining a linear layer and a Softmax activation function, and is used to restore the document encoding feature from a numerical representation to a text representation. Document restoration prediction is to input the document encoding feature into the restoration prediction layer to obtain a document in text form, that is, predicted document data.

[0121] In step S503 of some embodiments, calculate the loss value according to the distribution probability of the words corresponding to each hidden position in the predicted document data and the words corresponding to the hidden positions in the sample document data, that is, based on the loss function calculation method of randomly masking some words in the input sequence in the Masked Language Model (MLM), and then predicting these masked words, so as to improve the encoding ability and restoration prediction ability in the target document encoding model.

[0122] In step S104 of some embodiments, the preset document database includes multiple documents. The preset document database is the database to be retrieved. Each document in the preset document database is input into the target document encoding model to obtain the document encoding corresponding to each document. The set of all document encodings is the candidate document encoding.

[0123] Please refer to Figure 7 , in some embodiments, step S105 may include but is not limited to steps S701 to S704:

[0124] Step S701: Perform document embedding on the target retrieval formula to obtain the target embedding feature;

[0125] Step S702: Based on the target document encoding model, perform document encoding on the retrieval formula embedding feature to obtain the target encoding feature;

[0126] Step S703: Perform similarity retrieval according to the target encoding feature and the candidate document encoding to obtain the target document encoding;

[0127] Step S704: Screen the preset document database according to the target document encoding to obtain the target retrieval document.

[0128] Steps S701 to S704 illustrated in the embodiments of the present application perform document embedding on the target retrieval formula to obtain the target embedding feature, then perform document encoding on the retrieval formula embedding feature based on the target document encoding model to obtain the target encoding feature, then perform similarity retrieval according to the target encoding feature and the candidate document encoding to obtain the target document encoding, and finally screen the preset document database according to the target document encoding to obtain the target retrieval document, so as to realize the ability of the target document encoding model to perform context speculation and structural information reasoning based on incomplete information to encode the target retrieval formula, that is, to supplement information for the target retrieval formula to obtain the target document encoding, and then perform similarity retrieval according to the target document encoding and the candidate document encoding, so as to solve the problem of low accuracy in semantic matching retrieval between the retrieval formula and the target text in the financial scenario due to the characteristics of long financial text announcements, complex directory structures, and similar semantics in a large amount of text.

[0129] In step S701 of some embodiments, the document embedding principle of performing document embedding on the target retrieval formula is the same as that of performing document embedding on the sample document data, which will not be elaborated here.

[0130] In step S702 of some embodiments, encoding the retrieval embedded features based on the target document encoding model means inputting the retrieval embedded features into the target document encoding model, and the encoding layer of the target document encoding model encodes the retrieval embedded features, and the obtained numerical matrix is the target encoding feature.

[0131] In step S703 of some embodiments, similarity retrieval refers to calculating the similarity according to the target encoding feature and the candidate document encoding, such as cosine similarity, Euclidean distance, comparing the target encoding feature with the candidate document encoding, so as to find the encoding of the document most similar to the target encoding, that is, the target document encoding.

[0132] In step S704 of some embodiments, screening the preset document data according to the target document encoding means obtaining the document corresponding to the target document encoding in the preset document database to obtain the target retrieval document.

[0133] Please refer to Figure 8 , the embodiments of the present application further provide a document retrieval device, which can implement the above document retrieval method. The device includes:

[0134] A data acquisition module 801, configured to acquire sample document data;

[0135] A document embedding module 802, configured to perform document embedding on the sample document data to obtain document embedding features;

[0136] A model training module 803, configured to train a preset document encoding model according to the document embedding features and the sample document data to obtain a target document encoding model;

[0137] A document encoding module 804, configured to perform document encoding on a preset document database based on the target document encoding model to obtain candidate document encodings;

[0138] A document retrieval module 805, configured to obtain a target retrieval formula, and perform document retrieval on a preset document database based on the target document encoding model, the target retrieval formula, and the candidate document encoding.

[0139] The specific implementation manner of this document retrieval device is basically the same as the specific embodiments of the above document retrieval method, and will not be repeated here.

[0140] The embodiments of the present application further provide an electronic device. The electronic device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the above document retrieval method is implemented. The electronic device can be any intelligent terminal including a tablet computer, an in-vehicle computer, etc.

[0141] Please refer to Figure 9 , Figure 9Schematically shows the hardware structure of an electronic device according to another embodiment. The electronic device includes:

[0142] A processor 901, which can be implemented in ways such as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application;

[0143] A memory 902, which can be implemented in forms such as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 902 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 902 and are called by the processor 901 to execute the document retrieval method of the embodiments of the present application;

[0144] An input / output interface 903, which is used to implement information input and output;

[0145] A communication interface 904, which is used to implement communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.);

[0146] A bus 905, which transmits information between various components of the device (such as the processor 901, the memory 902, the input / output interface 903, and the communication interface 904);

[0147] Among them, the processor 901, the memory 902, the input / output interface 903, and the communication interface 904 are communicatively connected to each other inside the device through the bus 905.

[0148] The embodiments of the present application also provide a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above-mentioned document retrieval method is implemented.

[0149] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include memories remotely provided with respect to the processor, and these remote memories can be connected to the processor through a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0150] The document retrieval method, document retrieval device, electronic device, and storage medium provided by the embodiments of the present application obtain sample document data, then perform document embedding on the sample document data to obtain document embedding features, and then train a preset document encoding model according to the document embedding features and the sample document data to obtain a target document encoding model that can well perceive the potential connection between the document and the encoding, so that the target document encoding model can mine the semantic and structural information in the document data, thereby transforming the semantic and structural information of the document into the encoding of the document; further, perform document encoding on a preset document database based on the target document encoding model to obtain candidate document encodings, then obtain a target retrieval formula, and perform document retrieval on the preset document database based on the target document encoding model, the target retrieval formula, and the candidate document encodings, so as to match the target retrieval formula and the candidate document encodings based on the semantic and structural information of the document by the target document encoding model, and screen the information obtained by the matching from the preset document database, so as to improve the retrieval accuracy of financial texts.

[0151] The embodiments described in the embodiments of the present application are for more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art can know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems.

[0152] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than those shown in the figures, or combine certain steps, or different steps.

[0153] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0154] Those of ordinary skill in the art will understand that all or some of the steps in the methods disclosed above, and the functional modules / units in systems and devices, can be implemented as software, firmware, hardware, or a suitable combination thereof.

[0155] As used in the specification of this application and the above-mentioned drawings, the terms "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of this application described here can be implemented in an order different from those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0156] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or a similar expression means any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0157] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the above-mentioned unit division is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces, and the indirect coupling or communication connection of devices or units can be in an electrical, mechanical, or other form.

[0158] The units described above as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0159] In addition, the functional units in various embodiments of the present application may be integrated into one processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0160] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present application. The aforementioned storage medium includes: various media that can store programs such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.

[0161] The preferred embodiments of the embodiments of the present application have been described above with reference to the accompanying drawings. However, this does not limit the scope of the rights of the embodiments of the present application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall fall within the scope of the rights of the embodiments of the present application.

Claims

1. A document retrieval method, characterized in that: The method comprises: Get sample document data; Performing document embedding on the sample document data to obtain document embedding features; Training a preset document encoding model according to the document embedding features and the sample document data to obtain a target document encoding model; Performing document encoding on a preset document database based on the target document encoding model to obtain candidate document encodings; A target search formula is obtained, and a document search is performed on the preset document database based on the target document encoding model, the target search formula, and the candidate document encoding.

2. The method according to claim 1, characterized in that The step of embedding the sample document data to obtain document embedding features includes: Performing document segmentation on the sample document data to obtain a document word sequence; Performing paragraph segmentation on the sample document data to obtain a document paragraph sequence; Extracting the directory structure of the sample document data to obtain document structure information; Document embedding is performed according to the document word sequence, the document paragraph sequence and the document structure information to obtain the document embedding feature.

3. The method according to claim 2, characterized in that The step of embedding the document according to the document word sequence, the document paragraph sequence and the document structure information to obtain the document embedding feature includes: Performing text embedding on the document word sequence to obtain a text vector embedding; Perform directory tree position embedding on the document paragraph sequence according to the document structure information to obtain directory tree vector embedding; Index embedding the document structure information based on a preset directory index encoder to obtain a directory index embedding; Feature concatenation is performed according to the text vector embedding, the directory tree vector embedding and the directory index embedding to obtain the document embedding feature.

4. The method according to claim 3, characterized in that The document word sequence includes a target word and a target position of the target word in the document word sequence, and the text embedding of the document word sequence to obtain a text vector embedding includes: Embedding the target word into a word vector to obtain a target word vector; Performing position vector embedding on the target position to obtain a target position vector; The target word vector and the target position vector are feature summed to obtain the text vector embedding.

5. The method according to claim 1, characterized in that The step of training a preset document encoding model according to the document embedding feature and the sample document data to obtain a target document encoding model includes: Performing document masking processing on the document embedding feature to obtain a document mask feature; Performing document prediction and restoration on the document mask feature according to the preset document encoding model to obtain predicted document data; The parameters of the preset document coding model are optimized according to the predicted document data and the sample document data to obtain the target document coding model.

6. The method according to claim 5, characterized in that The preset document coding model includes a coding layer and a restoration prediction layer, and the document mask feature is subjected to document prediction restoration according to the preset document coding model to obtain predicted document data, including: Performing document encoding on the document mask feature according to the encoding layer to obtain a document encoding feature; Document restoration prediction is performed on the document encoding features according to the restoration prediction layer to obtain the predicted document data.

7. The method according to any one of claims 1 to 6, characterized in that: The performing document retrieval on the preset document database based on the target document encoding model, the target search formula, and the candidate document encoding includes: Performing document embedding on the target search formula to obtain a target embedding feature; Performing document encoding on the search formula embedding feature based on the target document encoding model to obtain a target encoding feature; Perform similarity retrieval based on the target coding feature and the candidate document coding to obtain the target document coding; The preset document database is screened according to the target document code to obtain a target retrieval document.

8. A document retrieval device, characterized in that: The device comprises: A data acquisition module is used to acquire sample document data; A document embedding module, used to embed the sample document data into a document to obtain document embedding features; A model training module, used to train a preset document encoding model according to the document embedding features and the sample document data to obtain a target document encoding model; A document encoding module, used to perform document encoding on a preset document database based on the target document encoding model to obtain candidate document encodings; The document retrieval module is used to obtain a target retrieval formula and perform document retrieval on the preset document database based on the target document encoding model, the target retrieval formula, and the candidate document encoding.

9. An electronic device, characterized in that: The electronic device comprises a memory and a processor, the memory stores a computer program, and the processor implements the document retrieval method according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the document retrieval method according to any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Intelligent document coding system and method based on rule driving

    CN121809403A