Knowledge recombination retrieval method based on unstructured document
By preprocessing and deep semantic coding of unstructured documents, combined with entity recognition, relationship extraction and deep learning models, the problem of inaccurate search results in the existing technology is solved, and efficient and accurate document retrieval and classification are achieved.
Patent Information
- Application Number
- CN202510434778.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-04-08
AI Technical Summary
When processing unstructured documents, it is difficult for the prior art to capture the deep semantic relationships and knowledge structures in the document, resulting in insufficient accuracy and correlation of the search results and inability to meet complex needs.
By obtaining unstructured document data, preprocessing and structure, entity recognition and relationship extraction tools are used to extract entity information and relational words, combined with BERT pre-trained model and graph neural network for deep semantic coding, and finally classification and retrieval through naive Bayes model.
It realizes efficient classification and retrieval of unstructured documents, improves the accuracy and relevance of search results, can quickly respond to user needs, and significantly improves user experience and retrieval efficiency.
Smart Images

Figure CN120179809A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of knowledge recombination retrieval, and particularly to a knowledge recombination retrieval method based on unstructured documents. Background Art
[0002] In today's digital age, the processing and retrieval of unstructured document data have become important challenges faced by various industries. With the explosive growth of information volume, traditional text processing and retrieval methods are difficult to meet the requirements of high efficiency and accuracy.
[0003] In addition, when existing retrieval technologies process unstructured documents, they often rely on keyword matching or simple semantic analysis, and it is difficult to capture the deep semantic relationships and knowledge structures in the documents. This results in insufficient accuracy and relevance of retrieval results, and cannot provide users with a satisfactory retrieval experience. Especially in the legal field, the complexity and professionalism of documents further increase the difficulty of retrieval, and traditional retrieval methods are difficult to quickly and accurately find documents highly relevant to users' needs.
[0004] Therefore, there is an urgent need for a technical solution that can effectively process unstructured documents, achieve knowledge recombination, and improve retrieval efficiency and accuracy to meet the complex needs in practical applications. Summary of the Invention
[0005] Aiming at the deficiencies of the prior art, the present invention provides a knowledge recombination retrieval method based on unstructured documents, which at least solves the problems of large retrieval difficulty, inaccurate output results, and low relevance in the prior art.
[0006] To achieve the above objectives, the present invention is realized through the following technical solutions: A knowledge recombination retrieval method based on unstructured documents, characterized by comprising: Step 1: Obtain unstructured document data, preprocess the unstructured document data, and extract text data for analysis to form structured text; Step 2: Identify each structured text through an entity recognition tool, and output an entity information text marked with entity types and lexical relationships; Step 3: Analyze each entity information text through a relationship extraction tool, extract the entity types in the entity information text and the relationship words between different entity types, and form a relationship analysis text; Step 4: Process the relationship analysis text through a BERT pre-trained model to obtain an embedding vector corresponding to each relationship analysis text; Step 5: Input the embedding vector for analysis through a graph neural network, and output an embedding representation vector; Step 6: Analyze the embedded representation vectors through the Naive Bayes model, output the posterior probability and class probability of each relationship analysis text, and classify and store all unstructured document data; Step 7: After inputting the retrieval prompt words in the legal field, the retrieval prompt words form retrieval document data. Through the data processing and analysis in Steps 1 to 6, obtain the posterior probability and class probability of the corresponding retrieval document data, and filter out the unstructured document data that meets the requirements according to the posterior probability and class probability as the retrieval result of this retrieval.
[0007] In the preferred solution of the above knowledge recombination retrieval method based on unstructured documents, perform text extraction on each file in the unstructured document data through OCR technology, convert the text in the document and picture data into text data, and then clean the text data through a text editor or an automated cleaning script group to generate standardized text; use a word segmentation tool to split the words in the standardized text into several words to generate structured text.
[0008] In the preferred solution of the above knowledge recombination retrieval method based on unstructured documents, input each entity information text into the relationship extraction model, extract the entity types in each text and the relationship words between different entity types, and form a relationship analysis text with the extracted entity types and corresponding relationship words.
[0009] In the preferred solution of the above knowledge recombination retrieval method based on unstructured documents, input the relationship analysis text into the BERT pre-trained model, map the entity types and lexical relationships in the relationship analysis text to the vocabulary of the model, corresponding to unique identifiers or indexes, to form an integer sequence T={t1,t2,…,tn}, input the integer sequence T into the embedding layer of the BERT pre-trained model to obtain the initial embedding vector E={e1,e2,…,en}, input the initial embedding vector E into the Transformer encoder of the BERT pre-trained model to output the attention score vector; input the attention score vector into the feed-forward neural network, and through non-linearly transforming the attention score vector, obtain the attention transformation vector; perform residual connection and layer normalization on the attention score vector and the attention transformation vector to obtain the embedding vector.
[0010] In the preferred solution of the above knowledge recombination retrieval method based on unstructured documents, the method for obtaining the attention score vector is: Input the initial embedding vector E into each Transformer encoder. For the input initial embedding vector E, through three different linear transformations, query vector Q, key vector K, and value vector V are respectively generated. The dimensions of the three vectors are the same, denoted as the key vector dimension dk. The formulas for calculating query vector Q, key vector K, and value vector V based on the initial embedding vector E are as follows: ; ; ; Among them, Q h is the query vector generated by the h th encoder; K h is the key vector generated by the h th encoder; V h is the value vector generated by the h th encoder; h is the serial number of the encoder; W Q h is the weight matrix of the query vector of the h th encoder; W K h is the weight matrix of the key vector of the h th encoder; W V h is the weight matrix of the value vector of the h th encoder; Calculate the attention scores of each query vector and all key vectors through each encoder. The formula is as follows: ; Among them, ZYL h is the attention score output by the h th encoder; Concatenate the output results of all encoders and obtain the attention score vector through a linear transformation. The formula is: .
[0011] In the preferred scheme of the above knowledge reorganization retrieval method based on unstructured documents, the calculation method of the attention conversion vector is: Input the attention score vector into the feed-forward neural network, and through non-linear transformation of the attention score vector, obtain the attention conversion vector. The formula is as follows: ; Among them, W1 is the conversion weight two, W2 is the conversion weight three, b1 is the bias parameter one, and b2 is the bias parameter two.
[0012] In a preferred embodiment of the above knowledge recombination retrieval method based on unstructured documents, the formula for obtaining the embedding vector is as follows: ; In a preferred embodiment of the above knowledge recombination retrieval method based on unstructured documents, the formula for calculating the embedding representation vector is as follows: ; Among them, QR is the embedding representation vector; H i is the weight coefficient of the embedding vector, i represents the serial number of the feature in the embedding vector, and the value ranges from 1 to n; n is a positive integer, and σ is the activation function.
[0013] In a preferred embodiment of the above knowledge recombination retrieval method based on unstructured documents, the calculation methods of the posterior probability and the class probability are as follows: Calculate the likelihood probability of all entity information texts' embedding representation vectors under the given prior probability. Specifically, the likelihood probability can be decomposed into the product of the probabilities of each feature in the embedding representation vector. The formula is as follows: ; Among them, QRi is the i-th feature in the embedding representation vector QR, and n is the number of features; Calculate the posterior probability according to Bayes' theorem. The formula is as follows: ; After traversing all posterior probabilities, the formula for selecting the maximum posterior probability as the class probability value is as follows: ; Select the class corresponding to the class probability value as the classification result, and classify and store all entity information texts according to the corresponding classification result C QR to form a classification database.
[0014] In a preferred embodiment of the above knowledge recombination retrieval method based on unstructured documents, select the corresponding classification database according to the class probability of the retrieved document data, traverse and compare the posterior probability with all posterior probabilities in this classification database, and screen out several posterior probabilities whose similarity index meets the requirements; The method for determining whether the similarity index meets the requirements is: Calculate the similarity between the posterior probability corresponding to the retrieved document data and each posterior probability in the classification database. The unstructured document data corresponding to the posterior probability whose similarity exceeds the similarity threshold meets the requirements.
[0015] The present invention provides a knowledge recombination retrieval method based on unstructured documents, which has the following beneficial effects: (1) Obtain unstructured document data, preprocess the unstructured document data, and extract the text data for analysis to form structured text, which can convert unstructured documents in various formats, such as PDFs, images, etc., into a unified structured text format, improving the usability of the data and the efficiency of subsequent analysis, laying a solid foundation for subsequent entity recognition and relationship extraction, and avoiding the problems that traditional unstructured document processing methods are difficult to efficiently convert a large amount of unstructured data into a structured form for analysis, resulting in low data utilization and low analysis efficiency.
[0016] (2) Identify the entity information in the structured text through an entity recognition tool, and analyze the relationships between entities through a relationship extraction tool. By combining entity recognition and relationship extraction, key information in the text can be more comprehensively extracted. The entity recognition tool can accurately label entity types and lexical relationships, while the relationship extraction tool further analyzes the relationships between entities to form a complete relationship analysis text. This combination not only improves the accuracy of information extraction but also better captures the semantic relationships in the document, providing rich information for subsequent knowledge graph construction and semantic understanding, and overcoming the problems in the prior art that entity recognition and relationship extraction are carried out independently and lack effective integration of the correlation between the two, resulting in incomplete and inaccurate information extraction.
[0017] (3) Obtain the embedding vector of the relationship analysis text through the BERT pre-training model, then analyze it using a graph neural network, and finally classify it through a naive Bayes model, which can perform deep semantic encoding on the relationship analysis text, obtain high-quality embedding vectors, and capture the semantic features in the text. The graph neural network further analyzes the embedding vectors and outputs an embedding representation vector, enhancing the understanding of the complex relationships between entities. The naive Bayes model classifies based on the embedding representation vector and outputs posterior probabilities and class probabilities, realizing efficient classification storage of unstructured documents. The comprehensive application of these deep learning models makes full use of the advantages of each model, significantly improving the accuracy and efficiency of text analysis and classification, and avoiding the limitations of traditional text analysis methods in dealing with complex semantic relationships and large-scale data, which are difficult to meet the requirements of deep semantic understanding and efficient classification.
[0018] (4) After entering the retrieval prompt words in the legal field, the posterior probability and category probability of the corresponding retrieved document data are obtained through data processing and analysis, and the unstructured document data that meets the requirements is screened out according to these probabilities. By comprehensively applying various advanced technologies in the foregoing steps, in-depth semantic understanding and analysis of the retrieval prompt words can be carried out. After obtaining the posterior probability and category probability of the retrieved document data, according to the set threshold or ranking strategy, the unstructured document data that meets the user's needs can be accurately screened out. This method not only improves the accuracy and relevance of the retrieval results, but also can quickly respond to the user's retrieval request, greatly improving the user experience and retrieval efficiency; it avoids the problem that when dealing with unstructured documents, relying only on keyword matching or simple semantic analysis makes it difficult to accurately screen out the documents highly relevant to the user's needs, resulting in insufficient accuracy and relevance of the retrieval results. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 It is a schematic diagram of the steps of a knowledge recombination retrieval method based on unstructured documents according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0020] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0021] Embodiment 1:
[0022] Please refer to Figure 1 , the present invention provides a knowledge recombination retrieval method based on unstructured documents, including: Step 1: Obtain unstructured document data, preprocess the unstructured document data, and extract the text data for analysis to form structured text.
[0023] Step 101: Obtain unstructured documents related to the legal field, including but not limited to laws and regulations, case judgments, contract texts, legal research reports, etc. For example, data on laws and regulations can be obtained from official government platforms such as the website of the National People's Congress of China, the website of the Chinese government, and the database of laws and regulations of the Ministry of Justice, or from commercial databases such as Pkulaw. Case judgments can be obtained from the China Judgments Online, the China Court Trial Live Network, etc., or from open data platforms and APIs, academic and industry resources such as the libraries of law schools. Contract texts and templates can be obtained from public resource platforms such as the SEC EDGAR database, the China Securities Regulatory Commission, or the Internet public contract library. Legal research reports and literature can be obtained from academic databases such as CNKI or open access platforms such as Google Scholar.
[0024] Step 102: Use OCR technology to extract text from each file in the unstructured document data, convert the text in the document and image data into editable text data, such as text files (TXT) or editable word / pdf files, and then clean the text data through a text editor or an automated cleaning script group. This mainly includes deleting garbled characters, correcting incorrect line breaks by merging redundant line breaks in paragraphs, replacing OCR recognition errors, and deleting meaningless repeated paragraphs, etc., to generate standardized text.
[0025] In this step, using OCR technology to extract text from unstructured document data can convert the text in pictures or scanned copies into editable text data, such as TXT files or editable Word / PDF files. This not only improves the processability of the data but also facilitates subsequent text cleaning and analysis. For example, through OCR technology, scanned legal regulations documents, case judgments, etc. can be converted into editable text, which is convenient for further processing and analysis, solving the problem that a large amount of text data in unstructured documents exists in the form of pictures or scanned copies, and traditional methods are difficult to directly extract and analyze this data. Especially in the legal field, many historical documents, scanned copies, etc. are stored in image form and cannot be directly processed for text. By cleaning the text data through a text editor or an automated cleaning script, this cleaning process can significantly improve the quality of the text data, reduce the interference of noisy data on subsequent analysis, improve the processing efficiency and accuracy, and avoid situations where the text data extracted from unstructured documents often has problems such as garbled characters, incorrect line breaks, OCR recognition errors, and repeated paragraphs, which affect the subsequent text analysis and processing efficiency.
[0026] Step 103: Use existing word segmentation tools such as the jieba library and the PKUSEG library to split the text in the standardized text into several words. Specifically, after splitting each word in the standardized text, a corresponding processed structured text will be generated.
[0027] In this step, splitting the text in the standardized text into several words by the word segmentation tool can accurately split the text in the standardized text into several words and generate the processed structured text. This not only provides a basis for subsequent entity recognition and relationship extraction, but also can better capture the professional terms and semantic information in legal texts, improving the accuracy and efficiency of analysis. It can solve the problem that in the legal field, there are many professional terms and complex sentence patterns in the text, and traditional word segmentation tools may not be able to accurately process these special words and sentence patterns, resulting in poor word segmentation effects and affecting subsequent entity recognition and relationship extraction.
[0028] Step 2: Use an entity recognition tool to recognize each structured text and output an entity information text marked with entity types and lexical relationships. Step 201: Train the entity recognition NER model tool. Specifically, select the BERT deep learning model. Before model training, add the words in the legal-related fields to the model's vocabulary. By expanding its vocabulary to include more legal terms, the expanded vocabulary can include the words in the legal dictionary and the legal words and lexical relationships in the legal knowledge base for pre-training. Then, input the legal text data marked with entity types and lexical relationships. For example, in the legal text data, in the sentence "Plaintiff Zhang San sues Defendant Li Si for infringement", mark "Zhang San" as "plaintiff", "Li Si" as "defendant", and the relationship between Zhang San and Li Si as "litigation" into the BERT deep learning model for NER task training, and then obtain the trained entity recognition NER model tool.
[0029] In this step, by selecting the BERT deep learning model and expanding its vocabulary before training to add legal-related domain vocabulary (such as vocabulary in legal dictionaries and legal vocabulary and vocabulary relationships in legal knowledge bases), the model can better understand and identify professional terms and entity relationships in legal texts. Training with legal text data annotated with entity types and vocabulary relationships further improves the accuracy and generalization ability of the model. For example, in the sentence "Plaintiff Zhang San sues Defendant Li Si for infringement", the model can accurately label "Zhang San" as the "plaintiff", "Li Si" as the "defendant", and identify the "litigation" relationship between them. This training method enables the entity recognition tool to adapt to the special needs of the legal field, provides accurate entity information for subsequent relationship extraction and knowledge graph construction, and solves the problem that existing entity recognition tools have a low recognition accuracy when dealing with professional terms and complex texts in the legal field and cannot meet the accurate recognition requirements for professional entities (such as case parties, court names, legal provisions, etc.) in legal documents. In addition, the entity relationships in legal texts are complex and diverse, and it is difficult for traditional entity recognition tools to capture these relationships.
[0030] It should be noted that the entity types can be case parties, court names, legal provisions, judgment results, etc.; for example, case parties such as plaintiffs and defendants; the court name is the name of the court that makes the judgment; the legal provision can be the legal provision used as the basis for adjudication; the judgment result is the final judgment result, such as the compensation amount is RMB XXX yuan; the relationship between entity types can be that there is a "litigation" relationship between the plaintiff and the defendant, and there is a "citation" relationship between the judgment result and the legal provision, etc.
[0031] Step 202: Input the entity information text of all unstructured document data into the entity recognition NER model tool, and the entity recognition NER model tool outputs the entity information text annotated with entity types and vocabulary relationships corresponding to each entity information text.
[0032] In this step, by applying the trained entity recognition NER model tool to all unstructured document data, it is possible to automatically and efficiently analyze each entity information text, and output the entity information text marked with entity types and lexical relationships. This automated processing method greatly improves the efficiency and accuracy of information extraction, can process large-scale legal document data, and provides rich entity information for subsequent knowledge recombination and retrieval. For example, for a large number of case judgments, the model can quickly identify entities such as case parties, court names, legal provisions, and judgment results, and mark the relationships between them, such as the "citation" relationship. This not only saves the time and cost of manual annotation, but also ensures the accuracy and integrity of entity information, providing reliable data support for subsequent analysis and retrieval; it avoids the problem that traditional methods are difficult to efficiently extract and annotate entity information when dealing with large-scale unstructured document data, resulting in incomplete and inaccurate information extraction and inability to meet the needs of rapid analysis of a large number of legal documents in practical applications.
[0033] Step 3: Analyze each entity information text through a relationship extraction tool, extract the entity types in the entity information text and the relationship words between different entity types, and form a relationship analysis text.
[0034] Step 301: Train the relationship extraction model; first extract relevant entity type vocabulary, relationship vocabulary, and concept vocabulary from the legal knowledge base and incorporate them into the model vocabulary to enhance the model's understanding and extraction ability of legal knowledge. This may include keywords in legal provisions, legal principles, legal procedures, etc. For example, add legal action words such as "prosecute", "judge", "compensate", legal entity words such as "plaintiff", "defendant", "court", and relationship words such as "according to", "result in", "involve"; use legal text data with relationship annotations, such as judgments, contracts, legal consultation records, etc., covering various legal fields and case types; perform preprocessing operations such as cleaning, word segmentation, and part-of-speech tagging on the annotated data to make it meet the requirements of model input, and divide the prepared dataset into a training set, a validation set, and a test set, and the ratio can be 7:2:1; input the data into a CNN convolutional neural network for training and testing to obtain a relationship extraction model.
[0035] In this step, the model's understanding and extraction capabilities of legal knowledge are enhanced by extracting relevant entity types, lexical relationships, and conceptual vocabulary from the legal knowledge base and incorporating them into the model vocabulary. Legal text data annotated with inclusion relationships is used for training, covering various legal fields and case types to ensure the generalization ability of the model. Data preprocessing operations such as cleaning, word segmentation, and part-of-speech tagging make the data meet the model input requirements. The dataset is divided into a training set, a validation set, and a test set in a ratio of 7:2:1 to ensure the sufficiency and reliability of model training. A CNN convolutional neural network is used for training and testing to obtain a relationship extraction model that can accurately extract entity relationships in legal texts. For example, the model can accurately identify the "basis" relationship in the text "According to Article 107 of the Contract Law" and the "sue" relationship in the text "Plaintiff Zhang San sues Defendant Li Si", significantly improving the accuracy and efficiency of relationship extraction; overcoming the problems of inaccurate and incomplete extraction when existing relationship extraction tools handle complex relationships in the legal field. The relationships in legal texts are diverse and complex, such as the "citation" relationship between legal provisions and judgment results, and the "litigation" relationship between the plaintiff and the defendant. Traditional methods are difficult to effectively capture these relationships.
[0036] Step 302: Input each entity information text into the relationship extraction model, extract the entity types in each text and the relationship words between different entity types, and form a relationship analysis text with the extracted entity types and corresponding relationship words.
[0037] In this step, by applying the trained relationship extraction model to each entity information text, the entity types and the relationship words between them in the text can be automatically and efficiently extracted to form a complete relationship analysis text. This automated processing method greatly improves the efficiency of relationship extraction, can process large-scale legal document data, and provides rich relationship information for subsequent knowledge graph construction and semantic understanding. For example, for a judgment document containing multiple case parties, the model can accurately extract the entity types of each party (such as plaintiff, defendant) and the relationships between them (such as litigation relationship), and integrate this information into a relationship analysis text, providing strong support for subsequent legal analysis and retrieval; overcoming the difficulty of traditional methods in efficiently extracting entity types and the relationship words between them when processing large-scale legal documents, resulting in incomplete and inaccurate relationship analysis and unable to meet the needs of in-depth analysis of legal documents in practical applications.
[0038] Step Four: Process the relationship analysis text through the BERT pre-trained model to obtain the embedding vector corresponding to each relationship analysis text.
[0039] Step 401: Construct a BERT pre-trained model containing an embedding layer and multiple Transformer encoders.
[0040] In this step, by constructing a BERT pre-training model that includes an embedding layer and multiple Transformer encoders, it is possible to perform deep semantic encoding on the relationship analysis text, capturing semantic features and context information in the text. This model structure can effectively handle professional terms and complex sentence patterns in the legal field, improving the understanding and analysis ability of legal texts.
[0041] Step 402: Add special tokens to the relationship analysis text, such as the start token ([CLS]), end token ([SEP]), padding token ([PAD]), etc., so that the model can correctly identify the start and end of a sentence and align texts of different lengths during batch processing. This can be achieved through string operations or list operations in a programming language, such as Python.
[0042] In this step, by adding special tokens such as the start token [CLS], end token [SEP], and padding token [PAD], the model can correctly identify the start and end of a sentence and align texts of different lengths during batch processing. This tokenization method ensures a unified input format for the model, improving the processing efficiency and accuracy of the model, and avoiding problems where the model has difficulty correctly identifying the start and end of a sentence and aligning texts of different lengths during batch processing.
[0043] Step 403: Generate a custom vocabulary by counting relevant words and concepts extracted from the legal knowledge base. For example, use the Counter class in the collections library to count word frequencies, and then generate a custom vocabulary sorted by word frequencies as the vocabulary for the BERT pre-training model. Input the relationship analysis text with special tokens added into the BERT pre-training model, and then map the entity types and word relationships in the relationship analysis text to the vocabulary of the model through the BertTokenizer in the transformers library, corresponding to unique identifiers or indices, forming an integer sequence T={t1,t2,…,tn}, where tn represents the nth word and n is the maximum number of words.
[0044] In this step, by generating a custom vocabulary and using it as the vocabulary of the BERT pre-trained model, the model can better understand and process the professional terms and concepts in legal texts. For example, use the Counter class in the collections library to count word frequencies, and then generate a custom vocabulary according to the sorted word frequencies to ensure that the model can accurately map the entity types and lexical relationships in the relationship analysis text to the model's vocabulary, avoiding the problem that the vocabulary of the existing pre-trained model may not contain professional terms and concept words in the legal field, resulting in the model being unable to accurately understand the semantics when processing legal texts. And by converting the relationship analysis text into an integer sequence, the model can directly process this sequence data. This conversion method not only preserves the semantic information of the text but also enables the model to perform batch processing efficiently.
[0045] Step 404: Input the integer sequence T into the embedding layer of the BERT pre-trained model. The embedding layer converts each word into a word embedding vector according to the input integer sequence T, such as using Word2Vec, etc., and then adds a position embedding vector to each word embedding vector. The position embedding can be generated in a deterministic way such as using sine and cosine functions, and then looked up and added according to the position index of the word in the sentence; add the word embedding vector and the position embedding vector element by element to obtain the initial embedding vector E = {e1, e2, …, en}, where en represents the nth vector, which is used as the initial input of the encoder of the Transformer model.
[0046] In this step, each word is converted into a word embedding vector through the embedding layer, and a position embedding vector is added to generate the initial embedding vector. This embedding method not only captures the semantic information of the word but also takes into account the position information of the word in the sentence, improving the model's ability to understand the text.
[0047] Step 405: Input the initial embedding vector E into each Transformer encoder layer. For the input initial embedding vector E, through three different linear transformations, query vector Q, key vector K, and value vector V are respectively generated. The dimensions of the three vectors are the same, denoted as the key vector dimension dk; the formulas for calculating the query vector Q, key vector K, and value vector V based on the initial embedding vector E are as follows: ; ; ; where, Q h is the query vector generated by the h th encoder; K h is the hThe key vectors generated by an encoder; V h is the h value vectors generated by an encoder; h is the serial number of the encoder; W Q h is the h weight matrix of the query vector of the W K h is the h weight matrix of the key vector of the W V h is the h weight matrix of the value vector of the
[0048] It should be noted that the initial values of the weight matrices are random. Usually, methods such as Xavier initialization or He initialization are used to set the range of the initial values according to the input and output dimensions of the weight matrices.
[0049] In this step, by generating query vectors, key vectors, and value vectors, the model can calculate the attention scores between words and capture the correlations between them. This mechanism enables the model to focus on the important information in the text and improve its ability to understand and analyze the text.
[0050] Step 406: Calculate the attention scores of each query vector with all key vectors through each encoder, according to the following formula: ; where, ZYL h is the h attention score vector output by the
[0051] It should be noted that the attention scores represent the correlations or matching degrees between the query vectors and each key vector. This score reflects the influence degree of other words on the current word when encoding it. To prevent the values from being too large during exponential operations, the dot product result is usually divided by the square root of the key vector dimension for scaling.
[0052] In this step, the attention scores reflect the correlations between the query vectors and each key vector, helping the model determine the influence degree of other words on the current word when encoding it. By scaling the dot product result, the problem of values being too large during exponential operations is avoided, improving the numerical stability of the model and solving the problem of how to quantify the correlations between words so that the model can better understand the semantic structure of the text.
[0053] Step 407: Concatenate the output results of all encoders and obtain the attention score vector through linear transformation. The formula is as follows: ; where Wo is the transformation weight matrix. Usually, methods such as Xavier initialization or He initialization are used to set the range of the initial value according to the input and output dimensions of the weight matrix, so as to ensure that the network can stably propagate gradients at the beginning of training and avoid the problems of gradient disappearance or explosion.
[0054] In this step, by concatenating the output results of all encoders and performing linear transformation, the model can generate a comprehensive attention score vector. This integration method fully utilizes the information of multiple encoders, improving the expression ability and accuracy of the model.
[0055] Step 408: Input the attention score vector into the feed-forward neural network, and obtain the attention transformation vector through non-linear transformation of the attention score vector. The formula is as follows: ; where, W1 is the transformation weight two, W2 is the transformation weight three, b1 is the bias parameter one, and b2 is the bias parameter two.
[0056] It should be noted that usually, methods such as Xavier initialization or He initialization are used to initialize the transformation weight two and the transformation weight three; the bias parameter one and the bias parameter two are usually initialized as zero vectors because the update of the bias term mainly depends on the data itself during training, and the initial value of zero will not have a negative impact on the training of the model.
[0057] Through the non-linear transformation of the attention score vector by the feed-forward neural network, the model can generate the attention transformation vector, further enhancing the ability to understand and analyze the text semantics. Perform residual connection and layer normalization on the attention score vector and the attention transformation vector to obtain the embedding vector. The formula is as follows: ; In this step, residual connection and layer normalization help to stabilize the training process of the model, accelerate the convergence speed, and improve the generalization ability of the model. By performing residual connection and layer normalization on the attention score vector and the attention transformation vector, the model can better capture the semantic information in the text, generate high-quality embedding vectors, accelerate the convergence speed, and improve the generalization ability of the model.
[0058] Step Five: Input the embedding vector into the graph neural network for analysis and output the embedding representation vector.
[0059] Specifically: The embedded vector EZ {EZ1, EZ2, …, EZn} is input into the graph neural network, and the embedded representation vector is output. The formula is as follows: ; where QR is the embedded representation vector; H i is the weight coefficient of the embedded vector, i represents the serial number of the feature in the embedded vector, with values ranging from 1 to n; n is a positive integer, and σ is the activation function.
[0060] In this solution, by representing the embedded vector as a graph structure, the relationships between entities can be intuitively displayed, such as the litigation relationships between the parties in a case, the citation relationships between legal provisions and judgment results, etc. This graph structure representation not only enriches the form of knowledge representation but also provides a basis for subsequent relationship analysis and knowledge reasoning; it overcomes the problem that traditional methods are difficult to effectively capture the complex relationships between entities when dealing with unstructured documents in the legal field, resulting in incomplete and inaccurate knowledge representation; by using an activation function (such as ReLU), nonlinearity is introduced in the update process of node embedded representation, and the complex relationships between nodes can be better captured. This nonlinear transformation improves the expressive power and generalization ability of the model, enabling the model to more accurately represent the complex knowledge structure in the legal field.
[0061] Step Six: Analyze the embedded representation vector through the Naive Bayes model, output the posterior probability and classification result of each relationship analysis text, and classify and store all unstructured document data.
[0062] Step 601: Define the category variable C, which is used to represent the category of the document. In the legal document classification task, the category can be "contract dispute", "traffic accident liability dispute", "intellectual property dispute", etc., which are represented by C1, C2, and C3 respectively, and there can be several, denoted as Cn.
[0063] In this step, by defining the category variable C, the category of the document is clearly represented, such as "contract dispute", "traffic accident liability dispute", "intellectual property dispute", etc., providing a clear category identifier for subsequent classification analysis and avoiding the problem of inaccurate classification results caused by the lack of a clear category definition and representation method in the legal document classification task.
[0064] Step 602: Calculate the prior probability of each category; that is, use the document data of the defined categories to train the Naive Bayes model and calculate the occurrence frequency of the stable data of each category used. For example, if there are 100 documents in the document data, and 30 of them belong to the "contract dispute" category, then P(C1 = contract dispute) = 0.3.
[0065] In this step, the Naive Bayes model is trained using the document data of defined categories, and the prior probability of each category is calculated. For example, if there are 100 documents in the document data and 30 of them belong to the "contract dispute" category, then P(C1 = contract dispute) = 0.3. This method of calculating the prior probability based on data can more accurately reflect the distribution of categories in the data, providing reliable prior information for subsequent classification and avoiding the problem that existing classification methods often rely on simple frequency statistics when calculating the prior probability and lack in-depth analysis of the data distribution.
[0066] Step 603: Calculate the likelihood probability of all entity information texts' embedding representation vectors under the given prior probability, specifically: The likelihood probability can be decomposed into the product of the probabilities of each feature in the embedding representation vector, and the formula is as follows: ; where QRi is the i-th feature in the embedding representation vector QR, and n is the number of features.
[0067] By decomposing the likelihood probability into the product of the probabilities of each feature in the embedding representation vector, the contribution of each feature to classification can be calculated more precisely. This decomposition method not only improves the calculation efficiency but also better captures the independence assumption between features, improving the classification accuracy of the model and avoiding the problem that traditional methods are difficult to effectively decompose and calculate the probabilities of each feature in the embedding representation vector when calculating the likelihood probability, resulting in inaccuracies.
[0068] Step 604: Calculate the posterior probability according to Bayes' theorem, and the formula is as follows: ; Through Bayes' theorem, the prior probability and the likelihood probability are combined to calculate the posterior probability of each category. This method can comprehensively consider the influence of prior knowledge and observed data, improving the accuracy and reliability of the classification results. Existing classification methods often ignore the combined influence of the prior probability and the likelihood probability when calculating the posterior probability, resulting in inaccurate classification results.
[0069] Step 605: After traversing all the posterior probabilities, select the maximum posterior probability as the category probability value, and select the category corresponding to the category probability value as the classification result. The formula is as follows: ; By traversing all the posterior probabilities, selecting the maximum value as the category probability value, and selecting the corresponding category as the classification result. This decision rule based on the maximum posterior probability can ensure the optimality of the classification results and improve the classification performance of the model; it overcomes the problem of lacking an effective decision rule to determine the final classification result in multi-category classification tasks.
[0070] Step 606: Classify and store all entity information texts according to the corresponding classification result C QR to form a classified database.
[0071] By classifying and storing all entity information texts according to the classification result, a structured classified database is formed. This classification and storage method not only improves the document management efficiency but also provides convenience for subsequent retrieval and analysis. For example, users can quickly locate the required documents by category, greatly improving the retrieval efficiency and solving the problems of chaotic document management and low retrieval efficiency in the existing methods when dealing with large-scale legal documents due to the lack of an efficient classification and storage mechanism.
[0072] Step Seven: After inputting the retrieval prompt words in the legal field, the retrieval prompt words form retrieval document data. Through the data processing and analysis in Steps One to Six, the posterior probability and classification result corresponding to the retrieval document data are obtained. According to the posterior probability and classification result, the unstructured document data that meets the requirements is screened out as the retrieval result of this retrieval.
[0073] Step 701: When inputting, use the legal language format for the retrieval prompt words. For example, "Plaintiff Zhang San sues Defendant Li Si for arrears of wages, the amount is 20,000, the court is xx court, query similar cases". After inputting the retrieval prompt words in the legal field, the retrieval prompt words form editable retrieval document data.
[0074] In this step, by converting the retrieval prompt words in the legal field into editable retrieval document data, it is ensured that the prompt words can be processed by the subsequent model. For example, when the user inputs "Plaintiff Zhang San sues Defendant Li Si for arrears of wages, the amount is 20,000, the court is xx court, query similar cases", the system converts it into structured retrieval document data, providing a basis for subsequent analysis and processing, converting the user's natural language query into a format that can be processed by the model, and making the retrieval result more accurate.
[0075] Step 702: Through the data processing and analysis in Steps One to Seven, the posterior probability and category probability corresponding to the retrieval document data are obtained.
[0076] In this step, the existing retrieval methods cannot make full use of the embedding vectors and classification information generated in the previous steps, resulting in insufficient relevance of the retrieval results. This solution ensures that the retrieval document data can be comprehensively analyzed through a complete data processing process from preprocessing, entity recognition, relationship extraction to embedding vector generation and classification. This enables the system to accurately understand the semantics and classification of the retrieval prompt words and provides an accurate basis for subsequent screening of similar documents.
[0077] Step 703: Select the corresponding classification database according to the category probability in Step 702, traverse and compare the posterior probability corresponding to the retrieved document data with all the posterior probabilities in this classification database, and screen out several posterior probabilities whose similarity index meets the requirements; the method for judging whether the similarity index meets the requirements is as follows: Calculate the similarity between the posterior probability corresponding to the retrieved document data and each posterior probability in the classification database through the well-known cosine similarity. The value range should be set in [0, 1]. The similarity threshold standard can be set to 0.8, that is, the unstructured document data corresponding to the posterior probability with a similarity exceeding 0.8 meets the standard. The similarity threshold can be adjusted according to the number of documents that meet the standard after retrieval. For example, if too many retrieved documents are output, the similarity threshold standard can be increased to 0.9.
[0078] In this step, by selecting the corresponding classification database and using the cosine similarity to calculate the similarity between the retrieved document and the documents in the database, the system can screen out the documents highly relevant to the retrieval prompt words. For example, setting the similarity threshold to 0.8, only the documents with a similarity exceeding 0.8 will be selected, ensuring the relevance and accuracy of the retrieval results. When screening similar documents by traditional methods, there is a lack of effective similarity calculation and screening mechanisms, resulting in inaccurate retrieval results.
[0079] Step 704: Use the unstructured document data corresponding to the posterior probability that meets the requirements as the output result of this retrieval.
[0080] In this step, by using the unstructured document data corresponding to the posterior probability that meets the requirements as the output result, the system can provide the user with directly available retrieval results. Such results not only accurately match the user's retrieval needs but also can be presented in a user-friendly manner, improving the user experience and retrieval efficiency. However, existing retrieval systems are difficult to convert the calculation results of the model into retrieval results understandable by users, resulting in a poor user experience.
[0081] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraint conditions of the technical solution.
[0082] The unit described as a separation component may or may not be physically separated. The component displayed as a unit may or may not be a physical unit, that is, it may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0083] As described above, the foregoing is only a specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application.
Claims
1. A knowledge reorganization retrieval method based on unstructured documents, characterized in that: include: Step 1: Obtain unstructured document data, preprocess the unstructured document data, and extract text data for analysis to form structured text; Step 2: Use entity recognition tools to identify each structured text and output entity information text annotated with entity types and vocabulary relationships; Step 3: Analyze each entity information text through the relationship extraction tool, extract the entity types and relationship words between different entity types in the entity information text, and form a relationship analysis text; Step 4: Process the relationship analysis text through the BERT pre-training model to obtain the embedding vector corresponding to each relationship analysis text; Step 5: Analyze the embedded vector input through the graph neural network and output the embedded representation vector; Step 6: Analyze the embedded representation vector through the naive Bayes model, output the posterior probability and category probability of each relational analysis text, and classify and store all unstructured document data; Step 7: After inputting the search prompt words in the legal field, the search prompt words form the search document data. Through the data processing and analysis of steps 1 to 6, the posterior probability and category probability of the corresponding search document data are obtained. According to the posterior probability and category probability, the unstructured document data that meets the requirements are screened out as the search results of this search.
2. The knowledge reorganization retrieval method based on unstructured documents according to claim 1 is characterized in that: In step one, text is extracted from each file in the unstructured document data using OCR technology, and the text in the document and image data is converted into text data. The text data is then cleaned using a text editor or an automated cleaning script group to generate standardized text. A word segmentation tool is used to split the text in the standardized text into several words to generate structured text.
3. The knowledge reorganization retrieval method based on unstructured documents according to claim 2 is characterized in that: In step three, each entity information text is input into the relation extraction model, the entity type in each text and the relation words between different entity types are extracted, and the extracted entity types and corresponding relation words are combined to form a relation analysis text.
4. The method for knowledge reorganization and retrieval based on unstructured documents according to claim 3 is characterized in that: In step 4, the relational analysis text is input into the BERT pre-trained model, and the entity types and lexical relations in the relational analysis text are mapped into the vocabulary of the model, corresponding to unique identifiers or indexes, to form an integer sequence T={t1,t2,…,tn}, and the integer sequence T is input into the embedding layer of the BERT pre-trained model to obtain the initial embedding vector E={e1,e2,…,en}, and the initial embedding vector E is input into the Transformer encoder of the BERT pre-trained model to output the attention score vector; the attention score vector is input into the feedforward neural network, and the attention conversion vector is obtained by performing a nonlinear transformation on the attention score vector; the attention score vector and the attention conversion vector are residually connected and layer-normalized to obtain the embedding vector.
5. The knowledge reorganization retrieval method based on unstructured documents according to claim 4 is characterized in that: The method for obtaining the attention score vector is: The initial embedding vector E is input into each Transformer encoder. For the input initial embedding vector E, three different linear transformations are used to generate the query vector Q, key vector K and value vector V respectively. The dimensions of the three vectors are the same, which is recorded as the key vector dimension dk. The formulas for calculating the query vector Q, key vector K and value vector V respectively through the initial embedding vector E are as follows: ; ; ; in, Q h It is h The query vector generated by the encoder; K h It is h The key vector generated by the encoder; V h It is h A vector of values generated by the encoder; h is the serial number of the encoder; W Q h It is h The weight matrix of the query vector for each encoder; W K h It is h The weight matrix of the key vectors of the encoders; W V h It is h The weight matrix of the encoder value vector; The attention score of each query vector and all key vectors is calculated by each encoder according to the following formula: ; in, ZL h It is h The attention scores output by the encoder; The output results of all encoders are concatenated and the attention score vector is obtained through linear transformation. The formula is: 。 6. The method for knowledge reorganization and retrieval based on unstructured documents according to claim 5, characterized in that: The attention transformation vector is calculated as: The attention score vector is input into the feedforward neural network, and the attention conversion vector is obtained by performing a nonlinear transformation on the attention score vector. The formula is as follows: ; in, W1 is the conversion weight two, W2 is the conversion weight three, b1 is the bias parameter one, and b2 is the bias parameter two.
7. The method for knowledge reorganization and retrieval based on unstructured documents according to claim 6, characterized in that: The formula for obtaining the embedding vector is as follows: 。 8. The method for knowledge reorganization and retrieval based on unstructured documents according to claim 7, characterized in that: The formula for calculating the embedding representation vector is as follows: ; Where QR is the embedding representation vector; H i is the weight coefficient of the embedding vector, i It represents the sequence number of the feature in the embedding vector, and its value ranges from 1 to n; n is a positive integer, and σ is the activation function.
9. The method for knowledge reorganization and retrieval based on unstructured documents according to claim 8, characterized in that: The calculation method of the posterior probability and class probability is: Calculate the likelihood probability of the embedded representation vector of all entity information texts under a given prior probability. Specifically, the likelihood probability can be decomposed into the probability product of each feature in the embedded representation vector. The formula is as follows: ; Where QRi is the i-th feature in the embedding representation vector QR, and n is the number of features; The posterior probability is calculated according to Bayes' theorem, and the formula is as follows: ; After traversing all posterior probabilities, the formula for selecting the maximum posterior probability as the category probability value is as follows: ; Select the category corresponding to the category probability value as the classification result, and classify all entity information texts according to the corresponding classification result C QR Carry out classified storage to form a classified database.
10. The knowledge reorganization retrieval method based on unstructured documents according to claim 9, characterized in that: In step 7, a corresponding classification database is selected according to the category probability of the retrieved document data, the posterior probability is traversed and compared with all posterior probabilities in the classification database, and several posterior probabilities whose similarity index meets the requirements are screened out; The method to determine whether the similarity index meets the requirements is: The similarity between the posterior probability corresponding to the retrieved document data and each posterior probability in the classification database is calculated, and the unstructured document data corresponding to the posterior probability whose similarity exceeds the similarity threshold meets the requirements.
Citation Information
Patent Citations
Knowledge graph construction method based on reliability of aircraft parts
CN116644192A
Information retrieval query method for legal data service platform
CN118981512A
Bayesian visual interactive search
US20170091319A1
Methods, apparatus, and systems for transforming unstructured natural language information into structured computer- processable data
WO2019050968A1
Knowledge graph-based question and answer method and apparatus, and storage medium
WO2021051558A1