Document noise processing method and device based on large language model, equipment and storage medium
By building a multi-source noise library and dynamically injecting controllable noise documents, combining sparse search and dense search algorithm scoring, identifying and eliminating trap documents, and optimizing document queue input, the problem of low document information processing efficiency in large language models in a noisy environment is solved, and the reliability and accuracy of generated results are improved.
Patent Information
- Application Number
- CN202510512522.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-07-11
AI Technical Summary
The existing search-enhanced generation system based on large language models is insufficient in the face of problems such as noise interference, difficult negative sample concealment misleading, and inefficient utilization of context position sensitivity, and difficult to effectively process document information in complex scenarios.
By building a multi-source noise library, dynamically inject controllable noise documents, combining sparse search and dense search algorithms for scoring, using a binary classification model to identify trap documents, and grouping and reordering documents to optimize document queue input to improve model robustness and efficiency.
It improves the document information processing efficiency of large language models in noisy environments, improves the reliability and accuracy of generated results, and improves the user experience.
Smart Images

Figure CN120296149A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and particularly relates to a method, device, equipment and storage medium for document noise processing based on a large language model. Background Art
[0002] Currently, the Retrieval-Augmented Generation (RAG) technology has significantly improved the accuracy and reliability of knowledge-intensive tasks by combining information retrieval and language model generation capabilities, and is widely used in scenarios such as open-domain question answering, real-time report generation, medical diagnosis assistance, etc.
[0003] However, with the complexity of application scenarios and the enhancement of knowledge dynamics, traditional RAG systems face the following key bottlenecks:
[0004] First, the contradiction between noise interference and model robustness: Existing systems usually directly input the retrieved documents, but irrelevant or misleading content (noise) will significantly reduce the generation quality. For example, in medical question answering, retrieving a fragment of a guideline that is related to the query keyword but has outdated content may lead the model to output incorrect conclusions. In addition, although completely deleting noisy documents can reduce interference, excessive cleaning of the context will cause the attention distribution of the language model to be too concentrated (entropy collapse), which will instead reduce its ability to understand complex semantics.
[0005] Second, the concealment and misguidance of difficult negative samples: Traditional retrieval systems are difficult to distinguish documents that are "semantically related but lack answers" (difficult negative samples). For example, when querying "the color of Napoleon's horse", a document about "the color of Napoleon's wife's horse" is retrieved. Such a document is highly relevant to the query but contains the wrong answer and is extremely likely to mislead the model generation.
[0006] Third, existing filtering methods (such as fixed threshold truncation) cannot effectively identify such pitfalls, resulting in insufficient reliability of the generated results.
[0007] Fourth, the inefficient utilization of context position sensitivity: Research shows that language models pay significantly more attention to the beginning and end positions of the context than the middle position (U-shaped attention distribution), but existing RAG systems usually arrange documents in descending order of retrieval scores, which may place key information in inefficient areas. For example, when highly relevant documents are concentrated in the middle section, the model may ignore their content. Although some studies have tried to alleviate this problem by modifying the model structure (such as optimizing the attention mechanism), such solutions require retraining the model, have poor compatibility, and high deployment costs.
[0008] As can be seen from the above, how to improve the processing efficiency of document information in the process of document noise processing based on a large language model is an urgent problem to be solved at present. Summary of the Invention
[0009] In view of this, the purpose of the present invention is to provide a method, device, equipment and storage medium for document noise processing based on a large language model, which can improve the efficiency of processing document information during the document noise processing based on the large language model. The specific scheme is as follows:
[0010] In a first aspect, the present application provides a method for document noise processing based on a large language model, including:
[0011] Performing probability distribution processing on the output result obtained after the large language model processes the document to be recognized, obtaining a confidence entropy value, and determining whether the confidence entropy value is greater than a preset entropy value threshold. If the confidence entropy value is not greater than the preset entropy value threshold, determining the number of documents based on the difference between the confidence entropy value and a preset target entropy value, and selecting the number of noise documents from a multi-source noise library; the multi-source noise library includes randomly generated text, cross-domain text, and semantic interference text;
[0012] Inserting the number of noise documents into the middle position of the context of the document to be recognized according to a preset insertion rule, obtaining a document to be processed, and scoring the document to be processed using a scoring model including a sparse retrieval algorithm and a dense retrieval algorithm. Then, determining whether the scoring result is greater than a preset scoring threshold. If it is greater, processing the document to be processed using a preset binary classification model to obtain a trap document probability; the trap document probability is the probability indicating that the document has a semantic trap;
[0013] Determining whether the trap document probability is greater than a preset probability threshold. If it is not greater, performing grouped determination on the document to be processed based on the scoring result to obtain a relevant group determination result, and determining a target queue based on the relevant group determination result and an initial queue, so as to input the target queue into the large language model for processing.
[0014] Optionally, determining the number of documents based on the difference between the confidence entropy value and a preset target entropy value, and selecting the number of noise documents from a multi-source noise library includes:
[0015] Determining the number of documents based on the entropy gain corresponding to the document to be recognized and the difference between the confidence entropy value and a preset target entropy value; the magnitude of the difference is positively correlated with the magnitude of the number of documents;
[0016] Generating meaningless text using a Markov chain, setting the meaningless text as randomly generated text, then collecting cross-domain text irrelevant to the target domain from websites using a preset collection rule, and generating semantic interference text related to the target keyword but with incorrect content using a preset text rewriting technique;
[0017] A multi-source noise library is established based on the randomly generated text, the cross-domain text, and the semantic interference text, and the number of noise documents equal to the number of the documents is selected from the multi-source noise library.
[0018] Optionally, scoring the to-be-processed document by using a scoring model including a sparse retrieval algorithm and a dense retrieval algorithm includes:
[0019] Determine whether the number of query words corresponding to the to-be-processed document is greater than a preset query word threshold. If the number of query words corresponding to the to-be-processed document is greater than the preset query word threshold, use the dense retrieval algorithm in the scoring model to score the to-be-processed document to obtain a corresponding scoring result; the dense retrieval algorithm is used to perform semantic matching scoring on the to-be-processed document;
[0020] If the number of query words corresponding to the to-be-processed document is not greater than the preset query word threshold, use the sparse retrieval algorithm in the scoring model to score the to-be-processed document to obtain a corresponding scoring result; the sparse retrieval algorithm is used to perform keyword matching scoring on the to-be-processed document.
[0021] Optionally, determine whether the scoring result is greater than a preset scoring threshold. If it is greater, use a preset binary classification model to process the to-be-processed document to obtain a trap document probability, including:
[0022] Determine whether the scoring result is greater than a preset scoring threshold. If the scoring result is greater than the preset scoring threshold, use the preset binary classification model to identify the lexical overlap degree of the to-be-processed document to obtain a corresponding overlap degree result; the overlap degree result is the overlap degree between the to-be-processed document and the query statement corresponding to the to-be-processed document;
[0023] Use an NER tool to identify entity conflicts in the to-be-processed document to obtain a corresponding conflict identification result; the conflict identification result is used to represent whether the query entity corresponding to the to-be-processed document does not match the attribute value;
[0024] Use a preset pattern matching rule to perform an answer missing marking operation on the to-be-processed document to obtain a corresponding missing marking result; the missing marking result is used to represent whether the to-be-processed document includes a preset keyword;
[0025] Determine the trap document probability based on the overlap degree result, the conflict identification result, and the missing marking result.
[0026] Optionally, determine whether the trap document probability is greater than a preset probability threshold. If it is not greater, perform a grouping determination on the to-be-processed document based on the scoring result to obtain a relevant group determination result, including:
[0027] Determine whether the probability of the trap document is greater than a preset probability threshold. If the probability of the trap document is greater than the preset probability threshold, mark the document to be processed corresponding to the probability of the trap document as a trap document and eliminate the trap document;
[0028] If the probability of the trap document is not greater than the preset probability threshold, determine whether the scoring result meets the preset high - relevant group conditions. If the scoring result meets the preset high - relevant group conditions, set the document to be processed corresponding to the scoring result as a high - relevant group document;
[0029] If the scoring result does not meet the preset high - relevant group conditions, determine whether the scoring result meets the preset medium - relevant group conditions. If the scoring result meets the preset medium - relevant group conditions, set the document to be processed corresponding to the scoring result as a medium - relevant group document;
[0030] If the scoring result does not meet the preset medium - relevant group conditions, set the document to be processed corresponding to the scoring result as a low - relevant group document.
[0031] Optionally, determining the target queue based on the relevant group determination result and the initial queue to input the target queue into the large language model for processing includes:
[0032] If the document to be processed is a high - relevant group document, determine whether the number of documents corresponding to the high - relevant group document is less than a preset document number threshold. If the number of documents corresponding to the high - relevant group document is not less than the preset document number threshold, allocate each of the high - relevant group documents to the initial queue according to the preset descending score alternating allocation rule to obtain the target queue, and input the target queue into the large language model for processing;
[0033] Or, determine whether the maximum scoring result among all the high - relevant group documents is greater than a preset score threshold. If the maximum scoring result among all the high - relevant group documents is not greater than the preset score threshold, allocate a preset percentage of the high - relevant group documents among all the high - relevant group documents to the end of the initial queue to obtain the target queue, and input the target queue into the large language model for processing.
[0034] Optionally, determining the target queue based on the relevant group determination result and the initial queue to input the target queue into the large language model for processing includes:
[0035] If the document to be processed is a medium-related group document, distribute each of the medium-related group documents to the initial queue corresponding to the large language model according to a preset uniform distribution rule to obtain a to-be-processed queue, and insert preset noise documents into the to-be-processed queue at a preset interval to obtain a target queue, so as to input the target queue into the large language model for processing;
[0036] If the document to be processed is a low-related group document, insert each of the low-related group documents into the initial queue according to a preset random insertion rule to obtain a target queue, so as to input the target queue into the large language model for processing.
[0037] In a second aspect, the present application provides a document noise processing device based on a large language model, including:
[0038] A noise document selection module, configured to perform probability distribution processing on the output result obtained after the large language model processes the document to be recognized, obtain a confidence entropy value, and determine whether the confidence entropy value is greater than a preset entropy value threshold. If the confidence entropy value is not greater than the preset entropy value threshold, determine the number of documents based on the difference between the confidence entropy value and a preset target entropy value, and select the number of noise documents from a multi-source noise library; the multi-source noise library includes randomly generated text, cross-domain text, and semantic interference text;
[0039] A trap document probability determination module, configured to insert the number of noise documents into the middle position of the context of the document to be recognized according to a preset insertion rule to obtain a to-be-processed document, and use a scoring model including a sparse retrieval algorithm and a dense retrieval algorithm to score the to-be-processed document, and then determine whether the scoring result is greater than a preset scoring threshold. If it is greater, use a preset binary classification model to process the to-be-processed document to obtain a trap document probability; the trap document probability is the probability that the document has a semantic trap;
[0040] A target queue determination module, configured to determine whether the trap document probability is greater than a preset probability threshold. If it is not greater, perform grouping determination on the to-be-processed document based on the scoring result to obtain a relevant group determination result, and determine a target queue based on the relevant group determination result and the initial queue, so as to input the target queue into the large language model for processing.
[0041] In a third aspect, the present application provides an electronic device, including:
[0042] A memory, configured to store a computer program;
[0043] A processor, configured to execute the computer program to implement the foregoing document noise processing method based on a large language model.
[0044] Fourthly, the present application provides a computer-readable storage medium for storing a computer program, wherein when the computer program is executed by a processor, the foregoing method for processing document noise based on a large language model is implemented.
[0045] As can be seen from the above, before performing document noise processing based on a large language model in the present application, it is necessary to perform probability distribution processing on the output result obtained after the large language model processes the document to be recognized, obtain a confidence entropy value, and determine whether the confidence entropy value is greater than a preset entropy threshold value. If the confidence entropy value is not greater than the preset entropy threshold value, the number of documents is determined based on the difference between the confidence entropy value and the preset target entropy value, and the number of noise documents is selected from the multi-source noise library; the multi-source noise library includes randomly generated text, cross-domain text, and semantic interference text; the number of noise documents is inserted into the middle position of the context of the document to be recognized according to a preset insertion rule to obtain a document to be processed, and a scoring model including a sparse retrieval algorithm and a dense retrieval algorithm is used to score the document to be processed, and then it is determined whether the scoring result is greater than a preset scoring threshold value. If it is greater, a preset binary classification model is used to process the document to be processed to obtain a trap document probability; it is determined whether the trap document probability is greater than a preset probability threshold value. If it is not greater, the document to be processed is grouped and determined based on the scoring result to obtain a relevant group determination result, and the target queue input to the large language model is determined based on the relevant group determination result and the initial queue.
[0046] Thus, it can be seen that the present application first needs to perform probability distribution processing on the output result obtained after the large language model processes the document to be recognized, obtain a confidence entropy value, and determine whether the confidence entropy value is greater than a preset entropy threshold value. If the confidence entropy value is not greater than the preset entropy threshold value, the number of documents is determined based on the difference between the confidence entropy value and the preset target entropy value, and the number of noise documents is selected from the multi-source noise library; subsequently, the number of noise documents is inserted into the middle position of the context of the document to be recognized according to a preset insertion rule to obtain a document to be processed, and a scoring model including a sparse retrieval algorithm and a dense retrieval algorithm is used to score the document to be processed, and then it is determined whether the scoring result is greater than a preset scoring threshold value. If it is greater, a preset binary classification model is used to process the document to be processed to obtain a trap document probability; finally, it is determined whether the trap document probability is greater than a preset probability threshold value. If it is not greater, the document to be processed is grouped and determined based on the scoring result to obtain a relevant group determination result, and the target queue input to the large language model is determined based on the relevant group determination result and the initial queue. In this way, the processing efficiency of document information is improved during the process of document noise processing based on a large language model, thereby enhancing the user experience. Description of the Drawings
[0047] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to the provided drawings.
[0048] Figure 1 Flowchart of a method for processing document noise based on a large language model disclosed in the present application;
[0049] Figure 2 Schematic structural diagram of a device for processing document noise based on a large language model disclosed in the present application;
[0050] Figure 3 Structural diagram of an electronic device disclosed in the present application. Specific embodiments
[0051] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0052] Currently, retrieval-augmented generation technology has significantly improved the accuracy and reliability of knowledge-intensive tasks by combining information retrieval and language model generation capabilities, and is widely used in scenarios such as open-domain question answering, real-time report generation, and medical diagnosis assistance. However, with the complexity of application scenarios and the enhancement of knowledge dynamics, traditional RAG systems face the following key bottlenecks: the contradiction between noise interference and model robustness, the hidden misguidance of difficult negative samples, and the inability of existing filtering methods (such as fixed-threshold truncation) to effectively identify such traps, resulting in insufficient reliability of the generated results and inefficient utilization of context position sensitivity. Therefore, the present application provides a method for processing document noise based on a large language model, which can improve the efficiency of processing document information during the process of processing document noise based on a large language model.
[0053] See Figure 1 As shown, the embodiments of the present invention disclose a method for processing document noise based on a large language model, including:
[0054] Step S11: Perform probability distribution processing on the output result obtained by processing the document to be recognized by the large language model to obtain a confidence entropy value, and determine whether the confidence entropy value is greater than a preset entropy value threshold. If the confidence entropy value is not greater than the preset entropy value threshold, determine the number of documents based on the difference between the confidence entropy value and the preset target entropy value, and select the number of noise documents from the multi-source noise library; the multi-source noise library includes randomly generated text, cross-domain text, and semantic interference text.
[0055] In this embodiment, since the generation quality of the large language model is closely related to the attention entropy value. When the entropy value is too low (e.g., approaching 0), the model may over-rely on a single document, resulting in incorrect generalization; when the entropy value is too high (e.g., exceeding 5 bits), the model may generate irrelevant content due to information chaos. However, if irrelevant documents are completely deleted, it may cause the attention distribution of the model to be too concentrated (entropy collapse), but retaining noise documents may introduce misleading information, especially "difficult negative samples" that are related to the query semantics but lack answers. Therefore, the embodiment of the present application injects controllable noise dynamically to stabilize the entropy value within the target range (3 - 5 bits), thereby improving the robustness of the model.
[0056] Further, in the process of dynamically injecting controllable noise, the embodiment of the present application first needs to construct a multi-source noise library, and the noise library contains randomly generated invalid text (such as scrambled words and sentences generated by a Markov chain), cross-domain documents (such as introducing science and technology news fragments in a medical Q&A system), and semantic interference fragments (such as replacing "ACE inhibitor" with "ACE stimulant"). Among them, the data in the noise library is classified according to domain labels and has a regular update function to maintain data diversity. In a specific implementation, in the financial report generation scenario, the noise library may contain fragments of stock historical data or social media comments irrelevant to the query and supports on-demand invocation.
[0057] Secondly, in the process of the large language model generating answers, the embodiment of the present application needs to calculate the entropy value of the output probability distribution in real time and determine whether the obtained threshold is greater than the preset threshold. If the entropy value is not greater than the preset threshold (e.g., 3 bits), it indicates that the model over-relies on a single document and may fall into an incorrect mode; if the entropy value is greater than the threshold, automatically select a preset number of documents from the multi-source noise library, where the preset number of documents can be dynamically adjusted according to the difference between the current entropy value and the target value (e.g., when the difference is 2 bits, inject 3 noise documents).
[0058] In a specific implementation, the formula for determining the confidence entropy value is as follows:
[0059] ;
[0060] Where, is the probability distribution of the output tokens, is the output token, is the number of documents in the document to be recognized. When the confidence entropy value is lower than the threshold T (such as 3 bits), the noise injection operation is triggered.
[0061] Among them, the process of the noise injection operation is as follows: First, the number of noise documents is dynamically adjusted according to the difference between the current entropy value and the target entropy value, and the determination formula is as follows:
[0062] ;
[0063] Among them, is the entropy gain of a single noise document, is the target entropy value, is the current entropy value.
[0064] Subsequently, several noise documents are randomly inserted into the middle section of the context (such as the 20%-80% position) to avoid interfering with the key information at the beginning and end, and the random distribution of the noise can disperse the model's attention and prevent it from over-focusing on local content.
[0065] Specifically, the number of documents is determined based on the difference between the confidence entropy value and the preset target entropy value, and the number of noise documents equal to the number of documents is selected from the multi-source noise library, which can include: determining the number of documents based on the entropy gain corresponding to the document to be recognized and the difference between the confidence entropy value and the preset target entropy value; the magnitude of the difference is positively correlated with the magnitude of the number of documents; generating meaningless text using a Markov chain and setting the meaningless text as randomly generated text, then collecting cross-domain text irrelevant to the target domain from the website using a preset collection rule, and generating semantic interference text related to the target keyword but with incorrect content using a preset text rewriting technique; establishing a multi-source noise library based on the randomly generated text, cross-domain text, and semantic interference text, and selecting the number of noise documents equal to the number of documents from the multi-source noise library.
[0066] Step S12: Insert the number of noise documents into the middle position of the context of the document to be recognized according to the preset insertion rule to obtain a document to be processed, and use a scoring model including a sparse retrieval algorithm and a dense retrieval algorithm to score the document to be processed, and then determine whether the scoring result is greater than a preset scoring threshold. If it is greater, use a preset binary classification model to process the document to be processed to obtain the probability of a trap document; the probability of a trap document is the probability that the document has a semantic trap.
[0067] In this embodiment, since threshold filtering based on retrieval scores cannot effectively identify documents that are semantically relevant but lack answers (for example, when querying "the color of Napoleon's horse", a document about "the color of Napoleon's wife's horse" is retrieved), therefore, the embodiment of the present application achieves precise filtering by performing a hybrid scoring operation and a semantic trap detection operation.
[0068] Among them, the first-layer filtering is dynamic filtering of hybrid retrieval scores: a preliminary screening is performed on the retrieval results, that is, the scoring results of BM25 (Okapi Best Matching 25) and dense retrieval (semantic matching) are combined. Among them, for short queries, such as "the effect of ACE inhibitors" (Angiotensin-Converting Enzyme Inhibitor), BM25 filtering is preferentially used; for complex long queries, such as "Highlights of the 2023 WHO (World Health Organization) Hypertension Guidelines Update", the weight of dense retrieval is increased to emphasize the use of dense retrieval.
[0069] Furthermore, after obtaining the scoring results, the embodiment of the present application needs to calculate the threshold: according to the distribution of the scores of the current batch of documents, an automatic adjustment and filtering operation is performed on the threshold. For example, if the scores of most documents are low, the threshold is lowered to retain more potentially relevant documents.
[0070] In a specific implementation, the scoring model is a model that combines sparse retrieval (BM25) and dense retrieval (such as Contriever), and when using the scoring model to determine the scoring results, if it is a long query, the scoring model tends to use the dense retrieval algorithm for processing, and if it is a short query, the scoring model tends to use the dense retrieval algorithm for processing.
[0071] Subsequently, a threshold is set to determine whether the current document needs to enter the second-layer judgment according to the threshold and the scoring results. Specifically, using a scoring model including a sparse retrieval algorithm and a dense retrieval algorithm to score the document to be processed may include: determining whether the number of query words corresponding to the document to be processed is greater than a preset query word threshold. If the number of query words corresponding to the document to be processed is greater than the preset query word threshold, the dense retrieval algorithm in the scoring model is used to score the document to be processed to obtain the corresponding scoring result; the dense retrieval algorithm is used to perform semantic matching scoring on the document to be processed; if the number of query words corresponding to the document to be processed is not greater than the preset query word threshold, the sparse retrieval algorithm in the scoring model is used to score the document to be processed to obtain the corresponding scoring result; the sparse retrieval algorithm is used to perform keyword matching scoring on the document to be processed.
[0072] In this embodiment, for the documents that have passed the preliminary screening, the embodiments of the present application use a lightweight classifier to perform secondary filtering on the documents. Among them, the classifier can determine whether a document is a "trap document" through the following features: Keyword conflict: Detect whether the document contains the entity in the query but the attributes do not match (such as for the query "Population of Paris", the document mentions "Area of Paris"); Answer missing mark: Identify explicit negative expressions such as "not mentioned" and "no data available" in the document; Context contradiction: Through semantic similarity calculation, determine whether the content of the document conflicts with the known correct answer. In a specific implementation, for the query "Color of Napoleon's horse", if the document describes "The horse of Napoleon's wife is black", although it contains the keywords "Napoleon", "horse", and "color", it is marked as a trap document and excluded due to the missing answer.
[0073] Among them, in a specific implementation, the input of the binary classification model is a query-document pair, and the output is the probability of a "trap document". And the features include: Lexical overlap, entity conflict, and answer missing mark. Among them, the lexical overlap is the result obtained after calculating the Jaccard (Jaccard Index) similarity between the query and the document; The entity conflict is the result obtained by using the NER (Named Entity Recognition) tool to detect whether the document contains the query entity but the attribute values do not match, such as for the query "Population of Paris", the document mentions "GDP of Paris"; The answer missing mark is the result obtained by detecting whether the document contains keywords such as "unknown" and "not mentioned" through pattern matching (such as regular expressions). After obtaining the trap probability document based on the above features, the embodiments of the present application need to perform real-time classification on the document, and mark the documents with a trap document probability exceeding 0.7 as trap documents and perform exclusion operations.
[0074] Specifically, determine whether the scoring result is greater than the preset scoring threshold. If it is greater, then use the preset binary classification model to process the document to be processed to obtain the trap document probability, which may include: Determine whether the scoring result is greater than the preset scoring threshold. If the scoring result is greater than the preset scoring threshold, then use the preset binary classification model to perform lexical overlap recognition on the document to be processed to obtain the corresponding overlap result; The overlap result is the overlap degree between the document to be processed and the query statement corresponding to the document to be processed; Use the NER tool to perform entity conflict recognition on the document to be processed to obtain the corresponding conflict recognition result; The conflict recognition result is used to represent whether the query entity corresponding to the document to be processed does not match the attribute value; Use the preset pattern matching rule to perform answer missing mark operation on the document to be processed to obtain the corresponding missing mark result; The missing mark result is used to represent whether the document to be processed includes the preset keyword; Determine the trap document probability based on the overlap result, conflict recognition result, and missing mark result.
[0075] Step S13: Determine whether the probability of the trap document is greater than a preset probability threshold. If it is not greater, perform grouped determination on the document to be processed based on the scoring result to obtain a relevant group determination result, and determine a target queue based on the relevant group determination result and the initial queue, so as to input the target queue into the large language model for processing.
[0076] In this embodiment, since the large language model's attention to context shows a "U-shaped distribution", that is, the attention to the beginning and end is higher than that to the middle. Therefore, in the embodiment of the present application, the documents are graded by hierarchical re-ranking to place the highly relevant documents in the high-attention area, and at the same time use noise to fill the low-attention area. In a specific implementation, the document grading strategy is as follows: The top 20% of the documents with the highest numerical values in the scoring result are high-relevant group documents, the documents with numerical values from 20% to 60% of the highest are medium-relevant group documents, and the bottom 40% of the documents with the lowest numerical values are low-relevant group documents. Among them, the high-relevant group documents are core documents including answers or strongly relevant evidence (such as the original text of the latest medical guidelines); the medium-relevant group includes documents with auxiliary information (such as historical version guidelines or relevant research papers); the low-relevant group documents include noise documents and weakly relevant content.
[0077] Specifically, to determine whether the probability of the trap document is greater than a preset probability threshold. If it is not greater, perform grouped determination on the document to be processed based on the scoring result to obtain a relevant group determination result, which may include: Determine whether the probability of the trap document is greater than a preset probability threshold. If the probability of the trap document is greater than the preset probability threshold, mark the document to be processed corresponding to the probability of the trap document as a trap document and eliminate the trap document; if the probability of the trap document is not greater than the preset probability threshold, determine whether the scoring result meets the preset high-relevant group condition. If the scoring result meets the preset high-relevant group condition, set the document to be processed corresponding to the scoring result as a high-relevant group document; if the scoring result does not meet the preset high-relevant group condition, determine whether the scoring result meets the preset medium-relevant group condition. If the scoring result meets the preset medium-relevant group condition, set the document to be processed corresponding to the scoring result as a medium-relevant group document; if the scoring result does not meet the preset medium-relevant group condition, set the document to be processed corresponding to the scoring result as a low-relevant group document.
[0078] It is worth mentioning that after the relevant group judgment of the document, the embodiment of the present application can allocate the position of the queue according to the relevant group corresponding to the document. In the first specific implementation, if the number of documents within the group corresponding to the high-relevant group documents is not less than 2, the above documents are alternately assigned to the head and tail positions of the queue in descending order according to the retrieval score. For example, the first document is placed at the top, the second document is placed at the bottom, the third document is placed second from the top, and so on. In addition, if the retrieval quality of the document is lower than 0.6, 70% of the high-relevant documents are assigned to the tail of the queue to utilize the position advantage of adjacent queries.
[0079] Specifically, determining a target queue based on the relevant group determination result and the initial queue, and inputting the target queue into a large language model for processing may include: if the document to be processed is a high-relevance group document, determining whether the number of documents corresponding to the high-relevance group document is less than a preset document number threshold; if the number of documents corresponding to the high-relevance group document is not less than the preset document number threshold, distributing each high-relevance group document to the initial queue according to a preset descending score alternating distribution rule to obtain a target queue, and inputting the target queue into the large language model for processing; or, determining whether the largest scoring result among the scoring results of each high-relevance group document is greater than a preset score threshold; if the largest scoring result among the scoring results of each high-relevance group document is not greater than the preset score threshold, distributing a preset percentage of the high-relevance group documents in each high-relevance group document to the tail of the initial queue to obtain a target queue, and inputting the target queue into the large language model for processing.
[0080] In the second specific implementation manner, if the document is a medium-relevance group document, the above-mentioned document is evenly inserted into the middle section of the queue, such as the 20%-80% position, and one noise document is added after every two medium-relevance documents are inserted into the queue to prevent the phenomenon of attention dispersion. In the third specific implementation manner, if the document is a low-relevance group document, the above-mentioned document is randomly inserted into the queue to ensure the discretization of the noise distribution. Specifically, determining a target queue based on the relevant group determination result and the initial queue, and inputting the target queue into a large language model for processing may include: if the document to be processed is a medium-relevance group document, distributing each medium-relevance group document to the initial queue corresponding to the large language model according to a preset uniform distribution rule to obtain a queue to be processed, and inserting preset noise documents into the queue to be processed at a preset interval to obtain a target queue, and inputting the target queue into the large language model for processing; if the document to be processed is a low-relevance group document, inserting each low-relevance group document into the initial queue according to a preset random insertion rule to obtain a target queue, and inputting the target queue into the large language model for processing.
[0081] As can be seen from the above, in the embodiment of the present application, it is first necessary to perform probability distribution processing on the output result obtained by the large language model after processing the document to be recognized, obtain the confidence entropy value, and determine whether the confidence entropy value is greater than the preset entropy threshold. If the confidence entropy value is not greater than the preset entropy threshold, the number of documents is determined based on the difference between the confidence entropy value and the preset target entropy value, and the number of noise documents is selected from the multi-source noise library; subsequently, the number of noise documents is inserted into the middle position of the context of the document to be recognized according to the preset insertion rule to obtain the document to be processed, and the document to be processed is scored by a scoring model including a sparse retrieval algorithm and a dense retrieval algorithm, and then it is determined whether the scoring result is greater than the preset scoring threshold. If it is greater, the document to be processed is processed by a preset binary classification model to obtain the trap document probability; finally, it is determined whether the trap document probability is greater than the preset probability threshold. If it is not greater, the document to be processed is grouped and determined based on the scoring result to obtain the relevant group determination result, and the target queue input to the large language model is determined based on the relevant group determination result and the initial queue. In this way, the efficiency of processing document information is improved during the document noise processing based on the large language model.
[0082] Correspondingly, as shown in Figure 2 the present application also provides a document noise processing device based on a large language model, including:
[0083] A noise document selection module, configured to perform probability distribution processing on the output result obtained by the large language model after processing the document to be recognized, obtain the confidence entropy value, and determine whether the confidence entropy value is greater than the preset entropy threshold. If the confidence entropy value is not greater than the preset entropy threshold, the number of documents is determined based on the difference between the confidence entropy value and the preset target entropy value, and the number of noise documents is selected from the multi-source noise library; the multi-source noise library includes randomly generated text, cross-domain text, and semantic interference text;
[0084] A trap document probability determination module, configured to insert the number of noise documents into the middle position of the context of the document to be recognized according to the preset insertion rule to obtain the document to be processed, and score the document to be processed by a scoring model including a sparse retrieval algorithm and a dense retrieval algorithm, and then determine whether the scoring result is greater than the preset scoring threshold. If it is greater, the document to be processed is processed by a preset binary classification model to obtain the trap document probability; the trap document probability is the probability that the document has a semantic trap;
[0085] A target queue determination module, configured to determine whether the probability of the trap document is greater than a preset probability threshold. If not, group and determine the to-be-processed document based on the scoring result to obtain a relevant group determination result, and determine a target queue based on the relevant group determination result and an initial queue, so as to input the target queue into the large language model for processing.
[0086] As can be seen from the above, before performing document noise processing based on the large language model in the embodiment of the present application, it is first necessary to perform probability distribution processing on the output result obtained after the large language model processes the to-be-identified document to obtain a confidence entropy value, and determine whether the confidence entropy value is greater than a preset entropy value threshold. If the confidence entropy value is not greater than the preset entropy value threshold, determine the number of documents based on the difference between the confidence entropy value and the preset target entropy value, and select the number of noise documents from the multi-source noise library; subsequently, insert the number of noise documents into the middle position of the context of the to-be-identified document according to a preset insertion rule to obtain a to-be-processed document, and score the to-be-processed document using a scoring model including a sparse retrieval algorithm and a dense retrieval algorithm, and then determine whether the scoring result is greater than a preset scoring threshold. If it is greater, process the to-be-processed document using a preset binary classification model to obtain the probability of the trap document; finally, determine whether the probability of the trap document is greater than a preset probability threshold. If not, group and determine the to-be-processed document based on the scoring result to obtain a relevant group determination result, and determine the target queue input to the large language model based on the relevant group determination result and the initial queue. In this way, the efficiency of processing document information is improved during the document noise processing based on the large language model.
[0087] In some specific embodiments, the noise document selection module 11 may specifically include:
[0088] A document quantity determination unit, configured to determine the number of documents based on the entropy gain corresponding to the to-be-identified document and the difference between the confidence entropy value and the preset target entropy value; the magnitude of the difference is positively correlated with the magnitude of the number of documents;
[0089] A text generation unit, configured to generate meaningless text using a Markov chain, set the meaningless text as randomly generated text, then collect cross-domain text irrelevant to the target domain from a website using a preset collection rule, and generate semantic interference text related to the target keyword but with incorrect content using a preset text rewriting technique;
[0090] A multi-source noise library establishment unit, configured to establish a multi-source noise library based on the randomly generated text, the cross-domain text, and the semantic interference text, and select the number of noise documents from the multi-source noise library.
[0091] In some specific embodiments, the trap document probability determination module 12 may specifically include:
[0092] A query term quantity judgment unit, configured to judge whether the quantity of query terms corresponding to the document to be processed is greater than a preset query term threshold. If the quantity of query terms corresponding to the document to be processed is greater than the preset query term threshold, then use the dense retrieval algorithm in the scoring model to score the document to be processed to obtain a corresponding scoring result; the dense retrieval algorithm is used to perform semantic matching scoring on the document to be processed;
[0093] A scoring result determination unit, configured to if the quantity of query terms corresponding to the document to be processed is not greater than the preset query term threshold, then use the sparse retrieval algorithm in the scoring model to score the document to be processed to obtain a corresponding scoring result; the sparse retrieval algorithm is used to perform keyword matching scoring on the document to be processed.
[0094] In some specific embodiments, the trap document probability determination module 12 may specifically include:
[0095] An overlap degree result determination unit, configured to judge whether the scoring result is greater than a preset scoring threshold. If the scoring result is greater than the preset scoring threshold, then use a preset binary classification model to identify the lexical overlap degree of the document to be processed to obtain a corresponding overlap degree result; the overlap degree result is the overlap degree between the document to be processed and the query statement corresponding to the document to be processed;
[0096] A conflict recognition result determination unit, configured to use a NER tool to perform entity conflict recognition on the document to be processed to obtain a corresponding conflict recognition result; the conflict recognition result is used to represent whether the query entity corresponding to the document to be processed does not match the attribute value;
[0097] A missing mark result determination unit, configured to perform an answer missing mark operation on the document to be processed using a preset pattern matching rule to obtain a corresponding missing mark result; the missing mark result is used to represent whether the document to be processed includes a preset keyword;
[0098] A trap document probability determination subunit, configured to determine the trap document probability based on the overlap degree result, the conflict recognition result, and the missing mark result.
[0099] In some specific embodiments, the target queue determination module 13 may specifically include:
[0100] A trap document probability judgment unit, configured to judge whether the trap document probability is greater than a preset probability threshold. If the trap document probability is greater than the preset probability threshold, mark the document to be processed corresponding to the trap document probability as a trap document, and remove the trap document;
[0101] A first scoring result judgment unit, configured to, if the trap document probability is not greater than the preset probability threshold, judge whether the scoring result meets the preset high correlation group condition. If the scoring result meets the preset high correlation group condition, set the document to be processed corresponding to the scoring result as a high correlation group document;
[0102] A second scoring result judgment unit, configured to, if the scoring result does not meet the preset high correlation group condition, judge whether the scoring result meets the preset medium correlation group condition. If the scoring result meets the preset medium correlation group condition, set the document to be processed corresponding to the scoring result as a medium correlation group document;
[0103] A document setting unit, configured to, if the scoring result does not meet the preset medium correlation group condition, set the document to be processed corresponding to the scoring result as a low correlation group document.
[0104] In some specific embodiments, the target queue determination module 13 may specifically include:
[0105] A document quantity judgment unit, configured to, if the document to be processed is a high correlation group document, judge whether the document quantity corresponding to the high correlation group document is less than a preset document quantity threshold. If the document quantity corresponding to the high correlation group document is not less than the preset document quantity threshold, allocate each of the high correlation group documents to the initial queue according to a preset descending score alternating allocation rule to obtain a target queue, so as to input the target queue into the large language model for processing;
[0106] A document allocation unit, configured to judge whether the maximum scoring result value among each of the high correlation group documents is greater than a preset score threshold. If the maximum scoring result value among each of the high correlation group documents is not greater than the preset score threshold, allocate a preset percentage of the high correlation group documents among each of the high correlation group documents to the tail of the initial queue to obtain a target queue, so as to input the target queue into the large language model for processing.
[0107] In some specific embodiments, the target queue determination module 13 may specifically include:
[0108] A first document insertion unit, configured to, if the document to be processed is a medium-related group document, distribute each of the medium-related group documents to an initial queue corresponding to the large language model according to a preset uniform distribution rule to obtain a to-be-processed queue, and insert preset noise documents into the to-be-processed queue at a preset interval to obtain a target queue, so as to input the target queue into the large language model for processing;
[0109] A second document insertion unit, configured to, if the document to be processed is a low-related group document, insert each of the low-related group documents into the initial queue according to a preset random insertion rule to obtain a target queue, so as to input the target queue into the large language model for processing.
[0110] Furthermore, an embodiment of the present application also discloses an electronic device. Figure 3 It is a structural diagram of an electronic device 20 shown according to an exemplary embodiment, and the content in the figure should not be considered as any limitation on the scope of use of the present application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. Among them, the memory 22 is used to store a computer program, and the computer program is loaded and executed by the processor 21 to implement the relevant steps in the document noise processing method based on the large language model disclosed in any of the foregoing embodiments. In addition, the electronic device 20 in this embodiment may specifically be an electronic computer.
[0111] In this embodiment, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows is any communication protocol applicable to the technical solution of the present application, and no specific limitation is imposed thereon here; the input / output interface 25 is used to obtain external input data or output data to the outside, and its specific interface type can be selected according to specific application requirements, and no specific limitation is made here.
[0112] In addition, as a carrier for resource storage, the memory 22 may be a read-only memory, a random access memory, a magnetic disk, or an optical disc, etc., and the resources stored thereon may include an operating system 221, a computer program 222, etc., and the storage method may be temporary storage or permanent storage.
[0113] Among them, the operating system 221 is used to manage and control each hardware device and computer program 222 on the electronic device 20, and it can be Windows Server, Netware, Unix, Linux, etc. In addition to the computer program that can be used to complete the large language model-based document noise processing method executed by the electronic device 20 disclosed in any of the foregoing embodiments, the computer program 222 can further include computer programs that can be used to complete other specific tasks.
[0114] Furthermore, the present application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the large language model-based document noise processing method disclosed above. For the specific steps of this method, reference can be made to the corresponding content disclosed in the foregoing embodiments, and details will not be repeated here.
[0115] In this specification, the various embodiments are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and reference can be made to the method part for the relevant parts.
[0116] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0117] The steps of the method or algorithm described in combination with the embodiments disclosed in this article can be directly implemented by hardware, software modules executed by a processor, or a combination of the two. The software module can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, register, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field.
[0118] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the said element.
[0119] The technical solutions provided in this application have been introduced in detail above. Specific examples are used in this text to elaborate on the principles and implementation manners of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application. At the same time, for those of ordinary skill in the art, according to the idea of this application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to this application.
Claims
1. A method for processing document noise based on large language models, characterized in that, Including: Performing probability distribution processing on the output result obtained after the large language model processes the document to be recognized, obtaining a confidence entropy value, and determining whether the confidence entropy value is greater than a preset entropy value threshold. If the confidence entropy value is not greater than the preset entropy value threshold, determining the number of documents based on the difference between the confidence entropy value and the preset target entropy value, and selecting the number of noise documents from the multi-source noise library; the multi-source noise library includes randomly generated text, cross-domain text, and semantic interference text; Inserting the number of noise documents into the middle position of the context of the document to be recognized according to a preset insertion rule, obtaining a document to be processed, and using a scoring model including a sparse retrieval algorithm and a dense retrieval algorithm to score the document to be processed, and then determining whether the scoring result is greater than a preset scoring threshold. If it is greater, using a preset binary classification model to process the document to be processed, obtaining a trap document probability; the trap document probability is the probability characterizing that there is a semantic trap in the document; Determining whether the trap document probability is greater than a preset probability threshold. If it is not greater, performing a grouping determination on the document to be processed based on the scoring result, obtaining a relevant group determination result, and determining a target queue based on the relevant group determination result and the initial queue, so as to input the target queue into the large language model for processing.
2. The method for processing document noise based on a large language model according to claim 1, wherein The determining the number of documents based on the difference between the confidence entropy value and the preset target entropy value, and selecting the number of noise documents from the multi-source noise library includes: Determining the number of documents based on the entropy gain corresponding to the document to be recognized and the difference between the confidence entropy value and the preset target entropy value; the magnitude of the difference is positively correlated with the magnitude of the number of documents; Generating meaningless text using a Markov chain, setting the meaningless text as randomly generated text, then collecting cross-domain text irrelevant to the target domain from websites using a preset collection rule, and generating semantic interference text related to the target keyword but with incorrect content using a preset text rewriting technique; Establishing a multi-source noise library based on the randomly generated text, the cross-domain text, and the semantic interference text, and selecting the number of noise documents from the multi-source noise library.
3. The method for processing document noise based on a large language model according to claim 1, wherein The using a scoring model including a sparse retrieval algorithm and a dense retrieval algorithm to score the document to be processed includes: Determining whether the number of query words corresponding to the document to be processed is greater than a preset query word threshold. If the number of query words corresponding to the document to be processed is greater than the preset query word threshold, using the dense retrieval algorithm in the scoring model to score the document to be processed, obtaining a corresponding scoring result; the dense retrieval algorithm is used to perform semantic matching scoring on the document to be processed; If the number of query words corresponding to the document to be processed is not greater than the preset query word threshold, using the sparse retrieval algorithm in the scoring model to score the document to be processed, obtaining a corresponding scoring result; the sparse retrieval algorithm is used to perform keyword matching scoring on the document to be processed.
4. The method for processing document noise based on a large language model according to claim 1, characterized in that, Determine whether the judgment scoring result is greater than a preset scoring threshold. If it is greater, use a preset binary classification model to process the document to be processed to obtain the probability of a trap document, including: Determine whether the judgment scoring result is greater than a preset scoring threshold. If the scoring result is greater than the preset scoring threshold, use a preset binary classification model to identify the lexical overlap degree of the document to be processed to obtain the corresponding overlap degree result; the overlap degree result is the overlap degree between the document to be processed and the query statement corresponding to the document to be processed; Use the NER tool to identify entity conflicts in the document to be processed to obtain the corresponding conflict identification result; the conflict identification result is used to represent whether the query entity corresponding to the document to be processed does not match the attribute value; Use a preset pattern matching rule to perform an answer missing marking operation on the document to be processed to obtain the corresponding missing marking result; the missing marking result is used to represent whether the document to be processed includes a preset keyword; Determine the probability of a trap document based on the overlap degree result, the conflict identification result, and the missing marking result.
5. The method for processing document noise based on a large language model according to any one of claims 1 to 4, characterized in that, Determine whether the probability of the trap document is greater than a preset probability threshold. If it is not greater, perform a grouping determination on the document to be processed based on the scoring result to obtain a relevant group determination result, including: Determine whether the probability of the trap document is greater than a preset probability threshold. If the probability of the trap document is greater than the preset probability threshold, mark the document to be processed corresponding to the probability of the trap document as a trap document, and remove the trap document; If the probability of the trap document is not greater than the preset probability threshold, determine whether the scoring result meets the preset high relevant group condition. If the scoring result meets the preset high relevant group condition, set the document to be processed corresponding to the scoring result as a high relevant group document; If the scoring result does not meet the preset high relevant group condition, determine whether the scoring result meets the preset medium relevant group condition. If the scoring result meets the preset medium relevant group condition, set the document to be processed corresponding to the scoring result as a medium relevant group document; If the scoring result does not meet the preset medium relevant group condition, set the document to be processed corresponding to the scoring result as a low relevant group document.
6. The method for processing document noise based on a large language model according to claim 5, wherein Determine the target queue based on the relevant group determination result and the initial queue, and input the target queue into the large language model for processing, including: If the document to be processed is a high relevant group document, determine whether the number of documents corresponding to the high relevant group document is less than a preset document number threshold. If the number of documents corresponding to the high relevant group document is not less than the preset document number threshold, allocate each high relevant group document to the initial queue according to the preset descending order of scores alternating allocation rule to obtain the target queue, and input the target queue into the large language model for processing; Alternatively, determine whether the maximum scoring result among the documents in each of the highly relevant groups is greater than a preset score threshold. If the maximum scoring result among the documents in each of the highly relevant groups is not greater than the preset score threshold, then allocate a preset percentage of the highly relevant group documents in each of the highly relevant groups to the tail of the initial queue to obtain a target queue, so as to input the target queue into the large language model for processing.
7. The method for processing document noise based on a large language model according to claim 5, wherein Determining the target queue based on the relevant group determination result and the initial queue, and inputting the target queue into the large language model for processing, includes: If the document to be processed is a medium relevant group document, then distribute each of the medium relevant group documents to the initial queue corresponding to the large language model according to a preset uniform distribution rule to obtain a queue to be processed, and insert preset noise documents into the queue to be processed at a preset interval to obtain a target queue, so as to input the target queue into the large language model for processing; If the document to be processed is a low relevant group document, then insert each of the low relevant group documents into the initial queue according to a preset random insertion rule to obtain a target queue, so as to input the target queue into the large language model for processing.
8. A document noise processing device based on a large language model, characterized in that, Includes: A noise document selection module, configured to perform probability distribution processing on the output result obtained after the large language model processes the document to be recognized to obtain a confidence entropy value, and determine whether the confidence entropy value is greater than a preset entropy threshold. If the confidence entropy value is not greater than the preset entropy threshold, then determine the number of documents based on the difference between the confidence entropy value and a preset target entropy value, and select the number of noise documents from a multi-source noise library; the multi-source noise library includes randomly generated text, cross-domain text, and semantic interference text; A trap document probability determination module, configured to insert the number of noise documents into the middle position of the context of the document to be recognized according to a preset insertion rule to obtain a document to be processed, and score the document to be processed using a scoring model including a sparse retrieval algorithm and a dense retrieval algorithm, and then determine whether the scoring result is greater than a preset scoring threshold. If it is greater, then process the document to be processed using a preset binary classification model to obtain a trap document probability; the trap document probability is the probability indicating that the document has a semantic trap; A target queue determination module, configured to determine whether the trap document probability is greater than a preset probability threshold. If it is not greater, then perform grouping determination on the document to be processed based on the scoring result to obtain a relevant group determination result, and determine a target queue based on the relevant group determination result and the initial queue, so as to input the target queue into the large language model for processing.
9. An electronic device, characterized in that, Includes: A memory, configured to store a computer program; A processor, configured to execute the computer program to implement the method for processing document noise based on a large language model according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, For storing a computer program, wherein the computer program, when executed by a processor, implements the method for processing document noise based on a large language model according to any one of claims 1 to 7.