Large model-based structured xml standard document intelligent query assistance system

By using a large-scale structured XML standard document intelligent query assistance system, which combines precise and fuzzy matching to calculate the relevance representation value and dynamically adjusts system parameters, the system solves the problem of low query efficiency in existing technologies and achieves a comprehensive and accurate improvement in document query.

CN121434158BActive Publication Date: 2026-05-05CHINA NAT INST OF STANDARDIZATION
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHINA NAT INST OF STANDARDIZATION
Filing Date
2025-10-30
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Existing technologies have failed to effectively optimize document search systems based on user feedback, resulting in low search efficiency and difficulty in quickly and accurately finding relevant documents.

Method used

A structured XML standard document intelligent query assistance system based on a large model is adopted. The system classifies documents through a document comparison module, determines relevance through a data partitioning module, generates index tags through a data analysis module, and calculates the relevance representation value by combining precise and fuzzy matching. The system parameters are dynamically adjusted to optimize document query.

Benefits of technology

It improves the comprehensiveness and accuracy of document retrieval, reduces information omissions caused by inconsistent keywords, dynamically adjusts system parameters to adapt to different types and sizes of document data, and improves query efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121434158B_ABST
    Figure CN121434158B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of document query assistance, in particular to a structured XML standard document intelligent query assistance system based on a large model, which comprises a data receiving module, a document comparison module, a data division module, a data analysis module, a text generation module, a data statistics module and a data processing module, whether the processing of each document is qualified is determined based on a click probability, when it is determined that the processing of each document is abnormal, a preset reference quantity is adjusted to a corresponding value, whether the operation of the document generation module is qualified is determined based on the average push quantity of the average single output push document of the text generation module, when it is determined that the operation of the document generation module is abnormal, a preset category quantity used for dividing the relevance of single batch documents is adjusted to a corresponding value. The feedback of the user is analyzed, the system is optimized according to the use behavior of the user, and the query efficiency of the document is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of document query assistance technology, and in particular to an intelligent query assistance system for structured XML standard documents based on a large model. Background Technology

[0002] In large enterprises, various departments generate a large number of structured XML standard documents, such as technical specifications, business process descriptions, project reports, etc.

[0003] A single section may store a large number of files related to a single event. If a user searches for highly relevant files using only partial keywords from the document name, the number of related files pushed to the system will be large, making it difficult to filter files specifically and causing users to find the documents they need quickly and accurately.

[0004] Chinese Patent Publication No. CN105027115A discloses a document querying and indexing method, including generating a document index based on a document set and using it to identify documents that match one or more queries. A tree with nodes corresponding to each object in the document is generated for each document. The nodes of the generated tree are merged or combined to generate a document index, which is itself a tree. Additionally, an inverted index is generated for each node of this index, identifying the one or more trees from which the node originates. When a query is received, the query is first executed against the document index tree. During execution, the correct set operation is applied to the inverted index associated with the node matching the query. The resulting set identifies documents that can match the query. The query is then executed on the identified documents. However, the above technical solution has the following problem: it does not consider analyzing user feedback, making it impossible to optimize the system based on user behavior, thus affecting document query efficiency. Summary of the Invention

[0005] To address this issue, the present invention provides an intelligent query assistance system for structured XML standard documents based on a large model, which overcomes the problem in existing technologies that do not consider the analysis of user feedback, thus failing to optimize the system based on user behavior and affecting document query efficiency.

[0006] To achieve the above objectives, this invention provides an intelligent query assistance system for structured XML standard documents based on a large model, comprising:

[0007] The document comparison module is used to classify each document based on the comparison results obtained from each document and to obtain the number of categories;

[0008] A data segmentation module, which is connected to the document comparison module, is used to determine the relevance of a single batch of documents based on the number of categories of the received single batch of documents.

[0009] The data analysis module is connected to the data segmentation module and the document comparison module respectively, and generates index tags for each document. When the correlation of a single batch of documents is determined to be strong, the preset number of references used for document classification is adjusted.

[0010] The text generation module is used to generate several push documents in descending order of relevance based on the search data input by the user;

[0011] The data processing module, which is connected to the document comparison module and the data segmentation module respectively, is used to determine whether the processing of each document is qualified based on the click probability of the top-ranked push document clicked by users, and to adjust the preset reference quantity to the corresponding value when it is determined that the processing of each document is abnormal. It also determines whether the operation of the document generation module is qualified based on the average number of push documents output by the text generation module per batch, and to adjust the preset number of categories used to classify the relevance of a single batch of documents to the corresponding value when it is determined that the operation of the document generation module is abnormal.

[0012] Furthermore, the text generation module is used to generate exclusive search text based on the user-input search data, and to perform a search to generate several push documents arranged in descending order of relevance representation values, including:

[0013] It is used to match the keywords of the exclusive search text with the index tags of each document in the database, and the product of the number of exactly matched keywords and the preset precision coefficient is determined as the first association value of the corresponding search tag;

[0014] It is used to perform semantic analysis and fuzzy matching on keywords and index tags of the exclusive search text, so as to determine the second association value of the corresponding search tag by multiplying the number of fuzzy matched keywords with the preset fuzzy coefficient.

[0015] The sum of the first association value and the second association value is used to determine the association characteristic value of the corresponding search tag;

[0016] This is used to compare the retrieved data with each index tag, and to identify the documents corresponding to each index tag whose selected relevance value is greater than the preset relevance comparison value as the push documents.

[0017] Furthermore, the document comparison module is used to classify each document based on the comparison results of each document, and to obtain the number of categories, including:

[0018] This is used to obtain the high-frequency words of a single document. The high-frequency words are sorted in descending order according to their frequency of occurrence. A preset number of high-frequency words are selected as the representative words for a single document.

[0019] Used to compare the characteristic words of each document in a single batch of documents received by the data receiving module;

[0020] This is used to group documents with the same signature terms into the same category.

[0021] Furthermore, the data segmentation module is used to determine the relevance of a single batch of documents based on the number of categories in the single batch of documents received by the data receiving module, including:

[0022] If the number of categories is less than or equal to the preset number of categories, the relevance of a single batch of documents is determined to be strong, and the preset reference number used for document classification is adjusted to the corresponding value based on the number of categories.

[0023] If the number of categories is greater than the preset number of categories, the correlation of a single batch of documents is determined to be weak, and the data comparison module is controlled to continue running using the current operating parameters.

[0024] Furthermore, the data analysis module is used to generate index tags for each document based on the relevance of a single batch of documents, including:

[0025] If the correlation of a single batch of documents is determined to be strong, the preset reference number used for document classification will be adjusted to the corresponding value based on the number of documents in the single batch.

[0026] If the relevance of a single batch of documents is determined to be weak, then each characteristic term corresponding to the document is determined as the index label of the corresponding document.

[0027] Furthermore, the data analysis module is used to adjust the preset reference quantity for document classification to a corresponding value based on the number of documents in a single batch, wherein,

[0028] The increase in the number of preset references is positively correlated with the number of documents in a single batch.

[0029] The data analysis module is used to control the document comparison module to redetermine the characteristic words of each document after adjusting the preset number of references.

[0030] The data analysis module is used to determine the redefined key terms corresponding to a document as the index tags for that document.

[0031] Furthermore, the data processing module is used to determine whether the processing of each document is qualified based on the click probability, provided that the running time of the data receiving module reaches an integer multiple of the preset running time. This includes:

[0032] If the click probability is less than or equal to the preset click probability, it is determined that the processing of each document is abnormal. Based on the click probability, the preset reference quantity is adjusted to the corresponding value. Based on the number of push documents output by the text generation module, it is determined whether the operating parameters of the document generation module are qualified.

[0033] The increase in the preset reference quantity is negatively correlated with the click probability.

[0034] Furthermore, if the click probability is greater than the preset click probability, then the processing of each document is deemed qualified, and the current parameters are used to continue processing each document.

[0035] Furthermore, the data processing module is used to determine whether the operating parameters of the document generation module are qualified based on the average number of pushes, including:

[0036] If the average number of pushes is greater than the preset number of pushes, it is determined that the document generation module is malfunctioning. Based on the average number of pushes, the preset number of categories used to classify the relevance of a single batch of documents is adjusted to the corresponding value.

[0037] The increase in the number of preset categories is positively correlated with the average number of push notifications.

[0038] Furthermore, the data processing module, after adjusting the number of preset categories, continuously monitors the click probability, and when the running time of the data receiving module reaches an integer multiple of the preset running time again;

[0039] The data processing module determines whether to correct the operating parameters for the text generation module based on historically changed parameters, including:

[0040] It is used to draw a probability time-domain curve based on the click probability within each preset runtime in the acquired historical data, and the slope of the calculated curve at the current time node is determined as the historical change parameter.

[0041] If the historical change parameters are greater than the preset change parameters, the operating parameters of the text generation module are deemed to be qualified.

[0042] If the historical change parameter is less than or equal to the preset change parameter, the preset accuracy coefficient is adjusted to the corresponding value based on the historical change parameter;

[0043] The increase in the preset accuracy coefficient is negatively correlated with the historical changes in parameters.

[0044] Compared with existing technologies, the advantages of this invention are as follows: First, the user-input search data is converted into exclusive search text. Then, through precise and fuzzy matching of the search text keywords and document index tags, a first correlation value and a second correlation value are calculated respectively. These two values ​​are added together to obtain a correlation characterization value. Documents with a correlation characterization value greater than a preset correlation comparison value are selected as push documents. Fuzzy matching compensates for the shortcomings of precise matching, finding semantically relevant documents and improving the comprehensiveness of the query. By setting preset precision coefficients and preset fuzzy coefficients, the importance of precise and fuzzy matching keywords in calculating the correlation characterization value is differentiated. The correlation characterization value represents the correlation between the search text and each document. Quantified correlation determines documents with strong correlation to the search tags as push documents. Combining precise and fuzzy matching comprehensively determines the user's query intent, thus comprehensively identifying relevant documents and avoiding the omission of important information due to incomplete keyword consistency. Documents are sorted according to the correlation characterization value, prioritizing the push of the most relevant documents to the user, improving document query efficiency.

[0045] Furthermore, the document comparison module acquires high-frequency words from individual documents, sorts them by frequency, and selects the top preset reference words as character terms. It then compares the character terms of a batch of documents, grouping documents with identical character terms into the same category. Within a single department's professional field, there may be semantically similar but substantively different professional terms. Classifying documents based solely on identical character terms improves the accuracy of category classification, ensuring the system can accurately categorize documents across different language environments. High-frequency words reflect the document's theme and content; comparing character terms allows for rapid document classification, facilitating subsequent management and retrieval. The preset reference number determines how many high-frequency words are selected as character terms for each document. This number directly affects the accuracy of document classification and the amount of data the text generation module needs to process during user searches. Based on character term comparison, the document comparison module quickly and accurately classifies a large number of documents, improving the efficiency of batch document management. By not relying on semantic similarity, it avoids interference from semantically similar but different words in professional fields, improving classification accuracy and consequently, document retrieval efficiency.

[0046] Furthermore, the data segmentation module determines the relevance of a single batch of documents based on the number of categories received by the data receiving module. When the number of categories is less than or equal to a preset number of categories, the number of categories in a single batch of documents is small, and the relevance between documents is strong; conversely, the relevance is weak. The data analysis module performs targeted processing based on the relevance of the single batch of documents. When a strong relevance is determined, increasing the preset reference number allows for more detailed document classification and improves the accuracy of index tags; if a weak relevance is determined, the characteristic words are directly used as index tags to simplify the processing flow. The parameters for determining index tags are dynamically adjusted according to the strength of document relevance, enabling the system to better adapt to document data of different types and sizes, improving document processing efficiency, and thus improving document query efficiency.

[0047] Furthermore, when the data receiving module's runtime reaches an integer multiple of the preset runtime, the document processing qualification is determined based on the click probability. The click probability reflects the accuracy of the pushed documents and the user's attention to the top-ranked pushed document. When the click probability is greater than the preset click probability, the click probability is high, and the system-pushed documents meet user needs, indicating qualified document processing. When the click probability is less than or equal to the preset click probability, processing is abnormal. In this case, the preset reference quantity is adjusted to holistically regulate the amount of index tag data, improving retrieval accuracy. Simultaneously, the document generation module's operating parameters are judged based on the average push quantity, which reflects the redundancy of the pushed documents output by the text generation module. When the average push quantity is greater than the preset push quantity, too many documents are pushed, making it difficult for users to find the information they need among numerous documents. Additionally, when the number of pushed files is too large, the interface needs to render more content, requiring more computational resources and time for the layout, formatting, and link generation of each pushed document, leading to display delays on the push interface. In this case, the preset category quantity used to classify the relevance of a single batch of documents is adjusted to the corresponding value to redefine the criteria for classifying relevance. Improving document matching accuracy reduces the number of searches required by users, thereby reducing overall search time. By monitoring and analyzing user click behavior and the number of documents pushed, system parameters are dynamically adjusted to continuously optimize system performance, thus improving document retrieval efficiency.

[0048] Furthermore, based on historical change parameters, the system determines whether to adjust the operating parameters of the text generation module. These historical change parameters reflect the trend of click probability over time. When the historical change parameters are less than or equal to the second preset change comparison value and greater than the first preset change comparison value, it indicates that the click probability is increasing or stabilizing due to the adjustment of the preset category quantity. When the historical change parameters are less than or equal to the preset change parameters, the click probability is decreasing or showing little improvement. In this case, an anomaly exists in the matching between the document and the user's search data, resulting in the failure to accurately prioritize the documents requested by the user. In this situation, the operating parameters of the text generation module should be adjusted. The preset accuracy coefficient is adjusted based on the historical change parameters, and the weights of precise matching and fuzzy matching are adjusted in a timely manner according to the trend of click probability changes. This allows the system to continuously adapt to changes in documents and user needs, maintaining a good push effect and further improving document query efficiency. Attached Figure Description

[0049] Figure 1 This is a block diagram of the intelligent query assistance system for structured XML standard documents based on a large model, as described in an embodiment of the present invention.

[0050] Figure 2 This is a block diagram of the data segmentation module in an embodiment of the present invention, which determines the relevance of a single batch of documents based on the number of categories.

[0051] Figure 3 The block diagram of the data analysis module of this invention generates index tags for each document based on the relevance of a single batch of documents;

[0052] Figure 4 This is a block diagram of the data processing module in an embodiment of the present invention, which determines whether the processing of each document is qualified based on the click probability. Detailed Implementation

[0053] To make the objectives and advantages of the present invention clearer, the present invention will be further described below with reference to embodiments; it should be understood that the specific embodiments described herein are merely for explaining the present invention and are not intended to limit the present invention.

[0054] Preferred embodiments of the present invention will now be described with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are merely illustrative of the technical principles of the present invention and are not intended to limit the scope of protection of the present invention.

[0055] Furthermore, it should be noted that, in the description of this invention, unless otherwise explicitly specified and limited, the term "connected" should be interpreted broadly. For example, it can refer to a fixed connection, a detachable connection, or an integral connection; it can refer to a mechanical connection or an electrical connection; it can refer to a direct connection or an indirect connection through an intermediate medium; and it can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.

[0056] Please see Figure 1 The diagram shown is a block diagram of the intelligent query assistance system for structured XML standard documents based on a large model, according to an embodiment of the present invention. The system of the present invention includes:

[0057] The data receiving module is used to acquire several documents stored by the user;

[0058] The document comparison module, which is connected to the data receiving module, is used to classify each document based on the comparison results of each document and obtain the number of categories.

[0059] A data segmentation module, which is connected to the document comparison module and the data receiving module respectively, is used to determine the relevance of a batch of documents based on the number of categories of a batch of documents received by the data receiving module.

[0060] The data analysis module is connected to the data segmentation module and the document comparison module respectively. It is used to generate index tags for each document based on the relevance of a single batch of documents. When the relevance of a single batch of documents is determined to be strong, the preset reference number used for document classification is adjusted to the corresponding value based on the number of documents in the single batch.

[0061] The text generation module is used to generate personalized search text based on the search data input by the user, and to perform a search to generate several push documents arranged in descending order of relevance.

[0062] The data statistics module, which is connected to the text generation module, is used to calculate the click probability of the first pushed document clicked by the user;

[0063] The data processing module, which is connected to the data statistics module, the document comparison module, and the data segmentation module, is used to determine whether the processing of each document is qualified based on the click probability, and to adjust the preset reference quantity to the corresponding value when the processing of each document is determined to be abnormal. It also determines whether the operation of the document generation module is qualified based on the average number of pushed documents output by the text generation module per batch, and to adjust the preset number of categories used to classify the relevance of a single batch of documents to the corresponding value when the operation of the document generation module is determined to be abnormal.

[0064] Specifically, the data receiving module further includes a data format verification unit, which determines whether the document is an XML document after receiving it, and reports an error if the document does not conform to the XML format.

[0065] Specifically, the text generation module is used to generate exclusive search text based on user-input search data, and to perform a search to generate several push documents arranged in descending order of relevance representation values, including:

[0066] It is used to match the keywords of the exclusive search text with the index tags of each document in the database, and the product of the number of exactly matched keywords and the preset precision coefficient is determined as the first association value of the corresponding search tag;

[0067] It is used to perform semantic analysis and fuzzy matching on keywords and index tags of the exclusive search text, so as to determine the second association value of the corresponding search tag by multiplying the number of fuzzy matched keywords with the preset fuzzy coefficient.

[0068] The sum of the first association value and the second association value is used to determine the association characteristic value of the corresponding search tag;

[0069] This is used to compare the retrieved data with each index tag, and to identify the documents corresponding to each index tag whose selected relevance value is greater than the preset relevance comparison value as the push documents.

[0070] Specifically, the preset correlation comparison value is selected within the range [0.5W0, 1.1W0], where W0 is the number of keywords in the exclusive search text. Those skilled in the art can select the preset correlation comparison value themselves. It can be understood that it can achieve the filtering of documents corresponding to search tags with strong correlation.

[0071] Specifically, the user-input search data is first converted into personalized search text. Then, through precise and fuzzy matching of the search text keywords with document index tags, a first and second correlation value are calculated respectively. These two values ​​are added together to obtain a correlation characterization value. Documents with a correlation characterization value greater than a preset correlation comparison value are selected as push documents. Fuzzy matching compensates for the shortcomings of precise matching, finding semantically relevant documents and improving the comprehensiveness of the query. By setting preset precision and fuzzy coefficients, the importance of precise and fuzzy matching keywords in calculating the correlation characterization value is differentiated. The correlation characterization value represents the correlation between the search text and various documents. Quantified correlation determines documents with strong relevance to the search tags as push documents. Combining precise and fuzzy matching comprehensively determines the user's query intent, thus comprehensively identifying relevant documents and avoiding the omission of important information due to incomplete keyword matching. Documents are sorted according to the correlation characterization value, prioritizing the push of the most relevant documents to the user, improving document search efficiency.

[0072] Specifically, the preset precision coefficient is greater than the preset fuzziness coefficient. In this embodiment, preferably, the preset precision coefficient is 1.21 and the preset fuzziness coefficient is 0.79. The sum of the preset precision coefficient and the preset fuzziness coefficient is 2.

[0073] Specifically, the text generation module can perform fuzzy matching using a natural language processing tool. The selection of the natural language processing tool is not limited and can include WordNet.

[0074] Specifically, the text generation module is used to generate personalized search text based on user-input search data, including:

[0075] Used to segment the search data entered by the user;

[0076] Used to delete function words in each word segment that have no actual retrieval meaning;

[0077] Used to identify the keywords of each word segment after removing function words as the keywords of the generated exclusive search text;

[0078] Keywords are used to identify synonyms of each segment after removing function words as keywords for the generated exclusive search text.

[0079] Specifically, the document comparison module is used to classify each document based on the comparison results obtained from each document, and to obtain the number of categories, including:

[0080] This is used to obtain the high-frequency words of a single document. The high-frequency words are sorted in descending order according to their frequency of occurrence. A preset number of high-frequency words are selected as the representative words for a single document.

[0081] Used to compare the characteristic words of each document in a single batch of documents received by the data receiving module;

[0082] This is used to group documents with the same signature terms into the same category.

[0083] Specifically, the document comparison module is also used to remove stop words in the document before determining high-frequency words;

[0084] Specifically, stop words can be stop words from a stop word list, which may include the Harbin Institute of Technology stop word list and the NLTK stop word list.

[0085] Specifically, the document comparison module acquires high-frequency words from individual documents, sorts them by frequency, and selects the top preset reference words as character terms. It then compares the character terms of a batch of documents, grouping documents with identical character terms into the same category. Within a single department's professional field, there may be semantically similar but substantively different professional terms. Classifying documents based solely on identical character terms improves the accuracy of category classification, ensuring the system can accurately categorize documents across different language environments. High-frequency words reflect the document's theme and content; comparing character terms allows for rapid document classification, facilitating subsequent management and retrieval. The preset reference number determines how many high-frequency words are selected as character terms for each document. This number directly affects the accuracy of document classification and the amount of data the text generation module needs to process during user searches. Based on character term comparison, the document comparison module quickly and accurately classifies a large number of documents, improving the efficiency of batch document management. By not relying on semantic similarity, it avoids interference from semantically similar but different words in professional fields, improving classification accuracy and consequently, document retrieval efficiency.

[0086] Please see Figure 2 The diagram shown is a block diagram of a data segmentation module in an embodiment of the present invention that determines the relevance of a single batch of documents based on the number of categories. The data segmentation module of the present invention is used to determine the relevance of a single batch of documents based on the number of categories of the single batch of documents received by the data receiving module, and includes:

[0087] If the number of categories is less than or equal to the preset number of categories, the relevance of a single batch of documents is determined to be strong, and the preset reference number used for document classification is adjusted to the corresponding value based on the number of categories.

[0088] If the number of categories is greater than the preset number of categories, the correlation of a single batch of documents is determined to be weak, and the data comparison module is controlled to continue running using the current operating parameters.

[0089] Specifically, the number of preset categories is selected within the range [0.43N0, 0.61N0], where N0 is the total number of documents in a single batch. Those skilled in the art can select and determine the number of preset categories themselves. Tests can be conducted on document datasets of different sizes. The number of preset categories is determined based on the accuracy of the document relevance judgment under different thresholds. It can be understood that it can achieve the classification of the relevance of each document.

[0090] Please see Figure 3 The diagram shown is a block diagram of the data analysis module of this invention, which generates index tags for each document based on the relevance of a single batch of documents. The data analysis module of this invention is used to generate index tags for each document based on the relevance of a single batch of documents, and includes:

[0091] If the correlation of a single batch of documents is determined to be strong, the preset reference number used for document classification will be adjusted to the corresponding value based on the number of documents in the single batch.

[0092] If the relevance of a single batch of documents is determined to be weak, then each characteristic term corresponding to the document is determined as the index label of the corresponding document.

[0093] Specifically, index tags include several index sub-tags, which are the characteristic words of the corresponding documents.

[0094] Specifically, the data segmentation module determines the relevance of a single batch of documents based on the number of categories received by the data receiving module. When the number of categories is less than or equal to a preset number, the number of categories in a single batch of documents is small, indicating a strong correlation between documents; conversely, the correlation is weak. The data analysis module performs targeted processing based on the relevance of the single batch of documents. When a strong correlation is determined, increasing the preset reference number allows for more detailed document classification, improving the accuracy of index tags; if a weak correlation is determined, the characteristic words are directly used as index tags to simplify the processing flow. The parameters for determining index tags are dynamically adjusted according to the strength of document relevance, enabling the system to better adapt to document data of different types and sizes, improving document processing efficiency, and consequently, document query efficiency.

[0095] Specifically, the data analysis module is used to adjust the preset reference quantity for document classification to a corresponding value based on the number of documents in a single batch, wherein,

[0096] The increase in the number of preset references is positively correlated with the number of documents in a single batch.

[0097] The data analysis module is used to control the document comparison module to redetermine the characteristic words of each document after adjusting the preset number of references.

[0098] The data analysis module is used to determine the redefined key terms corresponding to a document as the index tags for that document.

[0099] In this embodiment, optionally,

[0100] Compare the number of documents in a single batch with the first preset document count comparison value and the second preset document count comparison value;

[0101] If the number of documents in a single batch is less than or equal to the first preset document quantity comparison value, the preset reference quantity will be adjusted to 1.13 times the initial preset reference quantity.

[0102] If the number of documents in a single batch is less than or equal to the second preset document quantity comparison value and greater than the first preset document quantity comparison value, then the preset reference quantity will be adjusted to 1.21 times the initial preset reference quantity.

[0103] If the number of documents in a single batch is greater than the second preset document quantity comparison value, the preset reference quantity will be adjusted to 1.29 times the initial preset reference quantity.

[0104] The first preset document quantity comparison value is 1.3J0, and the second preset document quantity comparison value is 2.1J0, where J0 is the average number of documents in each historical batch.

[0105] Specifically, the larger the number of documents in a single batch, the greater the diversity and complexity of the documents. To classify these documents more accurately, the number of preset references is increased to cover more feature information, thereby distinguishing different documents more finely.

[0106] Please see Figure 4 The diagram shown is a block diagram of the data processing module of this invention, which determines whether the processing of each document is qualified based on the click probability. The data processing module of this invention is used to determine whether the processing of each document is qualified based on the click probability when the running time of the data receiving module reaches an integer multiple of a preset running time. This includes:

[0107] If the click probability is less than or equal to the preset click probability, it is determined that the processing of each document is abnormal. Based on the click probability, the preset reference quantity is adjusted to the corresponding value. Based on the number of push documents output by the text generation module, it is determined whether the operating parameters of the document generation module are qualified.

[0108] If the click probability is greater than the preset click probability, then the processing of each document is deemed qualified, and the current parameters are used to continue processing each document.

[0109] Specifically, the click probability is the ratio of the number of times the user clicks the top-ranked push document within a preset runtime to the number of times the user enters search data.

[0110] Specifically, the preset click probability is selected within the range [0.34, 0.44]. Those skilled in the art can select and determine the preset click probability according to the actual use scenario. It can be determined by combining the expected user behavior in the business scenario. It can be based on historical data to analyze the number of times the user clicks the first pushed document in different time periods. The preset click probability can be determined by statistical analysis of a large amount of historical data.

[0111] Specifically, the data processing module is used to adjust a preset reference quantity to a corresponding value based on the click probability, wherein,

[0112] The increase in the preset reference quantity is negatively correlated with the click probability.

[0113] In this embodiment, optionally,

[0114] Compare the click probability with the first preset click probability comparison value and the second preset click probability comparison value;

[0115] If the click probability is less than or equal to the first preset click probability comparison value, then the preset reference quantity will be adjusted to 1.25 times the current preset reference quantity;

[0116] If the click probability is less than or equal to the second preset click probability comparison value and greater than the first preset click probability comparison value, then the preset reference quantity will be adjusted to 1.17 times the current preset reference quantity.

[0117] If the click probability is greater than the second preset click probability comparison value, then the preset reference quantity will be adjusted to 1.11 times the current preset reference quantity;

[0118] The first preset click probability comparison value is 0.83G0, and the second preset click probability comparison value is 0.91G0, where G0 is the preset click probability.

[0119] Specifically, when the click probability is low, the system's current document categorization and push mechanism has problems. Increasing the preset number of references can make document categorization more detailed, generate more accurate index tags, thereby improving the relevance of pushed documents and increasing the probability of user clicks.

[0120] Specifically, the data processing module is used to determine whether the operating parameters of the document generation module are qualified based on the average number of pushes, including:

[0121] If the average number of pushes is less than or equal to the preset number of pushes, the document generation module's operating parameters are deemed qualified, and the document generation module is controlled to continue running using the current operating parameters.

[0122] If the average number of pushes is greater than the preset number of pushes, it is determined that the document generation module is malfunctioning. Based on the average number of pushes, the preset number of categories used to classify the relevance of a single batch of documents is adjusted to the corresponding value.

[0123] The increase in the number of preset categories is positively correlated with the average number of push notifications.

[0124] Specifically, when the data receiving module's runtime reaches an integer multiple of the preset runtime, the document processing qualification is determined based on the click probability. The click probability reflects the accuracy of the pushed documents and the user's attention level to the top-ranked pushed document. When the click probability is greater than the preset click probability, the click probability is high, and the system-pushed documents meet user needs, indicating qualified document processing. When the click probability is less than or equal to the preset click probability, processing is abnormal, and the preset reference quantity is adjusted to holistically regulate the amount of index tag data to improve retrieval accuracy. Simultaneously, the document generation module's operating parameters are judged based on the average push quantity, which reflects the redundancy of the pushed documents output by the text generation module. When the average push quantity is greater than the preset push quantity, too many documents are pushed, making it difficult for users to find the information they need among numerous documents. Furthermore, when the number of pushed files is too large, the interface needs to render more content, requiring more computing resources and time for the layout, formatting, and display of generated links for each pushed document, leading to display delays on the push interface. In this case, the preset number of categories used to classify the relevance of a single batch of documents is adjusted to the corresponding value to redefine the criteria for classifying relevance. Improving document matching accuracy reduces the number of searches required by users, thereby reducing overall search time. By monitoring and analyzing user click behavior and the number of documents pushed, system parameters are dynamically adjusted to continuously optimize system performance, thus improving document retrieval efficiency.

[0125] Specifically, the preset number of pushes is selected within the range [5, 8]. Those skilled in the art can select the preset number of pushes according to the actual application scenario. Experiments can be conducted on document datasets of different sizes and types. The preset number of pushes is determined based on the display delay time of pushes under different pushes. It can be understood that it can ensure that there are enough documents for users to choose from without causing the display delay time to be too long.

[0126] Specifically, the average number of push documents output by the text generation module in a single run is the average number of push documents output by the text generation module in each run within a preset runtime.

[0127] In this embodiment, optionally,

[0128] Compare the average number of pushes with the first preset push comparison value and the second preset push comparison value;

[0129] If the average number of pushes is less than or equal to the first preset push comparison value, the number of preset categories will be adjusted to 1.11 times the initial number of preset categories.

[0130] If the average number of pushes is less than or equal to the second preset push comparison value and greater than the first preset push comparison value, then the number of preset categories will be adjusted to 1.18 times the initial number of preset categories.

[0131] If the average number of pushes is greater than the second preset push comparison value, the number of preset categories will be adjusted to 1.21 times the initial number of preset categories.

[0132] The first preset push comparison value is 1.31T0, and the second preset push comparison value is 2.21T0, where T0 is the preset push quantity.

[0133] Specifically, because the current document classification method is too broad, it misses some strongly related batches of documents, resulting in inaccurate document push matching. In this case, the number of preset categories is increased to classify documents more finely and make document push more accurate.

[0134] Specifically, the data processing module, after adjusting the number of preset categories, continuously monitors the click probability, and when the running time of the data receiving module reaches an integer multiple of the preset running time again;

[0135] The data processing module determines whether to correct the operating parameters for the text generation module based on historically changed parameters, including:

[0136] It is used to draw a probability time-domain curve based on the click probability within each preset runtime in the acquired historical data, and the slope of the calculated curve at the current time node is determined as the historical change parameter.

[0137] If the historical change parameters are greater than the preset change parameters, the operating parameters of the text generation module are deemed to be qualified.

[0138] If the historical change parameter is less than or equal to the preset change parameter, the preset accuracy coefficient is adjusted to the corresponding value based on the historical change parameter;

[0139] The increase in the preset accuracy coefficient is negatively correlated with the historical changes in parameters.

[0140] Specifically, the preset change parameters are selected within the range [0.042, 0.051]. Those skilled in the art can select the preset change parameters themselves and conduct experiments on document datasets of different sizes and types. The preset change parameters are determined based on the historical change parameters under different conditions. It can be understood that it is possible to divide the situation into whether the click probability has improved.

[0141] In this embodiment, optionally,

[0142] Compare the historical change parameters with the first preset change comparison value and the second preset change comparison value;

[0143] If the historical change parameter is less than or equal to the first preset change comparison value, the preset accuracy coefficient will be adjusted to 1.27 times the initial preset accuracy coefficient.

[0144] If the historical change parameter is less than or equal to the second preset change comparison value and greater than the first preset change comparison value, then the preset accuracy coefficient will be adjusted to 1.22 times the initial preset accuracy coefficient.

[0145] If the historical change parameter is greater than the second preset change comparison value, the preset accuracy coefficient will be adjusted to 1.16 times the initial preset accuracy coefficient.

[0146] The first preset change comparison value is 0.61Y0, and the second preset change comparison value is 0.82Y0, where Y0 is the preset change parameter.

[0147] Specifically, the system determines whether to adjust the operating parameters of the text generation module based on historical change parameters, which reflect the trend of click probability over time. When the historical change parameters are less than or equal to the second preset change comparison value and greater than the first preset change comparison value, it indicates that the click probability is increasing or stabilizing due to the adjustment of the preset category quantity. When the historical change parameters are less than or equal to the preset change parameters, the click probability is decreasing or showing little improvement. In this case, an anomaly exists in the matching between the document and the user's search data, resulting in the failure to accurately prioritize the documents requested by the user. In this situation, the operating parameters of the text generation module should be adjusted. The preset accuracy coefficient is adjusted based on the historical change parameters, and the weights of exact matching and fuzzy matching are adjusted in a timely manner according to the trend of click probability changes. This allows the system to continuously adapt to changes in documents and user needs, maintaining good push performance and further improving document query efficiency.

[0148] The technical solution of the present invention has been described above with reference to the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the scope of protection of the present invention.

[0149] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A structured XML standard document intelligent query assistance system based on a large model, characterized in that, include: The document comparison module is used to classify each document based on the comparison results obtained from each document and to obtain the number of categories; A data segmentation module, which is connected to the document comparison module, is used to determine the relevance of a single batch of documents based on the number of categories of the received single batch of documents. The data analysis module is connected to the data segmentation module and the document comparison module respectively, and generates index tags for each document. When the correlation of a single batch of documents is determined to be strong, the preset number of references used for document classification is adjusted. The text generation module is used to generate several push documents in descending order of relevance based on the search data input by the user; The data processing module, connected to both the document comparison module and the data segmentation module, is used to determine whether the processing of each document is qualified based on the click probability of the top-ranked push document clicked by users, and to adjust the preset reference quantity to the corresponding value when the processing of each document is determined to be abnormal. It also determines whether the operation of the text generation module is qualified based on the average number of push documents output by the text generation module per batch, and adjusts the preset number of categories used to classify the relevance of a single batch of documents to the corresponding value when the operation of the text generation module is determined to be abnormal. The document comparison module is used to classify each document based on the comparison results obtained, and to obtain the number of categories, including: This is used to obtain the high-frequency words of a single document. The high-frequency words are sorted in descending order according to their frequency of occurrence. A preset number of high-frequency words are selected as the representative words for a single document. Used to compare the characteristic words of each document in a single batch of documents received by the data receiving module; This is used to group documents with the same key terms into the same category; The data segmentation module is used to determine the relevance of a single batch of documents based on the number of categories in the single batch of documents received by the data receiving module, including: If the number of categories is less than or equal to the preset number of categories, the relevance of a single batch of documents is determined to be strong, and the preset reference number used for document classification is adjusted to the corresponding value based on the number of categories. If the number of categories is greater than the preset number of categories, the correlation of a single batch of documents is determined to be weak, and the data comparison module is controlled to continue running using the current operating parameters; The data analysis module is used to generate index tags for each document based on the relevance of a single batch of documents, including: If the correlation of a single batch of documents is determined to be strong, the preset reference number used for document classification will be adjusted to the corresponding value based on the number of documents in the single batch. If the relevance of a single batch of documents is determined to be weak, then each characteristic term corresponding to the document is determined as the index label of the corresponding document; The data analysis module is used to adjust the preset reference quantity for document classification to a corresponding value based on the quantity of documents in a single batch. The increase in the number of preset references is positively correlated with the number of documents in a single batch. The data analysis module is used to control the document comparison module to redetermine the characteristic words of each document after adjusting the preset number of references. The data analysis module is used to determine the redefined key terms corresponding to a document as the index tags for that document.

2. The intelligent query assistance system for structured XML standard documents based on a large model according to claim 1, characterized in that, The text generation module is used to generate personalized search text based on user-input search data, and to perform a search to generate several push documents arranged in descending order of relevance representation values, including: It is used to match the keywords of the exclusive search text with the index tags of each document in the database, and the product of the number of exactly matched keywords and the preset precision coefficient is determined as the first association value of the corresponding search tag; It is used to perform semantic analysis and fuzzy matching on keywords and index tags of the exclusive search text, so as to determine the second association value of the corresponding search tag by multiplying the number of fuzzy matched keywords with the preset fuzzy coefficient. The sum of the first association value and the second association value is used to determine the association characteristic value of the corresponding search tag; This is used to compare the retrieved data with each index tag, and to identify the documents corresponding to each index tag whose selected relevance value is greater than the preset relevance comparison value as the push documents.

3. The intelligent query assistance system for structured XML standard documents based on a large model according to claim 2, characterized in that, The data processing module is used to determine whether the processing of each document is qualified based on the click probability, provided that the running time of the data receiving module reaches an integer multiple of the preset running time. This includes: If the click probability is less than or equal to the preset click probability, it is determined that the processing of each document is abnormal. Based on the click probability, the preset reference quantity is adjusted to the corresponding value. Based on the number of push documents output by the text generation module, it is determined whether the operating parameters of the text generation module are qualified. The increase in the preset reference quantity is negatively correlated with the click probability.

4. The intelligent query assistance system for structured XML standard documents based on a large model according to claim 3, characterized in that, If the click probability is greater than the preset click probability, then the processing of each document is deemed qualified, and the current parameters are used to continue processing each document.

5. The intelligent query assistance system for structured XML standard documents based on a large model according to claim 4, characterized in that, The data processing module is used to determine whether the operating parameters of the text generation module are qualified based on the average number of pushes, including: If the average number of pushes is greater than the preset number of pushes, it is determined that the text generation module is malfunctioning. Based on the average number of pushes, the preset number of categories used to classify the relevance of a single batch of documents is adjusted to the corresponding value. The increase in the number of preset categories is positively correlated with the average number of push notifications.

6. The intelligent query assistance system for structured XML standard documents based on a large model according to claim 5, characterized in that, The data processing module continuously monitors the click probability after adjusting the number of preset categories, and when the running time of the data receiving module reaches an integer multiple of the preset running time again; The data processing module determines whether to correct the operating parameters for the text generation module based on historically changed parameters, including: It is used to draw a probability time-domain curve based on the click probability within each preset runtime in the acquired historical data, and the slope of the calculated curve at the current time node is determined as the historical change parameter. If the historical change parameters are greater than the preset change parameters, the operating parameters of the text generation module are deemed to be qualified. If the historical change parameter is less than or equal to the preset change parameter, the preset accuracy coefficient is adjusted to the corresponding value based on the historical change parameter; The increase in the preset accuracy coefficient is negatively correlated with the historical changes in parameters.

Citation Information

Patent Citations

  • Query and index over documents

    CN105027115A

  • Resilient document queries

    US20030088550A1