Data element retrieval and archiving system and method for natural language understanding of text
By designing a data element retrieval and archive system for text natural language understanding, the problems of low text retrieval efficiency and semantic information ignorance in the existing technology are solved, and personalized and intelligent search strategies are realized, which improves the accuracy of text recommendations and user satisfaction.
Patent Information
- Application Number
- CN202510107310.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-27
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art is inefficient in retrieval when processing large-scale text data, has high index construction and maintenance costs, and ignores the semantic information of the text, resulting in a deviation from the user's intention and a lack of a personalized sorting mechanism.
A data element retrieval and archiving system for text natural language understanding was designed. Through multi-source text preprocessing, text complexity evaluation and archiving, keyword expansion strategies and data element tracking, a personalized and intelligent search strategy was realized, and secondary recommendations were made through user behavior analysis and text association analysis.
It improves the accuracy and user satisfaction of text recommendations, optimizes information processing efficiency and accuracy, enhances user experience, and provides more accurate and valuable text recommendations for corporate decision-making.
Smart Images

Figure CN120045641A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data element retrieval and archiving, and specifically to a data element retrieval and archiving system and method for natural language understanding of text. Background Art
[0002] With the rapid development of information technology, the generation and accumulation of a large amount of text data have made information retrieval and archiving an important research topic. Enterprises generate and receive a large amount of text data every day, including reports, emails, market research, social media feedback, etc. When formulating strategic and operational decisions, enterprises need to quickly obtain relevant information and insights. Effective text retrieval and archiving can help decision-makers find the required materials more quickly and improve the quality of decisions.
[0003] Traditional information retrieval methods are unable to cope when dealing with large-scale data sets, resulting in low retrieval efficiency. Existing indexing technologies may not be able to effectively support complex query requirements, the cost of index construction and maintenance is too high, and many existing systems only perform retrieval based on keyword matching, ignoring the semantic information of the text, resulting in a large deviation between the retrieval results and the actual intention of the user's query; at the same time, the sorting of retrieval results often only depends on simple relevance scores, lacking a sorting mechanism that takes into account the personalized needs and preferences of users, and the feedback of retrieval results is often ignored, and it is unable to optimize and adjust itself according to user feedback.
[0004] Therefore, in view of the above problems, there is an urgent need for a data element retrieval and archiving system and method for natural language understanding of text. Summary of the Invention
[0005] Aiming at the deficiencies of the prior art, the present invention provides a data element retrieval and archiving system and method for natural language understanding of text, which solves the problem of how to efficiently process multi-source text data, and improves the accuracy of text recommendation and user satisfaction through personalized and intelligent retrieval strategies, as well as a secondary recommendation mechanism based on user behavior analysis and text association analysis.
[0006] To achieve the above objectives, the present invention is implemented through the following technical solutions: A data element retrieval and archiving system for natural language understanding of texts, including a data acquisition and preprocessing module, a text complexity evaluation and archiving module, a retrieval recommendation module, and a secondary recommendation module, where: The data acquisition and preprocessing module is used to acquire multi-source texts, preprocess the multi-source texts, and establish an initial text library; The text complexity evaluation and archiving module is used to analyze the complexity of texts based on text data element information and text structure information, determine the complexity level of the texts, and then, according to the archiving strategy corresponding to the complexity level of the texts, combine with the initial text library to archive texts of different complexity levels respectively; The retrieval recommendation module is used to receive keywords input by users, perform text retrieval based on keyword expansion strategies and data element tracking, and then generate a text recommendation folder; The secondary recommendation module is used to analyze the user access data of each text in the text recommendation folder, mark the texts of interest to users in the text recommendation folder, and then perform text screening based on the correlation analysis of the texts of interest to users in their respective archives and the usability information analysis of each text in their respective archives to generate a secondary text recommendation folder.
[0007] Further, the specific analysis of acquiring multi-source texts, preprocessing the multi-source texts, and establishing an initial text library is as follows: Obtain the sources of multi-source texts, and the sources of multi-source texts specifically include enterprise websites, enterprise databases, enterprise document management systems, and enterprise-related social media; Obtain the multi-source texts from the sources of multi-source texts, and perform data preprocessing on the multi-source texts. The data preprocessing specifically includes detecting and deleting duplicate text content, uniformly converting the text formats of the multi-source texts, performing word segmentation on the multi-source texts to generate a set of words for each multi-source text, performing part-of-speech analysis and marking on the set of words for each multi-source text, and removing stop words and common words from the set of words for each multi-source text; Convert each preprocessed multi-source text into a feature vector respectively, and then store the multi-source texts in the constructed database and provide a multi-source text index. The specific table structure of the constructed database includes multi-source text ID, original multi-source text, and the field of the converted multi-source text.
[0008] Further, the specific analysis of analyzing the complexity of texts based on text data element information and text structure information is as follows: Obtain the text element information and text structure information of the texts; The text element information specifically includes the proportion of complex words in the text and the information density of the text. Normalize the proportion of complex words in the text and the information density of the text, and then take the average of the proportion of complex words in the text and the information density of the text as the complexity of the text data elements; The text structure information specifically includes the number of text sentence types and the number of topic supporting sentences. Normalize the number of text sentence types and the number of topic supporting sentences, and then take the average of the number of text sentence types and the number of topic supporting sentences as the complexity of the text structure.
[0009] Further, the specific analysis for determining the complexity level of the text is as follows: obtain the upper and lower thresholds of the complexity of text data elements and the upper and lower thresholds of the complexity of text structure; the upper and lower thresholds of the complexity of text data elements specifically include the upper threshold of the complexity of text data elements and the lower threshold of the complexity of text data elements; the upper and lower thresholds of the complexity of text structure specifically include the upper threshold of the complexity of text structure and the lower threshold of the complexity of text structure; when the complexity of text data elements is less than or equal to the lower threshold of the complexity of text data elements and the complexity of text structure is less than or equal to the lower threshold of the complexity of text structure, mark the complexity level of the text as text first-level complexity. When the complexity of text data elements is greater than or equal to the upper threshold of the complexity of text data elements, the complexity of text structure is greater than the lower threshold of the complexity of text structure and less than the upper threshold of the complexity of text structure, or the complexity of text structure is less than or equal to the lower threshold of the complexity of text structure, mark the complexity level of the text as text second-level complexity; when the complexity of text data elements is greater than the lower threshold of the complexity of text data elements and less than the upper threshold of the complexity of text data elements, the complexity of text structure is greater than or equal to the upper threshold of the complexity of text structure, or the complexity of text structure is greater than the lower threshold of the complexity of text structure and less than the upper threshold of the complexity of text structure, or the complexity of text structure is less than or equal to the lower threshold of the complexity of text structure, mark the complexity level of the text as text second-level complexity; when the complexity of text data elements is less than or equal to the lower threshold of the complexity of text data elements, the complexity of text structure is greater than or equal to the upper threshold of the complexity of text structure, or the complexity of text structure is greater than the lower threshold of the complexity of text structure and less than the upper threshold of the complexity of text structure, mark the complexity level of the text as text second-level complexity; when the complexity of text data elements is greater than or equal to the upper threshold of the complexity of text data elements and the complexity of text structure is greater than or equal to the upper threshold of the complexity of text structure, mark the complexity level of the text as text third-level complexity.
[0010] Further, the specific analysis for archiving texts with different complexity levels respectively in combination with the initial text library according to the archiving strategy corresponding to the complexity level of the text is as follows: for the text corresponding to the text first-level complexity, the archiving strategy is specifically: store and archive it using a lightweight database, and set the access permission to open access; for the text corresponding to the text second-level complexity, the archiving strategy is specifically: store and archive it using a relational database, and set the access permission to restricted access; for the text corresponding to the text third-level complexity, the archiving strategy is specifically: perform structured storage and archiving using a document management system, and provide backup and recovery functions, set the access permission to restricted access, and perform secondary restriction through an access key.
[0011] Furthermore, receiving keywords input by users, performing text retrieval according to keyword expansion strategy and data element tracking, and then generating a specific analysis of text recommendation folders are as follows: the keyword expansion strategy specifically includes synonym and near-synonym expansion, related word expansion and semantic expansion; obtaining keyword expansion information based on the keywords input by users, the keyword expansion information specifically includes synonyms of keywords, near-synonyms of keywords, keyword-related words and keyword semantic meaning words; generating an extended search set based on the keyword expansion information, and determining data elements according to keywords, the data elements specifically include author, source and subject tags; searching in the text library using the extended search set and data elements, and then marking the retrieved texts; screening the marked texts according to the complexity level and timestamp of the marked texts, screening out texts with a complexity level of text level one complexity and a creation time outside the defined time range, and then combining the remaining marked texts into a text recommendation folder, and extracting the text summary and tags of each remaining marked text to generate a text outline.
[0012] Furthermore, the user access data of each text in the text recommendation folder is analyzed, and the user-focused texts in the text recommendation folder are marked. Then, text screening is performed based on the association analysis of the user-focused texts in the corresponding archives and the usability information analysis of each text in the corresponding archives. The specific analysis to generate the secondary text recommendation folder is as follows: obtaining the user access data of each text in the text recommendation folder, the user access data specifically includes click-through rate, browsing time, download status, and reprint status; obtaining the user behavior pattern threshold, the user behavior pattern threshold specifically includes the click-through rate threshold and the browsing time threshold; comparing the click-through rate and browsing time of each text in the text recommendation folder with the click-through rate threshold and the browsing time threshold, and analyzing in combination with the download status and the reprint status; when the click-through rate of the text in the text recommendation folder is greater than the click-through rate threshold , and the browsing time is greater than the browsing time threshold, and there are both downloading and reprinting situations, marking the text as the user's attention text; extracting the text outline of the user's attention text, using the text outline of the user's attention text to search in the corresponding archives, identifying the association information between the text outline of the user's attention text and each text in the corresponding archives, the association information is specifically the text similarity; using the text similarity to filter out the text in the corresponding archives whose text similarity is greater than the text similarity threshold, and then obtaining the usability information of the filtered text, the usability information specifically includes the complexity level and timestamp of the text, filtering out the text with a complexity level of the first-level complexity of the text and a creation time outside the defined time range, and then combining the remaining texts into a secondary text recommendation folder, and extracting the text summary and label of each remaining text to generate a text outline.
[0013] A method for retrieving and archiving data elements for text natural language understanding, which applies the above-mentioned system for retrieving and archiving data elements for text natural language understanding, includes the following steps: obtaining multi-source texts, preprocessing the multi-source texts, and establishing an initial text library; analyzing the complexity of the texts based on text data element information and text structure information, determining the complexity level of the texts, and then, according to the archiving strategy corresponding to the complexity level of the texts, combining with the initial text library, archiving texts with different complexity levels respectively; receiving keywords input by the user, performing text retrieval based on the keyword expansion strategy and data element tracking, and then generating a text recommendation folder; analyzing the user access data of each text in the text recommendation folder, marking the texts concerned by the user in the text recommendation folder, and then screening the texts based on the correlation analysis of the texts concerned by the user in the affiliated archives and the usability information analysis of each text in the affiliated archives, and generating a secondary text recommendation folder.
[0014] The present invention has the following beneficial effects:
[0015] The system and method for retrieving and archiving data elements for text natural language understanding preprocess multi-source texts and establish an initial text library, ensuring the consistency and processability of data and laying a solid foundation for subsequent analysis and archiving. Based on the analysis of text complexity, it can select archiving strategies targeted, avoiding the "one-size-fits-all" processing method, thus improving the efficiency and accuracy of information processing; through the keyword expansion strategy and data element tracking, it can more accurately understand the user's query intention and retrieve relevant texts from the huge text library accordingly. The personalized retrieval method not only improves the retrieval accuracy but also enhances the user experience. The generated text recommendation folder further reflects the intelligent characteristics and helps users quickly find the content they are interested in; through the analysis of user access data, it can identify the texts concerned by the user, and then screen the texts based on the correlation analysis of these texts in the affiliated archives and the usability information analysis, generating a secondary text recommendation folder, which not only helps to optimize the organizational structure of text resources, improve resource utilization efficiency, but also provides more accurate and valuable text recommendations for users, provides strong support for enterprise decision-making, promotes business innovation and development, and at the same time helps to build a perfect knowledge base system, enhancing the enterprise's knowledge management level and competitiveness. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 It is a flowchart of the structure for retrieving and archiving data elements for text natural language understanding according to the present invention.
[0017] Figure 2 It is a flowchart of the method for retrieving and archiving data elements for text natural language understanding according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0018] Embodiments of the present application implement an information retrieval and archiving system and method for text natural language understanding, which can efficiently extract information from multi-source text data. Through intelligent retrieval and personalized recommendation, accurate and valuable text content is provided for users.
[0019] The general idea for the problems in the embodiments of the present application is as follows:
[0020] First, collect text data from multiple sources to ensure the diversity and comprehensiveness of the data. Preprocess the multi-source text collected, and integrate the preprocessed text data into a unified initial text library, providing a basis for subsequent analysis and archiving. Based on the text data element information and text structure information, analyze the complexity of the text. According to the analysis results, divide the text into different complexity levels. For texts with different complexity levels, select corresponding archiving strategies for archiving to achieve the orderly organization and efficient management of text resources. Receive the keywords input by the user, and based on the keyword expansion strategy, generate more comprehensive query conditions. Use the data element tracking technology to retrieve the texts related to the query conditions in the initial text library. According to the retrieval results, generate a text recommendation folder to provide personalized text recommendation services for users. Analyze the user access data of each text in the text recommendation folder to understand the user's preferences and behavior patterns. Based on the user access data, mark the texts that the user is interested in. According to the correlation analysis of the texts that the user is interested in in their respective archives and the usability information analysis of each text in the respective archives, perform text screening to generate a secondary text recommendation folder to provide more accurate and valuable text recommendations for users.
[0021] Please refer to Figure 1 Embodiments of the present invention provide a technical solution: A data element retrieval and archiving system for text natural language understanding, including a data acquisition and preprocessing module, a text complexity evaluation and archiving module, a retrieval and recommendation module, and a secondary recommendation module, where: The data acquisition and preprocessing module is used to acquire multi-source text, preprocess the multi-source text, and establish an initial text library; The text complexity evaluation and archiving module is used to analyze the complexity of the text based on the text data element information and text structure information, determine the complexity level of the text, and then, according to the archiving strategy corresponding to the complexity level of the text, combine with the initial text library to archive texts with different complexity levels respectively; The retrieval and recommendation module is used to receive the keywords input by the user, perform text retrieval based on the keyword expansion strategy and data element tracking, and then generate a text recommendation folder; The secondary recommendation module is used to analyze the user access data of each text in the text recommendation folder, mark the texts that the user is interested in in the text recommendation folder, and then perform text screening according to the correlation analysis of the texts that the user is interested in in their respective archives and the usability information analysis of each text in the respective archives to generate a secondary text recommendation folder.
[0022] Specifically, the specific analysis of obtaining multi-source text, preprocessing the multi-source text, and establishing an initial text library is as follows: Obtain the sources of multi-source text, which specifically include enterprise websites, enterprise databases, enterprise document management systems, and enterprise-related social media; Obtain the multi-source text from the sources of multi-source text, and perform data preprocessing on the multi-source text. The data preprocessing specifically includes detecting and deleting duplicate text content, uniformly converting the text format of the multi-source text, segmenting the multi-source text to generate a set of words for each multi-source text, performing part-of-speech analysis and tagging on the set of words for each multi-source text, and removing stop words and common words from the set of words for each multi-source text; Convert each preprocessed multi-source text into a feature vector respectively, and then store the multi-source text in the constructed database and provide a multi-source text index. The specific table structure of the constructed database includes multi-source text ID, original multi-source text, and the field of the multi-source text after conversion.
[0023] In this implementation plan, to detect and delete duplicate text content, a text deduplication algorithm is specifically used, such as a method based on a hash function or a comparison based on text similarity to delete duplicate text; The uniform conversion of text format is specifically: converting texts in different formats into the same format, for example, converting HTML text into plain text format; To segment and generate a set of words, a segmentation tool (such as the jieba Chinese segmentation library or the NLTK English segmentation library) is specifically used to segment the text and generate a set of words for each text; The part-of-speech analysis and tagging is specifically: performing part-of-speech tagging on the segmented words for subsequent semantic analysis and information extraction; To remove stop words and common words, it is specifically based on a stop word list, and common words include verbs, nouns, adjectives, and adverbs; The conversion of feature vectors is specifically achieved through the Bag of Words model and TF-IDF (Term Frequency-Inverse Document Frequency).
[0024] In the constructed database table structure, the multi-source text ID is the unique identifier for each text, the original multi-source text is the original text content without processing; The field of the multi-source text after conversion is the text representation after preprocessing and feature vector conversion.
[0025] An index is established in the database to speed up the text retrieval speed, and it is considered to use a full-text search index or an inverted index to support fast text retrieval operations.
[0026] Specifically, the specific analysis of analyzing the complexity of text based on text data element information and text structure information is as follows: Obtain the text element information and text structure information of the text; The text element information specifically includes the proportion of complex words in the text and the text information density. Normalize the proportion of complex words in the text and the text information density, and then take the average value of the proportion of complex words in the text and the text information density as the complexity of the text data elements; The text structure information specifically includes the number of text sentence types and the number of topic supporting sentences. Normalize the number of text sentence types and the number of topic supporting sentences, and then take the average value of the number of text sentence types and the number of topic supporting sentences as the text structure complexity.
[0027] The specific analysis of determining the complexity level of the text is as follows: Obtain the upper and lower thresholds of the text data element complexity and the upper and lower thresholds of the text structure complexity; The upper and lower thresholds of the text data element complexity specifically include the upper threshold of the text data element complexity and the lower threshold of the text data element complexity; The upper and lower thresholds of the text structure complexity specifically include the upper threshold of the text structure complexity and the lower threshold of the text structure complexity; When the text data element complexity is less than or equal to the lower threshold of the text data element complexity, and the text structure complexity is less than or equal to the lower threshold of the text structure complexity, mark the complexity level of the text as text first-level complexity. When the text data element complexity is greater than or equal to the upper threshold of the text data element complexity, the text structure complexity is greater than the lower threshold of the text structure complexity and less than the upper threshold of the text structure complexity, or the text structure complexity is less than or equal to the lower threshold of the text structure complexity, mark the complexity level of the text as text second-level complexity; When the text data element complexity is greater than the lower threshold of the text data element complexity and less than the upper threshold of the text data element complexity, the text structure complexity is greater than or equal to the upper threshold of the text structure complexity, or the text structure complexity is greater than the lower threshold of the text structure complexity and less than the upper threshold of the text structure complexity, or the text structure complexity is less than or equal to the lower threshold of the text structure complexity, mark the complexity level of the text as text second-level complexity; When the text data element complexity is less than or equal to the lower threshold of the text data element complexity, the text structure complexity is greater than or equal to the upper threshold of the text structure complexity, or the text structure complexity is greater than the lower threshold of the text structure complexity and less than the upper threshold of the text structure complexity, mark the complexity level of the text as text second-level complexity; When the text data element complexity is greater than or equal to the upper threshold of the text data element complexity, and the text structure complexity is greater than or equal to the upper threshold of the text structure complexity, mark the complexity level of the text as text third-level complexity.
[0028] In this implementation plan, according to the filing strategies corresponding to the complexity levels of the text, combined with the initial text library, the specific analysis of filing texts with different complexity levels is as follows: For texts corresponding to the first-level complexity of the text, the filing strategy is specifically: using a lightweight database for storage and filing, and setting the access permission to open access; for texts corresponding to the second-level complexity of the text, the filing strategy is specifically: using a relational database for storage and filing, and setting the access permission to restricted access; for texts corresponding to the third-level complexity of the text, the filing strategy is specifically: using a document management system for structured storage and filing, and providing backup and recovery functions, setting the access permission to restricted access, and performing secondary restriction through an access key.
[0029] The proportion of complex words in the text represents the ratio of the number of complex words (such as technical terms and long words) in the text to the total number of words; the information density of the text represents the ratio of the amount of information conveyed in the text to the length of the text, reflecting the compactness and effectiveness of the information; the number of text sentence types reflects the diversity of the text; the number of topic support sentences represents the number of core sentences that support the topic, usually the key sentences in each paragraph or section, reflecting the logic and structure of the text.
[0030] The specific way to obtain the proportion of complex words in the text is: using a pre-defined list of complex words to calculate the definition of complex words, and then counting the number of complex words and comparing it with the total number of words to obtain the proportion of complex words; the information density of the text is evaluated by calculating the average amount of information in each sentence, and the entropy formula in information theory can also be used to quantify the information density of the text; perform normalization processing (such as Min-Max normalization) on the proportion of complex words and information density in the text, and convert their values to the same scale (for example, between 0 and 1); the specific way to obtain the number of text sentence types is: counting the number of different types of sentences (such as declarative sentences, interrogative sentences, exclamatory sentences, etc.); the number of topic support sentences is determined by using a topic model to count the number of sentences related to the text topic.
[0031] The upper and lower thresholds of the complexity of text data elements and the upper and lower thresholds of the complexity of text structure are used to define the boundaries of different complexity levels, helping to evaluate the complexity of the text.
[0032] Use visualization tools (such as Matplotlib or D3.js) to display the results of text complexity analysis, helping to better understand the complexity distribution of the text. Machine learning models (such as support vector machines and decision trees) can also be used to predict the text complexity level and train using existing labeled data.
[0033] The access key can be specifically generated using a random number generator, or a secure key can be generated using an encryption library (such as the `secrets` module in Python or `SecureRandom` in Java). For texts of different complexity levels, the lightweight database used in the text first-level complexity archiving strategy is suitable for storing small text data, enabling fast querying and access. Open access ensures that all users can access and view the text, and all users can directly access it through a public URL. The review mechanism for the text content sets regular reviews of the text content to ensure the accuracy and update of the information, and allows users to submit feedback and modification suggestions; the relational database used in the text second-level complexity archiving strategy is suitable for storing structured data, supports complex queries, and utilizes the relationships between tables to help maintain data consistency. Restricted access specifically requires user authentication to access, ensuring that only authorized users can view the text. A role management system is used to assign different access permissions according to user roles; the document management system used in the text third-level complexity archiving strategy provides comprehensive document management functions, including version control and metadata management, supports document classification and retrieval, improves information management efficiency, and sets strict access permissions for restricted access and the access key. Only specific user groups can access, and two-factor authentication is carried out through the access key to ensure the high security of the data. Regular audits of access records are set for the text to ensure compliance with regulations.
[0034] Quantifying text complexity through specific metrics makes the evaluation process more objective and systematic; archiving texts according to complexity levels can facilitate users to obtain information at different levels, improving the availability and retrieval efficiency of information; based on different settings of complexity levels and access permissions, it is ensured that users can access texts suitable for their comprehension abilities, thereby optimizing the user experience; using different types of databases and storage strategies can adapt to different application scenarios and requirements, facilitating the expansion and maintenance of the system.
[0035] Specifically, the specific analysis of receiving the keywords input by the user, performing text retrieval based on the keyword expansion strategy and data element tracking, and then generating a text recommendation folder is as follows: The keyword expansion strategy specifically includes synonym and near-synonym expansion, related word expansion, and semantic expansion; obtaining keyword expansion information based on the keywords input by the user, where the keyword expansion information specifically includes synonyms of the keywords, near-synonyms of the keywords, keyword-related vocabulary, and keyword semantic meaning vocabulary; generating an extended retrieval set based on the keyword expansion information, and determining data elements based on the keywords, where the data elements specifically include author, source, and topic tags; using the extended retrieval set and data elements to perform retrieval in the text library, and then marking the retrieved text; screening the marked text according to the complexity level and timestamp of the marked text, screening out the text with a complexity level of first-level complexity and a creation time outside the defined time range, and then combining the remaining marked text into the text recommendation folder, and extracting the text summary and tags of each remaining marked text to generate a text outline.
[0036] In this implementation plan, synonyms and near-synonyms use a thesaurus (such as WordNet or a Chinese synonym dictionary) to obtain synonyms and near-synonyms of the keywords for expansion; related words are identified through context analysis or statistical correlation calculation (such as a co-occurrence matrix) to expand related vocabulary with the keywords; semantic expansion specifically uses natural language processing (NLP) models (such as BERT, Word2Vec) to obtain the semantic meaning of the keywords and expand relevant semantic vocabulary; data element extraction specifically extracts metadata from the documents in the text library to obtain the author, source, and topic tags of each document to ensure the accuracy of the data.
[0037] The click-through rate threshold is obtained through historical data analysis and is usually a percentile of the click-through rate. For example, the top 70% or 80% of the click-through rates are selected as the threshold; the browsing duration threshold is also determined by analyzing historical data to determine the average browsing duration, and a certain percentile higher than the average is selected as the threshold.
[0038] By performing synonym, near-synonym, related word, and semantic expansion on the keywords input by the user, the retrieval range can be expanded, thereby increasing the chance of finding relevant texts and ensuring that users can obtain more content related to their needs; based on the specific input and preferences of the user, the recommended content generated is more personalized and can better meet the needs of the user, thereby improving user satisfaction and engagement; screening the retrieved text according to the complexity level can ensure the quality and adaptability of the recommended content, avoiding presenting texts that are not suitable for the user's reading level or comprehension ability to the user, and optimizing the user's reading experience; by extracting text summaries and tags, it not only provides users with an information structure that is easy to understand but also facilitates users to quickly browse and select the content they are interested in, improving the efficiency of information management.
[0039] Specifically, analyze the user access data of each text in the text recommendation folder, mark the texts concerned by users in the text recommendation folder, and then screen the texts based on the correlation analysis of the texts concerned by users in their respective archives and the usability information analysis of each text in their respective archives. The specific analysis for generating the secondary text recommendation folder is as follows: Obtain the user access data of each text in the text recommendation folder, and the user access data specifically includes click-through rate, browsing duration, download situation, and reposting situation; Obtain the user behavior pattern thresholds, and the user behavior pattern thresholds specifically include click-through rate threshold and browsing duration threshold; Compare the click-through rate and browsing duration of each text in the text recommendation folder with the click-through rate threshold and browsing duration threshold, and analyze in combination with the download situation and reposting situation. When the click-through rate of the text in the text recommendation folder is greater than the click-through rate threshold, the browsing duration is greater than the browsing duration threshold, and there are both download situation and reposting situation, mark the text as a text concerned by users; Extract the text outline of the text concerned by users, use the text outline of the text concerned by users to search in their respective archives, and identify the correlation information between the text outline of the text concerned by users and each text in their respective archives. The correlation information is specifically text similarity; Use the text similarity for screening, screen out the texts in the respective archives with text similarity greater than the text similarity threshold, and then obtain the usability information of the screened texts. The usability information specifically includes the complexity level and timestamp of the text, screen out the texts with a complexity level of text first-level complexity and a creation time outside the defined time range, and then combine the remaining texts into the secondary text recommendation folder, and extract the text summary and tags of each remaining text to generate a text outline.
[0040] In this implementation, the click-through rate is calculated by recording the ratio of the number of times a user clicks on a piece of text to the number of times that text is displayed; the browsing duration is obtained by analyzing the time a user spends on a text page, usually starting to count when the page loads and stopping when the user leaves the page; the download situation is counted through the download function of the text management system, recording the number of times a user downloads a text; the reprint situation is obtained by tracking the reprint situation of the text on social media or other platforms; the click-through rate threshold is obtained through historical data analysis, usually the percentile of the click-through rate, for example, selecting the click-through rates of the top 70% or 80% as the threshold; the browsing duration threshold is also determined by analyzing historical data to determine the average browsing duration and selecting a certain percentile higher than the average as the threshold; the text similarity is obtained by using cosine similarity to represent the text as a vector and calculating its cosine value, or by calculating the ratio of the number of shared words between two texts to the number of words in the union using Jaccard similarity. It is also possible to vectorize the text based on term frequency and inverse document frequency through TF-IDF and then calculate the similarity; the text similarity threshold is determined by analyzing the similarity distribution of similar text pairs from historical similarity calculation data, selecting the similarity at a certain high percentile (such as 80%) as the threshold, or performing cross-validation to determine the optimal threshold.
[0041] By analyzing user access data, more text recommendations that better meet user needs can be provided, enhancing user satisfaction; using user behavior data to mark texts of user interest can make the recommendations more accurate and avoid information overload; as user behavior changes, the recommendation system can quickly adjust the recommendation strategy to ensure the timeliness of the recommended content; through association analysis, the relationships between texts of user interest and other texts can be discovered, promoting deeper information mining and recommendations; by combining multiple access metrics, a comprehensive evaluation of the texts can be conducted to improve the scientific nature of the recommendations.
[0042] A method for retrieving and archiving data elements of text natural language understanding, which applies the above-mentioned system for retrieving and archiving data elements of text natural language understanding, includes the following steps: obtaining multi-source texts, preprocessing the multi-source texts, and establishing an initial text library; analyzing the complexity of the texts based on text data element information and text structure information to determine the complexity level of the texts, and then, according to the archiving strategy corresponding to the complexity level of the texts, combining with the initial text library to archive texts of different complexity levels respectively; receiving keywords input by the user, performing text retrieval based on the keyword expansion strategy and data element tracking, and then generating a text recommendation folder; analyzing the user access data of each text in the text recommendation folder, marking the texts of user interest in the text recommendation folder, and then screening the texts based on the association analysis of the texts of user interest in their respective archives and the usability information analysis of each text in the respective archives to generate a secondary text recommendation folder.
[0043] In summary, the present application has at least the following effects:
[0044] By integrating multi-source documents and establishing an initial text library, it is possible to capture and analyze information more comprehensively, improve data utilization efficiency and retrieval accuracy; introducing an evaluation mechanism for text complexity, classifying texts into different complexity levels according to their structural and content characteristics, and formulating corresponding filing strategies for texts of different complexities can optimize the filing process and improve the flexibility and accuracy of the system; considering the keyword expansion strategy and data element tracking, it is possible to retrieve and recommend relevant documents more intelligently, enhancing the user experience; through the analysis of user access data, it is possible to provide real-time feedback on the user's interest in documents, thereby optimizing the relevance of recommendations and enhancing the personalized recommendation ability; the analysis of the relevance of documents in the recommended folder can recommend relevant documents again based on the user's access, and this multi-level recommendation strategy can further improve the utilization rate of documents and user satisfaction.
[0045] Those skilled in the art should understand that the embodiments of the present invention can be provided as methods and systems. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program code.
[0046] The present invention is described with reference to the flowcharts and structural diagrams of methods and systems according to the embodiments of the present invention. It should be understood that each process and module combination in the flowcharts and structural diagrams can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one process or multiple processes and structures Figure 1 one module or multiple modules.
[0047] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one process or multiple processes and structures Figure 1 one module or multiple modules.
[0048] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are performed on the computer or other programmable apparatus to produce a computer-implemented process, thereby providing instructions for implementing the process Figure 1 a process or processes and architectures Figure 1 steps for specifying functions in one or more modules.
[0049] Although the preferred embodiments of the present invention have been described, additional changes and modifications can be made by those skilled in the art once they learn the basic creative concept. Therefore, the appended claims are intended to be construed to cover the preferred embodiments as well as all changes and modifications falling within the scope of the present invention.
[0050] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.
Claims
1. A data element retrieval and archiving system for natural language understanding of text, characterized in that: It includes data acquisition and preprocessing module, text complexity evaluation and archiving module, retrieval recommendation module and secondary recommendation module, among which: The data acquisition and preprocessing module is used to acquire multi-source texts, preprocess the multi-source texts, and establish an initial text library; The text complexity evaluation and archiving module is used to analyze the complexity of the text based on the text data element information and the text structure information, determine the complexity level of the text, and then archive texts of different complexity levels in combination with the initial text library according to the archiving strategy corresponding to the complexity level of the text; The search recommendation module is used to receive keywords input by users, perform text search based on keyword expansion strategy and data element tracking, and then generate text recommendation folders; The secondary recommendation module is used to analyze the user access data of each text in the text recommendation folder, mark the user-focused texts in the text recommendation folder, and then perform text screening based on the association analysis of the user-focused texts in the corresponding archives and the usability information analysis of each text in the corresponding archives to generate a secondary text recommendation folder.
2. A data element retrieval and archiving system for natural language understanding of text according to claim 1, characterized in that: The specific analysis of obtaining multi-source texts, preprocessing them, and establishing the initial text library is as follows: Acquire multi-source text sources, wherein the multi-source text sources specifically include enterprise websites, enterprise databases, enterprise document management systems, and enterprise-related social media; Acquire multi-source texts from the multi-source text sources, and perform data preprocessing on the multi-source texts, wherein the data preprocessing specifically includes detecting and deleting repeated text content, converting the multi-source texts into a unified text format, segmenting the multi-source texts to generate word sets of each multi-source text, performing part-of-speech analysis and marking on the word sets of each multi-source text, and removing stop words and common words from the word sets of each multi-source text; The preprocessed multi-source texts are respectively converted into feature vectors, and then the multi-source texts are stored in a constructed database and a multi-source text index is provided. The specific table structure of the constructed database includes multi-source text IDs, original multi-source texts, and converted multi-source text fields.
3. A data element retrieval and archiving system for natural language understanding of text according to claim 1, characterized in that: The specific analysis of the complexity of analyzing the text based on the text data element information and the text structure information is as follows: Obtaining text element information and text structure information of the text; The text element information specifically includes the text complex vocabulary ratio and text information density, the text complex vocabulary ratio and text information density are normalized, and then the average value of the text complex vocabulary ratio and text information density is taken as the text data element complexity; The text structure information specifically includes the number of text sentence types and the number of topic supporting sentences. The number of text sentence types and the number of topic supporting sentences are normalized, and then the average value of the number of text sentence types and the number of topic supporting sentences is taken as the text structure complexity.
4. A data element retrieval and archiving system for natural language understanding of text according to claim 3, characterized in that: The specific analysis of determining the complexity level of the text is: Obtaining upper and lower thresholds of text data element complexity and upper and lower thresholds of text structure complexity; The upper and lower thresholds of the text data element complexity specifically include an upper threshold of the text data element complexity and a lower threshold of the text data element complexity; The upper and lower thresholds of the text structure complexity specifically include an upper threshold of text structure complexity and a lower threshold of text structure complexity; When the complexity of text data elements is less than or equal to the lower threshold of text data element complexity, and the text structure complexity is less than or equal to the lower threshold of text structure complexity, the complexity level of the marked text is the first level text complexity. When the complexity of text data elements is greater than or equal to the upper threshold of text data element complexity, the text structure complexity is greater than the lower threshold of text structure complexity and less than the upper threshold of text structure complexity, or the text structure complexity is less than or equal to the lower threshold of text structure complexity, the complexity level of the marked text is the second level of text complexity; When the complexity of text data elements is greater than the lower threshold of text data element complexity and less than the upper threshold of text data element complexity, the text structure complexity is greater than or equal to the upper threshold of text structure complexity, or the text structure complexity is greater than the lower threshold of text structure complexity and less than the upper threshold of text structure complexity, or the text structure complexity is less than or equal to the lower threshold of text structure complexity, the complexity level of the marked text is the second level of text complexity; When the complexity of the text data elements is less than or equal to the lower threshold of the text data elements complexity, the text structure complexity is greater than or equal to the upper threshold of the text structure complexity, or the text structure complexity is greater than the lower threshold of the text structure complexity and less than the upper threshold of the text structure complexity, the complexity level of the marked text is the second level of text complexity; When the complexity of text data elements is greater than or equal to the upper threshold of text data element complexity, and the text structure complexity is greater than or equal to the upper threshold of text structure complexity, the complexity level of the marked text is text level 3 complexity.
5. A data element retrieval and archiving system for natural language understanding of text according to claim 4, characterized in that: According to the archiving strategy corresponding to the complexity level of the text, the specific analysis of archiving texts of different complexity levels in combination with the initial text library is as follows: For the text corresponding to the first level of complexity, the archiving strategy is as follows: using a lightweight database for storage and archiving, and setting the access permission to open access; For the text corresponding to the second level of complexity of the text, the archiving strategy is specifically as follows: using a relational database for storage and archiving, and setting the access permission to restricted access; For the text corresponding to the third level of complexity, the archiving strategy is specifically: use the document management system for structured storage and archiving, and provide backup and recovery functions, set the access permission to restricted access, and perform secondary restrictions through access keys.
6. A data element retrieval and archiving system for natural language understanding of text according to claim 4, characterized in that: The specific analysis of receiving keywords input by users, performing text retrieval based on keyword expansion strategy and data element tracking, and then generating text recommendation folders is as follows: The keyword expansion strategy specifically includes synonym and near-synonym expansion, related word expansion and semantic expansion; Acquire keyword expansion information based on the keyword input by the user, the keyword expansion information specifically includes synonyms of the keyword, near-synonyms of the keyword, keyword-related words and keyword semantic meaning words; Generate an extended search set based on the keyword extension information, and determine data elements based on the keywords, wherein the data elements specifically include author, source, and subject tag; Searching the text library using the extended search set and data elements, and then marking the retrieved text; The marked texts are filtered according to their complexity levels and timestamps, and texts with a complexity level of text level 1 and a creation time outside the defined time range are filtered out. The remaining marked texts are then combined into text recommendation folders, and the text summaries and labels of each remaining marked text are extracted to generate a text outline.
7. A data element retrieval and archiving system for natural language understanding of text according to claim 6, characterized in that: The user access data of each text in the text recommendation folder is analyzed, and the user-focused texts in the text recommendation folder are marked. Then, the texts are screened based on the association analysis of the user-focused texts in the corresponding archives and the usability information analysis of each text in the corresponding archives. The specific analysis of generating the secondary text recommendation folder is as follows: Obtaining user access data of each text in the text recommendation folder, wherein the user access data specifically includes click rate, browsing time, download status, and reprint status; Obtaining a user behavior pattern threshold, wherein the user behavior pattern threshold specifically includes a click rate threshold and a browsing time threshold; Compare the click rate and browsing time of each text in the text recommendation folder with the click rate threshold and browsing time threshold, and analyze them in combination with the download and reprint situations. When the click rate of the text in the text recommendation folder is greater than the click rate threshold, and the browsing time is greater than the browsing time threshold, and there is both download and reprint situation, mark the text as the user's attention text; Extracting the text outline of the text that the user is interested in, searching the corresponding archives using the text outline of the text that the user is interested in, and identifying the association information between the text outline of the text that the user is interested in and each text in the corresponding archives, wherein the association information is specifically text similarity; The text similarity is used for screening, and the texts in the archives whose text similarity is greater than the text similarity threshold are screened out, and then the usability information of the screened texts is obtained, and the usability information specifically includes the complexity level and timestamp of the text, and the texts with the complexity level of the first-level text complexity and the creation time outside the defined time range are screened out, and then the remaining texts are combined into the secondary text recommendation folder, and the text summary and label of each remaining text are extracted to generate a text outline.
8. A method for retrieving and archiving data elements by natural language understanding of text, using a system for retrieving and archiving data elements by natural language understanding of text as claimed in any one of claims 1 to 7, characterized in that: The following steps are involved: Obtain multi-source texts, pre-process the multi-source texts, and establish an initial text library; Analyze the complexity of the text based on the text data element information and text structure information, determine the complexity level of the text, and then archive the texts of different complexity levels according to the archiving strategy corresponding to the complexity level of the text and the initial text library; Receive keywords input by users, perform text retrieval based on keyword expansion strategy and data element tracking, and then generate text recommendation folders; The user access data of each text in the text recommendation folder is analyzed, and the user-focused texts in the text recommendation folder are marked. Then, the texts are screened based on the correlation analysis of the user-focused texts in their respective archives and the usability information analysis of each text in the respective archives to generate a secondary text recommendation folder.
Citation Information
Cited By
Method for quickly changing business model definition into agent DB Description knowledge base
CN121807868A
A method to quickly convert business model definitions into an agent DBDescription knowledge base
CN121807868B