Traditional culture database construction method based on artificial intelligence

Through the traditional cultural database construction method based on artificial intelligence, the data collection frequency and search strategy are dynamically adjusted, and the cached content is optimized in combination with the literature popularity value, the problems of untimely data collection, low search efficiency and unreasonable cached content in the existing technology are solved, and more efficient data management and retrieval experience is achieved.

CN120216554AActive Publication Date: 2025-06-27SHANDONG POLYTECHNIC COLLEGE
View PDF 11 Cites 0 Cited by

Patent Information

Application Number
CN202510317751.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-06-27
Estimated Expiration
2045-03-18

AI Technical Summary

Technical Problem

The existing technology lacks flexibility in data collection, retrieval and cache management, and it is difficult to dynamically adjust strategies based on data source characteristics and user behavior, resulting in untimely data collection, low retrieval efficiency and unreasonable cache content.

Method used

Using a traditional cultural database construction method based on artificial intelligence, we analyze the update rules of data sources and user behavior, dynamically adjust the data collection frequency, build a query keyword collection to improve retrieval efficiency, and dynamically adjust the cache content according to the literature popularity value.

Benefits of technology

It realizes accurate matching of data collection frequency, improves retrieval efficiency and user experience, and ensures the timeliness and relevance of cached content, solving the problems of untimely data collection, low retrieval efficiency and unreasonable cached content in traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216554A_ABST
    Figure CN120216554A_ABST
Patent Text Reader

Abstract

The invention discloses a traditional culture database construction method based on artificial intelligence, and belongs to the technical field of database construction.The method comprises the steps that firstly, data source updating time points are analyzed to determine statistical duration, number source activeness and number source attention are calculated to obtain a frequency collection matching value, data source collection frequency is matched, and the problem that a traditional collection strategy is not flexible is solved; word segmentation is carried out on a user query statement, a query keyword set is constructed to determine a retrieval field, sorting is carried out according to literature matching values, and the retrieval efficiency is improved; and finally, collecting user behavior data, analyzing literature popularity values to determine cached literatures and cache points, caching the literatures and the cache points to a corresponding library, eliminating the literatures with low popularity when the space is insufficient, and updating the literatures every three days to adapt to literature popularity changes. Compared with the prior art, the method has the advantages that data source characteristics can be accurately mastered, acquisition strategies can be adjusted, valuable data can be quickly positioned, caches can be dynamically managed according to user behaviors and literature popularity, and timeliness, retrieval efficiency and user experience of databases can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of database construction, and particularly relates to a method for constructing a traditional culture database based on artificial intelligence. Background Art

[0002] In the era of digital information explosion, the dependence on literature data in academic research and other fields has increased dramatically. However, the literature is scattered in diverse data sources such as academic databases and professional forums, with significant differences in format, update frequency, and structure. In such an environment, many problems have emerged in the existing technologies;

[0003] Data collection level: Traditional data collection methods are difficult to flexibly adjust the collection strategy according to the characteristics of data sources;

[0004] Retrieval level: Due to the inability to predict complex query intentions, it takes a long time for users to screen results, and the response is slow;

[0005] Cache management level: Traditional cache mechanisms cannot dynamically adjust cache content according to the popularity of literature and user behavior. Therefore, we propose a method for constructing a traditional culture database based on artificial intelligence. Summary of the Invention

[0006] The purpose of the present invention is to provide a method for constructing a traditional culture database based on artificial intelligence to solve the problems raised in the above background art.

[0007] To achieve the above purpose, the present invention provides the following technical solutions: A method for constructing a traditional culture database based on artificial intelligence, including the following steps:

[0008] Step 1: Analyze the update time points of data sources to determine the statistical duration. Within the statistical duration, count the number of newly published documents, the updated quantity of existing documents, as well as the download volume, browsing duration, and collection times of documents, and calculate the data source activity and data source attention respectively; comprehensively analyze the statistical duration, data source activity, and data source attention to obtain a sampling frequency matching value, and match the collection frequency of data sources according to the sampling frequency matching value;

[0009] Step 2: Perform lexical segmentation on the user's query statement, construct a query keyword set by analyzing the frequency inverse value of the vocabulary, determine the retrieval field according to the query keyword set, analyze the document matching value of the documents within the retrieval field, and sort the documents within the retrieval field according to the document matching value;

[0010] Step 3: Collect the user's behavior data, analyze the literature popularity value according to the behavior data, analyze the cached literature according to the literature popularity value, determine the literature cache points by sending test data packets to each cache node, and cache the cached literature in the corresponding cache literature library of the literature cache points.

[0011] Preferably, in step one, the specific process of determining the statistical duration is as follows:

[0012] For each data source, first obtain the time points of the last six updates of the data source, calculate the time intervals between adjacent update time points, denoted as adjacent interval durations, and then use the standard deviation calculation formula to calculate the standard deviation value among all adjacent interval durations, denoted as adjacent standard deviation. Compare the adjacent standard deviation with the preset adjacent standard deviation threshold. If the adjacent standard deviation is greater than or equal to the corresponding threshold, recalculate the adjacent standard deviation in the way of adding two update time points in each batch until the update interval difference is less than the corresponding threshold and then stop; finally, add up all the adjacent update intervals corresponding to the operation when stopping calculating the adjacent standard deviation and divide by the number of intervals to obtain the average update duration of the data source;

[0013] Taking the time point when the data source last completed data update as the starting moment, and pushing forward the average update duration on this basis, the reached time point is the ending moment, and this duration is denoted as the statistical duration TC.

[0014] Preferably, in step one, the specific process of calculating the data source activity, data source attention, and sampling frequency matching value is as follows:

[0015] Obtain the number of new documents published by the data source and the number of updates of existing documents within the statistical duration, and assign different weight coefficients to both. After normalizing the number of new documents published by the data source and the number of updates of existing documents, multiply by the corresponding weight coefficients, and finally add them up to obtain the data source activity SH;

[0016] For each document in the data source, monitor the download volume, browsing duration, and collection times of the document within the statistical duration; then sum up the download volume, browsing duration, and collection times of all documents in the data source respectively to obtain the total document download volume WX, total document browsing duration WL, and total document collection times WS in the data source. After normalizing the total document download volume WX, total document browsing duration WL, and total document collection times WS in the data source, use the formula: GD = WX×a1 + WL×a2 + WS×a3 to obtain the data source attention GD, where a1, a2, and a3 are preset weight coefficients;

[0017] By using the formula with the statistical duration TC, data source activity SH, and data source attention GD: , the sampling frequency matching value CPZ is obtained, where p1, p2, and p3 are preset weight coefficients.

[0018] Preferably, in step one, the specific process of matching the data source's collection frequency according to the sampling frequency matching value is as follows:

[0019] Preset several frequency acquisition matching value intervals, and each frequency acquisition matching value interval corresponds to a data source acquisition frequency;

[0020] For each data source, by matching its corresponding frequency acquisition matching value with all frequency acquisition matching value intervals, output the acquisition frequency corresponding to the corresponding frequency acquisition matching value interval, and determine the literature data acquisition frequency of the data source through this acquisition frequency; preset an acquisition update adjustment factor k, where k is greater than 1, and use the formula: GT = TC × k to obtain the frequency acquisition update duration GT. Starting from the moment when the data source first performs data acquisition according to the determined acquisition frequency, when the elapsed interval duration since this starting point reaches the frequency acquisition update duration, immediately trigger the re-matching analysis process of the data source acquisition frequency.

[0021] Preferably, in step two, the specific process of analyzing the frequency inverse value of the vocabulary and constructing the query keyword set is as follows:

[0022] After the user inputs a query statement, split the query statement into several independent vocabulary words. For each split vocabulary word, count the number of times the vocabulary word appears in the query distance, and then divide the number of times the vocabulary word appears in the query distance by the total number of vocabulary words to obtain the vocabulary word frequency;

[0023] Organize all the literature data collected and obtained into a document set. For each vocabulary word, count the number of documents in the document set that include this vocabulary word, and organize it into a word-document set. By dividing the total number of documents in the document set by the total number of documents in the word-document set, obtain the inverse document frequency of the vocabulary word;

[0024] For each vocabulary word, multiply the vocabulary word frequency corresponding to the vocabulary word by the inverse document frequency of the vocabulary word to obtain the frequency inverse value. Compare the frequency inverse value corresponding to the vocabulary word with the preset keyword frequency inverse threshold. If the frequency inverse value corresponding to the vocabulary word is greater than or equal to the preset keyword frequency inverse threshold, mark the vocabulary word as a keyword corresponding to the query statement. By organizing the keywords corresponding to the user's query statement, form a keyword set, denoted as the query keyword set.

[0025] Preferably, in step two, the specific process of determining the retrieval field according to the query keyword set, analyzing the literature matching value of the literature in the retrieval field, and sorting the literature in the retrieval field according to the literature matching value is as follows:

[0026] Obtain each literature field involved in the document set. For each literature field, obtain the corresponding specific keyword set, denoted as the literature field keyword set;

[0027] For each literature field, by matching the set of query keywords with the set of literature field keywords, the number of matching keywords is output, denoted as the field matching value. The field matching value is compared with the preset field matching threshold. If the field matching value is greater than the preset field matching threshold, the literature field is marked as the retrieval field, and all retrieval fields are sorted to form a retrieval list;

[0028] For each retrieval field in the retrieval list, all the literatures in that field are sorted to form a field literature library. For each literature in the field literature library, the number of words in the literature that contain the words in the set of query keywords is counted, denoted as the literature matching value. All the literatures in the field literature library are sorted from high to low according to the size of the literature matching value.

[0029] Preferably, in step three, the specific process of analyzing the literature heat value is as follows:

[0030] Collect and record the query behavior data of each user. The query behavior data includes: query statement, query time, and corresponding retrieval results; the retrieval results include: the list of hit literatures, the retrieval fields where the literatures are located, the click time of the user on the literature, and the reading duration; at the same time, record the number of collections, comments, likes, and the literature release time of each literature in the literature database in real time;

[0031] For each user, obtain all the retrieval fields corresponding to all the query statements within the last three days, and generate an unordered list containing all the retrieval fields; then count the frequency of each different retrieval field in the list, denoted as the field duplicate check value. For each retrieval field, if the field duplicate check value is greater than the preset field duplicate check threshold, the retrieval field is marked as the cached field;

[0032] For each cached field, obtain the corresponding field literature library. For each literature in the field literature library, subtract the literature release time from the current time to obtain the literature release duration FT;

[0033] Try to obtain the most recent click time of the user on the literature. If the time can be obtained, subtract the most recent click time from the current time to obtain the literature idle duration. Preset the literature idle duration threshold. If the literature idle duration is greater than the preset literature idle duration threshold, then use the literature idle duration threshold as the literature idle duration; if there has never been a click record, directly use the literature idle duration threshold as the literature idle duration XT; obtain the comment volume PL, the number of collections SC, and the user reading duration YT of the literature;

[0034] After normalizing the comment volume PL, the number of collections SC, the user reading duration YT, the literature release duration FT, and the literature idle duration XT of the literature, use the formula: , the literature heat value WXD is obtained, where r1, r2, r3, r4, and r5 are preset weight coefficients.

[0035] Preferably, in step three, the cached literature is analyzed according to the literature heat value. The specific process of determining the literature cache point by sending test data packets to each cache node and caching the cached literature in the cache literature library corresponding to the literature cache point is as follows:

[0036] By comparing the literature heat value corresponding to the literature with a preset literature heat threshold, if the literature heat value is greater than the preset literature heat threshold, the literature is marked as cached literature;

[0037] Each cache node is set, and each cache node is correspondingly provided with a cache literature library. By sending test data packets to each cache node, the round-trip time from the user device to each cache node is measured, and the cache node with the shortest round-trip time is marked as the literature cache point, and all cached literatures are cached in the cache literature library corresponding to the literature cache point;

[0038] If the cache space of the cache literature library is insufficient, the priority elimination algorithm is adopted to sequentially eliminate the literatures with the lowest literature heat value;

[0039] A cache literature library update mechanism is set, and it is set to trigger an update operation every three days. When the update operation is triggered, all literatures in the cache literature library are deleted, and the cached literature is re-analyzed and stored in the cache literature library.

[0040] Preferably, a computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the above method are implemented.

[0041] Preferably, a computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.

[0042] Compared with the prior art, the beneficial effects of the present invention are:

[0043] (1). The method and system for constructing a traditional culture database based on artificial intelligence determine the average duration by calculating the adjacent duration and adjacent standard deviation of the data source, and then obtain the statistical duration, accurately grasping the update rule and stability of the data source; within the statistical duration, comprehensively consider the situation of the data source publishing new documents and updating old documents to measure the activity of the data source, and at the same time calculate the attention of the data source in combination with the download, browsing, and collection data of the documents by users; based on this, analyze the sampling frequency matching value to make the sampling frequency fit the actual update frequency of the data source, avoiding over-sampling or insufficient sampling; and set the sampling frequency update duration, recalculate the sampling frequency matching value when it expires, and adjust the sampling frequency in a timely manner to adapt to the dynamic changes of the data source, maintaining the timeliness and relevance of the database content, and solving the problem that it is difficult for traditional data collection methods to flexibly adjust the collection strategy according to the characteristics of the data source.

[0044] (2). The method and system for constructing a traditional culture database based on artificial intelligence, when a user inputs a query statement, segment the query statement, calculate the frequency inverse value of the vocabulary, construct a query keyword set according to the frequency inverse value, and understand the user's query intention; match the query keyword set with the domain keyword set to determine the retrieval field, and then sort the documents in the domain document library from high to low according to the document matching value between each document in the retrieval field and the query keyword set; this enables the newly constructed document database to display a retrieval field list when the user inputs a query statement. After clicking on a certain retrieval field, the documents in the corresponding domain document library will be presented to the user in descending order of the matching value, facilitating quick positioning to the most valuable materials for the user, greatly shortening the time-consuming for the user to screen the results, and effectively improving the retrieval efficiency and user experience, solving the problem that the user spends a long time screening the results and the response is slow due to the inability to predict complex query intentions.

[0045] (3). The method and system for constructing a traditional culture database based on artificial intelligence collect user query behavior data, generate an unordered list of all retrieval fields within the last three days, analyze the field duplicate detection value of each field, mark the cached fields after comparing with the preset threshold, and capture the fields that the user has frequently concerned about recently; for the documents in each cached field, comprehensively consider multi-dimensional factors such as the document release duration, user click and idle duration, comment volume, favorite times, and user reading duration to analyze the document heat value, and determine the documents that may be requested for access; determine the document cache point through the cache node positioning method based on network latency detection, store the documents that may be accessed in the corresponding cached document library, and the user can obtain data at a faster speed, significantly shortening the query response time; trigger the update operation of the cached document library every three days to adapt to the dynamic changes of the document heat, ensuring that the most likely documents to be accessed by the user are stored in the cached document library, and solving the problem that the traditional cache mechanism cannot dynamically adjust the cached content according to the document heat and user behavior. Brief Description of the Drawings

[0046] Figure 1 This is the flowchart of the present invention. Specific embodiments

[0047] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0048] Embodiment 1

[0049] Please refer to Figure 1 , the present invention provides a method for constructing a traditional culture database based on artificial intelligence, including: an intelligent data collection module, an intelligent retrieval module, and an intelligent distributed cache module;

[0050] The intelligent data collection module analyzes the update time points of the data sources to obtain the statistical duration. During the statistical duration, the number of newly published documents, the updated quantity of existing documents, as well as the download volume, browsing duration, and collection times of the documents are counted. The activity and attention of the data sources are calculated respectively. Finally, comprehensive analysis is performed based on the statistical duration, the activity of the data source, and the attention of the data source to obtain the acquisition frequency matching value, and the acquisition frequency of the data source is matched according to the acquisition frequency matching value. The specific process is as follows:

[0051] Communicate and connect with each data source through data acquisition technologies such as web crawler technology, data interface call technology, and sensor data acquisition technology;

[0052] For each data source, first obtain the time points of the last six updates of the data source, calculate the time interval between adjacent two update time points, denoted as the adjacent interval duration. Subsequently, use the standard deviation calculation formula to calculate the standard deviation value between all adjacent interval durations, denoted as the adjacent interval standard deviation. Preset the adjacent interval standard deviation threshold, and compare the adjacent interval standard deviation with the preset adjacent interval standard deviation threshold. If the adjacent interval standard deviation is greater than or equal to the corresponding threshold, it indicates that the update time interval fluctuates greatly and the stability is poor. Then, in the way of adding two update time points to each batch, recalculate the adjacent interval standard deviation until the update interval difference is less than the corresponding threshold and then stop; finally, add up all the adjacent update intervals corresponding to the operation of stopping calculating the adjacent interval standard deviation and divide by the number of intervals to obtain the average update duration of the data source;

[0053] Taking the time point when the data source last completed data update as the starting moment, and pushing the average update duration backward based on this, the reached time point is the ending moment. Denote this duration as the statistical duration TC;

[0054] Obtain the number of new documents published by the data source and the number of updates to existing documents within the statistical duration, and assign different weight coefficients to the two. After normalizing the number of new documents published by the data source and the number of updates to existing documents, multiply them by the corresponding weight coefficients, and finally add them together to obtain the data source activity SH;

[0055] For each document in the data source, monitor the download volume, browsing duration, and collection times of the document by users within the statistical duration; subsequently, sum up the download volume, browsing duration, and collection times of all documents in the data source respectively to obtain the total document download volume WX, the total document browsing duration WL, and the total document collection times WS in the data source. After normalizing the total document download volume WX, the total document browsing duration WL, and the total document collection times WS in the data source, use the formula: GD = WX×a1 + WL×a2 + WS×a3 to obtain the data source attention GD, where a1, a2, and a3 are preset weight coefficients. The larger the data source attention GD, the higher the attention of the documents presented in the data source by users;

[0056] By using the statistical duration TC, the data source activity SH, and the data source attention GD, with the formula: obtain the sampling frequency matching value CPZ, where p1, p2, and p3 are preset weight coefficients; the larger the sampling frequency matching value corresponding to the data source, the more frequent the update of the data source, and the higher the attention and activity of users to its content. It is necessary to perform data collection more frequently to ensure obtaining the latest and most concerned information and avoid data omission or obsolescence;

[0057] Preset several sampling frequency matching value intervals, and each sampling frequency matching value interval corresponds to a data source collection frequency. The larger the lower bound value and the upper bound value of the sampling frequency matching value interval, the smaller the corresponding collection frequency value, that is, the faster the collection frequency;

[0058] For each data source, by matching its corresponding sampling frequency matching value with all sampling frequency matching value intervals, output the collection frequency corresponding to the corresponding sampling frequency matching value interval, and determine the document data collection frequency of the data source through this collection frequency; preset a collection update adjustment factor k, where k is greater than 1, and use the formula: GT = TC×k to obtain the sampling frequency update duration GT. Starting from the moment when the data source first performs data collection according to the determined collection frequency, when the elapsed interval duration since this starting point reaches the sampling frequency update duration, immediately trigger the re-matching analysis process of the data source collection frequency.

[0059] It should be noted that by calculating the adjacent interval duration and the standard deviation of the adjacent intervals to determine the average duration, and then obtaining the statistical duration, the update rule and stability of the data source can be accurately grasped; based on this, the sampling frequency matching value is calculated, which can make the sampling frequency fit the actual update frequency of the data source, avoiding over-sampling or under-sampling; within the statistical duration, the activity of the data source is measured by comprehensively considering the situation of the data source publishing new documents and updating old documents, and at the same time, the attention of the data source is calculated by combining the download, browsing, and collection data of the documents by users, so that the sampling frequency integrates the self-update of the data source and the user's attention, ensuring that the collected documents are both new and popular, and improving the data value; by setting the sampling frequency update duration, the sampling frequency matching value is recalculated when it expires, and the sampling frequency is adjusted in a timely manner to adapt to the dynamic changes of the data source and maintain the timeliness and relevance of the database content.

[0060] The intelligent retrieval module performs lexical segmentation on the user's query statement, constructs a set of query keywords by analyzing the frequency-inverse value of the vocabulary, determines the retrieval field according to the set of query keywords, analyzes the document matching value of the documents in the retrieval field, and sorts the documents in the retrieval field according to the document matching value. The specific process is as follows:

[0061] When the user inputs a query statement, using natural language processing technology, the query statement is split into several independent vocabulary words. For each split vocabulary word, the number of times the vocabulary word appears in the query distance is counted. Subsequently, the number of times the vocabulary word appears in the query distance is divided by the total number of vocabulary words to obtain the word frequency of the vocabulary.

[0062] The document data obtained by the intelligent data acquisition module is sorted to form a document set. For each vocabulary word, the number of documents including the vocabulary word in the document set is counted and sorted into a word-document set. By dividing the total number of documents in the document set by the total number of documents in the word-document set, the inverse document frequency of the vocabulary is obtained.

[0063] For each vocabulary word, by multiplying the word frequency corresponding to the vocabulary word by the inverse document frequency of the vocabulary, the frequency-inverse value is obtained. The larger the frequency-inverse value of the vocabulary, the more critical the vocabulary is in the query statement and the stronger the indication of determining the theme of the statement. A preset keyword frequency-inverse threshold is set, and the frequency-inverse value corresponding to the vocabulary is compared with the preset keyword frequency-inverse threshold. If the frequency-inverse value corresponding to the vocabulary is greater than or equal to the preset keyword frequency-inverse threshold, the vocabulary is marked as a keyword corresponding to the query statement. By sorting the keywords corresponding to the user's query statement, a set of keywords is formed, denoted as the query keyword set.

[0064] Obtain the various literature fields involved in the document collection. For each literature field, obtain the corresponding specific keyword set, denoted as the literature field keyword set. The specific keyword set corresponding to the literature field can be obtained through the following methods: First, perform a full-text scan of the documents in this field, and use text mining technology to extract the words with relatively high frequencies of occurrence and having field representativeness. At the same time, refer to the terms in the professional field dictionary, authoritative academic papers, and the suggestions of field experts to screen and supplement the extracted words, and finally construct the specific keyword set corresponding to this literature field.

[0065] For each literature field, by matching the query keyword set with the literature field keyword set, output the number of matching keywords, denoted as the field matching value. Preset the field matching threshold, and compare the field matching value with the preset field matching threshold. If the field matching value is greater than the preset field matching threshold, mark this literature field as the retrieval field, and organize all the retrieval fields to form a retrieval list.

[0066] For each retrieval field in the retrieval list, organize all the documents in this field to form a field document library. For each document in the field document library, through text matching technology, count the number of words in the document that contain the words in the query keyword set, denoted as the document matching value. Sort the documents in the field document library from high to low according to the size of the document matching value. When the user inputs a detection statement, the retrieval list will be displayed first. When the user clicks on a certain retrieval field in the retrieval list, the documents in the field document library corresponding to the retrieval field will be sorted from high to low according to the size of the document matching value.

[0067] It should be noted that when the user inputs a query statement, the query statement will be segmented first, and then the frequency inverse value of the words will be calculated. According to the frequency inverse value of the words, analyze the query keyword set of the query statement. By segmenting, calculating the frequency inverse value, and constructing the keyword set, it is convenient to understand the user's query intention. Match the query keyword set with the literature field keyword set to determine the retrieval field. For each retrieval field, sort the documents in the field document library from high to low according to the size of the document matching value between each document in the retrieval field and the query keyword set. This can enable the newly constructed document database to achieve that when the user inputs a query statement, the retrieval field list will be displayed. When the user clicks on a certain retrieval field, the documents in the field document library corresponding to the retrieval field will be sorted from high to low according to the document matching value and presented to the user, which can facilitate quickly locating the most valuable materials for oneself, greatly shortening the time-consuming for the user to screen the results, and effectively improving the retrieval efficiency and user experience.

[0068] The intelligent distributed cache module collects the behavior data of users, analyzes the literature popularity value based on the behavior data, analyzes the cached literature according to the literature popularity value, determines the literature cache points by sending test data packets to each cache node, and caches the cached literature in the cache literature library corresponding to the literature cache points. The specific process is as follows:

[0069] Collect and record the query behavior data of each user. The query behavior data includes: query statements, query time, and corresponding retrieval results; the retrieval results include: the list of hit literatures, the retrieval fields where the literatures are located, the click time of the user on the literature, the reading duration, etc.; at the same time, record in real time the number of collections, the number of comments, the number of likes, and the literature release time of each literature in the literature database, etc.;

[0070] For each user, obtain all the retrieval fields corresponding to all the query statements within the last three days, and generate an unordered list containing all the retrieval fields; then, through programming counting logic, count the frequency of each different retrieval field appearing in the list, denoted as the field duplicate check value. Preset the field duplicate check threshold. For each retrieval field, if the field duplicate check value is greater than the preset field duplicate check threshold, then mark this retrieval field as a cached field;

[0071] For each cached field, obtain the corresponding field literature library. For each literature in the field literature library, by subtracting the literature release time from the current time, obtain the literature release duration FT;

[0072] Try to obtain the user's most recent click time on the literature. If this time can be obtained, subtract the most recent click time from the current time to obtain the literature idle duration. Preset the literature idle duration threshold. If the literature idle duration is greater than the preset literature idle duration threshold, then use the literature idle duration threshold as the literature idle duration; if there has never been a click record, directly use the literature idle duration threshold as the literature idle duration XT; obtain the number of comments PL, the number of collections SC, and the user reading duration YT of the literature;

[0073] After normalizing the number of comments PL, the number of collections SC, the user reading duration YT, the literature release duration FT, and the literature idle duration XT of the literature, use the formula: , to obtain the literature popularity value WXD, where r1, r2, r3, r4, r5 are preset weight coefficients. The larger the literature popularity value, the more attention and popularity the literature receives from users, which means that during the user's query process, the probability of this literature being requested for access is greater;

[0074] Preset the literature popularity threshold. By comparing the literature popularity value corresponding to the literature with the preset literature popularity threshold, if the literature popularity value is greater than the preset literature popularity threshold, then mark this literature as a cached literature;

[0075] Each cache node is correspondingly provided with a cached literature library. By sending test data packets to each cache node, the round-trip time from the user device to each cache node is measured. The cache node with the shortest round-trip time is marked as the literature cache point, and all cached literatures are cached in the cached literature library corresponding to the literature cache point to ensure that users can obtain data at the fastest speed when querying relevant literatures. If the cache space of the cached literature library is insufficient, the priority elimination algorithm is adopted to eliminate the literatures with the lowest literature heat value in turn.

[0076] A cache literature library update mechanism is set. It is set to trigger an update operation every three days. When the update operation is triggered, all literatures in the cached literature library are deleted, and the cached literatures are re-analyzed and stored in the cached literature library.

[0077] It should be noted that by collecting user query behavior data, an unordered list of all retrieval fields in the last three days is generated, and the field duplicate check values of each field are analyzed. After comparing with the preset threshold, if it is greater, it is marked as the cached field. In this way, it is convenient to capture the fields that users have frequently concerned about recently. For the literatures in each cached field, the literature heat value is analyzed by comprehensively considering multi-dimensional factors such as the publication duration of the literature, whether there is user click and idle duration, comment volume, favorite times, and user reading duration, so as to determine the literatures that may be requested to be accessed. Through the cache node positioning method based on network delay detection, the literature cache point is analyzed, and the literatures that may be requested to be accessed are stored in the cached literature library corresponding to the literature cache point. In this way, it is convenient for users to obtain data at a faster speed, significantly shorten the query response time, and greatly improve the user experience. An update operation of the cached literature library is triggered every three days. This mechanism can adapt to the dynamic changes of literature heat and ensure that the cached literature library stores the literatures that are most likely to be accessed by customers at present.

[0078] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirits of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for constructing a traditional cultural database based on artificial intelligence, characterized in that: The following steps are involved: Step 1: Analyze the update time of the data source and determine the statistical duration. During the statistical duration, count the number of new document releases, the number of existing document updates, the number of document downloads, the browsing duration, and the number of collections, and calculate the data source activity and data source attention respectively; comprehensively analyze the statistical duration, data source activity, and data source attention to obtain the sampling frequency matching value, and match the data source collection frequency according to the sampling frequency matching value; Step 2: Segment the user's query sentence, construct a query keyword set by analyzing the frequency inverse value of the vocabulary, determine the search field based on the query keyword set, analyze the document matching value of the documents in the search field, and sort the documents in the search field based on the document matching value; Step 3: Collect user behavior data, and analyze the document heat value based on the behavior data, analyze the cached documents based on the document heat value, determine the document cache point by sending a test data packet to each cache node, and cache the cached documents in the cache document library corresponding to the document cache point.

2. The method for constructing a traditional cultural database based on artificial intelligence according to claim 1, characterized in that: In step 1, the specific process of determining the statistical duration is as follows: For each data source, first obtain the time points of the six most recent updates of the data source, calculate the time interval between two adjacent update time points, record it as the inter-adjacent time length, then use the standard deviation calculation formula to calculate the standard deviation value between all inter-adjacent time lengths, record it as the inter-adjacent standard deviation, compare the inter-adjacent standard deviation with the preset inter-adjacent standard deviation threshold, if the inter-adjacent standard deviation is greater than or equal to the corresponding threshold, recalculate the inter-adjacent standard deviation by adding two update time points in each batch, until the update interval difference is less than the corresponding threshold and then stop; Finally, add up all the adjacent update time intervals corresponding to the stop operation of calculating the adjacent standard deviation, and divide by the number of intervals to get the average update time of the data source; The starting time is the time when the data source last completed data update. The time when the average update duration is moved forward is the end time. This period of time is recorded as the statistical duration TC.

3. The method for constructing a traditional cultural database based on artificial intelligence according to claim 2 is characterized in that: In step 1, the specific process of calculating data source activity, data source attention, and frequency matching value is as follows: Obtain the number of new documents published by the data source and the number of updated existing documents within the statistical period, and assign different weight coefficients to the two respectively. After normalizing the number of new documents published by the data source and the number of updated existing documents, multiply them by the corresponding weight coefficients, and finally add them up to obtain the data source activity SH; For each document in the data source, monitor the number of downloads, browsing time, and number of collections by users during the statistical period; Then, the download volume, browsing time and collection times of all documents in the data source are summed up respectively to obtain the total download volume WX, total browsing time WL and total collection times WS of documents in the data source. After normalizing the total download volume WX, total browsing time WL and total collection times WS of documents in the data source, the data source attention GD is obtained using the formula: GD=WX×a1+WL×a2+WS×a3, where a1, a2 and a3 are preset weight coefficients; By combining the statistical time TC, data source activity SH and data source attention GD, the formula is used: , and obtain the sampling frequency matching value CPZ, where p1, p2, and p3 are preset weight coefficients.

4. The method for constructing a traditional cultural database based on artificial intelligence according to claim 3 is characterized in that: In step 1, the specific process of matching the acquisition frequency of the data source according to the acquisition frequency matching value is as follows: A number of sampling frequency matching value intervals are preset, and each sampling frequency matching value interval corresponds to a data source sampling frequency; For each data source, by matching its corresponding frequency sampling matching value with all the frequency sampling matching value intervals, the collection frequency corresponding to the corresponding frequency sampling matching value interval is output, and the literature data collection frequency of the data source is determined through this collection frequency; the collection update adjustment factor k is preset, where k is greater than 1, and the frequency sampling update duration GT is obtained using the formula: GT=TC×k. The starting point is the moment when the data source first collects data based on the determined collection frequency. When the interval duration from the starting point reaches the frequency sampling update duration, the re-matching analysis process of the data source collection frequency is immediately triggered.

5. The method for constructing a traditional cultural database based on artificial intelligence according to claim 4 is characterized in that: In step 2, the specific process of analyzing the inverse frequency value of the vocabulary and constructing the query keyword set is as follows: When the user enters a query, the query is split into several independent words. For each split word, the number of times the word appears in the query distance is counted, and then the number of times the word appears in the query distance is divided by the total number of words to obtain the word frequency. All the collected document data are sorted into a document set. For each word, the number of documents that include the word in the document set is counted and sorted into a word-file set. The inverse document frequency of the word is obtained by dividing the total number of documents in the document set by the total number of documents in the word-file set. For each word, the inverse frequency value is obtained by multiplying the word frequency corresponding to the word with the inverse document frequency of the word. The inverse frequency value corresponding to the word is compared with the preset keyword inverse frequency threshold. If the inverse frequency value corresponding to the word is greater than or equal to the preset keyword inverse frequency threshold, the word is marked as a keyword of the corresponding query statement. The keywords corresponding to the user's query statement are organized into a keyword set, which is recorded as a query keyword set.

6. The method for constructing a traditional cultural database based on artificial intelligence according to claim 5 is characterized in that: In step 2, the specific process of determining the search field according to the query keyword set, analyzing the document matching values ​​of the documents in the search field, and sorting the documents in the search field according to the document matching values ​​is as follows: Obtain each document field involved in the document collection, and for each document field, obtain a corresponding specific keyword set, which is recorded as a document domain keyword set; For each document field, the query keyword set is matched with the text domain keyword set, the number of matched keywords is output, recorded as the field matching value, and the field matching value is compared with the preset field matching threshold. If the field matching value is greater than the preset field matching threshold, the document field is marked as the search field, and all search fields are sorted into a search list; For each search field in the search list, all the documents in the field are organized into a field document library. For each document in the field document library, the number of words in the query keyword set contained in the document is counted and recorded as the document matching value. All the documents in the field document library are sorted according to the size of the document matching value.

7. The method for constructing a traditional cultural database based on artificial intelligence according to claim 6 is characterized in that: In step 3, the specific process of analyzing the literature heat value is as follows: Collect and record each user's query behavior data, including query statements, query time, and corresponding search results; search results include: hit list of documents, search field of the document, user click time on the document, reading time; at the same time, record the number of collections, comments, likes, and publication time of each document in the document database in real time; For each user, all search fields corresponding to all query statements in the past three days are obtained, and an unordered list containing all search fields is generated; then the frequency of each different search field in the list is counted, recorded as the field duplicate check value, and for each search field, if the field duplicate check value is greater than the preset field duplicate check threshold, the search field is marked as a cache field; For each cached domain, obtain the corresponding domain document library. For each document in the domain document library, obtain the document publishing time FT by subtracting the document publishing time from the current time. Try to obtain the time when the user last clicked on the document. If the time can be obtained, subtract the time of the last click from the current time to get the document idle time. Preset the document idle time threshold. If the document idle time is greater than the preset document idle time threshold, use the document idle time threshold as the document idle time. If there has never been a click record, directly use the document idle time threshold as the document idle time XT. Obtain the number of comments PL, the number of collections SC, and the user reading time YT of the document. The formula is used after normalizing the number of comments PL, the number of collections SC, the user reading time YT, the document publishing time FT and the document idle time XT. , and obtain the document heat value WXD, where r1, r2, r3, r4, and r5 are preset weight coefficients.

8. The method for constructing a traditional cultural database based on artificial intelligence according to claim 7 is characterized in that: In step 3, the cached documents are analyzed according to the document heat value, and the document cache point is determined by sending a test data packet to each cache node, and the cached document is cached in the cache document library corresponding to the document cache point. The specific process is as follows: By comparing the document heat value corresponding to the document with the preset document heat threshold, if the document heat value is greater than the preset document heat threshold, the document is marked as a cached document; Each cache node is set up with a corresponding cache document library. By sending a test data packet to each cache node, the round-trip time from the user device to each cache node is measured, and the cache node with the shortest round-trip time is marked as a document cache point, and all cached documents are cached in the cache document library corresponding to the document cache point; If the cache space of the cache document library is insufficient, a priority elimination algorithm is used to eliminate documents with the lowest document heat value in turn; Set up a cached document library update mechanism, set it to trigger an update operation every three days. When the update operation is triggered, delete all documents in the cached document library, re-analyze the cached documents, and store them in the cached document library.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 8 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Content inquiry method and device based on multiple stages of cache modules

    CN104217019A

  • A power news data acquisition system

    CN109101597A

  • Science and technology intelligence data multi-level cache management method and system

    CN110472004A

  • Online literature induction and storage system based on document data analysis

    CN113239207A

  • Standard literature analysis management system and method applying big data technology

    CN115618014A