Object-based storage data capture path tracking system and method

By receiving user data search text in the stored data search tracking system, extracting topic keywords and reducing the index range, combining natural language processing and scenario analysis technology, the problem of low search and tracking efficiency and accuracy in the existing technology is solved, and more efficient and accurate data text crawling is achieved.

CN119988582AActive Publication Date: 2025-05-13辽宁天空云网络科技有限公司
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510062200.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-05-13
Estimated Expiration
2045-01-15

AI Technical Summary

Technical Problem

The prior art has problems with low search rates and low search rates in the search tracking of stored data, and lacks evaluation of the preliminary index tracking scope and scenario analysis of user search queries, resulting in a large number of data texts indexed, which increases the difficulty and complexity of index tracking.

Method used

By receiving the user's data, search text, preprocessing and extracting topic keywords, combining statistical feature technology and Boolean operators to reduce the index range, obtain a new data text candidate set, and use natural language processing technology to calculate semantic vector representation and cosine similarity, and perform correlation analysis based on the user's situational feature data to obtain the matching degree and recommend the corresponding data text to the user.

Benefits of technology

By conducting index range evaluation and scenario analysis on the data text candidate set, data text with low index matching correlation is reduced, and the crawling efficiency and accuracy are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988582A_ABST
    Figure CN119988582A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of data tracking, and provides an object-based storage data capture path tracking system and method.The method comprises the steps that a data text candidate set is screened out through comparison according to theme keywords, a retrieval range value is obtained through processing according to the attribution field and the text number of data texts, whether the index range needs to be reduced or not is judged, and if yes, the object-based storage data capture path is obtained; the index range is reduced in combination with Boolean operators, so that the index matching efficiency is improved; utilizing a natural language processing technology to obtain a semantic vector of the data text and a semantic vector of the data retrieval text, and processing to obtain similarity; according to the method, the feature data of the data text in the corresponding new data text candidate set is combined, correlation analysis is carried out to obtain the correlation value, the similarity between the data text and the data retrieval text is combined, the matching degree is obtained through processing, the corresponding data text is recommended to the user according to the matching degree, and the capture accuracy of the stored data is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data tracking, and in particular to a system and method for tracking a data capture path based on an object storage. Background Art

[0002] With the rapid development of information technology, the retrieval and tracking of massive storage data has become a major challenge. Traditional storage data retrieval methods, such as keyword retrieval, often have problems with low recall and precision when facing complex and diverse storage data. At the same time, users often need to combine specific contextual information during the retrieval and tracking process to improve the accuracy and relevance of retrieval and tracking. Therefore, it is particularly important to develop a crawling path tracking system and method based on object storage data.

[0003] In the prior art, the search and tracking of stored data is achieved by grouping the selected candidate texts, pre-processing them according to common features, and then adapting the corresponding similarity algorithm. However, the problem is that there is a lack of evaluation of the scope of its preliminary index tracking, such as the amount of text involved in the candidate data text and the field of attribution involved. The index scope is not analyzed and processed, resulting in a large number of indexed data texts, which increases the difficulty and complexity of index tracking. Secondly, semantic matching and crawling are achieved through a semantic similarity algorithm, but the context analysis of user search queries is lacking, and the search and tracking of data texts cannot be achieved more accurately.

[0004] To this end, the present invention provides a system and method for tracing a data crawling path based on an object storage. Summary of the invention

[0005] In order to make up for the deficiencies of the prior art, at least one technical problem raised in the background technology is solved.

[0006] The technical solution adopted by the present invention to solve the technical problem is: a method for tracing a crawling path based on object storage data, comprising:

[0007] Receive the data search text input by the user on its search platform and perform preprocessing, extract the subject keywords in the data search text based on the preprocessed data search text by using statistical feature technology, compare the subject keywords in the data search text with the subject keywords of the data text stored in the search platform, screen out the data texts that meet the subject keywords in the data search text as the data text candidate set, analyze and process the data texts in the data text candidate set according to the domain and the number of the data texts, and determine whether the index range needs to be reduced according to the search range value. If so, generate a reduction signal;

[0008] Based on the reduction signal, the index scope is reduced according to the subject keywords and combined with Boolean operators, and a new data text candidate set is obtained;

[0009] Based on the new data text candidate set, the data text in the new data text candidate set and the data retrieval text are converted into vector representations using natural language processing technology to obtain the semantic vectors of the data text and the semantic vectors of the data retrieval text, and the similarity between the data text and the data retrieval text is calculated using cosine similarity;

[0010] The user's situational feature data is obtained, and combined with the feature data of the data text in the corresponding new data text candidate set, and a correlation analysis is performed to obtain the association value between the user's situational feature data and the data text, and combined with the similarity between the data text and the data retrieval text, the matching degree is obtained through processing and analysis, and the corresponding data text is recommended to the user based on the matching degree.

[0011] As a further technical solution of the present invention, the process of extracting the subject keywords includes:

[0012] Among all the words, obtain the word's part of speech, and filter out the subject keywords from all the words based on the word's part of speech and the word's key value GJ;

[0013] Specifically, if the part of speech of a word falls within the range of parts of speech commonly used to express a subject word, and the key value GJ of the word is greater than or equal to a preset key threshold, the word is marked as a subject keyword.

[0014] As a further technical solution of the present invention, the key value GJ is obtained in the following manner:

[0015] The obtained word frequency ratio PL, TF-IDF ratio TB and word span ratio KB are processed by the formula: The key value GJ is obtained, wherein s1, s2 and s3 are all preset proportional coefficients.

[0016] As a further technical solution of the present invention, the word occurrence frequency ratio PL is obtained by ratio processing the word occurrence frequency with the sum of the occurrence frequencies of all words, and the word TF-IDF ratio TB and word span ratio KB are obtained in the same way as the occurrence frequency ratio PL.

[0017] As a further technical solution of the present invention, the method for obtaining the search range value is:

[0018] Obtaining the attribution domain of each data text in the data text candidate set, removing duplicate attribution domains according to the types of attribution domains, obtaining the number of types of valid attribution domains, performing ratio processing on the number of types of valid attribution domains in the data text candidate set and the total number of types of attribution domains stored in the retrieval platform, and obtaining the effective domain ratio of the data text candidate set;

[0019] The number of data texts in each type of valid attribution domain in the data text candidate set is counted, and the sum is averaged to obtain the mean value of the domain text quantity, and the mean value of the domain text quantity is compared with the total number of data texts stored in the retrieval platform to obtain the text quantity ratio of the data text candidate set;

[0020] The effective domain ratio and the text quantity ratio of the data text candidate set are summed to obtain the retrieval range value.

[0021] As a further technical solution of the present invention, the method of converting the data text in the new data text candidate set and the data retrieval text into a vector representation by using natural language processing technology includes:

[0022] D1, build a vocabulary containing all unique words from the data text;

[0023] D2, use the pre-trained word embedding model BERT to map each word in the vocabulary into a high-dimensional vector space to obtain the vocabulary vector;

[0024] D3, for each data text, sum up the TF-IDF values ​​of all words to obtain the total vector weight of the word, and perform ratio processing on the TF-IDF value of the word and the total vector weight of the word to obtain the vector weight of the word;

[0025] D4, for each data text, all word vectors are multiplied by the vector weights of the corresponding words, and then the sum is averaged to obtain the semantic vector of the data text.

[0026] As a further technical solution of the present invention, the matching degree is obtained by summing up the association value between the user's scenario feature data and the data text and the similarity between the data text and the data retrieval text.

[0027] As a further technical solution of the present invention, the association value is obtained by averaging the scores between the feature data of the data text and the scenario feature data of the user, specifically:

[0028] All the scores obtained are marked as F1, F2, ..., Fz, where z represents the number of scores;

[0029] Average all the scores:

[0030] The associated value GL is obtained by the formula: GL=c1*F1+c2*F2+...cn*Fn, where c1, c2...cn are weight coefficients.

[0031] As a further technical solution of the present invention, the weight coefficient is obtained in the following manner:

[0032] Obtain the historical query records of the user within the historical query period. If the user's scenario feature data appears in a single historical query record, mark the single historical query record as a query record that appears. If the user's scenario feature data does not appear in a single historical query record, mark the single historical query record as a query record that does not appear.

[0033] Count the number of query records that appear, and compare it with the total number of historical query records to obtain the query record ratio, which is marked as CX.

[0034] Obtain the number of scenario feature data appearing in the query record, sum and average the number of scenario feature data in all query records, obtain the mean number of scenario feature data in the query record, and compare it with the total number of scenario feature data of the user to obtain the occurrence feature data ratio, which is marked as TZ;

[0035] The occurrence value CV of the scenario feature data is obtained by the formula: CV = g1*CX+g2*TZ, where g1 and g2 are both preset proportional coefficients;

[0036] The occurrence values ​​CV of all scenario feature data are summed up to obtain the total occurrence value, and the occurrence value CV of the scenario feature data is ratioed to the total occurrence value to obtain the weight coefficient of the corresponding score of the scenario feature data.

[0037] Based on object storage data crawling path tracking system, including:

[0038] Retrieval tracking and evaluation module: receiving the data retrieval text input by the user on its retrieval platform and performing preprocessing, based on the preprocessed data retrieval text, using statistical feature technology to extract the subject keywords in the data retrieval text, comparing the subject keywords in the data retrieval text with the subject keywords of the data text stored in the retrieval platform, screening out the data texts that meet the subject keywords in the data retrieval text as the data text candidate set, analyzing and processing the retrieval range value according to the belonging field and the number of texts of each data text in the data text candidate set, judging whether it is necessary to reduce the index range according to the retrieval range value, and if so, generating a reduction signal;

[0039] Retrieval tracking control module: Based on the reduction signal, the index scope is reduced according to the subject keywords and combined with Boolean operators, and a new data text candidate set is obtained;

[0040] Similarity analysis module: Based on the new data text candidate set, the data text in the new data text candidate set and the data retrieval text are converted into vector representation using natural language processing technology to obtain the semantic vector of the data text and the semantic vector of the data retrieval text, and the similarity between the data text and the data retrieval text is calculated using cosine similarity;

[0041] Association matching processing module: obtain the user's situational feature data, and combine it with the feature data of the data text in the corresponding new data text candidate set, and perform correlation analysis to obtain the association value between the user's situational feature data and the data text, and combine it with the similarity between the data text and the data retrieval text, process and analyze to obtain the matching degree, and recommend the corresponding data text to the user based on the matching degree.

[0042] The beneficial effects of the present invention are as follows:

[0043] 1. Receive the data retrieval text input by the user on its retrieval platform and perform preprocessing. Based on the preprocessed data retrieval text, use statistical feature technology to extract subject keywords in the data retrieval text, compare the subject keywords in the data retrieval text with the subject keywords of the data text stored in the retrieval platform, screen out the data texts that meet the subject keywords in the data retrieval text as the data text candidate set, analyze and process the retrieval range value according to the belonging field and the number of texts of each data text in the data text candidate set, judge whether it is necessary to reduce the index range according to the retrieval range value, and if so, generate a reduction signal; based on the reduction signal, reduce the index range according to the subject keywords and combined with Boolean operators, and obtain a new data text candidate set. The present invention evaluates the index range of its data text candidate set and reduces the index range using Boolean operators when the range is large, which is conducive to reducing data texts with low index matching relevance and improving crawling efficiency.

[0044] 2. Based on the new data text candidate set, natural language processing technology is used to convert the data text in the new data text candidate set and the data retrieval text into vector representations, and the semantic vector of the data text and the semantic vector of the data retrieval text are obtained, and the similarity between the data text and the data retrieval text is calculated using cosine similarity; the user's situational feature data is obtained, and combined with the feature data of the data text in the corresponding new data text candidate set, and correlation analysis is performed to obtain the association value between the user's situational feature data and the data text, and combined with the similarity between the data text and the data retrieval text, the matching degree is obtained through processing and analysis, and the corresponding data text is recommended to the user according to the matching degree. The present invention not only calculates the similarity, but also integrates the situational analysis, and performs data text crawling and tracking through the similarity obtained by semantic analysis and the association value obtained by situational analysis, thereby improving the crawling efficiency of stored data while improving the crawling accuracy of stored data. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] The present invention will be further described below in conjunction with the accompanying drawings.

[0046] Figure 1 is a flowchart of the steps of the object storage data crawling path tracking method according to an embodiment of the present invention;

[0047] Figure 2 It is a flowchart of the object storage data crawling path tracking system according to an embodiment of the present invention. DETAILED DESCRIPTION

[0048] In order to make the technical means, creative features, objectives and effects achieved by the present invention easy to understand, the present invention is further explained below in conjunction with specific implementation methods.

[0049] Example 1

[0050] like Figure 1 As shown, the object storage data crawling path tracking method described in the embodiment of the present invention includes:

[0051] Step 1: receiving the data search text input by the user on its search platform and preprocessing it, extracting the subject keywords in the data search text by using statistical feature technology based on the preprocessed data search text, comparing the subject keywords in the data search text with the subject keywords of the data text stored in the search platform, screening out the data texts that meet the subject keywords in the data search text as the data text candidate set, analyzing and processing the search range value according to the domain and number of the data texts in the data text candidate set, judging whether it is necessary to reduce the index range according to the search range value, and if so, generating a reduction signal;

[0052] Among them, the process of preprocessing the data retrieval text includes:

[0053] S1, Word Segmentation: For Chinese text, it is first necessary to perform word segmentation to split the sentence into independent words or phrases;

[0054] S2, Stop Word Removal: Stop words refer to words that frequently appear in the text but contribute little to the text theme, such as "de" (的), "le" (了), "zai" (在), etc.;

[0055] S3, Part-of-Speech Tagging: Perform part-of-speech tagging on the words in the text, including but not limited to: nouns, verbs;

[0056] The extraction of topic keywords includes:

[0057] A1, Word Frequency Statistics: Count the occurrence frequency of each word in the preprocessed data retrieval text;

[0058] A2, Calculate the TF-IDF value of the word using the TF-IDF algorithm, where TF (Term Frequency) represents the word frequency and IDF (Inverse Document Frequency) represents the inverse document frequency;

[0059] A3, Measure the word span of the word, where the word span represents the distance between the first occurrence and the last occurrence of the word in the text;

[0060] Based on the occurrence frequency, TF-IDF value, and word span of the word, calculate and obtain the occurrence frequency ratio PL, TF-IDF ratio TB, and word span ratio KB of the word respectively;

[0061] Perform data processing on the obtained occurrence frequency ratio PL, TF-IDF ratio TB, and word span ratio KB of the word through the formula: Obtain the key value GJ, where s1, s2, and s3 are all preset proportionality coefficients, s1 takes the value of 1.102, s2 takes the value of 1.134, and s3 takes the value of 1.24;

[0062] It should be noted that there are m groups of historical data, and each group of historical data includes the occurrence frequency ratio PL, TF-IDF ratio TB, and word span ratio KB of the word. Since the occurrence frequency ratio PL, TF-IDF ratio TB, and word span ratio KB of the word will all affect the key degree of the word, according to the key degree of the word, the staff uses an offline model to fit based on the occurrence frequency ratio PL, TF-IDF ratio TB, and word span ratio KB of the word to obtain the key value GJ matched by the occurrence frequency ratio PL, TF-IDF ratio TB, and word span ratio KB of the word;

[0063] Then, based on the frequency ratios PL, TF-IDF ratios TB, and word span ratios KB of multiple word groups, a linear model is used to fit the historical data of the frequency ratios PL, TF-IDF ratios TB, word span ratios KB, and key values ​​GJ of multiple word groups, and the prepared historical data is brought into the selected fitting model for fitting to obtain fitting coefficients s1, s2, and s3;

[0064] The word frequency ratio PL is obtained by comparing the word frequency with the sum of all word frequencies. The word TF-IDF ratio TB and word span ratio KB are obtained in the same way as the frequency ratio PL. For example, in the data retrieval text, the word appears N i , where N represents frequency, i represents word number, i=1, 2, 3, 4...n; the frequency ratio of the words is:

[0065] Among all the words, obtain the word's part of speech, and filter out the subject keywords from all the words based on the word's part of speech and the word's key value GJ;

[0066] Specifically, if the part of speech of a word falls within the range of parts of speech of commonly used expression keywords, and the key value GJ of the word is greater than or equal to the preset key threshold, the word is marked as a topic keyword;

[0067] If the word's part of speech is within the range of common expression keywords, and the word's key value GJ is less than the preset key threshold, the word is marked as a non-topic keyword;

[0068] If the part of speech of a word does not fall within the range of the parts of speech of commonly expressed subject words, and the key value GJ of the word is greater than or equal to the preset key threshold, the word is marked as a non-topic keyword;

[0069] If the part of speech of a word does not fall within the range of the parts of speech of commonly expressed subject words, and the key value GJ of the word is less than the preset key threshold, the word is marked as a non-topic keyword;

[0070] It should be noted that, generally speaking, common parts of speech used to express topics include nouns, verbs, proper nouns, etc. By obtaining the parts of speech of a word and matching it with the common parts of speech used to express the topic, topic keywords can be screened out;

[0071] It can be understood that the key value GJ is obtained by processing the frequency ratio PL of the word, the TF-IDF ratio TB and the word span ratio KB. This type of data belongs to statistical data. This type of data can be easily collected through existing statistical monitoring technology and combined with existing mathematical calculation rules to perform numerical operations on the collected data. According to the positive and negative influence of the data, the floating influence of multiple data can be synchronously collected to complete the data processing requirements of the current technical solution steps. In the process of data processing requirements, threshold comparison is involved. The thresholds of the corresponding data are all the experience values ​​of those in this field. Threshold comparison is a conventional data judgment method in the prior art, and the threshold set for the corresponding threshold comparison is the current data. In the real-time collection scenario, the empirical values ​​proposed by the personnel in this field for the current data monitoring, for example, the staff calculates the key value GJ based on multiple sets of data, and identifies the criticality of the words according to the key value GJ, and obtains a corresponding relationship between the key value GJ and the criticality of the words. The staff derives and divides the thresholds according to the corresponding relationship between multiple sets of key values ​​and the criticality of the words, thereby obtaining the thresholds. For example, if the key value exceeds the threshold B in the historical monitoring period, the criticality of the words is high, and the threshold in the next historical monitoring period can be set to B. On the contrary, when the threshold in the next period is lower than B, the criticality of the words is also high, and the current threshold continues to be lowered based on B;

[0072] The way to obtain the search range value is:

[0073] Obtaining the attribution domain of each data text in the data text candidate set, removing duplicate attribution domains according to the types of attribution domains, obtaining the number of types of valid attribution domains, performing ratio processing on the number of types of valid attribution domains in the data text candidate set and the total number of types of attribution domains stored in the retrieval platform, and obtaining the effective domain ratio of the data text candidate set;

[0074] The number of data texts in each type of valid attribution domain in the data text candidate set is counted, and the sum is averaged to obtain the mean value of the domain text quantity, and the mean value of the domain text quantity is compared with the total number of data texts stored in the retrieval platform to obtain the text quantity ratio of the data text candidate set;

[0075] The effective domain ratio and the text quantity ratio of the data text candidate set are summed to obtain the search range value;

[0076] In some preferred implementations, the search range value is compared with a preset search range threshold, and the specific comparison process is as follows:

[0077] If the search range value is greater than or equal to the preset search range threshold, it means that the number of data texts indexed by the current subject keyword is large and the fields are diverse, and the search range is large, which is not conducive to subsequent specific index matching. The index range needs to be reduced, and a reduction signal is generated;

[0078] If the search range value is less than the preset search range threshold, it means that the number of data texts indexed by the current subject keyword is small and the field types are not diverse, so the search range is small, which is conducive to subsequent specific index matching and there is no need to reduce the index range;

[0079] It can be understood that the effective domain ratio and text quantity ratio of the data text candidate set are both relevant text data in the actual scene, for example, including the number of types and the number of data texts that obtain the effective belonging field. Such data can be easily collected through existing sensor technology or network monitoring technology and combined with existing mathematical calculation rules to perform numerical operations on the collected data, and the floating influence of multiple data can be synchronously collected according to the positive and negative influence of the data to complete the data processing requirements of the current technical solution steps. In the process of data processing requirements, threshold comparison is involved, and the thresholds of the corresponding data are all empirical values ​​of personnel in this field. Threshold comparison is a conventional data judgment method in the prior art, and the threshold set for the corresponding threshold comparison is the empirical value proposed by personnel in this field for current data monitoring in the real-time collection scenario of the current data, such as the search range value. If the search range value exceeds the threshold A in the historical monitoring period, it means that the search range is large, then the threshold can be set to A in the next historical monitoring period. On the contrary, if the threshold is lower than A in the next period and there is also a large search range, the current threshold will continue to be lowered based on A.

[0080] Step 2: Based on the reduction signal, reduce the index range according to the subject keywords and combine with Boolean operators, and obtain a new data text candidate set;

[0081] Specifically, related subject keywords are combined using Boolean operators to obtain combined subject keywords. For example, the related words "apple" and "mobile phone" are combined into "apple" AND "mobile phone", where AND is a Boolean operator.

[0082] Re-comparison is performed using the combined subject keywords with the subject keywords of the data text stored in the retrieval platform to obtain a new data text candidate set;

[0083] It should be noted that the search range value corresponding to the new data text candidate set is smaller than the preset search range threshold;

[0084] Step 3: Based on the new data text candidate set, the data text in the new data text candidate set and the data retrieval text are converted into vector representations using natural language processing technology to obtain the semantic vector of the data text and the semantic vector of the data retrieval text, and the similarity between the data text and the data retrieval text is calculated using cosine similarity;

[0085] Specifically, natural language processing technology is used to convert the data text in the new data text candidate set and the data retrieval text into vector representation, including:

[0086] D1, build a vocabulary containing all unique words from the data text;

[0087] D2, use the pre-trained word embedding model BERT to map each word in the vocabulary into a high-dimensional vector space to obtain the vocabulary vector;

[0088] D3, for each data text, sum up the TF-IDF values ​​(term frequency - inverse document frequency) of all words to obtain the total vector weight of the word, and perform ratio processing on the TF-IDF value of the word and the total vector weight of the word to obtain the vector weight of the word (v1, v2, v3...vn, where n represents the number of words);

[0089] D4, for each data text, all word vectors are multiplied by the vector weights of the corresponding words, and then the sum is averaged to obtain the semantic vector of the data text;

[0090] Specifically, the semantic vector of the data retrieval text is obtained in the same way as the semantic vector of the data text;

[0091] The cosine similarity is used to calculate the similarity between the data text and the data retrieval text, specifically:

[0092] By formula: The similarity cosθ between the data text and the data retrieval text is obtained, where A represents the semantic vector of the data retrieval text, B represents the semantic vector of the data text, and ||A|| and ||B|| represent the modulus length of vectors A and B respectively (i.e., the length of the vector, which can be obtained by finding the square root of the sum of the squares of the elements of the vector);

[0093] Step 4: Obtain the user's scenario feature data, and combine it with the feature data of the data text in the corresponding new data text candidate set, and perform correlation analysis to obtain the correlation value between the user's scenario feature data and the data text, and combine it with the similarity between the data text and the data search text, process and analyze to obtain the matching degree, and recommend the corresponding data text to the user according to the matching degree;

[0094] Specifically, the user's context feature data includes but is not limited to the query time and location, and the corresponding feature data of the data text in the new data text candidate set includes but is not limited to the data text creation time and the included location;

[0095] The specific process of correlation analysis is as follows:

[0096] For the feature data of the data text and the scenario feature data of the user, the scores between the feature data are calculated respectively, and all the scores are averaged to obtain their association values;

[0097] Exemplary:

[0098] Use the time window to compare the data text creation time and the user query time, and combine the linear function to calculate the score, specifically:

[0099] First, calculate the time difference (in days) between the data text creation time and the user query time;

[0100] Set a time window (e.g. 7 days);

[0101] Calculate the score:

[0102] Use linear function: score = 1-(time difference / time window length);

[0103] Note that if the time difference is greater than the time window length, the score can be 0 or a small negative number, indicating complete irrelevance or negative correlation;

[0104] Exemplary:

[0105] Get the distance between the data text location and the user's location;

[0106] Calculate the score:

[0107] Use a linear function: score = 1-(distance / maximum distance threshold), where the maximum distance threshold is an empirical value set according to the application scenario;

[0108] All the scores obtained are marked as F1, F2, ..., Fz, where z represents the number of scores;

[0109] Average all the scores:

[0110] The correlation value GL is obtained by the formula: GL = c1*F1+c2*F2+......cn*Fn, where c1, c2...cn are weight coefficients;

[0111] The weight coefficient is obtained as follows:

[0112] Obtain the historical query records of the user within the historical query period. If the user's scenario feature data appears in a single historical query record, mark the single historical query record as a query record that appears. If the user's scenario feature data does not appear in a single historical query record, mark the single historical query record as a query record that does not appear.

[0113] Count the number of query records that appear, and compare it with the total number of historical query records to obtain the query record ratio, which is marked as CX.

[0114] Obtain the number of scenario feature data appearing in the query record, sum and average the number of scenario feature data in all query records, obtain the mean number of scenario feature data in the query record, and compare it with the total number of scenario feature data of the user to obtain the occurrence feature data ratio, which is marked as TZ;

[0115] The appearance value CV of the scenario feature data is obtained by the formula: CV = g1*CX+g2*TZ, where g1 and g2 are both preset proportional coefficients, where g1 is 1.064 and g2 is 1.031;

[0116] It should be noted that there are m groups of historical data, each group of historical data includes query record ratio CX and occurrence feature data ratio TZ. Since query record ratio CX and occurrence feature data ratio TZ will affect the appearance of scenario feature data, according to the appearance degree of scenario feature data, the staff uses the offline model to fit according to the query record ratio CX and occurrence feature data ratio TZ to obtain the appearance value CV of the scenario feature data matched by the query record ratio CX and occurrence feature data ratio TZ;

[0117] Based on multiple sets of query record ratios CX and occurrence feature data ratios TZ, a linear model is used to fit the historical data of multiple sets of query record ratios CX, occurrence feature data ratios TZ, and occurrence value CV of scenario feature data, and the prepared historical data is brought into the selected fitting model for fitting to obtain fitting coefficients g1 and g2;

[0118] Sum up the occurrence values ​​CV of all scenario feature data to obtain the total occurrence value, perform ratio processing on the occurrence value CV of the scenario feature data and the total occurrence value to obtain the weight coefficient of the corresponding score of the scenario feature data;

[0119] The obtained association value between the user's scenario feature data and the data text and the similarity between the data text and the data search text are summed to obtain the matching degree;

[0120] Recommend corresponding data texts to users in descending order of matching degree;

[0121] The technical solution of the embodiment of the present invention is: receiving the data retrieval text input by the user on its retrieval platform and performing preprocessing, based on the preprocessed data retrieval text, using statistical feature technology to extract the subject keywords in the data retrieval text, comparing the subject keywords in the data retrieval text with the subject keywords of the data text stored in the retrieval platform, screening out the data texts that meet the subject keywords in the data retrieval text as the data text candidate set, analyzing and processing the retrieval range value according to the belonging field and the number of texts of each data text in the data text candidate set, judging whether it is necessary to reduce the index range according to the retrieval range value, and if so, generating a reduction signal; based on the reduction signal, reducing the index range according to the subject keywords and in combination with Boolean operators, and obtaining a new data text candidate set. The present invention evaluates the index range of its data text candidate set and reduces the index range using Boolean operators when its range is large, which is conducive to reducing the number of index matching with low relevance. According to the text, it is beneficial to improve the crawling efficiency; based on the new data text candidate set, the data text in the new data text candidate set and the data retrieval text are converted into vector representations using natural language processing technology to obtain the semantic vector of the data text and the semantic vector of the data retrieval text, and the cosine similarity is used to calculate the similarity between the data text and the data retrieval text; the user's situational feature data is obtained, and combined with the feature data of the data text in the corresponding new data text candidate set, and a correlation analysis is performed to obtain the correlation value between the user's situational feature data and the data text, and combined with the similarity between the data text and the data retrieval text, the matching degree is obtained through processing and analysis, and the corresponding data text is recommended to the user according to the matching degree. The present invention not only calculates the similarity, but also integrates the situational analysis, and crawls and tracks the data text through the similarity obtained by semantic analysis and the correlation value obtained by situational analysis, thereby improving the crawling efficiency of the stored data and the crawling accuracy of the stored data.

[0122] Example 2

[0123] like Figure 2 As shown, the object storage data crawling path tracking system according to the embodiment of the present invention includes:

[0124] Retrieval tracking and evaluation module: receiving the data retrieval text input by the user on its retrieval platform and performing preprocessing, based on the preprocessed data retrieval text, using statistical feature technology to extract the subject keywords in the data retrieval text, comparing the subject keywords in the data retrieval text with the subject keywords of the data text stored in the retrieval platform, screening out the data texts that meet the subject keywords in the data retrieval text as the data text candidate set, analyzing and processing the retrieval range value according to the belonging field and the number of texts of each data text in the data text candidate set, judging whether it is necessary to reduce the index range according to the retrieval range value, and if so, generating a reduction signal;

[0125] Retrieval tracking control module: Based on the reduction signal, the index scope is reduced according to the subject keywords and combined with Boolean operators, and a new data text candidate set is obtained;

[0126] Similarity analysis module: Based on the new data text candidate set, the data text in the new data text candidate set and the data retrieval text are converted into vector representation using natural language processing technology to obtain the semantic vector of the data text and the semantic vector of the data retrieval text, and the similarity between the data text and the data retrieval text is calculated using cosine similarity;

[0127] Association matching processing module: obtain the user's situational feature data, and combine it with the feature data of the data text in the corresponding new data text candidate set, and perform correlation analysis to obtain the association value between the user's situational feature data and the data text, and combine it with the similarity between the data text and the data retrieval text, process and analyze to obtain the matching degree, and recommend the corresponding data text to the user based on the matching degree.

[0128] The above shows and describes the basic principles, main features and advantages of the present invention. It should be understood by those skilled in the art that the present invention is not limited to the above embodiments. The above embodiments and descriptions are only for explaining the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention may have various changes and improvements, which fall within the scope of the present invention. The scope of protection of the present invention is defined by the attached claims and their equivalents.

Claims

1. A method for tracing a crawling path based on object storage data, characterized in that: include: Receive the data search text input by the user on its search platform and perform preprocessing, extract the subject keywords in the data search text based on the preprocessed data search text by using statistical feature technology, compare the subject keywords in the data search text with the subject keywords of the data text stored in the search platform, screen out the data texts that meet the subject keywords in the data search text as the data text candidate set, analyze and process the data texts in the data text candidate set according to the domain and the number of the data texts, and determine whether the index range needs to be reduced according to the search range value. If so, generate a reduction signal; Based on the reduction signal, the index scope is reduced according to the subject keywords and combined with Boolean operators, and a new data text candidate set is obtained; Based on the new data text candidate set, the data text in the new data text candidate set and the data retrieval text are converted into vector representations using natural language processing technology to obtain the semantic vectors of the data text and the semantic vectors of the data retrieval text, and the similarity between the data text and the data retrieval text is calculated using cosine similarity; The user's situational feature data is obtained, and combined with the feature data of the data text in the corresponding new data text candidate set, and a correlation analysis is performed to obtain the association value between the user's situational feature data and the data text, and combined with the similarity between the data text and the data retrieval text, the matching degree is obtained through processing and analysis, and the corresponding data text is recommended to the user based on the matching degree.

2. The object storage data crawling path tracking method according to claim 1 is characterized in that: The extraction process of the subject keywords includes: Among all the words, obtain the word's part of speech, and filter out the subject keywords from all the words based on the word's part of speech and the word's key value GJ; Specifically, if the part of speech of a word falls within the range of parts of speech commonly used to express a subject word, and the key value GJ of the word is greater than or equal to a preset key threshold, the word is marked as a subject keyword.

3. The object storage data crawling path tracking method according to claim 2 is characterized in that: The key value GJ is obtained as follows: The obtained word frequency ratio PL, TF-IDF ratio TB and word span ratio KB are processed by the formula: The key value GJ is obtained, wherein s1, s2 and s3 are all preset proportional coefficients.

4. The object storage data crawling path tracking method according to claim 3 is characterized in that: The word occurrence frequency ratio PL is obtained by performing ratio processing on the word occurrence frequency and the sum of the occurrence frequencies of all words. The word TF-IDF ratio TB and word span ratio KB are obtained in the same way as the occurrence frequency ratio PL.

5. The object storage data crawling path tracking method according to claim 1 is characterized in that: The method for obtaining the search range value is as follows: Obtaining the attribution domain of each data text in the data text candidate set, removing duplicate attribution domains according to the types of attribution domains, obtaining the number of types of valid attribution domains, performing ratio processing on the number of types of valid attribution domains in the data text candidate set and the total number of types of attribution domains stored in the retrieval platform, and obtaining the effective domain ratio of the data text candidate set; The number of data texts in each type of valid attribution domain in the data text candidate set is counted, and the sum is averaged to obtain the mean value of the domain text quantity, and the mean value of the domain text quantity is compared with the total number of data texts stored in the retrieval platform to obtain the text quantity ratio of the data text candidate set; The effective domain ratio and the text quantity ratio of the data text candidate set are summed to obtain the retrieval range value.

6. The object storage data crawling path tracking method according to claim 1 is characterized in that: The method of converting the data text in the new data text candidate set and the data retrieval text into a vector representation by using natural language processing technology includes: D1, build a vocabulary containing all unique words from the data text; D2, use the pre-trained word embedding model BERT to map each word in the vocabulary into a high-dimensional vector space to obtain the vocabulary vector; D3, for each data text, sum up the TF-IDF values ​​of all words to obtain the total vector weight of the word, and perform ratio processing on the TF-IDF value of the word and the total vector weight of the word to obtain the vector weight of the word; D4, for each data text, all word vectors are multiplied by the vector weights of the corresponding words, and then the sum is averaged to obtain the semantic vector of the data text.

7. The object storage data crawling path tracking method according to claim 1 is characterized in that: The matching degree is obtained by summing up the association value between the user's scenario feature data and the data text and the similarity between the data text and the data search text.

8. The object storage data crawling path tracking method according to claim 7 is characterized in that: The association value is obtained by averaging the scores between the feature data of the data text and the scenario feature data of the user, specifically: All the scores obtained are marked as F1, F2, ..., Fz, where z represents the number of scores; Average all the scores: The associated value GL is obtained by the formula: GL=c1*F1+c2*F2+...cn*Fn, where c1, c2...cn are weight coefficients.

9. The object storage data crawling path tracking method according to claim 8 is characterized in that: The weight coefficient is obtained as follows: Obtain the historical query records of the user within the historical query period. If the user's scenario feature data appears in a single historical query record, mark the single historical query record as a query record that appears. If the user's scenario feature data does not appear in a single historical query record, mark the single historical query record as a query record that does not appear. Count the number of query records that appear, and compare it with the total number of historical query records to obtain the query record ratio, which is marked as CX. Obtain the number of scenario feature data appearing in the query record, sum and average the number of scenario feature data in all query records, obtain the mean number of scenario feature data in the query record, and compare it with the total number of scenario feature data of the user to obtain the occurrence feature data ratio, which is marked as TZ; The occurrence value CV of the scenario feature data is obtained by the formula: CV = g1*CX+g2*TZ, where g1 and g2 are both preset proportional coefficients; The occurrence values ​​CV of all scenario feature data are summed up to obtain the total occurrence value, and the occurrence value CV of the scenario feature data is ratioed to the total occurrence value to obtain the weight coefficient of the corresponding score of the scenario feature data.

10. A system for tracing a crawling path based on object storage data, the system being used to implement the tracing method according to any one of claims 1 to 9, characterized in that: include: Retrieval tracking and evaluation module: receiving the data retrieval text input by the user on its retrieval platform and performing preprocessing, based on the preprocessed data retrieval text, using statistical feature technology to extract the subject keywords in the data retrieval text, comparing the subject keywords in the data retrieval text with the subject keywords of the data text stored in the retrieval platform, screening out the data texts that meet the subject keywords in the data retrieval text as the data text candidate set, analyzing and processing the retrieval range value according to the belonging field and the number of texts of each data text in the data text candidate set, judging whether it is necessary to reduce the index range according to the retrieval range value, and if so, generating a reduction signal; Retrieval tracking control module: based on the reduction signal, it reduces the index scope according to the subject keywords and combines Boolean operators, and obtains a new data text candidate set; Similarity analysis module: Based on the new data text candidate set, the data text in the new data text candidate set and the data retrieval text are converted into vector representation using natural language processing technology to obtain the semantic vector of the data text and the semantic vector of the data retrieval text, and the similarity between the data text and the data retrieval text is calculated using cosine similarity; Association matching processing module: obtain the user's situational feature data, and combine it with the feature data of the data text in the corresponding new data text candidate set, and perform correlation analysis to obtain the association value between the user's situational feature data and the data text, and combine it with the similarity between the data text and the data retrieval text, process and analyze to obtain the matching degree, and recommend the corresponding data text to the user based on the matching degree.

Citation Information

Patent Citations

  • Search processing device, electronic equipment and search processing method

    CN102402525A

  • Intelligent remote-sensing image control point data search method

    CN103235810A

  • Text matching method and device

    CN110287396A

  • Privacy-protected lightweight double-layer filtering close contact person screening method

    CN113747424A

  • System and method for search space reduction for identifying an item

    US20240020858A1