Intelligent retrieval efficiency optimization method based on computer data
By calculating the similarity between text data in data retrieval and using the DBSCAN algorithm for clustering, and adjusting the radius parameters based on the continuous degree of sample similarity and differences of the core objects, the problem of inefficient traditional data retrieval is solved, and fine clustering and efficient retrieval of text data is achieved.
Patent Information
- Application Number
- CN202510549808.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-29
AI Technical Summary
When traditional data retrieval methods process large-scale data and complex information, the search range is too large and there is too much unrelated data, resulting in inaccurate retrieval efficiency, especially when using the DBSCAN algorithm, due to the Euclidean distance, the clustering results are not refined enough, resulting in inaccurate retrieval.
By obtaining text data and keywords in computer data, the degree of similarity between text data is calculated, and the text data is clustered using the DBSCAN algorithm. According to the core objects during the clustering process, the difference in the sample radius interval and the number of sample points contained are selected, the continuity and difference of sample similarity are calculated, and the radius parameters of the core objects are adjusted to improve the fineness of the clustering method.
By performing similarity measurement and application of DBSCAN algorithm on text data, determining the core object and adjusting the radius parameters, the fine division of text data of different similarity degrees is achieved, the accuracy of clustering is improved, and thus the efficiency of data retrieval is improved.
Smart Images

Figure CN120067283A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly relates to a method for optimizing the intelligent retrieval efficiency of computer data. Background Art
[0002] In traditional data retrieval methods, when dealing with large-scale data and complex information, the search scope is too large and there are too many irrelevant data, which is not conducive to the rapid progress of accurate retrieval. Therefore, in order to improve the retrieval efficiency and optimize the user experience, the DBSCAN algorithm is usually used to improve the retrieval efficiency at present. This algorithm classifies data before retrieval to achieve more targeted search for relevant data during retrieval, thereby improving the retrieval efficiency. However, the DBSCAN algorithm uses the Euclidean distance as the default distance metric method, which not only cannot comprehensively consider the characteristic relationships between data, but also causes the problem that the clustering results are not finely divided due to the large differences between data. For example, when retrieving text data, on the one hand, the text data cannot use the Euclidean distance for distance measurement. At the same time, the content information contained in different text data is rich and diverse, which leads to insufficiently fine division when using the DBSCAN algorithm for clustering, and thus inaccurate retrieval occurs during retrieval. Summary of the Invention
[0003] The present invention provides a method for optimizing the intelligent retrieval efficiency of computer data to solve the above problems.
[0004] A method for optimizing the intelligent retrieval efficiency of computer data according to the present invention adopts the following technical solutions: On the one hand, an embodiment of the present invention provides a method for optimizing the intelligent retrieval efficiency of computer data, and the method includes the following steps: Obtain the text data in the computer data and the keywords of the text data; Obtain the similarity degree between text data according to the number of occurrences of keywords in the text data and the similarity value of the keyword text vectors; Use the text data as sample points and use the DBSCAN algorithm to cluster the sample points to obtain the core objects during the clustering process, and the metric distance used during the clustering process is the similarity degree between text data; Obtain the sample similarity continuity degree of the core objects according to the sample radius interval screened out by the core objects and the quantity difference of the two types of sample points included in the sample radius interval; Divide the core objects into a first sample and a second sample according to the screened sample radius interval. The distance between the sample points in the first sample and the core object is greater than the sample radius interval, and the distance between the sample points in the second sample and the core object is less than or equal to the sample radius interval; Obtain the difference between the first sample and the second sample of the core object based on the selected sample radius interval, the shortest distance between the first sample and the second sample, and the distances from each sample point in the first sample and the second sample to the core object; Obtain the necessity of adjusting the radius parameter of the core object based on the degree of continuity of sample similarity of the core object and the difference between the first sample and the second sample; Adjust the radius parameter of the core object according to the necessity of adjusting the radius parameter of the core object; Classify the text data according to the adjusted radius parameter of the core object, and complete the text data retrieval according to the classification result of the text data.
[0005] Furthermore, the similarity degree between the text data includes the following specific steps: Calculate the product of the average weight of the keywords in the th combination of the keywords of any two text contents and the cosine similarity value of the keyword text vectors in the th combination, denoted as the first product of the th combination. Take the mean of the first products of all combinations as the similarity degree between any two text data.
[0006] Furthermore, the metric distance used in the clustering process is , where exp( ) represents the exponential function with the natural constant as the base, represents the similarity degree between any two text data.
[0007] Furthermore, the method for obtaining the degree of continuity of sample similarity of the core object based on the selected sample radius interval of the core object and the difference in the number of two types of sample points included in the sample radius interval includes the following specific steps: The distance from each sample point in the neighborhood of the core object to the core object is the sample point radius; Sort all the sample point radii in the neighborhood of the core object in ascending order, and calculate the difference between adjacent sample point radii after sorting; Select B largest adjacent sample point radius differences from the adjacent sample point radius differences, denoted as the sample radius interval; For the B sample radius intervals selected for the i-th core object, calculate the difference in the number of sample points included in the two sample points corresponding to the b-th sample radius interval selected for the i-th core object of the absolute value ratio as the first ratio of the b-th sample radius interval. Take the inverse normalization value of the mean of the first ratios of all sample radius intervals as the adjustment coefficient of the largest sample point radius in the neighborhood of the i-th core object; Obtain the sample similarity continuity degree of the core object according to the adjustment coefficient of the maximum sample point radius within the neighborhood of the core object and the maximum sample point radius within the neighborhood of the core object.
[0008] Further, the obtaining of the sample similarity continuity degree of the core object according to the adjustment coefficient of the maximum sample point radius within the neighborhood of the core object and the maximum sample point radius within the neighborhood of the core object includes the following specific steps: Take the normalized value of the product of the maximum sample point radius within the neighborhood of the i-th core object and the adjustment coefficient of the maximum sample point radius within the neighborhood of the i-th core object as the sample similarity continuity degree of the i-th core object.
[0009] Further, the obtaining of the difference between the first sample and the second sample of the core object according to the selected sample radius interval, the shortest distance between the first sample and the second sample, and the distances from each sample point in the first sample and the second sample to the core object includes the following specific steps: At any selected sample radius interval of the i-th core object, calculate the mean value of all the shortest distances in the shortest distance sequence between the first sample and the second sample, and take the average distance from each sample point in the first sample to the i-th core object and the standard deviation of the distances from each sample point in the first sample to the core object The product is denoted as the second product. Take the average distance from each sample point in the second sample to the core object and the standard deviation of the distances from each sample point in the second sample to the core object The product is denoted as the second product. Take the sum value of the first product and the second product as the first sum value, and take the ratio of the mean value of all the shortest distances to the first sum value as the second ratio; Take the normalized value of the maximum value among the second ratios at all the selected sample radius intervals of the i-th core object as the difference between the first sample and the second sample of the i-th core object.
[0010] Further, the specific obtaining method of the shortest distance sequence between the first sample and the second sample is as follows: Traverse the shortest distances between any two sample points in the first sample and the second sample. The any two sample points are not selected repeatedly. Arrange the shortest distances between any two sample points from low to high to obtain the shortest distance sequence between the first sample and the second sample.
[0011] Further, the obtaining of the necessity for adjusting the radius parameter of the core object according to the sample similarity continuity degree of the core object and the difference between the first sample and the second sample includes the following specific steps: Calculate the difference between 1 and the sample similarity continuity degree of the i-th core object, and denote the product of the difference and the difference between the first sample and the second sample of the i-th core object as the adjustment necessity of the radius parameter of the i-th core object.
[0012] Furthermore, the adjustment of the radius parameter of the core object according to the adjustment necessity of the radius parameter of the core object includes the following specific steps: Calculate the difference between 1 and the adjustment necessity of the radius parameter of the i-th core object that needs to be adjusted, and denote the product of the difference and the preset radius as the radius parameter of the adjusted core object.
[0013] Furthermore, the classification of text data according to the radius parameter of the adjusted core object and the completion of text data retrieval according to the classification result of the text data include the following specific contents: Cluster the text data through the radius parameter of the adjusted core object to obtain the text data of each cluster after final clustering, and label them; Traverse the text data of each labeled cluster in sequence, retrieve the keywords of the text data in each cluster of text data. If the text data in the cluster does not contain the keyword to be retrieved, then retrieve the text data in the next cluster.
[0014] The beneficial effects of the technical solution of the present invention are as follows: By performing similarity measurement on text data and applying the DBSCAN algorithm, the core objects of the text data can be determined. Based on the continuity of sample similarity within the preset neighborhood of the core object and the analysis of the differences between high and low similarity samples of the core object, the necessity of adjusting the radius parameter of the core object can be determined. According to this analysis result, the radius parameter of the core object can be adaptively adjusted, thereby improving the text data clustering method using the Euclidean distance as the distance metric. This method realizes the fine division of text data with different similarity degrees, improves the accuracy of clustering, and further improves the efficiency of data retrieval. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0016] Figure 1 It is a step flowchart of a method for optimizing the intelligent retrieval efficiency of computer data according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0017] To further elaborate on the technical means and effects adopted by the present invention to achieve the intended invention purpose, the following specifically describes, with reference to the accompanying drawings and preferred embodiments, a method for optimizing the intelligent retrieval efficiency of computer data based on the present invention, including its specific implementation manner, structure, features, and effects. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, the specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.
[0018] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs.
[0019] The following specifically describes the specific solution of a method for optimizing the intelligent retrieval efficiency of computer data provided by the present invention with reference to the accompanying drawings.
[0020] Please refer to Figure 1 , which shows a flowchart of the steps of a method for optimizing the intelligent retrieval efficiency of computer data provided by an embodiment of the present invention. The method includes the following steps: Step S001: Obtain the text data in the computer data.
[0021] There are various types of computer data, such as text, audio, video, etc. In this embodiment, the intelligent retrieval optimization of text data is taken as an example. The text data is obtained by accessing the database. The Oracle database is used in this embodiment. The text data stored in the database is news data. In other embodiments, other databases or other text data can be used, such as Sql databases and paper data, etc. The text data includes the text title and the text content.
[0022] Step S002: Obtain the similarity degree between the text data.
[0023] For the text title, keywords in the text title are extracted through TextRank. Considering that some titles may have multiple keywords, the top two nodes with the highest ranking are selected as keywords according to TextRank. At the same time, for the extracted keywords, the keywords are converted into text vectors through the word embedding method. This method can obtain the semantic and syntactic relationships between the keywords, which helps to improve the accuracy of subsequent analysis of the similarity degree. Among them, TextRank and the word embedding method are well-known technologies, and the specific methods are not elaborated in this embodiment.
[0024] Since there may be multiple keywords in the text title, the measurement of the similarity degree between text data in this embodiment is analyzed by considering the weights and similarities between keywords. Specifically, the keywords extracted from the text titles of the text data are used to measure the similarity between the text data. The similarity degree between the text data is characterized by the cosine similarity of the text vectors, and the feature weights between the keywords are also considered. For each keyword in the text title, the number of times it appears in the text content is counted, and the ratio of the number of times the current keyword appears to the sum of the number of times all keywords appear is used as the weight corresponding to each keyword in the text title.
[0025] This embodiment provides a method for obtaining the similarity degree between text data as follows: Calculate the product of the average weight of the keywords in the m-th combination of the keywords of any two text contents and the cosine similarity value of the text vectors of the keywords in the n-th combination, and denote it as the first product of the m-th combination. The average value of the first products of all combinations is used as the similarity degree between any two text data. and the -th combination, and record it as the first product of the -th combination. The average value of the first products of all combinations is used as the similarity degree between any two text data.
[0026] It should be noted that: the calculation formula for the similarity degree between any two text data is: In the formula: represents the similarity degree between any two text data; M represents the number of combinations of keywords in the text titles of any two text data (pairwise combination of keywords in different text titles); represents the average weight of the keywords in the m-th combination of the keywords of any two text titles; represents the cosine similarity value of the text vectors of the keywords in the m-th combination of the keywords of any two text titles.
[0027] This calculation method pairwise combines the keywords of any two text titles, and uses the average weight of the keywords in the combination to adjust the cosine similarity value corresponding to the current combination. Since the weights of different keywords are different, the confidence levels corresponding to the similarity degrees analyzed in the combination are also different. When the weights of the two keywords in the combination are larger, it means that the keywords in the current combination are more representative for the two text data, and the confidence level of the similarity degree of the text data analyzed therefrom is greater. The closer the value of the cosine similarity is to 1, the greater the corresponding similarity degree.
[0028] In another embodiment, keywords in the text content are used to calculate the similarity degree between any two text data. The specific method is to count the number of occurrences of keywords in the text content, and use the ratio of the number of occurrences of the current keyword to the sum of the number of occurrences of all keywords as the weight corresponding to each keyword in the text content. A method for obtaining the similarity degree between text data is given as follows: In the formula: represents the similarity degree between any two text data; represents the number of combinations of keywords in the text content of any two text data (pairwise combination of keywords in different text contents); represents the th combination of keywords in any two text contents, and the average weight of the keywords; represents the th combination of keywords in any two text contents, and the cosine similarity value of the keyword text vectors;
[0029] Thus, the similarity degree between any two text data is obtained.
[0030] Step S003: Determine core objects for the text data according to the DBSCAN algorithm and the similarity degree between the text data, and obtain the sample similarity continuity degree of the core objects and the difference of the high and low similarity samples.
[0031] Regarding each text data as a sample point, all sample points are clustered by the DBSCAN algorithm. During the clustering process, the metric distance between any two text data is ; The core points (core objects) in the DBSCAN algorithm refer to the points that contain enough (greater than or equal to the minimum number of points MinPts) sample points within a given radius eps (i.e., within the neighborhood of the sample points). These points are considered as internal points of the cluster in the DBSCAN algorithm and are density-reachable. The principle of the DBSCAN algorithm and the concept of the core points in the DBSCAN algorithm are well-known, and the specific implementation process is not described in detail in this embodiment. In this embodiment, subsequent analysis of the text data is carried out according to the core points in the DBSCAN algorithm. In this embodiment, the preset neighborhood radius size eps = 6 and the minimum neighborhood sample number threshold MinPts = 5. In other embodiments, the numerical values can be set according to specific situations, and this embodiment does not make specific limitations.
[0032] Analyze the distribution continuity of text data for the neighborhood corresponding to the preset radius eps of the core object. The more continuous the text data distribution is, the better the similarity and continuity performance between the data, and the greater the possibility of corresponding to the same type of text data; the less continuous the distribution is, the more likely the data may have a break, indicating that there are text data of different categories in the neighborhood of the current core object. At the same time, consider the difference between high-similarity and low-similarity samples in the neighborhood of the core object based on the similarity and continuity. The farther the distance between the two and the denser their respective distributions are, the greater the possibility that there are text data of different categories in the neighborhood of the current core object.
[0033] It should be noted that the similarity and continuity between data are reflected by the sample similarity and continuity degree of the core object. Specifically, taking the distance from each sample point in the neighborhood of the core object to the core point as the sample point radius, the same sample point radius may correspond to multiple sample points (these sample points are all on the circle with the current core point as the center and the radius as the sample point radius). Statistically sort the sample point radii in ascending order (repeated ones are also sorted), and at the same time calculate the difference between adjacent sample point radii. In this embodiment, for the sorted sample point radii, calculate the difference between adjacent sample point radii by subtracting the previous one from the next one, where the last one does not participate in the calculation. From the obtained differences between adjacent sample point radii, select B largest differences between adjacent sample point radii, denoted as the sample radius interval. In this embodiment, B = 10 is taken as an example for description, and other values can be used in other embodiments, which are not specifically limited in this embodiment. A method for obtaining the sample similarity and continuity degree of the core object is given as follows: Calculate the b-th sample radius interval selected for the i-th core object The difference in the number of sample points included in the two sample points corresponding to the b-th sample radius interval The absolute value ratio is used as the first ratio of the b-th sample radius interval, and the inverse normalization value of the mean of the first ratios of all sample radius intervals is used as the adjustment coefficient of the maximum sample point radius in the neighborhood of the i-th core object; The normalized value of the product of the maximum sample point radius in the neighborhood of the i-th core object and the adjustment coefficient of the maximum sample point radius in the neighborhood of the i-th core object is used as the sample similarity and continuity degree of the i-th core object.
[0034] It should be noted that: The calculation formula for the sample similarity and continuity degree of the i-th core object is: In the formula: Represents the adjustment coefficient of the maximum sample point radius in the neighborhood of the i-th core object.
[0035] In the formula: Represents the sample similarity continuity degree of the $i$-th core object; Represents the sample point radius within the neighborhood of the $i$-th core object; Represents the maximum sample point radius within the neighborhood of the $i$-th core object; $B$ represents $B$ sample radius intervals selected for the $i$-th core object; Represents the $b$-th selected sample radius interval; the $b$-th sample radius interval corresponds to two sample points, and there are multiple sample points on the circle (the circle corresponding to the sample point radius with the $i$-th core as the center) where each of these two sample points is located. Denote the difference in the number of sample points on the circles corresponding to the two sample points as ; In this embodiment, Is briefly recorded as the difference in the number of sample points contained in the two sample points corresponding to the $b$-th sample radius interval, where Adding 1 in it is to avoid the denominator being 0; $\max( )$ represents taking the maximum value; $| |$ represents taking the absolute value; Represents the linear normalization function.
[0036] This calculation method reflects the maximum continuous value of the current core object through The larger its value, the greater the possible sample similarity continuity degree in the end. Because the overall formula is adjusted on the basis of
[0037] In the formula, Is used to determine the adjustment coefficient of the overall formula, Can only reflect the maximum continuous value of the current core object, but cannot determine the continuous situation during the process of reaching the maximum continuous value. Therefore, in this embodiment, multiple maximum sample radius intervals are selected to analyze the continuous situation at each sample interval. The more continuous at each sample interval, the more continuous the whole, and the larger the adjustment coefficient value; Represents the $b$-th selected sample radius interval. The larger this value, the less continuous the sample point distribution here (because the sample points are sorted, so the sample radius interval value here can reflect the continuous situation); The continuity is further reflected by The smaller the difference in the number of the two types of samples that make up the current sample radius interval, the more it can indicate that there is a fault here as a whole, that is, the sample point distribution here is less continuous; The larger it is, the less continuous it is. By Adjusting the corresponding relationship, so the continuous situation is reflected through each sample radius interval. The larger the value, the better the continuous situation, and the larger the adjustment coefficient value.
[0038] Further, according to the B sample radius intervals screened above, the core objects are divided into high - similarity samples (denoted as the first samples) and low - similarity samples (denoted as the second samples). Specifically, for the B sample radius intervals screened according to the i - th core point, for the b - th sample radius interval among the B sample radius intervals, when the distance between any other core point and the i - th core point is greater than the b - th sample radius interval, this core point is a high - similarity sample point of the i - th core point, and all high - similarity sample points constitute the high - similarity sample of the i - th core point, that is, the high - similarity sample of the i - th core object; when the distance between any other core point and the i - th core point is less than or equal to the b - th sample radius interval, this core point is a low - similarity sample point of the i - th core point, and all low - similarity sample points constitute the low - similarity sample of the i - th core point, that is, the low - similarity sample of the i - th core object.
[0039] Traverse the shortest distances between any two sample points in the high - and low - similarity samples (the points taken are not repeated, and the shortest distance is taken as the criterion), and arrange them from low to high to obtain a shortest - distance sequence of high - and low - similarity samples. For example: Calculate the distance between sample point qa and sample point qb in the high - and low - similarity samples. Next time, calculate the distance between any two sample points among the remaining sample points except sample point qa and sample point qb until all sample points in the high - and low - similarity samples have participated in the calculation.
[0040] Under any one of the sample radius intervals screened for the i - th core object, calculate the mean value of all the shortest distances in the shortest - distance sequence of the first samples and the second samples, and multiply the average distance from each sample point in the first samples to the i - th core object by the standard deviation of the distances from each sample point in the first samples to the core object to obtain a second product. Multiply the average distance from each sample point in the second samples to the core object by the standard deviation of the distances from each sample point in the second samples to the core object to obtain a second product. Take the sum of the first product and the second product as the first sum value, and take the ratio of the mean value of all the shortest distances to the first sum value as the second ratio; Take the normalized value of the maximum value among the second ratios under all the sample radius intervals screened for the i - th core object as the difference between the first samples and the second samples of the i - th core object.
[0041] It should be noted that: In this embodiment, a calculation method for obtaining the difference between high - and low - similarity samples is given as follows: In the formula: It represents the difference between the high- and low-similarity samples of the $i$-th core object, that is, the difference between the first sample and the second sample; $B$ represents the preset number $B$ of sample radius intervals selected for the $i$-th core object; $K$ represents the number of the shortest distances in the shortest distance sequence of the high- and low-similarity samples (the first sample and the second sample). It represents the $k$-th shortest distance in the shortest distance sequence of the high- and low-similarity samples (the first sample and the second sample). It represents the average distance from each sample point in the first sample (high-similarity sample) to the $i$-th core object. It represents the standard deviation of the distances from each sample point in the first sample (high-similarity sample) to the core object. It represents the average distance from each sample point in the second sample (low-similarity sample) to the core object. It represents the standard deviation of the distances from each sample point in the second sample (high-similarity sample) to the core object; $\max()$ represents taking the maximum value. It represents the linear normalization function.
[0042] This calculation method is through It represents the distance between the high- and low-similarity samples, which is mainly determined by averaging the accumulation of the shortest distances corresponding to the high- and low-similarity sample ranges. When the shortest distance is far enough, the actual distance is naturally farther. It represents the comprehensive density of the high- and low-similarity samples, where It represents the density of the high-similarity samples. Taking the core object as a reference, the smaller the standard deviation, the more consistent the distances from each sample point to the core object, but it cannot explain the specific distance degree, so it is reflected by the average distance. Combining the two together reflects the density of the high-similarity samples; for Similarly, it should be added that the analysis of this method mainly considers the overall density and does not consider the density in the internal local areas. Finally, the two are accumulated to determine the comprehensive density of the high- and low-similarity samples; through It reflects the difference between the high- and low-similarity samples corresponding to one sample radius interval. When the distance between the high- and low-similarity samples is larger, the comprehensive density of the high- and low-similarity samples is smaller, indicating that the difference between the high- and low-similarity samples corresponding to the current sample radius interval is larger. The maximum difference between the high- and low-similarity samples is selected through $\max()$ as the difference between the high- and low-similarity samples of the current core object.
[0043] So far, the sample similarity continuity of the sample data core object and the difference between the high- and low-similarity samples of the core object are obtained.
[0044] Step S004: Obtain the necessity of adjusting the radius parameter of the core object according to the continuity degree of sample similarity of the core object and the difference between high and low similarity samples, and adjust the radius parameter of the core object according to the necessity of adjustment.
[0045] The above process analyzes the continuity degree of sample similarity of the core object and the difference between high and low similarity samples, uses the continuity degree of sample similarity of the core object as a weight to adjust the difference between its high and low similarity samples, so as to judge the necessity of adjusting the radius parameter of the core object.
[0046] Calculate the difference between 1 and the continuity degree of sample similarity of the i-th core object, and denote the product of the difference and the difference between the first sample and the second sample of the i-th core object as the necessity of adjusting the radius parameter of the i-th core object.
[0047] Specifically, the calculation formula for the necessity of adjustment is as follows: In the formula: represents the necessity of adjusting the radius parameter of the i-th core object; represents the continuity degree of sample similarity of the i-th core object; represents the difference between high and low similarity samples of the i-th core object.
[0048] This calculation method shows that the smaller the continuity degree of sample similarity of the core object, the greater the possibility that there is a fault in the sample distribution, the greater the corresponding necessity of adjusting the final radius parameter, and the greater the determined weight; and at the same time, it is adjusted in combination with the difference between high and low similarity samples of the core object.
[0049] So far, obtain the necessity of adjusting the radius parameter of the core object. For those greater than or equal to the threshold, adjust their radius parameters, otherwise do not adjust. In this embodiment, the threshold is set to 0.6, and other implementation manners can select other values according to the actual situation, which are not limited in this embodiment. One adjustment method given in this embodiment is: Calculate the difference between 1 and the necessity of adjusting the radius parameter of the i-th core object that needs to be adjusted, and denote the product of the difference and the preset radius as the radius parameter of the adjusted core object.
[0050] Specifically, the calculation formula for the radius parameter of the adjusted core object is as follows: In the formula, represents the radius parameter of the adjusted core object; represents the preset radius; represents the necessity of adjusting the radius parameter of the i-th core object that needs to be adjusted.
[0051] Step S005: Perform final clustering on the text data according to the radius parameter of the adjusted core object, and perform retrieval based on the clustering.
[0052] Adjust the radius parameter of the core object through the above steps (the radius parameter of the core object that has not been adjusted is equal to the preset radius ), and accordingly perform clustering to obtain the text data of each cluster after final clustering, and number them.
[0053] When performing text data retrieval subsequently, traverse the text data of each numbered cluster in sequence, retrieve the keywords of the text data in the text data of each cluster. If there are relevant keywords to be retrieved within the cluster, perform text data retrieval within the current cluster, and do not perform text data retrieval in other clusters. If there are no relevant keywords within the current cluster, then perform screening of the next cluster. This method can achieve more targeted search and screening of relevant data within a small range, improving the retrieval efficiency of the data.
[0054] In other embodiments, after retrieving the text data within the current cluster, it is also possible to continue with the screening retrieval of the next cluster, and finally multiple text data of different cluster classes can be retrieved simultaneously.
[0055] Thus, this embodiment is completed.
[0056] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for optimizing the efficiency of intelligent retrieval of computer data, characterized in that: The method comprises the following steps: Obtaining text data and keywords of the text data in computer data; Obtain the similarity between text data according to the number of times the keyword appears in the text data and the similarity value of the keyword text vector; The text data is used as sample points and the DBSCAN algorithm is used to cluster the sample points to obtain the core objects in the clustering process, where the metric distance used in the clustering process is the similarity between the text data; The sample similarity continuity degree of the core object is obtained according to the sample radius interval screened out by the core object and the quantity difference of the two sample points included in the sample radius interval; The core object is divided into a first sample and a second sample according to the screened sample radius interval, wherein the distance between the sample point and the core object in the first sample is greater than the sample radius interval, and the distance between the sample point and the core object in the second sample is less than or equal to the sample radius interval; Obtain the difference between the first sample and the second sample of the core object according to the screened sample radius interval, the shortest distance between the first sample and the second sample, and the distance from each sample point in the first sample and the second sample to the core object; Obtaining the necessity of adjusting the radius parameter of the core object according to the degree of continuity of sample similarity of the core object and the difference between the first sample and the second sample; Adjusting the radius parameter of the core object according to the necessity of adjusting the radius parameter of the core object; The text data is classified according to the adjusted radius parameter of the core object, and the text data is retrieved according to the classification result of the text data.
2. According to claim 1, a method for optimizing computer data intelligent retrieval efficiency is characterized in that: The similarity between the text data includes the following specific steps: Calculate the first keyword of any two text contents The average weight of the keywords in the combination is The product of the cosine similarity values of the keyword text vectors in the combinations is recorded as The first product of all combinations is taken as the average of the first products of all combinations as the similarity between any two text data.
3. According to claim 1, a method for optimizing computer data intelligent retrieval efficiency is characterized in that: The metric distance used in the clustering process is , exp( ) represents an exponential function with a natural constant as the base, Indicates the similarity between any two text data.
4. According to claim 1, a method for optimizing computer data intelligent retrieval efficiency is characterized in that: The specific steps of obtaining the sample similarity continuity degree of the core object according to the sample radius interval screened out by the core object and the quantity difference of the two sample points included in the sample radius interval are as follows: The distance from each sample point in the neighborhood of the core object to the core object is the sample point radius; Sort the radius of all sample points in the neighborhood of the core object in ascending order, and calculate the radius difference of adjacent sample points after sorting; A preset number B of maximum radius differences of adjacent sample points are selected from the radius differences of adjacent sample points and recorded as the sample radius interval; For the B sample radius intervals filtered out by the ith core object, calculate the bth sample radius interval filtered out by the ith core object The difference in the number of sample points contained in the two sample points corresponding to the b-th sample radius interval The ratio of the absolute values of is taken as the first ratio of the b-th sample radius interval, and the inversely proportional normalized value of the mean of the first ratios of all sample radius intervals is taken as the adjustment coefficient of the maximum sample point radius in the neighborhood of the ith core object; The sample similarity continuity degree of the core object is obtained according to the adjustment coefficient of the maximum sample point radius in the core object neighborhood and the maximum sample point radius in the core object neighborhood.
5. According to claim 4, a method for optimizing computer data intelligent retrieval efficiency is characterized in that: The step of obtaining the sample similarity continuity degree of the core object according to the adjustment coefficient of the maximum sample point radius in the core object neighborhood and the maximum sample point radius in the core object neighborhood includes the following specific steps: The normalized value of the product of the maximum sample point radius in the neighborhood of the i-th core object and the adjustment coefficient of the maximum sample point radius in the neighborhood of the i-th core object is used as the sample similarity continuity degree of the i-th core object.
6. According to claim 1, a method for optimizing computer data intelligent retrieval efficiency is characterized in that: The specific steps of obtaining the difference between the first sample and the second sample of the core object according to the screened sample radius interval, the shortest distance between the first sample and the second sample, and the distance from each sample point in the first sample and the second sample to the core object are as follows: Under any sample radius interval selected by the i-th core object, calculate the average of all the shortest distances in the shortest distance sequence between the first sample and the second sample, and convert the average distance from each sample point in the first sample to the i-th core object into The standard deviation of the distance from each sample point in the first sample to the core object The product of is recorded as the second product, and the average distance from each sample point in the second sample to the core object is The standard deviation of the distance from each sample point in the second sample to the core object The product of is recorded as the second product, the sum of the first product and the second product is recorded as the first sum, and the ratio of the mean of all the shortest distances to the first sum is recorded as the second ratio; The normalized value of the maximum value of the second ratios of all sample radius intervals screened out from the i-th core object is used as the difference between the first sample and the second sample of the i-th core object.
7. According to claim 6, a method for optimizing computer data intelligent retrieval efficiency, characterized in that: The specific method of obtaining the shortest distance sequence between the first sample and the second sample is as follows: The shortest distances between any two sample points in the first sample and the second sample are traversed, and the any two sample points are not selected repeatedly. The shortest distances between any two sample points are arranged from low to high to obtain the shortest distance sequence between the first sample and the second sample.
8. According to claim 1, a method for optimizing computer data intelligent retrieval efficiency is characterized in that: The necessity of adjusting the radius parameter of the core object according to the degree of continuity of sample similarity of the core object and the difference between the first sample and the second sample is obtained, and the specific steps include the following: The difference between 1 and the sample similarity continuity of the ith core object is calculated, and the product of the difference and the difference between the first sample and the second sample of the ith core object is recorded as the necessity of adjusting the radius parameter of the ith core object.
9. The method for optimizing computer data intelligent retrieval efficiency according to claim 1, characterized in that: The step of adjusting the radius parameter of the core object according to the necessity of adjusting the radius parameter of the core object includes the following specific steps: The difference between 1 and the necessity of adjusting the radius parameter of the i-th core object that needs to be adjusted is calculated, and the product of the difference and the preset radius is recorded as the radius parameter of the adjusted core object.
10. The method for optimizing computer data intelligent retrieval efficiency according to claim 1, characterized in that: The text data is classified according to the radius parameter of the adjusted core object, and the text data retrieval is completed according to the classification result of the text data, including the specific contents as follows: Clustering the text data by adjusting the radius parameter of the core object, obtaining the text data of each cluster after the final clustering, and labeling them; Each cluster of labeled text data is traversed in turn, and the keywords of the text data are searched in each cluster of text data. If the text data in the cluster does not contain the keywords to be searched, the text data in the next cluster is searched.
Citation Information
Patent Citations
An improved algorithm for carrying out abnormal mining on density irregular data based on DBSCAN
CN109669990A
Density-based text clustering method, device and equipment, and storage medium
CN112528025A
Detection optimization method for similar duplicate records based on density clustering algorithm DBSCAN algorithm
CN116451675A
Intelligent screening method and system for approval workflow data based on machine learning
CN118312656A
Ultrasonic data intelligent management method and system
CN118364324A