Method and System for Expanding and Correcting Agricultural Knowledge Graph Data
By applying the word frequency calculation model and fluctuation calculation model in the agricultural knowledge graph, the inefficiency and data quality problems in the data expansion and correction process are solved, and more efficient and accurate data integration is achieved, and agricultural production and scientific and technological innovation are improved.
Patent Information
- Application Number
- CN202510100782.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-01-22
AI Technical Summary
In the process of data expansion and correction, agricultural knowledge graphs have problems such as inefficiency, susceptibility to subjective factors, data redundancy and irrelevance, which makes it difficult to guarantee the comprehensiveness and accuracy of the knowledge graph.
The method based on word frequency calculation model and fluctuation calculation model is adopted to obtain word frequency indicators of various categories of words, and filter out word frequency indicators to be processed through sorting and fluctuation analysis, formulate expansion strategies, correct the location of the data set, and ensure the comprehensiveness and accuracy of data expansion.
It improves the comprehensiveness and accuracy of data expansion, avoids the introduction of redundant and irrelevant data, and enables new data sets to be integrated into relevant categories in the knowledge graph more quickly, improving agricultural production efficiency and agricultural scientific and technological innovation.
Smart Images

Figure CN119557459B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly to a method and system for expanding and correcting agricultural knowledge graph data. Background Art
[0002] Today, with the rapid development of agricultural informatization, the agricultural knowledge graph, as an important carrier for agricultural knowledge management and services, its construction and update are of great significance for improving agricultural production efficiency and promoting agricultural scientific and technological innovation. However, the operation and maintenance of the agricultural knowledge graph, especially the expansion and correction, face many technical challenges and practical problems.
[0003] Most of the traditional methods for expanding agricultural knowledge graph data rely on manual screening and manual addition. This process is not only inefficient but also easily affected by subjective factors, making it difficult to guarantee the comprehensiveness and accuracy of data expansion. With the continuous emergence of new knowledge in the agricultural field, some automated expansion methods have emerged, but these methods often lack a fine screening mechanism, easily leading to redundancy and irrelevance of the expanded data, thus affecting the overall quality and application value of the knowledge graph. How to quickly and accurately integrate this new knowledge into the existing knowledge graph has become a key problem to be solved urgently; at the same time, the sorting of categories to be expanded in the agricultural knowledge graph has an important impact on user retrieval and use of knowledge. Traditional sorting methods may lead to chaotic category sorting due to the lack of a scientific basis and a dynamic adjustment mechanism, making it difficult for users to quickly locate the required information and reducing the practicality and user experience of the knowledge graph. Summary of the Invention
[0004] In view of the defects in the prior art, the present invention provides a method and system for expanding and correcting agricultural knowledge graph data.
[0005] A method for expanding and correcting agricultural knowledge graph data, comprising: obtaining multiple different categories to be expanded based on the agricultural knowledge graph, extracting category words corresponding to the categories to be expanded, and obtaining a new dataset, where the new dataset includes text data; obtaining word frequency indicators corresponding to each category word based on a word frequency calculation model and the new dataset, sorting the multiple word frequency indicators in descending order, and obtaining a word frequency sequence; obtaining a decline degree indicator of adjacent word frequency indicators in the word frequency sequence based on a fluctuation calculation model and the word frequency sequence, and determining whether there is a decline degree indicator greater than a first threshold; if there is no decline degree indicator greater than the first threshold, adding the new dataset to the multiple different categories to be expanded in sequence, and correcting the position of the new dataset in the category to be expanded according to the size of the word frequency indicator of the new dataset in the category to be expanded; if there is a decline degree indicator greater than the first threshold, screening out the first decline degree indicator greater than the first threshold in the order of the word frequency sequence, and using the adjacent word frequency indicators corresponding to the decline degree indicator as two word frequency indicators to be processed; determining whether the word frequency indicators to be processed are less than a second threshold, obtaining an expansion strategy according to the determination result, and correcting the position of the new dataset in the category to be expanded according to the size of the word frequency indicator of the new dataset in the category to be expanded.
[0006] Optionally, obtaining an expansion strategy according to the determination result includes: if the word frequency indicator to be processed is less than the second threshold, obtaining all the word frequency indicators before the word frequency indicator to be processed in the word frequency sequence, finding the corresponding category to be expanded according to the category words corresponding to the obtained word frequency indicators, and adding the new dataset to the obtained categories to be expanded in sequence.
[0007] Optionally, obtaining an expansion strategy according to the determination result further includes: if the word frequency indicator to be processed is not less than the second threshold, screening out the second decline degree indicator greater than the first threshold in the order of the word frequency sequence, and using the adjacent word frequency indicators corresponding to the decline degree indicator as two new word frequency indicators to be processed, and continuing to determine whether the new word frequency indicators to be processed are less than the second threshold.
[0008] Optionally, obtaining word frequency indicators corresponding to each category word based on a word frequency calculation model and the new dataset is expressed as: ; where is the word frequency indicator corresponding to the i-th category word, is the number of occurrences of the i-th category word in the text data of the new dataset, is the total number of words in the text data of the new dataset.
[0009] Optionally, obtaining the decline degree index of adjacent word frequency metrics in the word frequency sequence based on the fluctuation calculation model and the word frequency sequence includes: obtaining the i-th word frequency metric and the (i + 1)-th word frequency metric in the word frequency sequence; obtaining the decline value according to the i-th word frequency metric and the (i + 1)-th word frequency metric; and obtaining the decline degree index according to the decline value and the i-th word frequency metric.
[0010] Optionally, obtaining the decline degree index according to the decline value and the i-th word frequency metric includes: ; where is the decline degree index of the i-th word frequency metric and the (i + 1)-th word frequency metric, is the i-th word frequency metric, is the (i + 1)-th word frequency metric.
[0011] Optionally, correcting the position of the newly added dataset under the category to be expanded according to the word frequency metric size of the newly added dataset in the category to be expanded includes: obtaining the association relationship between the position range of the current category to be expanded and the word frequency metric based on the agricultural knowledge graph; obtaining the position range to be corrected of the newly added dataset in the current category to be expanded according to the association relationship and the word frequency metric of the newly added dataset, and correcting the position of the newly added dataset according to the position range to be corrected.
[0012] Optionally, the system includes: an obtaining module, configured to obtain a plurality of different categories to be expanded based on the agricultural knowledge graph, extract category words corresponding to the categories to be expanded, and obtain a newly added dataset, where the newly added dataset includes text data; a calculation module, configured to obtain word frequency metrics corresponding to each category word based on the word frequency calculation model and the newly added dataset, sort the plurality of word frequency metrics in descending order, and obtain a word frequency sequence; a calculation and judgment module, configured to obtain the decline degree index of adjacent word frequency metrics in the word frequency sequence based on the fluctuation calculation model and the word frequency sequence, and judge whether there is a decline degree index greater than a first threshold; a first judgment execution module, configured to, when there is no decline degree index greater than the first threshold, sequentially add the newly added dataset to the plurality of different categories to be expanded, and correct the position of the newly added dataset according to the word frequency metric size of the newly added dataset under the category to be expanded; a second judgment execution module, configured to, when there is a decline degree index greater than the first threshold, screen out the first decline degree index greater than the first threshold in the order of the word frequency sequence, and use the adjacent word frequency metric corresponding to the decline degree index as two word frequency metrics to be processed; a third judgment execution module, configured to judge whether the word frequency metric to be processed is less than a second threshold, obtain an expansion strategy according to the judgment result, and correct the position of the newly added dataset according to the word frequency metric size of the newly added dataset under the category to be expanded.
[0013] Optionally, the third judgment execution module is further configured to: if the word frequency index to be processed is less than the second threshold, obtain all the word frequency indexes before the word frequency index to be processed in the word frequency sequence, find the corresponding category to be expanded according to the category words corresponding to the obtained word frequency indexes, and add the new data set to the obtained category to be expanded in sequence.
[0014] Optionally, the third judgment execution module is further configured to: if the word frequency index to be processed is not less than the second threshold, screen out the second decline degree index greater than the first threshold in the order of the word frequency sequence, and use the adjacent word frequency indexes corresponding to the decline degree index as two new word frequency indexes to be processed, and continue to judge whether the new word frequency indexes to be processed are less than the second threshold.
[0015] The beneficial effects of the present invention are reflected in:
[0016] In the entire agricultural knowledge graph data expansion and correction method, first, the word frequency indexes of various category words are calculated based on the new data set and sorted to form a word frequency sequence, and then the significant fluctuation points in the word frequency sequence are identified through a fluctuation calculation model, and the word frequency indexes to be processed are screened out accordingly; further, the second threshold is used to further subdivide the importance of the word frequency indexes to be processed, so as to formulate a more reasonable expansion strategy, improve the comprehensiveness and accuracy of data expansion, avoid the introduction of redundant and irrelevant data, enable the new data set to be more quickly integrated into the relevant categories in the knowledge graph, and at the same time ensure that these categories are arranged in an orderly manner according to the degree of association with the new data set, thereby improving agricultural production efficiency, promoting agricultural scientific and technological innovation, and providing strong support for agricultural knowledge management and services. Description of the Drawings
[0017] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the specific embodiments or the prior art. In all the drawings, similar elements or parts are generally identified by similar reference numerals. In the drawings, the elements or parts are not necessarily drawn to scale.
[0018] Figure 1 It is a schematic diagram of the steps of the agricultural knowledge graph data expansion and correction method of the present invention;
[0019] Figure 2 It is a partial schematic diagram of the steps of S6 in the agricultural knowledge graph data expansion and correction method of the present invention;
[0020] Figure 3 It is a partial schematic diagram of the steps of S3 in the agricultural knowledge graph data expansion and correction method of the present invention;
[0021] Figure 4This is a schematic diagram of some steps of S4 in the method for expanding and correcting agricultural knowledge graph data of the present invention. Detailed implementation manners
[0022] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Usually, the components of the embodiments of the present invention described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations.
[0023] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed present invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0024] It should be noted that: similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. In addition, the terms "first", "second", etc. are only used for descriptive distinction and cannot be understood as indicating or implying relative importance.
[0025] As Figure 1 shown, a method for expanding and correcting agricultural knowledge graph data is provided, including:
[0026] S1. Obtain a plurality of different categories to be expanded based on the agricultural knowledge graph, extract the category words corresponding to the categories to be expanded, and obtain a new data set, where the new data set includes text data;
[0027] S2. Obtain the word frequency indexes corresponding to each category word based on the word frequency calculation model and the new data set, sort the plurality of word frequency indexes in descending order, and obtain a word frequency sequence;
[0028] S3. Obtain the descending degree indexes of adjacent word frequency indexes in the word frequency sequence based on the fluctuation calculation model and the word frequency sequence, and determine whether there is a descending degree index greater than a first threshold;
[0029] S4. If there is no descending degree index greater than the first threshold, then sequentially add the new data set to a plurality of different categories to be expanded, and correct the position where the new data set is located according to the size of the word frequency index of the new data set under the category to be expanded;
[0030] S5. If there is a degree-of-descent index greater than the first threshold, then screen out the first degree-of-descent index greater than the first threshold in the order of the word-frequency sequence, and use the adjacent word-frequency index corresponding to this degree-of-descent index as the two word-frequency indices to be processed.
[0031] S6. Determine whether the word-frequency index to be processed is less than the second threshold, obtain an expansion strategy based on the judgment result, and correct the position where the new dataset is located according to the word-frequency index size of the new dataset under the category to be expanded.
[0032] In this embodiment, it should be noted that in S1, first, it is necessary to identify and extract multiple different categories to be expanded from the already constructed agricultural knowledge graph. These categories to be expanded are predefined in the knowledge graph and are used to classify and organize various knowledge and information in the agricultural field. For example, the categories to be expanded may include "rice planting techniques", "pest control", "agricultural machinery and equipment", etc. Each category to be expanded corresponds to a category word. For example, the category word for "rice planting techniques" is "rice planting". By traversing the hierarchical structure and node information of the knowledge graph, these categories to be expanded can be automatically identified, laying a foundation for subsequent data expansion and correction work.
[0033] Next, it is necessary to obtain a new dataset and use it as the raw material for data expansion and correction. The new dataset can come from various channels, such as the latest research results released by agricultural research institutions, the empirical data accumulated in agricultural production practices, the update of agricultural policies and regulations, etc. These datasets usually exist in the form of text data and contain rich agricultural knowledge and information. After obtaining the new dataset, it is necessary to preprocess it, including operations such as deduplication, cleaning, and formatting. At the same time, it is also necessary to extract the category words related to the categories to be expanded from the new dataset. These category words are the key links connecting the new dataset and the categories to be expanded. For example, if the new dataset contains an article about "new rice planting techniques", then it is necessary to extract the category word "rice planting" from it so as to associate and expand it with the "rice planting techniques" category in the knowledge graph in the future.
[0034] In S2, it is necessary to use a word frequency calculation model to process the newly added dataset and calculate the word frequency indicators corresponding to each category of words. The word frequency calculation model is a model based on text data analysis. It can count the frequency of a specific word in the text, thereby reflecting the relevance of the word to the text data. In S2, then count the number of times each category of word appears in the newly added dataset, that is, the word frequency indicator corresponding to each category of word. For example, the category words corresponding to the category to be expanded in S1 include "rice planting", "pest control", and "agricultural machinery". In the newly added dataset, the word frequency of "rice planting" is relatively high, followed by "pest control", and the least is "agricultural machinery". According to the magnitude of the word frequency, sort the word frequency indicators of multiple category words to form a word frequency sequence from large to small. This sequence reflects the importance and relevance of each category of word in the newly added dataset.
[0035] Suppose the number of times each category of word appears is counted in the newly added dataset. It is found that the word frequency indicator of the word "rice planting" in the text data is 0.03; the word frequency indicator of the word "pest control" in the text data is 0.01; the word frequency indicator of the word "agricultural machinery" in the text data is 0.005. Then, according to these word frequency data, the three category words "rice planting", "pest control", and "agricultural machinery" will be arranged in descending order of word frequency to form a word frequency sequence.
[0036] In S3, it is necessary to further analyze the fluctuation of adjacent word frequency indicators in the word frequency sequence based on the fluctuation calculation model and the obtained word frequency sequence. The fluctuation calculation model is a tool used to quantify the degree of difference between adjacent elements in a data sequence. Here, it is used to evaluate the degree of decline indicator of adjacent word frequency indicators. The specific operation is to calculate the difference between each pair of adjacent word frequency indicators in the word frequency sequence, and then obtain the degree of decline indicator that reflects the change range of the word frequency indicator from one category to the next through these differences. Then, compare these differences with a preset first threshold to determine whether there is a significant fluctuation. The first threshold is a critical value determined according to experience or statistical methods, used to distinguish normal word frequency changes and significant fluctuations.
[0037] For example, suppose the word frequency sequence is "rice planting" (0.03), "pest control" (0.01), and "agricultural machinery" (0.005), and the differences of adjacent word frequency indicators are 0.03-0.01=0.02 and 0.01-0.005=0.005 respectively. If the preset first threshold is 0.001, then the first difference 0.02 and the second difference 0.005 are both greater than the first threshold, indicating that there is a significant fluctuation in the word frequency indicators between "rice planting" and "pest control", and there is also a significant fluctuation in the word frequency indicators between "pest control" and "agricultural machinery". In this case, it is considered that the correlation between the newly added data set and the "rice planting" category is significantly higher than that of the "pest control" category, so different expansion strategies may be needed to handle these two categories. On the contrary, if the differences of all adjacent word frequency indicators are less than the first threshold, it is considered that the correlation between the newly added data set and all the categories to be expanded is relatively uniform, or the difference is not enough to serve as a basis for classification.
[0038] In S4, based on the previous fluctuation analysis results, if it is determined that there is no decline degree indicator greater than the first threshold, that is, the fluctuations of adjacent word frequency indicators in the word frequency sequence are kept within the normal range, this indicates that the correlation between the new data set and all categories to be expanded is relatively uniform, without obvious preference or significant difference. At this time, a comprehensive expansion strategy will be adopted to add the new data set to multiple different categories to be expanded in turn. During the adding process, the position of the new data set will be corrected according to the size of the word frequency indicator of the new data set under the category to be expanded, ensuring that categories with higher word frequency indicators can give priority to displaying relevant new data. Doing so not only ensures the comprehensiveness of data expansion, but also improves the retrieval efficiency and user experience of the knowledge graph through word frequency sorting.
[0039] For example, suppose that the word frequency sequence analyzed in S3 is "rice planting" (such as 0.03), "pest control" (such as 0.028), and "agricultural machinery" (such as 0.027), and the difference between adjacent word frequency indicators does not exceed the preset first threshold (such as 0.002). This means that the correlation between the newly added data set and the three categories to be expanded, "rice planting", "pest control" and "agricultural machinery" is relatively close, without significant fluctuations. Therefore, in step S4, the newly added data set will be fully expanded to these three categories. At the same time, according to the size of the word frequency index, it will ensure that the newly added data set under the "rice planting" category is relatively advanced, followed by "pest control" and finally "agricultural machinery". In this way, when users search or browse related knowledge, they can first see the newly added content most related to "rice planting", thereby improving the practicality and user satisfaction of the knowledge graph.
[0040] In S5, it has been determined through the fluctuation calculation model that there is a degree-of-descent index greater than the first threshold in the word frequency sequence, which means that the correlation between the new dataset and certain categories to be expanded is significantly higher than that of other categories. Therefore, it is necessary to screen out the first degree-of-descent index greater than the first threshold in the order of the word frequency sequence, and use the adjacent word frequency index corresponding to this degree-of-descent index as the two word frequency indicators to be processed. Suppose the word frequency sequence is "rice cultivation" (0.03), "pest control" (0.01), "agricultural machinery" (0.005), and the preset first threshold is 0.015. By calculating the difference between adjacent word frequency indicators, it is found that the difference from "rice cultivation" to "pest control" is 0.02 (greater than the first threshold), while the difference from "pest control" to "agricultural machinery" is 0.005 (less than the first threshold). At this time, the first degree-of-descent index greater than the first threshold, that is, 0.02, will be screened out, and the word frequency indicators corresponding to the two category words "rice cultivation" and "pest control" will be used as the word frequency indicators to be processed.
[0041] In S6, the introduction of the second threshold is indeed to help more precisely judge the importance of the word frequency indicators to be processed and the subsequent word frequency indicators, and accordingly formulate a more reasonable expansion strategy. Specifically, the second threshold plays the role of a "screening threshold" here. After determining the word frequency indicators to be processed, it will be checked whether these indicators are all less than the second threshold. If all the word frequency indicators to be processed are less than the second threshold, this usually means that in the word frequency sequence, starting from the word frequency indicators to be processed, the subsequent word frequency indicators are also relatively small, that is, the correlation between the new dataset and these subsequent categories is not strong. Therefore, it may be decided not to consider these subsequent categories anymore, but to focus mainly on the categories associated with the word frequency indicators to be processed, and a relatively conservative expansion strategy will be adopted, such as expanding with a lower priority or within a small range. On the contrary, if any one or both of the word frequency indicators to be processed are greater than or equal to the second threshold, this indicates that the correlation between the new dataset and these categories is not only significantly higher than that of other categories, but also reaches a relatively high level. At this time, a more aggressive expansion strategy may be adopted, such as adding the new dataset to these categories with a higher priority, and considering further screening the subsequent word frequency indicators in the word frequency sequence so as to add the new dataset to more relevant categories. This can ensure the comprehensiveness and accuracy of the knowledge graph, and at the same time improve the efficiency of users' retrieval and use of knowledge.
[0042] In summary, in the entire method for expanding and correcting agricultural knowledge graph data, first, word frequency indicators for various categories of words are calculated based on the newly added data set and sorted to form a word frequency sequence. Then, significant fluctuation points in the word frequency sequence are identified through a fluctuation calculation model, and the word frequency indicators to be processed are screened out based on this. Further, the second threshold is used to further subdivide the importance of the word frequency indicators to be processed, thereby formulating a more reasonable expansion strategy, improving the comprehensiveness and accuracy of data expansion, avoiding the introduction of redundant and irrelevant data, enabling the newly added data set to be more quickly integrated into the relevant categories in the knowledge graph, and at the same time ensuring that these categories are arranged in an orderly manner according to their degree of association with the newly added data set, thereby improving agricultural production efficiency, promoting agricultural scientific and technological innovation, and providing strong support for agricultural knowledge management and services.
[0043] As Figure 2 shown, in one embodiment, obtaining the expansion strategy according to the judgment result in S6 includes:
[0044] S61. If the word frequency indicator to be processed is less than the second threshold, obtain all the word frequency indicators before the word frequency indicator to be processed in the word frequency sequence, find the corresponding categories to be expanded according to the category words corresponding to the obtained word frequency indicators, and sequentially add the newly added data set to the obtained categories to be expanded.
[0045] In this embodiment, it should be noted that in S61, a further detailed analysis of the word frequency indicators identified in S6 will be carried out. When the word frequency indicator to be processed is less than the second threshold, this usually means that starting from this word frequency indicator to be processed, the subsequent word frequency indicators have relatively small values in the word frequency sequence, reflecting that the degree of association between the newly added data set and these subsequent categories is not strong. Then, the word frequency sequence will be traced back to obtain all the word frequency indicators before the word frequency indicator to be processed. The category words corresponding to these word frequency indicators actually represent those categories to be expanded with the highest degree of association with the newly added data set. Subsequently, the corresponding categories to be expanded in the agricultural knowledge graph will be found according to these category words, and the newly added data set will be sequentially added to these screened categories to be expanded with a relatively high degree of association with the newly added data set.
[0046] Assume that "rice cultivation" (word frequency index 0.03) and "pest control" (word frequency index 0.01) have been determined as the word frequency indexes to be processed, and the second threshold is set to 0.015. If the word frequency index of "pest control" is less than the second threshold, then in step S61, the word frequency sequence will be traced back to obtain all the word frequency indexes before the word frequency index of "rice cultivation", which is the word frequency index to be processed. In this example, since "rice cultivation" is the first word frequency index in the word frequency sequence, the only word frequency index obtained is "rice cultivation". Next, according to the category word "rice cultivation", the corresponding "rice cultivation technology" category to be expanded will be found in the agricultural knowledge graph, and the new dataset will be added to this category to be expanded with a relatively high degree of association with the new dataset. In this way, it can be ensured that the new dataset is accurately integrated into the relevant categories in the knowledge graph, avoiding the introduction of redundant and irrelevant data, and improving the accuracy and efficiency of data expansion.
[0047] As Figure 2 shown, in one embodiment, obtaining the expansion strategy according to the judgment result in S6 further includes:
[0048] S62. If the word frequency index to be processed is not less than the second threshold, then the second decline degree index greater than the first threshold will be screened out in the order of the word frequency sequence, and the adjacent word frequency indexes corresponding to this decline degree index will be used as two new word frequency indexes to be processed, and it will continue to be judged whether the new word frequency indexes to be processed are less than the second threshold.
[0049] In this embodiment, it should be noted that in S62, when it is judged that the word frequency index to be processed is not less than the preset second threshold, it will be reviewed one by one according to the word frequency sequence (that is, the arrangement order of each keyword from high to low or from low to high in terms of appearance frequency) to find the second index whose decline degree exceeds the first threshold. Once such a decline degree index is found, this index and its immediately adjacent next word frequency index (that is, the word frequency after the decline) will be jointly marked as new word frequency indexes to be processed. Subsequently, it will be evaluated again whether these two new word frequency indexes to be processed are still not less than the second threshold. If so, the above logic will continue to be iteratively screened until a pair of word frequency indexes that meet the conditions is found, or the word frequency sequence is completely traversed. In this way, it can be ensured that the new dataset is integrated into more relevant categories in the knowledge graph that meet the requirements, and even if there are large fluctuations in adjacent word frequency indexes, it will not cause the omission of the expansion of the new dataset.
[0050] Taking actual operation as an example, assume that in a text analysis task, the word frequency sequence is [50, 45, 15, 10, 8], the first threshold is set to 20, and the second threshold is set to 10. Initially, the word frequency index to be processed is 50 (assuming it is the first value in the sequence). Since 50 is much greater than the second threshold 10, it enters step S62. First, it is found that the degree of decrease from 45 to 15 (30) exceeds the first threshold 20, so 15 and the subsequent 10 are used as the new word frequency indices to be processed. Then, it is judged whether both of these two new indices (15 and 10) are not less than the second threshold 10. It is found that both conditions are met. However, since the degree of decrease from 10 to 8 (2) does not exceed the first threshold, the iteration stops. At this time, 50, 45, 15, and 10 will be selected as the basis to find the category words and categories to be expanded corresponding to these word frequency indices, which not only reflect a significant decrease in word frequency but also meet the condition of not being lower than the second threshold.
[0051] In one embodiment, in S2, obtaining the word frequency index corresponding to each category word based on the word frequency calculation model and the new dataset is expressed as:
[0052] ; where
[0053] is the word frequency index corresponding to the i-th category word, is the number of occurrences of the i-th category word in the text data of the new dataset, is the total number of words in the text data of the new dataset.
[0054] In this embodiment, it should be noted that is used to quantify the number of occurrences of each category word in the new dataset. The expression defines how to calculate the word frequency index of the i-th category word. Among them, is the word frequency index of the i-th category word, which reflects the relative frequency of this word in the new dataset; is the actual number of occurrences of the i-th category word in the text data of the new dataset. It is an absolute value, indicating the specific occurrence of this word in the dataset; N is the total number of words in the text data of the new dataset, that is, all the words in the dataset, which is a denominator for normalization to ensure the comparability of word frequency indices between datasets of different scales.
[0055] Suppose the new dataset contains 1000 words, and the word "machine learning" appears 50 times, and "artificial intelligence" appears 30 times. For "machine learning", = 0.05; for "artificial intelligence", = 0.03.
[0056] That is to say, in the newly added dataset, the word frequency index of "machine learning" is 0.05, while the word frequency index of "artificial intelligence" is 0.03.
[0057] As Figure 3 shown, in one embodiment, obtaining the decline degree index of adjacent word frequency indexes in the word frequency sequence based on the fluctuation calculation model and the word frequency sequence in S3 includes:
[0058] S31. Obtain the i-th word frequency index and the (i + 1)-th word frequency index in the word frequency sequence;
[0059] S32. Obtain the decline value according to the i-th word frequency index and the (i + 1)-th word frequency index;
[0060] S33. Obtain the decline degree index according to the decline value and the i-th word frequency index.
[0061] In this embodiment, it should be noted that in S31, first, two adjacent word frequency indexes need to be obtained in sequence from the word frequency sequence that has been sorted according to the word frequency size, that is, the i-th word frequency index and the (i + 1)-th word frequency index. Here, i is a variable that starts from 1 and increases one by one until the entire word frequency sequence is traversed. During each iteration, the current two adjacent word frequency indexes will be recorded to prepare for calculating the decline degree index in the next step. If the word frequency sequence is [0.03, 0.01, 0.005], then in the first iteration, the i-th word frequency index is 0.03, and the (i + 1)-th word frequency index is 0.01.
[0062] In S32, the decline value between the i-th word frequency index and the (i + 1)-th word frequency index obtained in step S31 will be calculated. The calculation method of the decline value is to directly subtract the (i + 1)-th word frequency index from the i-th word frequency index, that is, decline value = the i-th word frequency index - the (i + 1)-th word frequency index. This decline value reflects the change amplitude of the word frequency index from the i-th category word to the (i + 1)-th category word, that is, the correlation degree difference between the newly added dataset and these two category words. Continuing with the above example, the decline value is 0.03 - 0.01 = 0.02.
[0063] In S33, the degree-of-descent indicator is jointly determined based on the descent value calculated in step S32 and the i-th word frequency indicator. The calculation method of the degree-of-descent indicator can be diverse, but the core idea is to compare or normalize the descent value with the i-th word frequency indicator in some form, so as to obtain an indicator that can reflect the magnitude of the degree of descent relative to the i-th word frequency indicator. This is done to make the degree-of-descent indicator comparable among different word frequency sequences. A possible calculation method is to divide the descent value by the i-th word frequency indicator, that is, degree-of-descent indicator = descent value / the i-th word frequency indicator. In this way, if the descent value is relatively large compared to the i-th word frequency indicator, then the degree-of-descent indicator will also be large, and vice versa. Continuing with the above example, if this calculation method is adopted, then the degree-of-descent indicator is 0.02 / 0.03 ≈ 0.67, indicating that the degree of decline in the word frequency indicator from "rice planting" to "pest control" is approximately 67%.
[0064] In one implementation, obtaining the degree-of-descent indicator based on the descent value and the i-th word frequency indicator in S33 includes:
[0065] ; where
[0066] is the degree-of-descent indicator between the i-th word frequency indicator and the (i + 1)-th word frequency indicator, is the i-th word frequency indicator, is the (i + 1)-th word frequency indicator.
[0067] In this implementation, it should be noted that this expression quantifies the degree of decline by calculating the difference between adjacent two word frequency indicator values, that is, the descent value, and dividing it by the previous word frequency indicator. The result is a value between 0 and 1 (TF(i) > TF(i + 1), because the entire word frequency sequence sorts multiple word frequency indicators in descending order), representing the relative decline ratio from the i-th word frequency indicator to the (i + 1)-th word frequency indicator.
[0068] As Figure 4 shown, in one implementation, the modification of the position where the new dataset is located according to the word frequency indicator size of the new dataset under the category to be expanded in S4 includes:
[0069] S41. Obtain the association relationship between the position range of the current category to be expanded and the word frequency indicator based on the agricultural knowledge graph;
[0070] S42. Obtain the position range to be corrected of the new dataset in the current category to be expanded according to the association relationship and the word frequency indicator of the new dataset, and correct the position where the new dataset is located according to the position range to be corrected.
[0071] In this embodiment, it should be noted that in S41, it is necessary to obtain the correlation between the position range and the word frequency index of the currently to-be-expanded category based on the agricultural knowledge graph. The core of this step lies in understanding the organizational structure of each category in the knowledge graph and their relative importance. The agricultural knowledge graph is a hierarchical or network structure that contains a large amount of knowledge in the agricultural field. These knowledges are organized into different categories, and each category has its specific position range in the graph. By analyzing the structure and existing data of the knowledge graph, a mapping relationship between the category position range and the word frequency index can be established. This relationship may be expressed as a function or rule that describes how the word frequency index affects the position of the category in the knowledge graph. For example, categories with higher word frequency indexes may be placed in more core or prominent positions in the graph, while categories with lower word frequency indexes may be located at the edge or less important positions in the graph. Through such a correlation, the fit between the new dataset and the existing knowledge graph can be understood more accurately, providing a basis for subsequent position correction.
[0072] In S42, according to the correlation established in the previous step and the word frequency index of the new dataset, obtain the to-be-corrected position range of the new dataset in the currently to-be-expanded category. Specifically, the word frequency index of the new dataset will be substituted into the previously established correlation, and the position range that the new dataset should be in the knowledge graph will be obtained through calculation. This position range may be a specific coordinate range or a relative position description, such as "under a subcategory of a certain category" or "adjacent to a certain category". After obtaining the to-be-corrected position range, the actual position of the new dataset in the knowledge graph will be corrected according to this position range. This may include moving the new dataset to a more appropriate position, adjusting its relationship with surrounding categories, or adding appropriate labels and descriptions to it. Through such correction, it can be ensured that the new dataset is accurately integrated into the knowledge graph, forming an effective association and complementarity with the existing knowledge, thereby improving the integrity and accuracy of the knowledge graph.
[0073] An agricultural knowledge graph data expansion and correction system is also provided. The system includes:
[0074] An acquisition module, configured to obtain multiple different to-be-expanded categories based on the agricultural knowledge graph, extract category words corresponding to the to-be-expanded categories, and obtain a new dataset, where the new dataset includes text data;
[0075] A calculation module, configured to obtain the word frequency index corresponding to each category word based on the word frequency calculation model and the new dataset, sort multiple word frequency indexes in descending order, and obtain a word frequency sequence;
[0076] A calculation and judgment module, configured to obtain a decline degree index of adjacent word frequency indexes in a word frequency sequence based on a fluctuation calculation model and the word frequency sequence, and judge whether there is a decline degree index greater than a first threshold;
[0077] A first judgment execution module, configured to, when there is no decline degree index greater than the first threshold, sequentially add the new data set to multiple different categories to be expanded, and correct the position where the new data set is located according to the word frequency index size of the new data set under the category to be expanded;
[0078] A second judgment execution module, configured to, when there is a decline degree index greater than the first threshold, screen out the first decline degree index greater than the first threshold in the order of the word frequency sequence and use the adjacent word frequency index corresponding to the decline degree index as two word frequency indexes to be processed;
[0079] A third judgment execution module, configured to judge whether the word frequency index to be processed is less than a second threshold, obtain an expansion strategy according to the judgment result, and correct the position where the new data set is located according to the word frequency index size of the new data set under the category to be expanded.
[0080] In one embodiment, the third judgment execution module is further configured to: if the word frequency index to be processed is less than the second threshold, obtain all the word frequency indexes before the word frequency index to be processed in the word frequency sequence, find the corresponding category to be expanded according to the category word corresponding to the obtained word frequency index, and sequentially add the new data set to the obtained category to be expanded.
[0081] In one embodiment, the third judgment execution module is further configured to: if the word frequency index to be processed is not less than the second threshold, screen out the second decline degree index greater than the first threshold in the order of the word frequency sequence and use the adjacent word frequency index corresponding to the decline degree index as two new word frequency indexes to be processed, and continue to judge whether the new word frequency index to be processed is less than the second threshold.
[0082] In this embodiment, it should be noted that regarding the above-mentioned agricultural knowledge graph data expansion and correction system, the specific manner of performing operations has been described in detail in the embodiments of the agricultural knowledge graph data expansion and correction method, and will not be elaborated here.
[0083] The preferred embodiments of the present disclosure have been described in detail above with reference to the accompanying drawings. However, the present disclosure is not limited to the specific details in the above embodiments. Within the technical concept scope of the present disclosure, various simple modifications can be made to the technical solutions of the present disclosure, and these simple modifications all fall within the protection scope of the present disclosure.
[0084] In addition, it should be noted that, in the above specific embodiments, the various specific technical features described can be combined in any appropriate manner without contradiction. To avoid unnecessary repetition, the present disclosure will not separately describe various possible combination manners.
[0085] Furthermore, any combinations can be made among the various different embodiments of the present disclosure as long as they do not violate the idea of the present disclosure, and the same should be regarded as the content disclosed by the present disclosure.
[0086] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the various embodiments of the present invention, and they should all be covered by the scope of the claims and the description of the present invention.
Claims
1. A method for expanding and correcting agricultural knowledge graph data, characterized in that: include: Based on the agricultural knowledge graph, a plurality of different categories to be expanded are obtained, and category words corresponding to the categories to be expanded are extracted, and a new data set is obtained, wherein the new data set includes text data; Based on the word frequency calculation model and the newly added data set, the word frequency index corresponding to each category word is obtained, and multiple word frequency indexes are sorted in descending order to obtain the word frequency sequence; Among them, the word frequency index corresponding to each category word is obtained based on the word frequency calculation model and the newly added data set as follows: ;in, is the word frequency index corresponding to the i-th category word, is the number of occurrences of the i-th category word in the text data of the newly added dataset, The total number of words in the text data of the newly added dataset; Obtaining a decrease degree index of adjacent word frequency indexes in the word frequency sequence based on the fluctuation calculation model and the word frequency sequence, and determining whether there is a decrease degree index greater than a first threshold; Wherein, the step of obtaining the decline degree indicator of adjacent word frequency indicators in the word frequency sequence based on the fluctuation calculation model and the word frequency sequence includes: obtaining the i-th word frequency indicator and the i+1-th word frequency indicator in the word frequency sequence; obtaining the decline value according to the i-th word frequency indicator and the i+1-th word frequency indicator; obtaining the decline degree indicator according to the decline value and the i-th word frequency indicator; The step of obtaining the decline degree index according to the decline value and the i-th word frequency index includes: ;in, is the decrease degree index of the i-th word frequency index and the i+1-th word frequency index, is the i-th word frequency index, is the i+1th word frequency index; If there is no decrease degree index greater than the first threshold, the newly added data set is sequentially added to a plurality of different categories to be expanded, and the position of the newly added data set is corrected according to the word frequency index size of the newly added data set under the category to be expanded; If there is a decrease degree index greater than the first threshold, the first decrease degree index greater than the first threshold is selected in the order of the word frequency sequence and the adjacent word frequency indexes corresponding to the decrease degree index are used as two word frequency indexes to be processed; It is determined whether the word frequency index to be processed is less than the second threshold, and an expansion strategy is obtained according to the determination result, and the position of the newly added data set is corrected according to the word frequency index size of the newly added data set under the category to be expanded.
2. The agricultural knowledge graph data expansion and correction method according to claim 1, characterized in that: The obtaining of the expansion strategy according to the judgment result includes: If the word frequency index to be processed is less than the second threshold, all word frequency indexes before the word frequency index to be processed in the word frequency sequence are obtained, and the corresponding category to be expanded is found according to the category words corresponding to the obtained word frequency index, and the newly added data sets are added to the obtained category to be expanded in sequence.
3. The agricultural knowledge graph data expansion and correction method according to claim 2, characterized in that: The obtaining of the expansion strategy according to the judgment result further includes: If the word frequency index to be processed is not less than the second threshold, the second decrease degree index greater than the first threshold is selected in the order of the word frequency sequence and the adjacent word frequency indicators corresponding to the decrease degree index are used as two new word frequency indicators to be processed, and the new word frequency indicators to be processed are further judged whether they are less than the second threshold.
4. The agricultural knowledge graph data expansion and correction method according to claim 1, characterized in that: The step of correcting the position of the newly added data set according to the word frequency index size of the newly added data set under the category to be expanded includes: Based on the agricultural knowledge graph, the correlation between the location range of the current category to be expanded and the word frequency index is obtained; The position range of the newly added data set to be corrected in the current category to be expanded is obtained according to the association relationship and the word frequency index of the newly added data set, and the position of the newly added data set is corrected according to the position range to be corrected.
5. An agricultural knowledge graph data expansion and correction system, characterized in that: The system comprises: An acquisition module, used to acquire a plurality of different categories to be expanded based on the agricultural knowledge graph and extract the category words corresponding to the categories to be expanded, and acquire a newly added data set, wherein the newly added data set includes text data; A calculation module is used to obtain the word frequency index corresponding to each category word based on the word frequency calculation model and the newly added data set, and sort the multiple word frequency indexes in order from large to small to obtain the word frequency sequence; Among them, the word frequency index corresponding to each category word is obtained based on the word frequency calculation model and the newly added data set as follows: ;in, is the word frequency index corresponding to the i-th category word, is the number of occurrences of the i-th category word in the text data of the newly added dataset, The total number of words in the text data of the newly added dataset; A calculation and judgment module, used to obtain a decrease degree index of adjacent word frequency indexes in a word frequency sequence based on a fluctuation calculation model and a word frequency sequence, and to judge whether there is a decrease degree index greater than a first threshold; Wherein, the step of obtaining the decline degree indicator of adjacent word frequency indicators in the word frequency sequence based on the fluctuation calculation model and the word frequency sequence includes: obtaining the i-th word frequency indicator and the i+1-th word frequency indicator in the word frequency sequence; obtaining the decline value according to the i-th word frequency indicator and the i+1-th word frequency indicator; obtaining the decline degree indicator according to the decline value and the i-th word frequency indicator; The step of obtaining the decline degree index according to the decline value and the i-th word frequency index includes: ;in, is the decrease degree index of the i-th word frequency index and the i+1-th word frequency index, is the i-th word frequency index, is the i+1th word frequency index; A first judgment execution module is used to add the newly added data set to a plurality of different categories to be expanded in sequence when there is no decrease degree index greater than the first threshold, and to correct the position of the newly added data set according to the word frequency index size of the newly added data set under the category to be expanded; The second judgment execution module is used for, when there is a decrease degree indicator greater than the first threshold, screening out the first decrease degree indicator greater than the first threshold in the order of the word frequency sequence and taking the adjacent word frequency indicators corresponding to the decrease degree indicator as two word frequency indicators to be processed; The third judgment execution module is used to judge whether the word frequency index to be processed is less than the second threshold, and obtain the expansion strategy according to the judgment result, and correct the position of the newly added data set according to the word frequency index size of the newly added data set under the category to be expanded.
6. The agricultural knowledge graph data expansion and correction system according to claim 5, characterized in that: The third judgment execution module is also used for: If the word frequency index to be processed is less than the second threshold, all word frequency indexes before the word frequency index to be processed in the word frequency sequence are obtained, and the corresponding category to be expanded is found according to the category words corresponding to the obtained word frequency index, and the newly added data sets are added to the obtained category to be expanded in sequence.
7. The agricultural knowledge graph data expansion and correction system according to claim 6, characterized in that: The third judgment execution module is also used for: If the word frequency index to be processed is not less than the second threshold, the second decrease degree index greater than the first threshold is selected in the order of the word frequency sequence and the adjacent word frequency indicators corresponding to the decrease degree index are used as two new word frequency indicators to be processed, and the new word frequency indicators to be processed are further judged whether they are less than the second threshold.
Citation Information
Patent Citations
Generation method, device and equipment of extended query word and storage medium
CN112925967A
Merchant text recognition method and device, equipment and storage medium
CN115618871A