A method and system for constructing a livestock and poultry germplasm resource database
By constructing a feature association diagram of the livestock and poultry germplasm resource database and optimizing the retrieval priority, the problem of low retrieval efficiency in the existing technology is solved, and a more efficient and accurate retrieval effect is achieved.
Patent Information
- Application Number
- CN202510790285.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2045-06-13
AI Technical Summary
The existing livestock and poultry germplasm resource database has too long a response time when searching based on a combination of multiple feature conditions, which affects the search efficiency and accuracy. The existing technology only adjusts the priority based on the search frequency and fails to effectively consider the intrinsic connection between population characteristics.
By constructing an initial feature association graph, analyzing the correlation and hierarchical deviation of population features, combining retrieval matching degree and comprehensive retrieval cost, optimizing feature combinations, establishing an optimized feature association graph, and using database federation technology to optimize retrieval priority.
It improves the retrieval efficiency of the livestock and poultry germplasm resources database, enhances the correlation between population characteristics at different levels, reduces the physical delay of full table scanning, and improves the accuracy and speed of retrieval.
Smart Images

Figure CN120319328B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of database construction, and in particular to a method and system for constructing a livestock and poultry germplasm resource database. Background Art
[0002] The livestock and poultry germplasm resource database is a systematic database of livestock and poultry species and their genetic resources. It is mainly used to aggregate and manage information on the varieties, subspecies, local varieties, genetic resources, and pedigrees of various livestock and poultry germplasms. It can help the livestock industry improve breeding efficiency and cultivate new varieties that better meet market demand. The livestock and poultry germplasm resource database generally contains multiple main databases and sub-databases, aiming to comprehensively record, manage, and analyze various data on different livestock and poultry germplasm resources. Currently, with the continuous increase in the amount of livestock and poultry germplasm resource data, the germplasm information contained in the database is becoming more and more abundant, and the different sub-databases are independent of each other and only contain the corresponding livestock and poultry feature data. Therefore, with such a large amount of data, if the user performs a combination search with multiple feature conditions, the system response time will be too long, affecting the search efficiency and accuracy.
[0003] In order to improve retrieval efficiency, existing technologies usually adjust retrieval priorities based on the frequency of occurrence of each search formula in the user's search history; however, the process of adjusting retrieval priorities based solely on the frequency of occurrence of the search formula does not take into account the intrinsic connection between population characteristics in the search formula. When the livestock and poultry germplasm resource database is more complex, its priority adjustment effect is more limited, thereby reducing retrieval efficiency. Summary of the Invention
[0004] In order to solve the technical problem that the prior art process of adjusting the search priority based solely on the frequency of occurrence of the search formula leads to low search efficiency in the livestock and poultry germplasm resource database, the purpose of this application is to provide a method and system for constructing a livestock and poultry germplasm resource database. The technical solutions adopted are as follows:
[0005] The first aspect of the present application provides a method for constructing a livestock and poultry germplasm resource database, comprising:
[0006] Acquire an initial livestock and poultry germplasm resource database of a multi-level classification storage architecture; obtain search results obtained by searching a search formula for each population characteristic combination in historical data; wherein the initial livestock and poultry germplasm resource database includes various population characteristics;
[0007] In the initial livestock and poultry germplasm resource database, the corresponding relevance is determined based on the correlation between each population characteristic and each other population characteristic and the hierarchical deviation; an initial characteristic correlation map is established based on the relevance between all population characteristics;
[0008] Determine the corresponding search matching degree based on the population feature matching between the search formula corresponding to each population feature combination and the corresponding search results; determine the corresponding comprehensive search cost based on the search frequency and search response of each search formula containing each population feature combination; determine the feature correlation degree of each population feature combination based on the search matching degree and the comprehensive search cost; and determine the optimized feature combination based on the feature correlation degree;
[0009] The initial feature association graph is optimized according to the population characteristics in the optimized feature combination to obtain an optimized feature association graph; and the retrieval priority of the initial livestock and poultry germplasm resource database is adjusted according to the optimized feature association graph.
[0010] Furthermore, the process of obtaining the relatability includes:
[0011] Each population feature is used as the target feature in turn; each population feature other than the target feature is used as the corresponding comparison feature;
[0012] Calculate the mutual information value between the target feature and the contrast feature; use the hierarchical path length of the target feature and the contrast feature in the initial livestock and poultry germplasm resources database as the corresponding inter-layer distance; normalize the product between the negative correlation mapping value of the inter-layer distance and the mutual information value to determine the relevance between the target feature and the contrast feature.
[0013] Furthermore, the process of obtaining the initial feature association graph based on the relatability among all population features includes:
[0014] Two population features whose corresponding associatability is greater than a preset association threshold are used as association features; an initial feature association graph is established based on the association features in all population features in a directed graph form.
[0015] Furthermore, the process of obtaining the search matching degree includes:
[0016] The number of population features corresponding to the intersection of all population features in the search formula corresponding to each population feature combination and all population features in the corresponding search results is positively correlated to determine the category matching degree; the negative correlation mapping value of the difference between the total number of population features in the search formula corresponding to each population feature combination and the total number of population features in the corresponding search results is used as the quantity matching degree;
[0017] The retrieval matching degree of each population feature combination is determined according to the product of the category matching degree and the quantity matching degree.
[0018] Furthermore, the process of obtaining the comprehensive search cost includes:
[0019] In a sampling time period, all search moments of all search formulas containing each population feature combination are arranged in chronological order to obtain a search moment time series sequence; in the search moment time series sequence, the time interval between each search moment and the previous search moment is used as the search time series interval of each search moment; and the search frequency of each population feature combination is determined based on the negative correlation mapping value of the mean of the search time series intervals of all search moments;
[0020] The ratio between the number of search moments corresponding to all search formulas containing each population feature combination in the sampling time period and the total number of search moments is used as the search demand degree of each population feature combination;
[0021] The average response time of all search expressions containing each population feature combination during the sampling period is used as the time cost;
[0022] The comprehensive search cost of each population feature combination is determined according to the product of the search frequency, the search demand and the time cost.
[0023] Furthermore, the process of obtaining the feature correlation degree includes:
[0024] The ratio between the search matching degree and the comprehensive search cost is normalized to determine the feature correlation degree of each population feature combination.
[0025] Furthermore, the process of obtaining the optimized feature combination includes:
[0026] The population feature combination with a feature correlation greater than a preset high correlation threshold is used as the optimized feature combination.
[0027] Furthermore, the process of obtaining the optimized feature association graph includes:
[0028] In the optimized feature combination, each population feature and other population features are used as associated features; the initial feature association graph is optimized in the form of a directed graph based on the associated features of all population features of each optimized feature combination to obtain an optimized feature association graph.
[0029] Furthermore, the process of adjusting the search priority of the initial livestock and poultry germplasm resources database according to the optimized feature association graph includes:
[0030] The database federation technology is used to map the physical table of the associated features in the optimized feature association graph into a unified logical view; when the user searches the initial livestock and poultry germplasm resource database, the search request is preferentially sent to the unified logical view for retrieval.
[0031] In a second aspect, the present application provides a system for constructing a livestock and poultry germplasm resource database, the system comprising:
[0032] The data acquisition and preprocessing module is used to obtain an initial livestock and poultry germplasm resource database with a multi-level classification storage architecture; obtain search results obtained by searching the historical data for each population characteristic combination using a search formula; wherein the initial livestock and poultry germplasm resource database includes various population characteristics;
[0033] An initial feature association graph construction module is used to determine the corresponding relevance based on the association between each population feature and each other population feature and the hierarchical deviation in the initial livestock and poultry germplasm resource database; and to establish an initial feature association graph based on the relevance between all population features;
[0034] An optimized feature combination determination module is used to determine the corresponding search matching degree based on the population feature matching between the search formula corresponding to each population feature combination and the corresponding search result; determine the corresponding comprehensive search cost based on the search frequency and search response of each search formula containing each population feature combination; determine the feature correlation degree of each population feature combination based on the search matching degree and the comprehensive search cost; and determine the optimized feature combination based on the feature correlation degree;
[0035] A retrieval priority adjustment module is used to optimize the initial feature association diagram according to the population characteristics in the optimized feature combination to obtain an optimized feature association diagram; and adjust the retrieval priority of the initial livestock and poultry germplasm resource database according to the optimized feature association diagram.
[0036] In a third aspect, the present application provides a computer device comprising a memory and a processor. The memory is configured to store computer program code, and the processor is configured to call and execute the computer program code from the memory to perform the method of the first aspect or any embodiment of the first aspect of the present application.
[0037] In a fourth aspect, the present application provides a computer program product, comprising a computer program code. When the computer program code is executed, the method of the first aspect or any embodiment of the first aspect of the present application is performed.
[0038] In a fifth aspect, the present application provides a computer-readable storage medium, which stores computer program code. When the computer program code is executed, it performs the method of the first aspect of the present application or any embodiment of the first aspect.
[0039] This application has the following beneficial effects:
[0040] First, an initial database is constructed based on the characteristic classification standards of livestock and poultry resources. Corresponding feature association rules, or the initial feature association graph, are established by analyzing the correlations between different livestock and poultry characteristics in the database. The difference between the user's search terms and the features contained in the search results is then continuously analyzed to determine the degree of feature association. The initial feature association graph is then dynamically adjusted based on the optimized feature combination obtained based on the feature association degree. This optimized feature association graph optimizes the search priority of the initial livestock and poultry germplasm resource database during the search process, strengthening the feature associations between population characteristics at different levels and improving search efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] In order to more clearly illustrate the technical solutions and advantages of the embodiments of the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the prior art descriptions. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0042] Figure 1 A flow chart of a method for constructing a livestock and poultry germplasm resource database provided by one embodiment of the present invention;
[0043] Figure 2 A schematic diagram of an initial feature association graph provided by one embodiment of the present invention;
[0044] Figure 3 A structural diagram of a livestock and poultry germplasm resource database construction system provided by one embodiment of the present invention;
[0045] Figure 4 The present invention provides a schematic diagram of a computer device structure according to an embodiment of the present invention. DETAILED DESCRIPTION
[0046] In order to further illustrate the technical means and effects adopted by the present invention to achieve the predetermined purpose of the invention, the following, in combination with the accompanying drawings and preferred embodiments, a method and system for constructing a livestock and poultry germplasm resource database proposed by the present invention, its specific implementation method, structure, characteristics and effects are described in detail as follows. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment, and the specific features, structures or characteristics in one or more embodiments may be combined in any suitable form. In addition, the terms "first" and "second" are used for descriptive purposes only and cannot be understood as implying or suggesting relative importance or implicitly indicating the number of technical features indicated. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features.
[0047] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.
[0048] The following describes in detail a method and system for constructing a livestock and poultry germplasm resource database provided by the present invention with reference to the accompanying drawings.
[0049] This application embodiment provides a method for constructing a livestock and poultry germplasm resource database. Figure 1 , which shows a flow chart of a method for constructing a livestock and poultry germplasm resource database provided by one embodiment of the present invention, the method comprising:
[0050] Step S101: obtaining an initial livestock and poultry germplasm resource database of a multi-level classification storage architecture; obtaining search results obtained by searching a search formula for each population feature combination in historical data; wherein the initial livestock and poultry germplasm resource database includes various population features.
[0051] General livestock and poultry germplasm resources include basic information, genetic information, breeding history, adaptability, production performance, and conservation measures for different types of livestock and poultry (such as cattle, sheep, pigs, chickens, ducks, and geese), as well as their breeds, subspecies, local breeds, and populations. Different data represent different aspects of livestock and poultry characteristics, so when constructing the initial database, multiple sub-databases need to be set up for separate storage, while also being divided into multiple main databases based on characteristic classification.
[0052] In one specific implementation of the present invention, the initial livestock and poultry germplasm resource database includes four main databases: a breed resource database, a genomic database, a genetic resource database, and a production performance database. The breed resource database is divided into population sub-databases based on population, specifically for different livestock and poultry populations, such as chickens, ducks, cattle, pigs, and geese. Each population sub-database contains population characteristics such as breed name, subspecies classification, local breed identifier, population distribution area, coat color, reproductive cycle, and age of sexual maturity. The genomic database directly includes population characteristics such as genome sequence, genetic markers, and genotype data. The genetic resource database is divided into genetic material sub-databases, specifically for different genetic materials, such as DNA, somatic cell, sperm, and tissue sample databases. Each genetic material sub-database is further divided into population sub-databases, specifically for different livestock and poultry populations, such as chickens, ducks, cattle, pigs, and geese. Each population sub-database contains population characteristics such as sample number, collection time, storage conditions, usage history, and extraction concentration. The final production performance database is first divided into sub-population databases based on the population, specifically for different livestock and poultry populations such as chickens, ducks, cattle, pigs, and geese. Each sub-population database includes population characteristics such as daily weight gain, feed conversion rate, slaughter rate, litter size (egg production), conception rate, and survival rate. It should be noted that the construction of sub-populations and the types of population characteristics in the initial livestock and poultry germplasm resource database can be adjusted according to the needs of the specific implementation environment and will not be further explained here.
[0053] In a specific implementation of an embodiment of the present invention, the sampling time range of historical data is set to within one month before the current moment, that is, the range of the sampling time period; and the population feature combination collected by the embodiment of the present invention is obtained through the type of search formula retrieved by the user. When the user searches, he usually performs a combined search with at least two population feature nouns as units. Each search formula usually corresponds to a different type of population feature. Therefore, this application uses the combination of population features corresponding to each search formula as the population feature combination; the search result is the result obtained when searching the initial livestock and poultry germplasm resource database with the search formula, and its content is also the combination of population features or a single population feature. It should be noted that there is an inclusion relationship between the search formulas. For example, the population feature combination corresponding to the included search formula also exists in the included search formula.
[0054] Step S102: In the initial livestock and poultry germplasm resource database, the corresponding relevance is determined based on the correlation between each population feature and each other population feature and the hierarchical deviation; and an initial feature association graph is established based on the relevance between all population features.
[0055] For the initial livestock and poultry germplasm database, the initial database construction process categorizes and stores the characteristics of different populations through multiple, independent sub-databases. Therefore, when users perform a combination search using multiple characteristics, the database system must perform a full table scan, increasing response time. Therefore, this step establishes corresponding feature association rules, or feature association graphs, based on the correlations between different livestock and poultry characteristics to improve retrieval efficiency; therefore, it is necessary to analyze the correlations between population characteristics.
[0056] Preferably, in some possible implementations of the embodiments of the present invention, the process of obtaining the associatability includes:
[0057] Each population characteristic is sequentially used as a target characteristic; each population characteristic other than the target characteristic is used as a corresponding comparison characteristic; the mutual information value between the target characteristic and the comparison characteristic is calculated; and the hierarchical path length between the target characteristic and the comparison characteristic in the initial livestock and poultry germplasm resource database is used as the corresponding inter-layer distance. It should be noted that the calculation of the mutual information value is a technical means well known to those skilled in the art and will not be further elaborated here. The mutual information value represents the degree of dependence between two population characteristics. The larger the corresponding mutual information value, the more correlated the two population characteristics are, that is, the mutual information value and the correlation are positively correlated. The larger the inter-layer distance, the more libraries the two population characteristics span, and the corresponding correlation relationship is generally greater, that is, the inter-layer distance and the correlation are negatively correlated. Therefore, the product of the negative correlation mapping value of the inter-layer distance and the mutual information value is further normalized to determine the correlation between the target characteristic and the comparison characteristic. Normalization makes the quantification of the correlation more intuitive, improving the accuracy and robustness of the subsequent correlation screening process.
[0058] In a specific implementation of the embodiment of the present invention, the process of obtaining the hierarchical path length includes:
[0059] A path diagram is constructed with each library (including the initial livestock and poultry germplasm resource database, the main database, and each sub-library) and each population characteristic as a node, and the inclusion relationship as a connecting line, and the length of all connecting lines is the same; on the path diagram, the number of libraries on the shortest path between each population characteristic to another population characteristic in the initial livestock and poultry germplasm resource database is used as the inter-layer distance; the statistics of the number of libraries include the main database and each sub-library; for example, if there are two population characteristics, one of which is a characteristic in sub-library B of main library A For example, if one population feature is C1 and the other is feature C2 from sub-library B of main library A, then the shortest path between the two population features is C1-B-C2. This shortest path only includes sub-library B, so the inter-layer distance between the two population features is 1. For another example, if there are two population features, one is feature C3 from sub-library B of main library A, and the other is feature D1 from sub-library E of main library A, then the shortest path between the two population features is C3-BAE-D1. This shortest path includes libraries B, A, and E, so the inter-layer distance between the two population features is 3. It should be noted that constructing a path graph can more intuitively obtain the shortest path. Implementers can choose methods other than constructing a path graph to obtain the shortest path based on the specific implementation environment. This will not be further elaborated here.
[0060] In a specific implementation of the embodiment of the present invention, the process of obtaining the associatability is expressed by the formula: ;in, For population characteristics and population characteristics The relevance between For population characteristics and population characteristics The mutual information value between ; For population characteristics and population characteristics The distance between layers; is the normalization function.
[0061] Further, feature association rules are constructed based on the relatability that characterizes the association relationship, that is, an initial feature association graph is established. Preferably, in some possible implementation methods of the embodiments of the present invention, the acquisition process of establishing the initial feature association graph based on the relatability between all population features includes: taking the two population features whose corresponding relatability is greater than a preset association threshold as association features; and establishing the initial feature association graph based on the association features among all population features in the form of a directed graph. In a specific implementation method of the embodiments of the present invention, the preset association threshold is set to 0.85, which can be adjusted according to the specific implementation environment.
[0062] It should be noted that directed graph is a technical means well known to those skilled in the art, please refer to Figure 2, which shows a schematic diagram of an initial feature association graph provided by an embodiment of the present invention, Figure 2 In the figure, a, b, c, d, e, f, g, and h all represent population characteristics. Since the characteristic influence between population characteristics, that is, the correlation, is bidirectional, the arrows in the initial characteristic correlation diagram obtained by the bidirectional diagram are also bidirectional. The bidirectional arrows indicate that the two population characteristics are correlated characteristics, that is, Figure 2 Specifically, hg, ga, gc, ag, ab, bd, be, ef, af, and ac correspond to associated feature relationships.
[0063] Step S103: Determine the corresponding search matching degree based on the population feature matching between the search formula corresponding to each population feature combination and the corresponding search result; determine the corresponding comprehensive search cost based on the search frequency and search response of each search formula containing each population feature combination; determine the feature correlation degree of each population feature combination based on the search matching degree and the comprehensive search cost; and determine the optimized feature combination based on the feature correlation degree.
[0064] The above method generates initial feature association graphs for different populations. Since this initial feature association graph is based on the correlation between population features, there is significant randomness between different features when users perform multi-feature combination searches. The initial feature association graph may not fully reflect the user's search habits, resulting in a significant improvement in search efficiency. Therefore, this step analyzes the user's search habits and the feature differences between the search and results to determine the target feature combination and dynamically adjust the initial feature association graph. Before adjustment, the differences in the user's search results are first analyzed. Generally, after a user searches, the system will provide the corresponding search results. If the feature differences between the user's search results and the search formula are large or the match is low, it indicates that the correlation between the population features in the search formula is highly likely not reflected in the initial feature association graph. Therefore, the feature differences between the corresponding search formula and the search results can be analyzed as the primary factor for updating the initial feature association graph. The greater the difference between the search formula and the search results, that is, the lower the search match, the greater the likelihood that the initial feature association graph does not contain the population feature association corresponding to the search formula. Consequently, the greater the need to update the initial feature association graph based on the population feature combination reflected by the search formula. Therefore, the search match is first calculated.
[0065] Preferably, in some possible implementations of the embodiments of the present invention, the process of obtaining the retrieval matching degree includes:
[0066] The number of population features corresponding to the intersection of all population features in the search formula corresponding to each population feature combination and all population features in the corresponding search results is positively mapped to determine the type matching degree; the negative correlation mapping value of the difference between the total number of population features in the search formula corresponding to each population feature combination and the total number of population features in the corresponding search results is used as the quantity matching degree; the search matching degree of each population feature combination is determined based on the product of the type matching degree and the quantity matching degree. First, for the search formula and the corresponding search results, the more consistent the type and number of corresponding population features are, the more similar the search formula and the search results are, and the corresponding matching degree should also be higher; therefore, the type matching degree is used to represent the degree of consistency in the type of population features; the quantity matching degree is used to represent the degree of consistency in the number of population features; the search matching degree is determined by combining the type matching degree and the quantity matching degree. It should be noted that the search formula corresponding to each population feature combination is a search formula that only contains the population features in this population feature combination.
[0067] In a specific implementation of the embodiment of the present invention, the process of obtaining the search matching degree is expressed by the formula: ;in, For the The retrieval matching degree of the combination of population characteristics; For the The number of population features corresponding to the intersection of all population features in the search formula corresponding to each population feature combination and all population features in the corresponding search results; For the The total number of population features in the search formula corresponding to each population feature combination; For the The total number of population features in the search results of the search formula corresponding to each population feature combination; For the The species matching degree of the search formula corresponding to the combination of population features is positively correlated with the number of population features in the intersection through the number of population features in the search formula, so as to prevent the search formulas with different numbers of species features from interfering with the species matching degree calculation; is an exponential function with a natural constant as its base; is the absolute value symbol; For the The quantitative matching degree of the search formula corresponding to each population feature combination; the smaller the calculated search matching degree is, the greater the need to optimize or adjust the initial feature association graph through the corresponding population feature combination.
[0068] The calculation of the search matching degree only considers whether the search results are accurate. However, for the population feature combination itself, the greater the cost of the corresponding population feature combination during the search, the more important the feature combination, i.e., the search formula, is to the user. Then, the correlation of the population features needs to be displayed on the initial feature correlation diagram. Therefore, further analysis related to the search cost is performed.
[0069] Preferably, in some possible implementations of the embodiments of the present invention, the process of obtaining the search cost includes:
[0070] In the sampling time period, all the search moments of all the search formulas containing each population feature combination are arranged in chronological order to obtain a search moment time series sequence; in the search moment time series sequence, the time interval between each search moment and the previous search moment is used as the search time series interval of each search moment; the search frequency of each population feature combination is determined according to the negative correlation mapping value of the mean of the search time series intervals of all the search moments; the ratio between the number of search moments corresponding to all the search formulas containing each population feature combination in the sampling time period and the total number of search moments is used as the search demand of each population feature combination; the mean of the response time of all the search formulas containing each population feature combination in the sampling time period when performing the search is used as the time cost.
[0071] First, if the search intervals for a search formula containing the corresponding population feature combination are smaller during the sampling period, and if the search formula containing this population feature combination appears more frequently during the entire sampling period, it indicates that users are more likely to adjust the search formula based on the corresponding population feature combination and perform multiple searches, and the importance of the corresponding population feature combination is higher. In other words, the greater the search frequency and the greater the search demand, the greater the comprehensive search cost representing the importance of the corresponding population feature combination. Furthermore, when users adjust the search formula, the search response time changes with the complexity of the search formula. The longer the overall response time, the more complex the search formula, and the higher the cost the user incurs to search for this population feature combination, the greater the corresponding search cost. In other words, the greater the time cost, the greater the comprehensive search cost. Finally, combining the various features representing the comprehensive search cost, the comprehensive search cost for each population feature combination is determined based on the product of search frequency, search demand, and time cost.
[0072] In some possible implementations of the embodiments of the present invention, the process of obtaining the comprehensive search cost is expressed by the formula: ;in, For the The comprehensive search cost of various population feature combinations; The sampling period includes The average response time of all search formulas for various population feature combinations during retrieval, that is, the time cost; The sampling period includes The number of all search moments of all search formulas for various population feature combinations, that is, the number of search moments in the corresponding search moment time series; The total number of search moments in the sampling period, that is, the total number of user searches in the sampling period; For the The first The time interval between the first retrieval moment and the previous retrieval moment, that is, the The retrieval time interval of each retrieval moment; For the How frequently combinations of population characteristics are retrieved; For the The retrieval demand for various population feature combinations.
[0073] The retrieval matching degree, which reflects relevance, and the comprehensive retrieval cost, which reflects importance, are further combined to comprehensively characterize the feature relevance of the population features in each population feature combination, thereby improving the accuracy of adjusting the initial feature relevance graph. Preferably, in some possible implementations of the present invention, the process of obtaining the feature relevance includes normalizing the ratio between the retrieval matching degree and the comprehensive retrieval cost to determine the feature relevance of each population feature combination. Since the comprehensive detection cost represents importance, and the smaller the retrieval matching degree, the greater the need to optimize the initial feature relevance graph using the corresponding population feature combination, the retrieval matching degree after negative correlation mapping is weighted by importance, i.e., the comprehensive retrieval cost, and the required feature relevance is determined as a ratio. The greater the feature relevance, the greater the need to associate the corresponding population features based on the initial feature relevance graph. Based on this characteristic, in one specific implementation of the present invention, the process of obtaining the optimized feature combination includes selecting the population feature combination with a feature relevance greater than a preset high relevance threshold as the optimized feature combination. In one specific implementation of the present invention, the preset high relevance threshold is set to 0.78 and can be adjusted according to the specific implementation environment.
[0074] In a specific implementation of the embodiment of the present invention, the process of obtaining the feature correlation degree is expressed by the formula: ;in, For the The degree of correlation between the characteristics of the population characteristics combination; For the The comprehensive search cost of various population feature combinations; For the The retrieval matching degree of various population feature combinations; in order to prevent the denominator from being meaningless due to 0, when the retrieval matching degree is 0, the denominator in the formula is replaced by 0.01 to calculate the feature correlation degree. The replacement value can be adjusted according to the specific implementation environment and will not be further limited or elaborated here.
[0075] Step S104: Optimizing the initial feature association graph according to the population features in the optimized feature combination to obtain an optimized feature association graph; adjusting the search priority of the initial livestock and poultry germplasm resources database according to the optimized feature association graph.
[0076] After obtaining the optimized feature combination for optimizing the initial feature association graph, the optimization process is further performed. Preferably, in some possible implementations of the embodiments of the present invention, the acquisition process of the optimized feature association graph includes:
[0077] In the optimized feature combination, each population feature is used as an association feature with each other population feature. The initial feature association graph is optimized based on the association features of all population features in each optimized feature combination in the form of a directed graph, thereby obtaining an optimized feature association graph representing the new feature association rule. It should be noted that the process of combining the association features with the directed graph here is the same as the process of obtaining the initial feature association graph. That is, on the basis of the original initial feature association graph, the newly added association features are combined through bidirectional arrows. Further explanation is not given here.
[0078] Finally, the search priority of the initial livestock and poultry germplasm resource database is adjusted based on the optimized feature association graph. Specifically, the database federation technology is used to map the physical table of associated features in the optimized feature association graph to a unified logical view; when the user searches the initial livestock and poultry germplasm resource database, the search request is preferentially sent to the unified logical view for retrieval. It should be noted that database federation technology is a technical means well known to those skilled in the art and will not be further defined or elaborated here. At this point, by giving priority to searching the optimized feature association graph that represents the association rules, the physical delay of traversing the huge amount of data during the search is reduced, making the retrieval efficiency of the constructed livestock and poultry germplasm resource database higher.
[0079] In summary, a method for constructing a livestock and poultry germplasm resource database first constructs an initial database based on the characteristic classification standards of livestock and poultry resources. Corresponding feature association rules, namely an initial feature association graph, are established by analyzing the correlations between different livestock and poultry features in the database. The difference between the features included in the user's search formula and the search results is then continuously analyzed to determine the feature association degree. The initial feature association graph is then dynamically adjusted based on the optimized feature combination obtained from the feature association degree. This optimized feature association graph optimizes the search priority of the initial livestock and poultry germplasm resource database during the search process, strengthens the feature associations between features at different levels of populations, and improves search efficiency.
[0080] This application also provides a system for constructing a livestock and poultry germplasm resource database. Figure 3 , which shows a structural diagram of a livestock and poultry germplasm resource database construction system provided by an embodiment of the present invention. The system includes: a data acquisition and preprocessing module 301, an initial feature association graph construction module 302, an optimized feature combination determination module 303 and a retrieval priority adjustment module 304.
[0081] The data acquisition and preprocessing module 301 is used to obtain an initial livestock and poultry germplasm resource database with a multi-level classification storage architecture; obtain search results obtained by searching the historical data for each population characteristic combination using a search formula; wherein the initial livestock and poultry germplasm resource database includes various population characteristics;
[0082] The initial feature association graph construction module 302 is used to determine the corresponding relevance in the initial livestock and poultry germplasm resource database based on the correlation between each population feature and each other population feature and the hierarchical deviation; and to establish the initial feature association graph based on the relevance between all population features;
[0083] The optimized feature combination determination module 303 is configured to determine the corresponding search matching degree based on the population feature matching between the search formula corresponding to each population feature combination and the corresponding search results; determine the corresponding comprehensive search cost based on the search frequency and search response of each search formula containing each population feature combination; determine the feature correlation degree of each population feature combination based on the search matching degree and the comprehensive search cost; and determine the optimized feature combination based on the feature correlation degree;
[0084] The search priority adjustment module 304 is used to optimize the initial feature association graph according to the population characteristics in the optimized feature combination to obtain an optimized feature association graph; and adjust the search priority of the initial livestock and poultry germplasm resource database according to the optimized feature association graph.
[0085] It should be noted that the system provided in the above embodiment is merely illustrative of the division of the aforementioned functional modules. In actual applications, the aforementioned functions can be assigned to different functional modules as needed, i.e., the internal structure of the computer device can be divided into different functional modules to complete all or part of the functions described above. Furthermore, the system for constructing a livestock and poultry germplasm resource database and the method for constructing a livestock and poultry germplasm resource database provided in the above embodiment are based on the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.
[0086] The present application also provides a computer device. Figure 4 , which shows a schematic diagram of the structure of a computer device provided by an embodiment of the present invention, the computer device includes a memory 401, a processor 402, and a computer program 403 stored in the memory 401 and running on the processor 402, wherein when the processor 402 executes the computer program 403, the computer device can execute any of the above-mentioned methods for constructing a livestock and poultry germplasm resources database.
[0087] An embodiment of the present application also provides a computer program product. When the computer program product is run on a computer device, the computer device can execute any of the methods for constructing a livestock and poultry germplasm resource database described above.
[0088] An embodiment of the present application also provides a computer-readable storage medium, which stores computer program code. When the computer program code runs on a computer device, the computer device can execute any of the methods for constructing a livestock and poultry germplasm resource database introduced above.
[0089] In the embodiments provided in the present application, it should be understood that the provided computer devices, computer program products, and computer-readable storage media are all used to execute the corresponding methods provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the methods provided above and will not be repeated here.
[0090] It should be noted that the order in which the embodiments of the present invention are described above is for illustrative purposes only and does not necessarily represent the superiority or inferiority of the embodiments. The processes depicted in the accompanying drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0091] The various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments.
Claims
1. A method for constructing a livestock and poultry germplasm resource database, characterized in that: The method comprises: Acquire an initial livestock and poultry germplasm resource database of a multi-level classification storage architecture; obtain search results obtained by searching a search formula for each population characteristic combination in historical data; wherein the initial livestock and poultry germplasm resource database includes various population characteristics; In the initial livestock and poultry germplasm resource database, the corresponding relevance is determined based on the correlation between each population characteristic and each other population characteristic and the hierarchical deviation; an initial characteristic correlation map is established based on the relevance between all population characteristics; Determine the corresponding search matching degree based on the population feature matching between the search formula corresponding to each population feature combination and the corresponding search results; determine the corresponding comprehensive search cost based on the search frequency and search response of each search formula containing each population feature combination; determine the feature correlation degree of each population feature combination based on the search matching degree and the comprehensive search cost; and determine the optimized feature combination based on the feature correlation degree; The initial feature association graph is optimized according to the population characteristics in the optimized feature combination to obtain an optimized feature association graph; and the retrieval priority of the initial livestock and poultry germplasm resource database is adjusted according to the optimized feature association graph.
2. The method for constructing a livestock and poultry germplasm resource database according to claim 1, characterized in that: The process of obtaining the relatability includes: Each population feature is used as the target feature in turn; each population feature other than the target feature is used as the corresponding comparison feature; Calculate the mutual information value between the target feature and the contrast feature; use the hierarchical path length of the target feature and the contrast feature in the initial livestock and poultry germplasm resources database as the corresponding inter-layer distance; normalize the product between the negative correlation mapping value of the inter-layer distance and the mutual information value to determine the relevance between the target feature and the contrast feature.
3. The method for constructing a livestock and poultry germplasm resource database according to claim 1, characterized in that: The process of obtaining the initial feature association graph based on the relatability among all population features includes: Two population features whose corresponding associatability is greater than a preset association threshold are used as association features; an initial feature association graph is established based on the association features in all population features in a directed graph form.
4. The method for constructing a livestock and poultry germplasm resource database according to claim 1, characterized in that: The process of obtaining the retrieval matching degree includes: The number of population features corresponding to the intersection of all population features in the search formula corresponding to each population feature combination and all population features in the corresponding search results is positively correlated to determine the category matching degree; the negative correlation mapping value of the difference between the total number of population features in the search formula corresponding to each population feature combination and the total number of population features in the corresponding search results is used as the quantity matching degree; The retrieval matching degree of each population feature combination is determined according to the product of the category matching degree and the quantity matching degree.
5. The method for constructing a livestock and poultry germplasm resource database according to claim 1, characterized in that: The process of obtaining the comprehensive search cost includes: In a sampling time period, all search moments of all search formulas containing each population feature combination are arranged in chronological order to obtain a search moment time series sequence; in the search moment time series sequence, the time interval between each search moment and the previous search moment is used as the search time series interval of each search moment; and the search frequency of each population feature combination is determined based on the negative correlation mapping value of the mean of the search time series intervals of all search moments; The ratio between the number of search moments corresponding to all search formulas containing each population feature combination in the sampling time period and the total number of search moments is used as the search demand degree of each population feature combination; The average response time of all search expressions containing each population feature combination during the sampling period is used as the time cost; The comprehensive search cost of each population feature combination is determined according to the product of the search frequency, the search demand and the time cost.
6. The method for constructing a livestock and poultry germplasm resource database according to claim 1, characterized in that: The process of obtaining the feature correlation degree includes: The ratio between the search matching degree and the comprehensive search cost is normalized to determine the feature correlation degree of each population feature combination.
7. The method for constructing a livestock and poultry germplasm resource database according to claim 1, characterized in that: The process of obtaining the optimized feature combination includes: The population feature combination with a feature correlation greater than a preset high correlation threshold is used as the optimized feature combination.
8. The method for constructing a livestock and poultry germplasm resource database according to claim 3, characterized in that: The process of obtaining the optimized feature association graph includes: In the optimized feature combination, each population feature and other population features are used as associated features; the initial feature association graph is optimized in the form of a directed graph based on the associated features of all population features of each optimized feature combination to obtain an optimized feature association graph.
9. The method for constructing a livestock and poultry germplasm resource database according to claim 8, characterized in that: The process of adjusting the search priority of the initial livestock and poultry germplasm resources database according to the optimized feature association graph includes: The database federation technology is used to map the physical table of the associated features in the optimized feature association graph into a unified logical view; when the user searches the initial livestock and poultry germplasm resource database, the search request is preferentially sent to the unified logical view for retrieval.
10. A system for constructing a livestock and poultry germplasm resource database, characterized in that: The system comprises: The data acquisition and preprocessing module is used to obtain an initial livestock and poultry germplasm resource database with a multi-level classification storage architecture; obtain search results obtained by searching the historical data for each population characteristic combination using a search formula; wherein the initial livestock and poultry germplasm resource database includes various population characteristics; An initial feature association graph construction module is used to determine the corresponding relevance based on the association between each population feature and each other population feature and the hierarchical deviation in the initial livestock and poultry germplasm resource database; and to establish an initial feature association graph based on the relevance between all population features; An optimized feature combination determination module is used to determine the corresponding search matching degree based on the population feature matching between the search formula corresponding to each population feature combination and the corresponding search result; determine the corresponding comprehensive search cost based on the search frequency and search response of each search formula containing each population feature combination; determine the feature correlation degree of each population feature combination based on the search matching degree and the comprehensive search cost; and determine the optimized feature combination based on the feature correlation degree; A retrieval priority adjustment module is used to optimize the initial feature association diagram according to the population characteristics in the optimized feature combination to obtain an optimized feature association diagram; and adjust the retrieval priority of the initial livestock and poultry germplasm resource database according to the optimized feature association diagram.
Citation Information
Patent Citations
Associated data retrieval method and system
CN114911826A
Standard literature analysis management system and method applying big data technology
CN115618014A