An intelligent comprehensive management method and system for orchid germplasm resources

By integrating multi-omics data processing and knowledge graph technology, and combining CNN and GNN for intelligent management of Cymbidium germplasm resources, the problems of information dispersion and inconsistent standards have been solved, achieving efficient and accurate germplasm resource management and recommendation, and promoting the intelligent development of the Cymbidium industry.

CN120744104BActive Publication Date: 2025-11-28ENVIRONMENTAL HORTICULTURE RES INST OF GUANGDONG ACADEMY OF AGRI SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511194891.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-26
Publication Date
2025-11-28
Estimated Expiration
2045-08-26

AI Technical Summary

Technical Problem

The fragmented information, inconsistent standards, and lack of intelligent analysis and prediction in the management of Cymbidium germplasm resources lead to low resource management efficiency and poor recommendation accuracy, which in turn affects the development of the industry.

Method used

By integrating multi-omics data processing, knowledge graph construction, convolutional neural network (CNN) semantic analysis, graph neural network (GNN) semantic representation, and data encryption and access control, intelligent management of germplasm resources can be achieved throughout the entire process from data collection to planting prediction.

Benefits of technology

It significantly improves the efficiency of germplasm resource management, the accuracy of recommendations, and the level of scientific decision-making, providing technical support for the digital and intelligent transformation of the Cymbidium orchid industry.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120744104B_ABST
    Figure CN120744104B_ABST
Patent Text Reader

Abstract

The application discloses a kind of orchid germplasm resource intelligent comprehensive management method and system, comprising: the multi-source heterogeneous data of collection orchid germplasm resource is standardized modeling and multidimensional label classification with data, constructs the orchid knowledge graph of fusing variety genetic characteristics, character index and environmental factor;According to user interactive behavior, dynamic semantic association is carried out in knowledge graph, and user feature correlation table is constructed;According to correlation table, relevant entity and relationship data in knowledge graph are extracted, and combined with flowering period, incense type index and market heat weight generation germplasm resource recommendation information, introduce weighted CNN prediction model to carry out accurate planting prediction to recommended variety.The application fuses multi-omics analysis, knowledge graph and deep learning technology, realizes the accurate recommendation of orchid germplasm resource, safety control and intelligent cultivation decision, significantly improves resource management efficiency and industrial application value.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of horticultural plant genetic resource management, and in particular to an intelligent comprehensive management method and system for orchid germplasm resources, which integrates multi-omics data analysis, knowledge graph construction and artificial intelligence algorithms. BACKGROUND

[0002] China has a long history of collecting, preserving and utilizing orchid germplasm resources, but there are still many problems in the systematic management of germplasm resources. For example, the information is scattered and lacks a unified naming system, the genetic information and multi-omics data records are not systematic; planting management relies on experience and lacks data-driven scientific decision-making; different institutions have different data standards and the degree of resource sharing is low. These factors limit the efficient protection, accurate evaluation and utilization of orchid germplasm resources, and affect the development and upgrading of the industry.

[0003] To solve the above problems, the present application proposes an intelligent comprehensive management method and system for orchid germplasm resources, which integrates multi-omics data (including genome, transcriptome, metabolome, proteome, phenotype data) and actual production management information. The method integrates resource integration, feature extraction, intelligent recommendation and growth prediction functions, and realizes accurate recommendation and scientific management of germplasm resources through artificial intelligence and knowledge graph technology, thereby significantly improving the evaluation accuracy, recommendation efficiency and scientificity of utilization, and providing technical support for the digitalization and intelligent transformation of the orchid industry. SUMMARY

[0004] The present application proposes an intelligent comprehensive management method and system for orchid germplasm resources to solve the problems of scattered information, non-uniform standards, lack of intelligent analysis and prediction in existing orchid germplasm resource management. The scheme integrates multi-omics data processing, knowledge graph construction, convolutional neural network (CNN) semantic analysis, graph neural network (GNN) semantic representation, data encryption and access control, and realizes the intelligent management of orchid germplasm resources from data collection, knowledge graph construction, user feature extraction, resource recommendation to planting prediction. Through the present application, the management efficiency, recommendation accuracy and scientific decision-making level of germplasm resources can be significantly improved, and technical support is provided for the digitalization and intelligentization of the orchid industry.

[0005] The present application provides an intelligent comprehensive management method for orchid germplasm resources, comprising the following steps:

[0006] S1: Collect multi-source heterogeneous data of national orchid germplasm resources, the multi-source heterogeneous data including genomic sequence, transcriptome expression profile, metabolome aroma component, phenotype image data and cultivation environment monitoring data; standardize modeling and multi-dimensional label classification on the multi-source heterogeneous data, and generate triple data based on genetic similarity of varieties, trait co-occurrence and environmental adaptability, and construct a knowledge graph fusing characteristics weight of national orchid based on the triple data;

[0007] S2: Real-time collection and analysis of user interaction behavior data at multiple time nodes, context semantic analysis and feature extraction of user behavior data based on CNN semantic model, obtaining user behavior features including flowering period preference, fragrance type preference and market heat; dynamic correlation analysis of related entity and relationship data in the knowledge graph based on the user behavior features, combined with the time correlation of each user behavior feature, generating a feature correlation table for each user;

[0008] S3: According to the feature correlation table, extract entity and relationship data matching user features from the knowledge graph, and combine multi-dimensional feature parameters such as flowering period, leaf type, and fragrance index of national orchid varieties to perform semantic conversion, and generate germplasm resource recommendation information;

[0009] S4: Based on the germplasm resource recommendation information and its data type selection matching encryption algorithm, configure access rights according to user roles and distribute independent keys, record access logs and recommendation data call records, and perform encrypted backup storage;

[0010] S5: Based on the germplasm resource recommendation information, obtain multiple recommended national orchid varieties, perform semantic representation on entity nodes of recommended varieties in the knowledge graph through a preset graph neural network, set semantic weights of recommended national orchid varieties, and combine historical and real-time growth sequences and semantic weights to introduce a weighted CNN prediction model to predict the growth state and flowering period of national orchids.

[0011] In the scheme, the S1 specifically includes:

[0012] The multi-source heterogeneous data includes:

[0013] Basic classification information, including variety category, flower type category, and leaf type category;

[0014] Geographical distribution information, including geographical coordinates, altitude, climate zone, and seasonal temperature and humidity changes of the original habitat and cultivation area;

[0015] Genetic and molecular information, including genomic sequence, transcriptome expression profile, and metabolome aroma compound spectrum;

[0016] Phenotype and trait information: leaf type, flower color, fragrance index, flowering period, and stress resistance index;

[0017] Data formatting, data cleaning, missing value filling and outlier processing are performed on multi-source heterogeneous data, and unit unification and threshold normalization preprocessing are performed according to the national standard of the industry;

[0018] After preprocessing, data modeling is performed based on a multi-dimensional label system, which includes genetic labels, trait labels, environmental adaptation labels and market heat labels;

[0019] Based on the labeled data, national entity extraction, attribute feature extraction and entity relationship construction are performed to obtain triple data, and the entity relationship includes genetic similarity relationship between varieties, trait co-occurrence relationship, environmental adaptability relationship and fragrance correlation relationship;

[0020] The triple data is introduced into the national characteristic weight factor for weighted processing to generate a knowledge graph that integrates the national characteristic weight;

[0021] Based on the knowledge graph structure, graphical visualization is performed, and each entity node and its relationship edge are interactively displayed through a user terminal to support variety comparison, feature retrieval and germplasm structure analysis.

[0022] In this scheme, the user terminal includes:

[0023] Mobile terminal: used for installing a special application program for national germplasm resource management, supporting shooting and uploading flower, leaf and root images, supporting uploading fragrance sensor data and geographic positioning information;

[0024] Computer terminal: with high-resolution map visualization and batch data analysis functions, supporting multi-window comparison of variety characteristics, guiding variety identification report and cultivation scheme;

[0025] Special collection device terminal: integrating fragrance component sensing module, environmental parameter collection module and RFID / QR code scanning module, used for on-site rapid collection of national variety fragrance spectrum, real-time climate parameters and germplasm identity information, and synchronous updating with the knowledge graph system;

[0026] Cloud interactive interface: supporting remote calling of knowledge graph data, online subscription of recommended variety updates and receiving of cultivation warning information pushed by prediction model.

[0027] In this scheme, S2 specifically includes:

[0028] Based on the interaction process between the user and the system, the behavior data of the user at multiple time nodes is collected in real time, including variety search records and click frequency, flower period selection preference and seasonal attention trend, fragrance category preference and historical evaluation data, leaf art type and flower color combination browsing time, market transaction and collection operation records;

[0029] A CNN semantic model combining the multi-modal features of the national lily is constructed, and the input of the semantic model includes a text description vector, a fragrance spectrum feature vector, and a variety image feature vector;

[0030] Through the semantic model, the entity data, relationship data, and attribute data in the knowledge graph are subjected to context semantic information extraction, obtaining graph document data containing three features of flowering period, fragrance, and market heat, and performing multi-dimensional semantic representation based on the document vector;

[0031] The semantic model is used to perform semantic analysis on user behavior data, generating user behavior features containing time-dependent features, and performing associated entity retrieval and relationship data analysis in the knowledge graph based on the user behavior features, obtaining retrieval entity data meeting the user's interest;

[0032] The retrieval entity data in the knowledge graph is labeled and an associated knowledge data set is generated, in which the similarity of each associated knowledge data based on variety features is calculated and numerically represented and sequenced by relationship strength, obtaining a first sequence;

[0033] Combined with the time series information of user behavior, the time correlation of each associated knowledge data with other associated knowledge data is calculated, and the time correlation is sequenced to obtain a second sequence, and the correlation coefficient of the first sequence and the second sequence is calculated based on the Pearson correlation coefficient, and the absolute value of the correlation coefficient is taken as the interest degree of the associated knowledge data;

[0034] According to the interest degree, all associated knowledge data is sorted, and the user behavior features and the associated knowledge data are stored correspondingly to form a feature association table maintained independently for each user.

[0035] In the scheme, the S3 specifically includes:

[0036] According to the interest degree in the feature association table, combined with the multi-dimensional feature weight of the national lily variety, a plurality of associated knowledge data meeting the user's preference is screened out, and the multi-dimensional feature weight includes:

[0037] Flowering period matching degree: calculated based on the climate conditions of the user's location and the historical flowering period prediction result of the national lily variety;

[0038] Fragrance matching degree: calculated based on the fragrance spectrum similarity and the user's fragrance preference label;

[0039] Leaf art and flower color combination preference degree: calculated based on the visual feature score of the variety leaf art category and the flower color category;

[0040] Market circulation heat: based on the comprehensive score of transaction records, collection quantity, and online discussion heat;

[0041] Environmental adaptability index: calculated based on the matching degree of the variety cultivation environment demand and the environmental parameters of the user's location;

[0042] Extract the entity and relationship data corresponding to the above-mentioned associated knowledge data from the knowledge graph, and perform semantic conversion to generate multi-modal germplasm resource recommendation information containing variety name, trait characteristics, flowering period prediction, fragrance index, cultivation suggestion and market reference price;

[0043] The recommendation information is prioritized according to the weighted interest degree, wherein the weighted interest degree is the weighted sum of the flowering period matching degree, the fragrance matching degree, the market circulation heat, and the environmental adaptability index parameters;

[0044] The recommendation information with the highest priority is displayed first on the user terminal interface, and secondary screening and sorting can be performed according to the flowering period, fragrance, and market heat conditions.

[0045] In this scheme, the S4 specifically includes:

[0046] According to the germplasm resource recommendation information generated by S3, the business scenario and the recommendation data type of the current recommendation process are determined;

[0047] The business scenario includes variety query scenario, cultivation management scenario, variety transaction and circulation scenario, variety copyright identification and infringement monitoring scenario;

[0048] According to the business scenario and the recommendation data type, the security level of the recommendation data is analyzed, and the corresponding encryption algorithm is matched, and the asymmetric encryption algorithm and the symmetric encryption algorithm are used for algorithm matching based on sensitive data and non-sensitive data respectively;

[0049] End-to-end encryption protocol is used for cross-platform data transmission;

[0050] According to the user role, access permission is configured and independent key is distributed, the user role includes ordinary user, registered merchant, breeding institution and platform administrator, and different roles have differentiated restrictions in access range, data export and secondary distribution permission;

[0051] Combined with the blockchain traceability mechanism, the unique identity of the national orchid variety, the transaction record, the copyright information and the call record of the recommendation data are stored on the chain;

[0052] In the user access process, the access log, the variety transaction record, the copyright certificate call record and the germplasm resource recommendation information access record are extracted through permission verification, and the data of the above-mentioned access process are stored in the disaster recovery backup system after encryption.

[0053] In this scheme, the S5 specifically includes:

[0054] According to the germplasm resource recommendation information, multiple recommended national orchid varieties are obtained, and entity and attribute data related to the varieties are extracted in the knowledge graph to construct a graph structure containing variety characteristics, genetic information, trait information, geographical information, and cultivation parameters;

[0055] In the graph structure, the relationship information of the five dimensions of variety, trait, geography, genetics, and quantitative index is analyzed, and the following national orchid industry-specific indicators are added to the relationship information:

[0056] Climate adaptability coefficient: calculated based on the similarity of meteorological data of the target cultivation area and the historical cultivation environment of the variety;

[0057] Flowering time window prediction value: based on historical flowering data and real-time environmental monitoring data, a time series analysis model is used to predict the most probable flowering start and end dates;

[0058] Fragrance stability index: based on the results of fragrance component detection over the years, the degree of fluctuation of the proportion of fragrance components of the variety under different environments is evaluated;

[0059] The entity nodes of the recommended national orchid varieties are semantically represented by a pre-set graph neural network, and the comprehensive semantic weight of each recommended national orchid variety is calculated in combination with the above industry-specific indicators;

[0060] The historical growth sequence and the current real-time growth sequence of multiple recommended national orchid varieties are input into a weighted CNN prediction model, and the comprehensive semantic weight is introduced as the authenticity of the growth sequence for sequence learning and prediction training;

[0061] In the prediction process, the output includes growth state level, flowering time window, and fragrance stability rating, and the prediction result is generated by setting the recommended cultivation measures and automatically pushed to the user terminal.

[0062] In the S5 described in the present scheme, the prediction result obtained by predicting the growth state and flowering period of national orchids is displayed in a multi-modal visualization manner through the user terminal, and the multi-modal visualization includes:

[0063] Flowering calendar view: the predicted flowering start and end dates are marked in the form of a time axis, and the flowering overlap degree of different varieties is dynamically displayed;

[0064] Fragrance component radar chart: shows the relative content and stability index of the main fragrance compounds;

[0065] Climate adaptability heat map: according to the climate adaptability coefficient, the suitable cultivation area is presented on the map, and zooming and area filtering are supported;

[0066] Growth state dynamic graph: based on historical and real-time growth sequences, a curve or animation is generated to show the change trend of plant height, leaf number, and flower bud differentiation index;

[0067] Cultivation suggestion panel: automatically generate corresponding fertilization, irrigation, temperature and humidity control and shading suggestions combined with prediction results, and export as a management schedule;

[0068] The visualization interface supports interactive operations, including filtering, comparison and collection by flowering period, fragrance type or climate conditions, and superimposed analysis of prediction information and historical cultivation records of selected varieties.

[0069] The second aspect of the application also provides an intelligent comprehensive management system for Cymbidium species resources, which comprises:

[0070] Memory: for storing the intelligent comprehensive management program of Cymbidium species resources and multi-source heterogeneous data sets, the data sets including genomic sequences, transcriptome expression profiles, metabolome aroma components, phenotype image data, cultivation environment monitoring data, user behavior data and market transaction records;

[0071] Processor: for executing the management program to realize the steps of the above-mentioned intelligent comprehensive management method for Cymbidium species resources;

[0072] Fragrance sensor interface module: for accessing gas chromatography-mass spectrometry or electronic nose aroma component collection equipment, and directly transmitting the collected fragrance spectrum data to the system database;

[0073] Climate and environmental data collection module: for real-time collection of light intensity, air temperature and humidity, soil temperature and humidity and carbon dioxide concentration, and dynamic linkage with the weighted CNN prediction model;

[0074] Blockchain node module: for storing unique identity of varieties, transaction contract, copyright certificate and recommended data call record;

[0075] User interaction terminal interface: for connecting mobile terminals, computer terminals and special collection equipment, for real-time display and interactive operation of flowering period calendar, fragrance radar chart, climate heat map and multi-modal prediction results.

[0076] The third aspect of the application provides a computer readable storage medium, the medium stores program instructions for executing the above-mentioned intelligent comprehensive management method for Cymbidium species resources.

[0077] The application discloses a kind of national orchid germplasm resource intelligent comprehensive management method and system, comprising: the multi-source heterogeneous data of collection national orchid germplasm resource are carried out data standardization modeling and multidimensional label classification, construct the national orchid knowledge graph that fuses variety genetic characteristics, character index and environmental factor;Dynamic semantic association is carried out from knowledge graph according to user interaction behavior, and user feature correlation table is constructed;According to the correlation table, relevant entity and relationship data in knowledge graph are extracted, and the weight such as flowering period, fragrance index and market heat is combined to generate germplasm resource recommendation information, introduce weighted CNN prediction model to carry out accurate planting prediction to recommended variety.The application fuses multiomics analysis, knowledge graph and deep learning technology, realizes the accurate recommendation of national orchid germplasm resource, safety control and intelligent cultivation decision-making, significantly improves resource management efficiency and industrial application value. BRIEF DESCRIPTION OF DRAWINGS

[0078] Figure 1 The flow chart of the national orchid germplasm resource intelligent comprehensive management method of the application is shown.

[0079] Figure 2 The module schematic diagram of the application is shown.

[0080] Figure 3 The containerization deployment module and the blockchain traceability module schematic diagram of the application are shown.

[0081] Figure 4 The running schematic diagram of the prediction credibility verification and scheduling module of the application is shown.

[0082] Figure 5 The block diagram of the national orchid germplasm resource intelligent comprehensive management system of the application is shown. DETAILED DESCRIPTION

[0083] In order to more clearly understand the above-mentioned purposes, features and advantages of the application, the application will be further described in detail below in combination with the drawings and specific embodiments. It should be noted that the embodiments of the application and the features in the embodiments can be combined with each other without conflict.

[0084] In the following description, many specific details are set forth in order to provide a thorough understanding of the application, but the application can also be implemented in other ways different from those described herein, therefore, the protection scope of the application is not limited by the specific embodiments disclosed below.

[0085] Figure 1 The flow chart of the national orchid germplasm resource intelligent comprehensive management method of the application is shown.

[0086] As Figure 1 shown, the first aspect of the application provides a kind of national orchid germplasm resource intelligent comprehensive management method, comprising:

[0087] S1: Collect multi-source heterogeneous data of national orchid germplasm resources, the multi-source heterogeneous data including genomic sequence, transcriptome expression profile, metabolome aroma component, phenotype image data and cultivation environment monitoring data; standardize modeling and multi-dimensional label classification are performed on the multi-source heterogeneous data, and ternary data is generated based on genetic similarity of varieties, trait co-occurrence and environmental adaptability, and a knowledge graph fusing national orchid characteristic weights is constructed through the ternary data;

[0088] S2: Real-time collection and analysis of user interaction behavior data at multiple time nodes, context semantic analysis and feature extraction of user behavior data based on a CNN semantic model, obtaining user behavior features including flowering period preference, fragrance preference and market heat; dynamic correlation analysis of related entity and relationship data in the knowledge graph based on the user behavior features, combined with the time correlation of each user behavior feature, generating a feature correlation table for each user;

[0089] S3: According to the feature correlation table, extracting entity and relationship data matched with user features from the knowledge graph, combining multi-dimensional characteristic parameters of national orchid variety flowering period, leaf type and fragrance index for semantic conversion, generating germplasm resource recommendation information;

[0090] S4: Based on the germplasm resource recommendation information and its data type, select a matching encryption algorithm, configure access rights according to the user role and distribute independent keys, record access logs and recommendation data call records, and perform encrypted backup storage;

[0091] S5: Based on the germplasm resource recommendation information, obtain multiple recommended national orchid varieties, perform semantic representation on entity nodes of the recommended varieties in the knowledge graph through a preset graph neural network (GNN), set semantic weights of the recommended national orchid varieties, combine historical and real-time growth sequences and semantic weights, and introduce a weighted CNN prediction model to predict the growth state and flowering period of the national orchid.

[0092] It should be noted that in the present application, the platform (or system) also includes a germplasm resource database, i.e. a system database, which is a digital management platform specially used for plant and animal germplasm information management, retrieval, archiving and sharing. Through the unified processing and structured classification (graph storage) of the present application, the visualization expression efficiency and user understanding ability of national orchid germplasm resource data and multi-source heterogeneous data are significantly improved, and the reusability and standardization level are higher in subsequent system calling and data interface development. The characteristic weight of national orchid includes the characteristic weight represented by each attribute and relationship data in the knowledge graph, such as geographical location, seasonal temperature and humidity, phenotype and trait class index, genetic similarity, trait co-occurrence, environmental adaptation characteristics and related factors and characteristic weights. Genetic similarity, trait co-occurrence and environmental adaptability are stored as attributes and relationship data.

[0093] Figure 2 The schematic diagram of the module of the application is shown;

[0094] It should be noted that the application comprises a generation module (the implementation process corresponds to S1), an acquisition module (the implementation process corresponds to S2, S3), an encryption module (the implementation process corresponds to S4), and an analysis and prediction module (the implementation process corresponds to S4), each module is connected through system integration at the software and hardware level to form a complete platform for comprehensive management and analysis and prediction of national orchid germplasm resource data, such as Figure 2 shown.

[0095] According to the embodiment of the application, the S1 is specifically:

[0096] The multi-source heterogeneous data comprises:

[0097] Basic classification information, including variety category, flower type category, and leaf type category;

[0098] Geographical distribution information, including geographical coordinates, altitude, climate zone, and seasonal temperature and humidity changes of the original habitat and cultivation area;

[0099] Genetic and molecular information, including genome sequence, transcriptome expression profile, and metabolome aroma compound spectrum;

[0100] Phenotype and trait information: leaf type, flower color, aroma type index, flowering period, and stress resistance index;

[0101] The multi-source heterogeneous data is subjected to data formatting, data cleaning, missing value filling, and abnormal value processing, and is subjected to unit unification and threshold normalization preprocessing according to national orchid industry standards;

[0102] After preprocessing, data modeling is performed based on a multi-dimensional label system, and the label system comprises genetic labels, trait labels, environmental adaptation labels, and market heat labels;

[0103] Based on the labeled data, national orchid entity extraction, attribute feature extraction, and entity relationship construction are performed to obtain triple data, and the entity relationship comprises genetic similarity relationship between varieties, trait co-occurrence relationship, environmental adaptability relationship, and aroma type correlation degree relationship;

[0104] The triple data is introduced into a national orchid characteristic weight factor for weighted processing to generate a knowledge graph fused with the national orchid characteristic weight;

[0105] Based on the knowledge graph structure, graphical visualization is performed, and each entity node and its relationship edge are interactively displayed through a user terminal to support variety comparison, feature retrieval, and germplasm structure analysis.

[0106] It should be noted that the entity is a node in the graph, each node represents a Cymbidium variety or sample, and each edge is a representation of relationship data, which represents the similarity, genetic association or commonality of growth conditions between different varieties. Based on the Cymbidium germplasm resources and their associated characteristics (such as varieties, geographical environment, growth cycle, climate factors, etc.), a knowledge graph is constructed, and the attribute data is the geographical distribution, planting characteristics and other information of the entity node. The knowledge graph is a kind of graph structure data, which can build a relationship based on the graph structure of different Cymbidium varieties, visualize the abstract germplasm data through graphics and image means, enhance the user's understanding of the resource structure and characteristics, improve the interactive experience and decision-making efficiency, and at the same time, the associated information of different Cymbidium varieties can be efficiently analyzed, which provides a data basis for subsequent germplasm resource recommendation.

[0107] The basic classification, geographical distribution, germplasm characteristics and label classification method in the multi-source heterogeneous data are as follows:

[0108] Basic classification: According to the growth habit and adaptability, the Cymbidium germplasm resources are divided into multiple variety types, such as spring orchid, elegant orchid, building orchid, ink orchid, cold orchid, lotus petal orchid, spring sword and bean petal orchid.

[0109] Geographical distribution: Cymbidium resources are particularly rich in Asia and Oceania, and the main producing areas include China, Japan, South Korea, Vietnam, Myanmar and Australia.

[0110] Germplasm characteristics: Different varieties have different genetic characteristics such as growth cycle, leaf morphology, flower type and color, fragrance type and stress resistance.

[0111] Classification method: Support multi-dimensional classification, including label management of germplasm resources according to variety, flower type, region, genotype, growth environment and other dimensions.

[0112] In the process of interaction and related data storage between the system (i.e. the platform of the application) and the user, the following data processing process can be realized: the system automatically records the user's query behavior on the platform, including keywords, access records and interaction behavior, which is used for user interest modeling, behavior characteristic analysis, etc. Based on the user's query frequency and click behavior, the user's preference trend for specific flower types or categories is extracted, such as the tendency of preference for unusual flower varieties. The system uses encryption algorithm to protect the key process and data, to ensure the user privacy and platform data security, and unauthorized users cannot read the original data. The system realizes user role permission management, and distinguishes between ordinary users, researchers and platform administrators, and different roles have different data access and operation permissions. The system defines access permissions, defines access levels of different roles according to security policies, and realizes effective isolation and permission control of core data resources.

[0113] According to the embodiment of the application, the user terminal comprises:

[0114] Mobile terminal: used for installing special application for national orchid germplasm management, supporting image shooting and uploading of flower, leaf and root, supporting uploading of fragrance sensor data and geographic positioning information;

[0115] Computer terminal: with high-resolution atlas visualization and batch data analysis functions, supporting multi-window comparison of variety characteristics, guiding variety identification report and cultivation scheme;

[0116] Special collection device terminal: integrating fragrance component sensing module, environmental parameter collection module and RFID / two-dimensional code scanning module, used for on-site rapid collection of national orchid variety fragrance spectrum, real-time climate parameters and germplasm identity information, and synchronous updating with knowledge graph system;

[0117] Cloud interaction interface: supporting remote calling of knowledge graph data, online subscription of recommended variety updates and receiving of cultivation warning information pushed by prediction model.

[0118] According to the embodiment of the present application, the S2, specifically:

[0119] Based on the interaction process between the user and the system, the behavior data of the user at multiple time nodes is collected in real time, and the behavior data includes variety search records and click frequency, flowering period selection preference and seasonal attention trend, fragrance category preference and historical evaluation data, browsing time length of leaf art type and flower color combination, market transaction and collection operation records;

[0120] A CNN semantic model combining national orchid multi-modal features is constructed, and the input of the semantic model includes text description vector, fragrance spectrum feature vector and variety image feature vector;

[0121] Through the semantic model, the context semantic information of entity data, relationship data and attribute data in the knowledge graph is extracted, the atlas document data containing flowering period-fragrance-market heat three features are obtained, and multi-dimensional semantic representation is performed based on document vector;

[0122] The semantic model is used for semantic analysis of user behavior data, user behavior features containing time-dependent features are generated, and based on the user behavior features, associated entity retrieval and relationship data analysis are performed in the knowledge graph, to obtain retrieval entity data meeting user interest;

[0123] The retrieval entity data is marked in the knowledge graph and an associated knowledge data set is generated, in the associated knowledge data set, the similarity based on variety characteristics is calculated for each associated knowledge data, and the relationship strength is numerically represented and serialized to obtain a first sequence;

[0124] The time correlation degree of each associated knowledge data with other associated knowledge data is calculated in combination with time sequence information of user behaviors, the time correlation degree is sequenced to obtain a second sequence, and the correlation coefficient of the first sequence and the second sequence is calculated based on a Pearson correlation coefficient, and the absolute value of the correlation coefficient is taken as the interest degree of the associated knowledge data.

[0125] All associated knowledge data is sorted according to the interest degree, and the user behavior features are stored in correspondence with the associated knowledge data to form a feature association table maintained independently for each user.

[0126] It should be noted that the CNN semantic model is a semantic analysis model based on a convolutional neural network, which uses semantic information representation of user behavior features to further retrieve corresponding associated entity (Guo Lan) data. In the retrieval of associated entities in the knowledge graph and the analysis of relationship data, corresponding retrieval entity data can be obtained, and dynamic association analysis of entities and relationship data in the knowledge graph can be performed to evaluate the context representation of behavior features in the knowledge graph and mark the corresponding entity data.

[0127] In the knowledge graph, the relationship between entities is generally determined by the geographical correlation between different Guo Lans, the growth correlation, the variety correlation, and the climate factor correlation. User behavior data includes user browsing records, search keywords, click heat, conversion behavior and other feature data, and has time sequence properties and context dependence characteristics. User portraits can be described through user behavior data to form user demographic information, behavior paths, interest preferences and other features.

[0128] The CNN semantic model is a semantic analysis and feature classification of feature data through convolution layers, pooling layers and fully connected layers.

[0129] Each user behavior data in the user behavior data (or user behavior feature vector) corresponds to a time node, each associated knowledge data also corresponds to a time node, and one user corresponds to one associated knowledge data set. The graph document data is a kind of context data, which is constructed based on entity data of the knowledge graph, and is used to retrieve associated entities of user behavior features in the application. Specifically, the retrieval entity data is marked with the corresponding associated entity, the associated knowledge data includes the retrieval entity data, the associated entity and the graph association information corresponding to the entity, the entity attribute information, etc.

[0130] The relationship strength is obtained through edge information in the graph, and the stronger the strength is, the closer the correlation between entities is, and the closer the distance between entities in the knowledge graph is. Herein, the variety characteristics are taken as the correlation dimension for analysis. The first sequence stores a plurality of correlation strength values, which are obtained by comparing one correlation knowledge data with other plurality of correlation knowledge data entities. The second sequence stores a plurality of time correlations, which are obtained through time correlation analysis. The time correlation is the distance between two time nodes, and the greater the distance is, the greater the time correlation is. The time correlation analysis can reflect the interaction characteristics and interaction similarity of the user in continuous time, so as to evaluate the effectiveness of the recommended data and mine the potential user resource data demand.

[0131] The feature correlation table stores the correlation state between the user behavior characteristics, the correlation knowledge data, the interest degree and the data.

[0132] It is worth mentioning here that in the traditional interaction analysis and recommendation of the germplasm resource platform, the user behavior analysis is often based on a single dimension, and the real-time recommendation effect is poor. Moreover, the retrieval efficiency of real-time correlation data in the high-frequency interaction user and the germplasm big data is low, it is difficult to adapt to the high-efficiency and dynamic recommendation under the condition of big data and multi-user high-frequency interaction, and there is a lack of recommendation correlation mode analysis means, which leads to the single dimension of the recommended data and the difficulty in mining the potential demand of the user.

[0133] In the present application, the user behavior data of the user at multiple time nodes is collected for behavior characteristic analysis, and before that, a multi-dimensional knowledge graph of the national orchid germplasm resource with context semantic information is constructed. Based on the semantic analysis form, the graph knowledge data retrieval analysis and entity correlation analysis of different behavior characteristics are carried out, a plurality of entity data are retrieved to construct the correlation knowledge data, the time correlation is introduced, the relationship between the corresponding correlation knowledge data of the user at different time nodes is analyzed, the interest degree of the recommended knowledge is set, and a feature correlation table is maintained for each user. In addition, based on the real-time update of the newly added user behavior characteristics or the knowledge graph, the correlation table can be updated in real time, the recommended data of different users can be efficiently and accurately generated, the recommendation mode combining the germplasm resource and the graph can be constructed, the correlation between the graph knowledge and the user characteristics can be effectively utilized, the user behavior characteristics can be effectively analyzed and the user portrait can be constructed based on the semantic model analysis, the correlation knowledge retrieval is carried out, the dynamic data recommendation for the user is realized, and the update and maintenance rules of the correlation table can be further set to improve the real-time and applicability of the recommendation, realize the dynamic and efficient recommendation analysis for more users and big data germplasm platform, dynamically predict the potential demand of the user in the future based on the correlation knowledge of the graph, and automatically match and push the model most suitable for the content or resource, thereby improving the user satisfaction and interaction efficiency.

[0134] According to the embodiment of the present application, the S3 is specifically:

[0135] According to the interest degree in the feature association table, combined with the multi-dimensional feature weight of the national orchid variety, a plurality of associated knowledge data meeting the user preference is screened out, and the multi-dimensional feature weight includes:

[0136] Flowering period matching degree: calculated based on the climate conditions of the user's region and the historical flowering period prediction result of the national orchid variety;

[0137] Fragrance matching degree: calculated based on the fragrance spectrum similarity and the user fragrance preference label;

[0138] Leaf art and flower color combination preference degree: calculated based on the visual feature score of the variety leaf art category and the flower color category;

[0139] Market circulation heat: based on the transaction record, the number of collections and the online discussion heat comprehensive score;

[0140] Environmental adaptability index: calculated based on the matching degree of the variety cultivation environment demand and the environmental parameters of the user's location;

[0141] The entity and relationship data corresponding to the above associated knowledge data are extracted from the knowledge graph, and semantic conversion is performed to generate multi-modal germplasm resource recommendation information including variety name, property characteristics, flowering period prediction, fragrance index, cultivation suggestion and market reference price;

[0142] The recommendation information is prioritized according to the weighted interest degree, wherein the weighted interest degree is the weighted sum result of the flowering period matching degree, the fragrance matching degree, the market circulation heat and the environmental adaptability index parameters;

[0143] The recommendation information with the highest priority is preferentially displayed on the user terminal interface, and secondary screening and sorting can be performed according to the flowering period, fragrance, market heat and other conditions.

[0144] In the feature association table, the interest degree can realize primary screening, the multi-dimensional feature weight can realize secondary screening, and the interest weighted screening is combined with the multi-dimensional evaluation standard.

[0145] According to the embodiment of the application, the S4 is specifically:

[0146] According to the germplasm resource recommendation information generated by S3, the business scenario and the recommendation data type in the current recommendation process are determined;

[0147] The business scenario includes variety query scenario, cultivation management scenario, variety transaction and circulation scenario, variety copyright identification and infringement monitoring scenario;

[0148] According to the business scenario and the recommendation data type, the security level of the recommendation data is analyzed, and the corresponding encryption algorithm is matched,

[0149] Symmetric encryption algorithm (AES, SM4, etc.) is used for general search data and other non-sensitive data.

[0150] Asymmetric encryption algorithm (RSA, SM2, etc.) is used for sensitive data such as transaction contracts and copyright certificates, combined with digital signature.

[0151] End-to-end encryption protocol is used for cross-platform data transmission.

[0152] Access permissions are configured according to user roles, including ordinary users, registered merchants, breeding institutions and platform administrators, and different roles have differentiated restrictions in access range, data export and secondary distribution permissions.

[0153] The unique identity of Guolandi varieties, transaction records, copyright information and call records of recommended data are stored on the chain in combination with the blockchain traceability mechanism, realizing data tamper-proofing and traceability.

[0154] During user access, access logs, variety transaction records, copyright certificate call records and germplasm resource recommendation information access records are extracted through permission verification, and the above access data is encrypted and stored in the disaster recovery backup system to prevent data loss and illegal leakage.

[0155] It should be noted that according to the business scenario and the type of recommended data, the permission level and security requirement level of the data in the corresponding recommended scenario can be analyzed, and different levels of encryption algorithms are further matched to ensure the security of the data. Market heat refers to market circulation heat.

[0156] The encryption process of S4 can ensure the auditability and disaster recovery capability of the system, and the encryption scenarios include but are not limited to resource retrieval, planting service reservation, variety selection and comparison, etc.

[0157] The recommended data type can be divided into structured content data (such as database fields), connection data (such as similarity correlation), image-text data (such as flower pattern + text description) and other multi-modal information. The encryption algorithm can be symmetric or asymmetric encryption algorithm, such as DES, AES, etc., with the advantages of fast encryption speed, high security, low implementation cost, etc., suitable for real-time protection in the recommendation process. User roles include but are not limited to administrators, auditors and ordinary users, and different roles have different system access permissions and operation boundaries.

[0158] According to the recommended scenario and data type, the appropriate encryption algorithm is matched, and the access permission is allocated according to the user role, which can realize the whole process data protection and permission control in the recommendation process. At the same time, user access records and resource call data are obtained and encrypted, further enhancing the overall data security, auditability and disaster recovery capability of the system.

[0159] According to the embodiment of the present application, the S5, in particular:

[0160] According to the germplasm resource recommendation information, a plurality of recommended national orchid varieties are obtained, and entities and attribute data related to the varieties are extracted in the knowledge graph to construct a graph structure containing variety characteristics, genetic information, trait information, geographical information, and cultivation parameters;

[0161] In the graph structure, the relationship information of the five dimensions of varieties, traits, geography, genetics, and quantitative indicators is analyzed, and the following national orchid industry-specific indicators are added to the relationship information:

[0162] Climate adaptability coefficient: based on the similarity of multi-year meteorological data (temperature, humidity, light, and precipitation) of the target cultivation area and the historical cultivation environment of the variety;

[0163] Flowering time window prediction value: based on historical flowering data and real-time environmental monitoring data, a time series analysis model is used to predict the most probable flowering start and end dates;

[0164] Fragrance stability index: based on the results of fragrance component detection over the years, the degree of fluctuation of the proportion of fragrance components of the variety under different environments is evaluated;

[0165] The entity nodes of the recommended national orchid varieties are semantically represented by a pre-set graph neural network, and the comprehensive semantic weight of each recommended national orchid variety is calculated in combination with the above industry-specific indicators;

[0166] The historical growth sequence and the current real-time growth sequence of the plurality of recommended national orchid varieties are input into a weighted CNN prediction model, and the comprehensive semantic weight is introduced as the authenticity of the growth sequence for sequence learning and prediction training;

[0167] In the prediction process, the growth state level, flowering time window, and fragrance stability rating are output, the recommended cultivation measures are set through the prediction output, the prediction result is generated, and is automatically pushed to the user terminal for assisting in variety selection and cultivation management decision-making.

[0168] It should be noted that in the sequence prediction training, the semantic weight is used as the authenticity of the historical growth sequence of different varieties, and based on the data authenticity, the sequence training of the CNN prediction model can be biased towards sequences with higher authenticity, and the sequence feature learning is targeted based on the weight, to construct a precise prediction model for the current germplasm resource recommendation information. In each prediction training process, the authenticity of different training sequences can be introduced in the loss calculation of the sequence prediction result to calculate the loss, so that the weighted CNN prediction model can be biased towards sequences with higher authenticity for feature learning. The prediction result includes the growth state level, the flowering time window, the fragrance stability rating, and the recommended cultivation measures.

[0169] The graph structure is a subgraph based on an S1 knowledge graph, and the relationship data is different. The graph neural network GNN specifically includes GCN, GAT, etc.

[0170] The specific data characteristics of the five dimensions of variety, trait, geography, genetics, entity, and quantitative indicators are as follows:

[0171] Variety information: such as national orchid name, scientific name, variety source, breeder information, parent material, etc.

[0172] Trait information: such as flower color, leaf type, fragrance type, flowering period, polysaccharide or sesquiterpene content, etc.

[0173] Geographical information: such as planting area, latitude and longitude, altitude, climate parameters (temperature, humidity, etc.);

[0174] Genetic information: such as genotype, phenotypic difference, allele frequency, heterozygosity;

[0175] Quantitative indicators: such as sample size, plant number, flower amount, pollination rate, etc.

[0176] The weighted CNN prediction model captures the time sequence characteristics and dependency relationships of the sequence at different time scales through the Conv1D one-dimensional convolution layer and the one-dimensional pooling layer. In the prediction process, the result is output through the full connection layer. In addition, before importing the historical growth sequence, the data needs to be preprocessed by corresponding sequence standardization. In planting prediction, the output result of the prediction sequence can obtain the flowering period classification (such as early flowering type, middle flowering type, and late flowering type) or the specific prediction date range and state (such as “predicted to bloom between late March and early April”, growth state prediction, etc.).

[0177] The preset graph neural network represents the semantic information of the entity nodes of the recommended national orchid variety, and analyzes the semantic weight of each recommended national orchid variety through the graph structure. Specifically, by recommending the semantic relationship of the entity of the national orchid variety in the graph structure, the context importance of each entity is evaluated. The context is the context formed by the semantic analysis of all entity relationship attribute information of the graph structure. Combined with the analysis of the key degree of the entity in the context by the graph structure, the context importance is obtained. In the graph structure, the correlation degree can be obtained based on the analysis of the relationship data between each entity. The correlation degree is determined by the number and strength of the entity edges. The number and strength of the entity edges are determined by the relationship of the five dimensions.

[0178] The context importance and the entity correlation degree can be used as semantic weights for analysis, further, the semantic weights can be obtained by weighted average based on the two values, optionally, on the basis of the semantic weights, secondary weight calculation can be carried out by adding national industry specific indicators, climate adaptability coefficient, flower period time window prediction value, fragrance stability index and the like for comprehensive semantic weight calculation, and based on the prediction demand and state analysis demand of national flowers, any type of weight index can be freely matched as a semantic weight for calculation, so that the adaptability of the platform to big data analysis is improved.

[0179] It is worth mentioning here that the flowering period of national flowers is affected by complex coupling of multiple factors such as genetic characteristics of varieties, environmental factors (temperature, light, humidity) and growth history. In the prior art, there are few means for planting prediction of recommended national flower planting resource data, and there is a lack of growth prediction analysis combining dimensions such as variety association, geographical association and genetic association. The traditional prediction method is difficult to handle the correlation between varieties and high-dimensional spatio-temporal data (growth sequence), resulting in poor generalization.

[0180] Based on this, the application constructs a sub-atlas for the recommended national flower variety, forms a graph structure, sets relationship data based on multi-dimensional national flower characteristics to obtain a complete graph structure, extracts semantic information of the recommended national flower variety in the graph structure by using a neural network and calculates semantic weights, learns sequence characteristics of the recommended national flower variety by the CNN prediction model through the semantic weights and the growth sequence, and finally performs accurate growth prediction of the recommended national flower variety by the CNN prediction model.

[0181] The application can realize heterogeneous data fusion prediction, fuse prediction training of the graph structure (variety association, geographical association and the like) and time sequence environment data, realize dynamic adaptability of recommended variety prediction, realize nonlinear response modeling of planting prediction and multi-dimensional environmental fluctuation. The application realizes semantic feature propagation and fusion learning across varieties by using a graph neural network, significantly improves the accuracy and generalization ability of national flower flowering time prediction, and especially exhibits good prediction effect in the scene with a large number of varieties and strong data heterogeneity.

[0182] In the training process of the CNN prediction model, a supervision mechanism or model structure based on reinforcement learning can be introduced for iterative optimization of hyperparameters.

[0183] In addition, when the graph structure is sparse or the sample data is insufficient, the random walk model (such as DeepWalk, node2vec) can be automatically switched to capture the semantic feature relationship, the path is sampled through the walking strategy, the semantic relationship of the entity node is learned, and then the time sequence is connected to perform lightweight prediction. The random walk model (Random Walk-based Model) is a method of constructing a sequence by simulating the walking behavior between nodes on a graph, and learning the node representation by using the sequence context, which can be used in the scene of sparse graph structure or difficult to train deep learning model.

[0184] According to the embodiment of the present application, in the S5, the prediction result obtained by predicting the growth state and flowering period of the orchid is displayed in a multi-modal visualization manner through the user terminal, and the multi-modal visualization includes:

[0185] Flowering calendar view: annotating the predicted flowering start and end dates in the form of a time axis, and dynamically displaying the flowering overlap degree of different varieties;

[0186] Radar chart of fragrance components: showing the relative content and stability index of main fragrance compounds (such as linalool, geraniol, phenylethanol, etc.);

[0187] Climate adaptability heat map: presenting the suitable cultivation area on the map according to the climate adaptability coefficient, and supporting zooming and area filtering;

[0188] Growth state dynamic graph: generating a curve or animation based on the historical and real-time growth sequence to show the change trend of plant height, leaf number, flower bud differentiation, etc.

[0189] Cultivation suggestion panel: automatically generating corresponding fertilization, irrigation, temperature and humidity control, and shading suggestions based on the prediction result, and supporting one-key export to a management schedule.

[0190] The visualization interface supports interactive operation, including filtering, comparing and collecting according to flowering period, fragrance type or climate condition, and allows superimposed analysis of the prediction information of the selected variety and the historical cultivation record to assist users in making variety selection and management decisions.

[0191] According to the embodiment of the present application, it further includes a containerization module, which includes:

[0192] Decomposition module: determining the system architecture of the orchid germplasm resource data management system, and decomposing the system architecture into a plurality of portable modules;

[0193] Selection module: obtaining the requirements of the portable module, and selecting a cloud service provider and a development infrastructure based on the requirements;

[0194] Packaging module: package the portable modules, cloud service providers and infrastructure according to the system architecture and write application code;

[0195] Starting module: create containers based on Docker technology and manage containers in real time, deploy the packaged modules to the containers and start them in the cloud, and ensure that the containers are correctly configured and compatible with the infrastructure.

[0196] It should be noted that the national orchid germplasm resource data management system is a software system for efficient management of national orchid germplasm resources, and the system architecture can include: user interface layer, business logic layer, data storage layer, etc. The national orchid germplasm resource data management system is the system of the present application, and the plurality of portable modules include but are not limited to:

[0197] Germplasm information management module: realizes the basic functions of adding, deleting, modifying and querying germplasm resources;

[0198] Data analysis and management module: provides statistical analysis tools, supports chart visualization, and assists decision-making;

[0199] Permission management and authentication module: provides user registration, login, access control, permission grading and other functions;

[0200] Data backup and recovery module: realizes automatic backup, manual recovery, and prevents data loss or damage.

[0201] The cloud service provider can be Ali Cloud, Huawei Cloud, Tencent Cloud, Amazon AWS, Microsoft Azure, etc., which supports elastic computing, storage, network and other resources. The infrastructure includes: container runtime environment, virtual machine, network topology, object storage, load balancer and other cloud resource configurations. Docker is a container technology platform that supports rapid building, delivery and deployment, which encapsulates application programs and their dependencies into images to achieve consistent operation across platforms. The present application uses the Docker technology platform to realize the portability of each module.

[0202] According to the embodiment of the present application, it further comprises a blockchain traceability module, which comprises:

[0203] Acquisition and recording module: acquires key data nodes and records hash values;

[0204] On-chain module: write verified data onto the chain using a consortium chain structure, and bind it with a smart contract;

[0205] Verification module: set traceability verification based on user role access strategy;

[0206] Evidence storage module: form data evidence using on-chain hash values and timestamp information;

[0207] It should be noted that the blockchain platform can be a controllable permission alliance chain platform such as Hyperledger Fabric, Fisco BCOS, and the smart contract language can be Go or Solidity.

[0208] Key data nodes such as resource registration, analysis records, recommendation results, and prediction labels. The verification module is used to ensure that the data obtained by the visitor is real, complete, and not tampered with. The evidence storage module is used to meet the trusted traceability requirements of national orchid germplasm resources in scientific research and trading scenarios. The blockchain traceability module application scenarios can include: variety right protection, breeding process traceability, circulation record evidence storage, user behavior log tamper-proofing, etc.

[0209] The blockchain traceability technology is combined with the containerized deployment mechanism in this embodiment, which not only guarantees the security, traceability and integrity of the key data in the national orchid germplasm resource management system, but also realizes the portability, easy deployment and high scalability of the system modules through technologies such as Docker. The containerized structure facilitates rapid deployment in different environments (local / edge / cloud), and the blockchain ensures the credibility and checkability of the whole process, providing a solid digital guarantee for scientific research, breeding, trading and other scenarios.

[0210] Figure 3 The containerized deployment module and the blockchain traceability module are shown in the schematic diagram of the application;

[0211] As Figure 3 The combination of the containerized deployment module and the blockchain traceability module is shown in the schematic diagram, which includes four portable modules in the upper left corner, the Docker container and the blockchain traceability part in the upper right corner, and the containerized deployment module based on the system architecture through decomposition, selection and packaging modules for modular deployment, and the data collected after deployment for request, chain, verification and other operations.

[0212] Figure 4 The running schematic diagram of the prediction credibility verification and scheduling module is shown in the schematic diagram of the application;

[0213] According to the embodiment of the application, it also includes a prediction credibility verification and scheduling module, which includes:

[0214] The first verification module: when the user initiates a prediction analysis request, a unique task identifier and input data hash digest are generated for each task, and they are recorded on the chain;

[0215] The second verification module: before the system executes the prediction model, the current container running environment fingerprint is obtained, and together with the prediction parameters, the digest is calculated again and written on the chain;

[0216] Smart contract module: The smart contract deployed on the blockchain platform automatically checks whether the task summary and the container environment match, and triggers prediction analysis after verification.

[0217] On-chain result recording module: The prediction output result is generated into a hash digest, and is signed by the system private key and written into the blockchain.

[0218] Heterogeneous container scheduling module: According to the label attribute of the prediction task, the most suitable instance is automatically scheduled in the multi-container node to execute the task.

[0219] Block query and audit module: Based on the user or administrator, the on-chain state of the prediction task, the call record, the input and output summary, and the model version are viewed.

[0220] It should be noted that the present application provides a prediction credibility verification and scheduling module based on the deep integration of blockchain and container deployment, which is used to enhance the data security, execution transparency and process traceability in the process of national germplasm resource prediction analysis.

[0221] The task identifier is a unique identification code generated by the user for each prediction request, which is used for on-chain task indexing and subsequent tracking; the input data hash digest is a summary of the key prediction input using an algorithm such as SHA-256, which ensures that the input parameters have not been tampered with. The task identifier includes UUID+timestamp.

[0222] The container running environment fingerprint includes but is not limited to: model image name and version, running container host information, GPU or CPU hardware identifier, call timestamp, container unique ID, etc.

[0223] The on-chain result recording module realizes the non-repudiation and post-verification of the result.

[0224] The smart contract is implemented in Solidity or Fabric chain code, which is used to perform parameter matching verification, task authorization verification, scheduling logic triggering and a series of trusted process control.

[0225] The heterogeneous container scheduling module schedules in multiple container instances by reading the task label and runtime state, ensures that the model version is correct, the authority is met, and the resources are adapted, to support the elastic execution of complex tasks. Label attributes such as model type, authority level, node load, etc.

[0226] The blockchain adopts private chain deployment (such as Hyperledger Fabric, Ethereum private chain, etc.), and all prediction-related records are written into blocks in the form of transactions to form a complete prediction process log chain for auditing and responsibility tracing.

[0227] The prediction credibility verification and scheduling module realizes the whole-process credible record from the prediction request initiation to the model execution, result output and audit backtracking, and improves the transparency, security and responsibility traceability of the system for the germplasm resource prediction, and is suitable for key scenes such as scientific research verification, intelligent agricultural management, variety certification and recordation and the like.

[0228] The prediction credibility verification and scheduling module can be seamlessly integrated with a Docker container, a Kubernetes orchestration platform, a blockchain node and a prediction analysis service (such as a deep learning model API), and provides services to the outside through a RESTful interface or a gRPC protocol.

[0229] The blockchain credible verification and container scheduling module greatly improves the performance of the national orchid germplasm resource prediction analysis platform in terms of safety, compliance and controllability, and provides a solid foundation for realizing large-scale multi-node collaborative prediction.

[0230] The application realizes traceable, verifiable and tamper-proof management of the whole process of the national orchid growth prediction task by combining blockchain and containerized deployment, greatly improving the credibility and transparency of the prediction results. The execution process of the prediction task, including input data, model version, running environment, output results and the like, can be recorded on the chain, ensuring that all behaviors are non-repudiable and facilitating later tracing and auditing. At the same time, the container fingerprint mechanism ensures the consistency of the model execution environment, avoids prediction bias caused by system differences, and improves the stability and reproducibility of the results. The system also supports elastic container scheduling based on heterogeneous resources, which can intelligently allocate computing tasks according to task priority and resource state, significantly improving overall response efficiency and resource utilization. In the multi-department or multi-agency collaborative use scene, the blockchain traceability mechanism effectively prevents data misuse, enhances the compliance and data security of the collaboration process, and, combined with on-chain encrypted summaries and container-level access control, can also resist illegal access and malicious tampering. In addition, containerized deployment supports rapid integration, flexible expansion and version rollback, providing technical support for the engineering landing and continuous iteration of the national orchid intelligent prediction platform, and promoting its large-scale application in scientific research, agricultural management and commercial services.

[0231] According to the embodiment of the application, further comprising:

[0232] In one interaction cycle, the user behavior characteristics and germplasm resource recommendation information of multiple interactive users are collected;

[0233] The user behavior characteristics are vectorized to obtain a behavior feature vector,

[0234] The text information of the knowledge graph is constructed through a Word2Vec model, and the germplasm resource recommendation information is subjected to semantic analysis based on a CNN semantic model to form a word vector in combination with the text information;

[0235] The feature vector of each user and the word vector are taken as clustering sample data, the users are clustered based on a kmeans clustering algorithm, and multiple user groups are obtained based on the clustering results;

[0236] The prediction results are grouped by the multiple user groups to obtain multiple prediction data groups;

[0237] Each prediction data group is encrypted and stored on a chain.

[0238] It should be noted that, by the present application, the users with similar prediction characteristics and behavior characteristics are clustered based on two-dimensional information (behavior characteristics and prediction characteristics), and the prediction data is classified and packaged on a chain, in the subsequent prediction result tracing and historical data query, the tracing and query efficiency of similar data can be effectively improved, and by the two-dimensional prediction data clustering analysis, the user prediction results with similar search and prediction characteristics can be effectively mined, and in the multi-user high-frequency interaction scene, efficient and stable data on-chain, encryption storage and tracing query process can be realized. In the clustering process of the feature vector and the word vector as the clustering sample data, the similarity between the clustering sample data is measured by the Euclidean distance mean of the two vectors.

[0239] Figure 5 A block diagram of an intelligent comprehensive management system for Cymbidium species germplasm resources is shown.

[0240] The second aspect of the present application also provides an intelligent comprehensive management system 5 for Cymbidium species germplasm resources, which comprises:

[0241] Memory: for storing the intelligent comprehensive management program for Cymbidium species germplasm resources and multiple source heterogeneous data sets, the data sets comprising genomic sequences, transcriptome expression profiles, metabolome aroma components, phenotype image data, cultivation environment monitoring data, user behavior data and market transaction records;

[0242] Processor: for executing the management program to realize the steps of the intelligent comprehensive management method for Cymbidium species germplasm resources as described above;

[0243] Fragrance sensor interface module: for accessing gas chromatography-mass spectrometry (GC-MS) or electronic nose and other aroma component collection equipment, and directly transmitting the collected fragrance spectrum data to the system database;

[0244] Climate and environment data collection module: for real-time collection of light intensity, air temperature and humidity, soil temperature and humidity, and carbon dioxide concentration, and dynamic linkage with the prediction model;

[0245] The blockchain node module is used for storing variety unique identity, transaction contract, copyright certificate and recommended data calling record, and realizes the decentralized storage and traceable management of data.

[0246] The user interaction terminal interface is used for connecting mobile terminals, computer terminals and special collection equipment, and supports real-time display and interactive operation of multi-modal prediction results such as flowering calendar, fragrance radar chart and climate heat map.

[0247] The kind of orchid germplasm resource intelligent comprehensive management system can realize any step of the kind of orchid germplasm resource intelligent comprehensive management method.

[0248] The third aspect of the application further provides a computer readable storage medium, the medium stores program instructions for executing any one of the orchid germplasm resource intelligent comprehensive management methods.

[0249] The application discloses an orchid germplasm resource intelligent comprehensive management method and system, which comprises the following steps: collecting basic information of orchid germplasm resources, generating a knowledge graph after standardization modeling and classification; analyzing user behavior data in real time, extracting features through a CNN semantic model and dynamically associating with the knowledge graph to generate a user feature association table; extracting graph data from the feature association table to generate germplasm resource recommendation information; matching an encryption algorithm and configuring role permissions to realize secure access and data backup of the recommendation information; and predicting planting effects by using a CNN model based on semantic representation and growth sequences of recommended varieties. The application combines graph and deep learning technologies, realizes accurate recommendation and data security control of germplasm resources, and improves the management efficiency and decision-making scientificity of germplasm resources.

[0250] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other manners. The above-described device embodiments are only schematic. For example, the division of the units is only a logical function division. There can be another division manner for the actual implementation, for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed coupling, or direct coupling or communication connection between the components can be indirect coupling or communication connection through some interfaces, devices or units, and can be electrical, mechanical or other forms.

[0251] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units; they can be located in one place, or distributed on multiple network units; and part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment.

[0252] In addition, each function unit in each embodiment of the present application can be integrated in one processing unit, or each unit can be separately as a unit, or two or more units can be integrated in one unit; the integrated unit can be realized in the form of hardware or in the form of hardware plus software function unit.

[0253] Those skilled in the art can understand that all or part of the steps of the above-mentioned method embodiments can be completed by program instruction related hardware, and the foregoing program can be stored in a computer readable storage medium, and the program executes the steps including the above-mentioned method embodiments when executed; and the foregoing storage medium includes a mobile storage device, a read-only memory (ROM), a random access memory (RAM), a magnetic disc or an optical disc, and various storage medium capable of storing program codes.

[0254] Alternatively, the integrated unit of the present application, if realized in the form of a software function module and sold or used as an independent product, can also be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the embodiments of the present application can be embodied in the form of a software product, and the computer software product is stored in a storage medium, includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in the embodiments of the present application. The foregoing storage medium includes a mobile storage device, a ROM, a RAM, a magnetic disc or an optical disc, and various storage medium capable of storing program codes.

[0255] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto, any person skilled in the art can easily think of changes or replacements within the technical range disclosed in the present application, which should be covered in the protection scope of the present application.

Claims

1. An intelligent comprehensive management method for Cymbidium germplasm resources, characterized in that, The method comprises the following steps: S1: Collecting multi-source heterogeneous data of national orchid germplasm resources, which includes genomic sequence, transcriptome expression profile, metabolome aroma component, phenotype image data and cultivation environment monitoring data; standardizing modeling and multi-dimensional label classification of multi-source heterogeneous data, and generating triple data based on genetic similarity of varieties, trait co-occurrence and environmental adaptability, and constructing knowledge graph integrating national orchid characteristic weight through the triple data; S2: Real-time collection and analysis of user interaction behavior data at multiple time nodes, context semantic analysis and feature extraction of user behavior data based on CNN semantic model, obtaining user behavior characteristics including flowering period preference, fragrance type preference and market heat; Based on the user behavior characteristics, the related entity and relationship data in the knowledge graph are dynamically associated and analyzed, and the time correlation of each user behavior characteristic is combined to generate a feature association table for each user; S3: According to the feature association table, extracting entity and relationship data matching user characteristics from the knowledge graph, combining multi-dimensional characteristic parameters of national orchid variety flowering period, leaf type, fragrance index for semantic conversion, and generating germplasm resource recommendation information; S4: Based on the germplasm resource recommendation information and its data type, select a matching encryption algorithm, configure access rights according to the user role and assign an independent key, record access logs and recommendation data call records, and perform encrypted backup storage; S5: Based on the germplasm resource recommendation information, obtain multiple recommended national orchid varieties, perform semantic representation of entity nodes of recommended varieties in the knowledge graph through a preset graph neural network, set semantic weights of recommended national orchid varieties, combine historical and real-time growth sequences and semantic weights, and introduce a weighted CNN prediction model to predict the growth state and flowering period of national orchids.

2. The intelligent comprehensive management method of Phalaenopsis germplasm resources according to claim 1, characterized in that, The S1 specifically comprises: The multi-source heterogeneous data comprises: Basic classification information, including variety category, flower type category, leaf type category; Geographical distribution information, including geographical coordinates, altitude, climate zone, seasonal temperature and humidity changes of the original habitat and cultivation area; Genetic and molecular information, including genomic sequence, transcriptome expression profile, and metabolome aroma compound spectrum; Phenotype and trait information: leaf type, flower color, fragrance index, flowering period, stress resistance index; Data formatting, data cleaning, missing value filling and outlier processing are performed on the multi-source heterogeneous data, and unit unification and threshold normalization preprocessing are performed according to national orchid industry standards; After preprocessing, data modeling is performed based on a multi-dimensional label system, which includes genetic label, trait label, environmental adaptability label and market heat label; Based on the labeled data, national orchid entity extraction, attribute feature extraction and entity relationship construction are performed to obtain triple data, and the entity relationship includes genetic similarity relationship between varieties, trait co-occurrence relationship, environmental adaptability relationship and fragrance correlation degree relationship; The triple data is introduced into the national orchid characteristic weight factor for weighted processing to generate a knowledge graph integrating the national orchid characteristic weight. Based on the knowledge graph structure, graphical visualization is performed, and each entity node and its relationship edge are interactively displayed through a user terminal to support variety comparison, feature retrieval, and germplasm structure analysis.

3. The intelligent comprehensive management method of Phalaenopsis germplasm resources according to claim 2, characterized in that, The user terminal comprises: Mobile terminal: used to install a special application program for national orchid germplasm resource management, support shooting flower, leaf, root image and uploading identification, support uploading fragrance sensor data and geographic positioning information; Computer terminal: with high-resolution map visualization and batch data analysis functions, support multi-window comparison of variety characteristics, guide variety identification report and cultivation scheme; Special collection device terminal: integrated with aroma component sensing module, environmental parameter collection module and RFID / two-dimensional code scanning module, used for on-site rapid collection of national orchid variety fragrance spectrum, real-time climate parameters and germplasm identity information, and synchronous update with knowledge graph system; Cloud interaction interface: support remote call knowledge graph data, online subscription recommended variety update and receive cultivation warning information pushed by prediction model.

4. The intelligent comprehensive management method of Phalaenopsis germplasm resources according to claim 1, characterized in that, The S2 specifically comprises: Based on the interaction process between the user and the system, the behavior data of the user at multiple time nodes is collected in real time, including variety search records and click frequency, flowering period selection preference and seasonal attention trend, fragrance category preference and historical evaluation data, leaf type and flower color combination browsing time, market transaction and collection operation records; A CNN semantic model combining national orchid multi-modal features is constructed, and the input of the semantic model includes text description vector, fragrance spectrum feature vector and variety image feature vector; Through the semantic model, the context semantic information of entity data, relationship data and attribute data in the knowledge graph is extracted, and the graph document data containing flowering period-fragrance-market heat characteristics is obtained, and the multi-dimensional semantic representation is performed based on the document vector; The semantic model is used for semantic analysis of user behavior data, and user behavior features containing time-dependent characteristics are generated, and based on the user behavior features, associated entity retrieval and relationship data analysis are performed in the knowledge graph, and the retrieval entity data meeting the user interest is obtained; The retrieval entity data is marked in the knowledge graph and an associated knowledge data set is generated, and in the associated knowledge data set, the similarity based on variety characteristics is calculated for each associated knowledge data, and the relationship strength is numerically represented and serialized to obtain a first sequence; Combined with the time series information of user behavior, the time correlation of each associated knowledge data with other associated knowledge data is calculated, the time correlation is serialized to obtain a second sequence, and the correlation coefficient of the first sequence and the second sequence is calculated based on the Pearson correlation coefficient, and the absolute value of the correlation coefficient is taken as the interest degree of the associated knowledge data; According to the interest degree, all the associated knowledge data are sorted, and the user behavior features and the associated knowledge data are stored correspondingly to form a feature association table maintained independently for each user.

5. The intelligent comprehensive management method of Cymbidium germplasm resources according to claim 4, characterized in that, The S3 specifically comprises: According to the interest degree in the feature association table, combined with the multi-dimensional feature weight of national orchid varieties, a plurality of associated knowledge data meeting the user preference are selected, and the multi-dimensional feature weight comprises: Flowering period matching degree: calculated based on the user's local climate conditions and the historical flowering period prediction results of the national orchid varieties; Fragrance type matching degree: calculated based on the similarity of fragrance spectrum and the user's fragrance preference label; Leaf art and flower color combination preference degree: calculated based on the visual feature scores of the variety's leaf art category and flower color category; Market circulation heat: comprehensive score based on transaction records, collection quantity, and online discussion heat; Environmental adaptability index: calculated based on the matching degree of the variety's cultivation environment requirements and the user's local environment parameters; Extract the entity and relationship data corresponding to the above related knowledge data from the knowledge graph, and perform semantic conversion to generate multi-modal germplasm resource recommendation information containing variety name, trait characteristics, flowering period prediction, fragrance index, cultivation suggestion, and market reference price; Set the priority of the recommendation information according to the weighted interest degree, where the weighted interest degree is the weighted sum of the flowering period matching degree, fragrance type matching degree, market circulation heat, and environmental adaptability index parameters; Display the recommendation information with the highest priority on the user terminal interface first, and support secondary screening and sorting according to the flowering period, fragrance type, and market heat conditions.

6. The intelligent comprehensive management method of Phalaenopsis germplasm resources according to claim 1, characterized in that, The S4 specifically includes: According to the germplasm resource recommendation information generated in S3, determine the business scenario and recommendation data type of the current recommendation process; The business scenarios include variety query scenario, cultivation management scenario, variety transaction and circulation scenario, and variety copyright identification and infringement monitoring scenario; According to the business scenario and recommendation data type, analyze the security level of the recommendation data, and match the corresponding encryption algorithm. Based on sensitive data and non-sensitive data, asymmetric encryption algorithm and symmetric encryption algorithm are used for algorithm matching respectively; End-to-end encryption protocol is used for cross-platform data transmission; According to the user role, configure access rights and assign independent keys. The user roles include ordinary users, registered merchants, breeding institutions, and platform administrators. Different roles have differentiated restrictions in access range, data export, and secondary distribution permissions; Combine the blockchain traceability mechanism to store the unique identity of the national orchid varieties, transaction records, copyright information, and call records of the recommendation data on the chain; During user access, access logs, variety transaction records, copyright certificate call records, and germplasm resource recommendation information access records are extracted through permission verification, and the data of the above access processes are encrypted and stored in the disaster recovery backup system.

7. The intelligent comprehensive management method of Cymbidium germplasm resources according to claim 1, characterized in that, The S5 specifically includes: According to the germplasm resource recommendation information, obtain multiple recommended national orchid varieties, and extract related entity and attribute data in the knowledge graph, construct a graph structure containing variety characteristics, genetic information, trait information, geographical information, and cultivation parameters; In the graph structure, analyze the relationship information of the five dimensions of variety, trait, geography, genetics, and quantity index, and add the following national orchid industry-specific indicators to the relationship information: Climate adaptability coefficient: calculated based on the similarity of meteorological data of the target cultivation area and the historical cultivation environment of the variety; Flowering period time window prediction value: based on historical flowering period data and real-time environmental monitoring data, use a time series analysis model to predict the most probable flowering start and end dates; Fragrance stability index: based on the detection results of fragrance components in previous years, the fluctuation degree of the proportion of fragrance components of the variety in different environments is evaluated; The entity nodes of the recommended national orchid varieties are semantically represented by a preset graph neural network, and the comprehensive semantic weight of each recommended national orchid variety is calculated by combining the above industry-specific indicators; The historical growth sequence and the current real-time growth sequence of multiple recommended national orchid varieties are input into the weighted CNN prediction model, and the comprehensive semantic weight is introduced as the authenticity of the growth sequence for sequence learning and prediction training; In the prediction process, the output includes growth state level, flowering time window, fragrance stability rating, and the prediction result is generated by setting the recommended cultivation measures, and is automatically pushed to the user terminal.

8. The intelligent comprehensive management method of Cymbidium germplasm resources according to claim 1, characterized in that, In the prediction result obtained by predicting the growth state and flowering period of the national orchid, the prediction result is displayed in a multi-modal visual manner through the user terminal, and the multi-modal visualization includes: Flowering prediction calendar view: mark the predicted flowering start and end dates in the form of a time axis, and dynamically display the flowering overlap degree of different varieties; Fragrance component radar chart: display the relative content and stability index of the main fragrance compounds; Climate adaptability heat map: according to the climate adaptability coefficient, the suitable cultivation area is presented on the map, and zooming and area filtering are supported; Growth state dynamic diagram: generate a curve or animation based on the historical and real-time growth sequence to show the change trend of plant height, leaf number, and flower bud differentiation index; Cultivation suggestion panel: automatically generate corresponding fertilization, irrigation, temperature and humidity control, and shading suggestions based on the prediction result, and export as a management plan; The visualization interface supports interactive operation, including filtering, comparing and collecting according to flowering period, fragrance or climate conditions, and superimposed analysis of the prediction information and historical cultivation records of the selected varieties.

9. An intelligent comprehensive management system for orchid germplasm resources, characterized in that, The system includes: Memory: for storing national orchid germplasm resource intelligent comprehensive management program and multi-source heterogeneous data set, the data set includes genome sequence, transcriptome expression profile, metabolome fragrance component, phenotype image data, cultivation environment monitoring data, user behavior data and market transaction record; Processor: for executing the management program to realize the steps of the national orchid germplasm resource intelligent comprehensive management method according to claim 1; Fragrance sensor interface module: for connecting gas chromatography-mass spectrometry or electronic nose fragrance component collection equipment, and directly transmitting the collected fragrance spectrum data to the system database; Climate and environmental data collection module: for real-time collection of light intensity, air temperature and humidity, soil temperature and humidity, and carbon dioxide concentration, and dynamic linkage with the weighted CNN prediction model; Blockchain node module: for storing variety unique identity, transaction contract, copyright certificate and recommended data call record; User interaction terminal interface: for connecting mobile terminal, computer terminal and special collection equipment, for real-time display and interactive operation of flowering calendar, fragrance radar chart, climate heat map and multi-modal prediction result.

10. A computer-readable storage medium, characterized in that, The medium stores program instructions for executing the national orchid germplasm resource intelligent comprehensive management method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Grape germplasm resource data integrated management system

    CN118350668A

  • Gene coding breeding prediction method and device based on graph clustering

    US20240119314A1