Malus spectabilis remote management and data modeling method based on cloud collaborative architecture
By employing a cloud-based collaborative architecture for data processing, the problems of inconsistent data formats and limited parsing in the Begonia Flower Management System were resolved, resulting in a highly adaptable remote management strategy that improved management efficiency and quality.
Patent Information
- Application Number
- CN202510944345.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-09
- Publication Date
- 2025-10-17
AI Technical Summary
The existing crabapple blossom management system suffers from problems such as inconsistent data formats, difficulty in integrating multi-source data, simplistic data parsing, lack of collaborative analysis, and poor adaptability of management strategies, resulting in low management efficiency and inconsistent quality.
A cloud-based collaborative architecture-based data processing method, including standardized preprocessing, multi-source data fusion, semantic-level collaborative disambiguation model, and knowledge base matching, is used to generate a remote management strategy for crabapple blossoms.
It establishes a unified foundation for data analysis, eliminates data ambiguity, generates precise management strategies, and is highly adaptable, suitable for large-scale gardens and home planting scenarios.
Smart Images

Figure CN120806508A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of plant management, in particular to a remote management and data modeling method for Chinese flowering crabapple based on a cloud collaborative architecture. BACKGROUND
[0002] At present, with the rapid development of smart agriculture, plant remote management has gradually become an important direction for the upgrading of agricultural technology. Chinese flowering crabapple, as a plant species with ornamental value and cultural connotation, precise regulation of its growth state has practical significance for landscape maintenance, home planting and other scenarios. However, in the existing management mode of Chinese flowering crabapple, there are still many problems to be solved. Traditional Chinese flowering crabapple management relies heavily on human experience. Management personnel need to regularly observe plant growth conditions and collect environmental data, and develop maintenance strategies based on personal experience. This approach not only consumes a lot of manpower and resources, but is also greatly influenced by subjective factors, making it difficult to achieve standardized and precise management. Especially in large-scale planting scenarios, manual management is inefficient, and differences in experience among different management personnel may lead to inconsistent management measures, thereby affecting the overall growth quality of Chinese flowering crabapple. With the application of sensor technology and the Internet of Things, some regions have begun to try to collect Chinese flowering crabapple growth environment data such as temperature, humidity, and light intensity through terminal devices, and transmit them to the management platform for analysis via the network. However, these systems often have limited data processing capabilities, and the collected data is mostly stored in raw form or subjected to simple statistical analysis, making it difficult to deeply mine the correlation between data. For example, the data collected by different sensors have different formats, and lack a unified standardized processing mechanism, making it difficult to fuse multi-source data and form an effective decision basis. Existing technologies for data analysis are mostly limited to a single dimension and fail to combine historical data and professional knowledge bases for collaborative analysis. When faced with complex growth environment data, data ambiguity may occur. For example, abnormal humidity data in a certain time period may be caused by soil properties, weather changes, irrigation system failures, and other factors, and existing systems are unable to accurately identify the specific cause, thereby leading to insufficient relevance of management strategies. The application of cloud technology in plant management is still in its early stages, and there are obvious shortcomings in the construction and collaborative application of cloud knowledge bases. Most systems lack a dynamic updating mechanism, and the correlation between different data sources is low, making it difficult to provide effective support for the precise matching of data entities. In terms of data entity disambiguation, existing methods often ignore the role of contextual information and only use simple keyword matching, resulting in low accuracy of disambiguation results and affecting the generation of subsequent management strategies. In the remote management strategy generation link, the prior art is mostly based on a single or a small amount of data indicators to formulate a scheme, and fails to comprehensively consider multi-dimensional factors of the growth of the Chinese flowering crabapple. For example, only the irrigation frequency is adjusted according to the soil humidity data, while the synergistic effects of factors such as light and temperature are ignored, so that the adaptability of the management strategy is poor and it is difficult to cope with the complex and changeable growth environment. SUMMARY
[0003] The purpose of the present application is to provide a Chinese flowering crabapple remote management and data modeling method based on a cloud collaborative architecture to solve the problems raised in the background art.
[0004] To achieve the above purpose, the present application provides a Chinese flowering crabapple remote management and data modeling method based on a cloud collaborative architecture, which comprises: S1: receiving Chinese flowering crabapple growth environment data collected by a terminal device; S2: standardizing and preprocessing the growth environment data and fusing and analyzing multi-source data to obtain a growth feature entity candidate list and a non-feature data segment; S3: extracting a first data entity to be modeled from the growth feature entity candidate list; S4: extracting a plurality of candidate data entity entries matched with the first data entity to be modeled from a cloud collaborative knowledge base; S5: inputting the plurality of candidate data entity entries matched with the first data entity to be modeled and the non-feature data segment into a collaborative disambiguation model to obtain a disambiguation data entity of the first data entity to be modeled, comprising: taking the non-feature data segment as context information, performing semantic-level collaborative aggregation analysis on the plurality of candidate data entity entries to obtain disambiguation data entity query response semantic aggregation encoding features; and performing semantic decoding on the disambiguation data entity query response semantic aggregation encoding features to obtain the disambiguation data entity of the first data entity to be modeled; S6: cyclically executing S3 to S5 to obtain a disambiguation data entity list; S7: generating a Chinese flowering crabapple remote management strategy based on the disambiguation data entity list.
[0005] Preferably, S2 comprises: noise filtering and cleaning of the growth environment data to obtain purified growth environment data; timestamp alignment processing of the purified growth environment data to obtain a set of growth environment data sequences; feature entity recognition and classification of each growth environment data sequence in the set of growth environment data sequences to obtain a growth feature entity candidate list and a non-feature data segment.
[0006] Preferably, S4 comprises: taking the first to-be-modeled data entity as an association keyword, performing semantic association query on the cloud collaborative knowledge base to obtain a plurality of candidate data entity entries matched with the first to-be-modeled data entity.
[0007] Preferably, the plurality of candidate data entity entries are subjected to semantic-level collaborative aggregation analysis with the non-feature data segment as context information to obtain a disambiguated data entity query response semantic aggregation encoding feature, comprising: The context information is subjected to semantic understanding with the non-feature data segment as context information to obtain a to-be-modeled data entity context information semantic encoding vector. The candidate data entity entries are respectively subjected to structured encoding using a candidate data entity entry embedding matrix to obtain a plurality of candidate data entity entry semantic embedding encoding vectors. The to-be-modeled data entity context information semantic encoding vector and the plurality of candidate data entity entry semantic embedding encoding vectors are subjected to data entity semantic collaborative aggregation encoding to obtain a disambiguated data entity query response semantic aggregation encoding vector as a disambiguated data entity query response semantic aggregation encoding feature.
[0008] Preferably, the context information is subjected to semantic understanding with the non-feature data segment as context information to obtain a to-be-modeled data entity context information semantic encoding vector, comprising: The to-be-modeled data entity context information is subjected to semantic embedding encoding using a to-be-modeled data entity embedding matrix to obtain a context sequence of the to-be-modeled data entity context information semantic embedding encoding vector. The context sequence of the to-be-modeled data entity context information semantic embedding encoding vector is subjected to context semantic encoding based on a bidirectional GRU model to obtain the to-be-modeled data entity context information semantic encoding vector.
[0009] Preferably, the to-be-modeled data entity context information semantic encoding vector and the plurality of candidate data entity entry semantic embedding encoding vectors are subjected to data entity semantic collaborative aggregation encoding to obtain a disambiguated data entity query response semantic aggregation encoding vector, comprising: The to-be-modeled data entity context information semantic encoding vector is subjected to feature enhancement based on an attention mechanism to obtain a to-be-modeled data entity context information semantic reinforced encoding vector. The to-be-modeled data entity context information semantic reinforced encoding vector and each of the plurality of candidate data entity entry semantic embedding encoding vectors are subjected to data entity monomer semantic query to obtain a set of to-be-modeled-candidate data entity monomer semantic query score encoding vectors. Based on the feature set independent distribution characteristics of the set of candidate data entity monomer semantic query score encoding vectors, the set of candidate data entity monomer semantic query score encoding vectors is subjected to adaptive aggregation analysis based on gating modulation to obtain a disambiguated data entity query response semantic aggregation encoding vector.
[0010] Preferably, the data entity monomer semantic query is performed on the to-be-modeled data entity context information semantic reinforcement encoding vector and each of the plurality of candidate data entity item semantic embedding encoding vectors to obtain a set of to-be-modeled-candidate data entity monomer semantic query score encoding vectors, including: The to-be-modeled data entity context information semantic reinforcement encoding vector and each of the plurality of candidate data entity item semantic embedding encoding vectors are respectively input into a data entity monomer semantic collaborative query unit to obtain a set of initial to-be-modeled-candidate data entity monomer semantic query score encoding vectors. The set of initial to-be-modeled-candidate data entity monomer semantic query score encoding vectors is subjected to feature dynamic optimization mapping to obtain a set of to-be-modeled-candidate data entity monomer semantic query score encoding vectors.
[0011] Preferably, based on the feature set independent distribution characteristics of the set of to-be-modeled-candidate data entity monomer semantic query score encoding vectors, the set of to-be-modeled-candidate data entity monomer semantic query score encoding vectors is subjected to adaptive aggregation analysis based on gating modulation to obtain a disambiguated data entity query response semantic aggregation encoding vector, including: Based on the feature set independent distribution characteristics of the set of to-be-modeled-candidate data entity monomer semantic query score encoding vectors, the monomer semantic matching degree of each to-be-modeled-candidate data entity monomer semantic query score encoding vector in the set of to-be-modeled-candidate data entity monomer semantic query score encoding vectors is determined to obtain a set of to-be-modeled-candidate data entity monomer semantic matching degrees. The set of to-be-modeled-candidate data entity monomer query semantic self-attention weights is input into a relationship gating agent module to obtain a set of to-be-modeled-candidate data entity monomer query semantic self-attention weights. Based on the set of to-be-modeled-candidate data entity monomer query semantic self-attention weights, the set of to-be-modeled-candidate data entity monomer semantic query score encoding vectors is aggregated to obtain a disambiguated data entity query response semantic aggregation encoding vector.
[0012] Preferably, the disambiguated data entity query response semantic aggregation encoding feature is subjected to semantic decoding to obtain a disambiguated data entity of the first to-be-modeled data entity, including: The semantic decoding based on the Transformer model is performed on the semantic aggregation coding vector of the disambiguation data entity query response to obtain the disambiguation data entity of the first to-be-modeled data entity.
[0013] Preferably, based on the disambiguation data entity list, a remote management strategy of the Malus halliana is generated, including: The disambiguation data entity list is input into a pre-trained management strategy generation model to obtain the remote management strategy of the Malus halliana.
[0014] Compared with the prior art, the method has the following advantages: In terms of data processing, the method first performs standardized preprocessing and multi-source data fusion analysis on the growth environment data collected by the terminal device, breaking down the format barriers between different data sources and enabling originally scattered and heterogeneous data to form a unified analysis basis. This processing method avoids the information silo problem caused by data format differences and enables various environmental parameters to work together, providing more comprehensive raw materials for subsequent data analysis. From the perspective of data entity extraction and matching, the first to-be-modeled data entity is extracted from the growth characteristic entity candidate list and matched with the candidate data entity entries in the cloud collaborative knowledge base, making full use of the massive historical data and professional knowledge stored in the cloud. This matching mechanism is not a simple surface information comparison, but a screening based on deep association of data entities, making the extracted candidate entries more consistent with the actual growth of the Malus halliana and reducing irrelevant information interference. The application of the collaborative disambiguation model is a highlight of the method. With non-feature data segments as context information, semantic-level collaborative aggregation analysis is performed on multiple candidate data entity entries, and then semantic decoding is performed to obtain the disambiguation data entity. This process fully considers the environmental background of the data and effectively eliminates data ambiguity. For example, when a humidity data is abnormal, combining with the non-feature data such as light, soil type and other non-feature data in the same period, the cause of the abnormality can be more accurately determined as improper irrigation or weather change, thereby providing more accurate data basis for subsequent management strategy. The extraction, matching and disambiguation steps are performed in a loop, and the disambiguation data entity list formed finally covers various key information in the growth process of the Malus halliana, and these information has been screened and verified layer by layer, with high accuracy and consistency. Based on such a list, a remote management strategy is generated, which can fully reflect the growth needs of the Malus halliana and avoid the strategy deviation caused by one-sided information in traditional management. The introduction of the cloud collaborative architecture makes the entire management process more scalable and adaptable. The cloud collaborative knowledge base can continuously accumulate new management experience and data cases, providing more abundant references for subsequent data analysis through continuous updates. At the same time, the collaborative architecture supports data access and sharing of multiple terminal devices. Whether it is multiple monitoring points in a large garden or a single sensor in a home garden, it can be integrated into a unified management system, achieving efficient management in different scale planting scenarios. The remote management strategy generated by this method is not a fixed instruction, but a dynamically adjusted scheme based on real-time data and historical knowledge. When the growth environment of Chinese flowering crabapple changes, new data collected by the terminal device will be uploaded to the cloud in a timely manner. After processing and analysis, the management strategy will also be updated accordingly, ensuring that the management measures are matched with the plant growth status in real time, so that Chinese flowering crabapple can obtain suitable maintenance conditions in different growth stages. BRIEF DESCRIPTION OF DRAWINGS
[0015] Figure 1 The working principle diagram of the Chinese flowering crabapple remote management and data modeling method based on the cloud collaborative architecture described in the present application; Figure 2 The flowchart of growth environment data preprocessing; Figure 3 The flowchart of semantic level collaborative aggregation analysis; Figure 4 The flowchart of context information semantic coding; Figure 5 The flowchart of data entity semantic collaborative aggregation coding. DETAILED DESCRIPTION
[0016] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0017] Please refer to Figures 1-5 The present application provides a Chinese flowering crabapple remote management and data modeling method based on a cloud collaborative architecture, and the specific implementation steps are as follows: S1: Receive the Chinese flowering crabapple growth environment data collected by the terminal device. The terminal device can include temperature sensors, humidity sensors, light sensors, soil nutrient sensors, etc. The collected data includes air temperature, air humidity, light intensity, soil pH value, soil moisture content, and other environmental parameters related to the growth of Chinese flowering crabapple.
[0018] S2: Standardize the growth environment data and perform multi-source data fusion analysis to obtain a list of growth feature entity candidates and non-feature data segments. Standardization aims to unify data formats and scales, and multi-source data fusion analysis integrates different types of data to distinguish between entity candidates that may reflect the growth characteristics of the plant and data segments that are not directly related to the growth characteristics.
[0019] S3: Extract a first data entity to be modeled from the list of growth feature entity candidates. The first data entity to be modeled is a data entity in the list of growth feature entity candidates that needs to be prioritized for modeling processing, such as soil moisture data entity within a certain time period.
[0020] S4: Extract multiple candidate data entity entries matching the first data entity to be modeled from the cloud collaborative knowledge base. The cloud collaborative knowledge base stores a large amount of historical data, expert knowledge, and growth model parameters related to the growth of the plant, and through matching, reference data entries associated with the first data entity to be modeled can be obtained.
[0021] S5: Input the multiple candidate data entity entries matching the first data entity to be modeled and the non-feature data segments into a collaborative disambiguation model to obtain a disambiguated data entity of the first data entity to be modeled. The specific process is as follows: use the non-feature data segments as context information, perform semantic-level collaborative aggregation analysis on the multiple candidate data entity entries to obtain disambiguated data entity query response semantic aggregation encoding features; then perform semantic decoding on the encoding features to obtain the disambiguated data entity of the first data entity to be modeled. The disambiguated data entity is a data entity that has been processed to eliminate ambiguity and accurately reflect the true meaning of the data entity.
[0022] S6: Recursively execute S3 to S5 to obtain a list of disambiguated data entities. All data entities to be modeled in the list of growth feature entity candidates are sequentially extracted, matched, and disambiguated to obtain a list of disambiguated data entities.
[0023] S7: Based on the list of disambiguated data entities, generate a remote management strategy for the plant. By analyzing the data in the list of disambiguated data entities and combining the growth patterns of the plant, a remote management strategy including irrigation adjustment, fertilization plan, light control, etc. is developed.
[0024] Embodiment 1: Noise filtering and cleaning of growth environment data to obtain purified growth environment data. During the collection process, the growth environment data may be mixed with various types of noise data due to the precision limitations of the sensor itself, sudden disturbances from the external environment, or signal loss during data transmission. These noise data appear as abnormal values that do not conform to the normal data distribution rules. For example, a temperature sensor may record a temperature value much higher than the actual environment when it is suddenly exposed to direct sunlight. A soil moisture sensor may produce a sudden increase in humidity data if it comes into contact with water droplets. Noise filtering and cleaning requires a pre-set algorithm to detect each point of the original data and identify these abnormal data. For a single isolated abnormal value, interpolation of adjacent data can be used for replacement, such as using the average value of multiple normal data before and after the abnormal value. For a continuous abnormal data segment, it needs to be determined whether it is caused by sensor failure. If confirmed, the data segment can be removed or filled with historical data under similar conditions. After such processing, the growth environment data originally containing noise is purified to form purified growth environment data.
[0025] Time stamp alignment processing of purified growth environment data to obtain a set of growth environment data sequences. Different terminal devices may have slight differences in their internal clocks when collecting data, resulting in data collected at the same actual time point being marked with different time stamps. In addition, different types of sensors may have different collection frequencies, such as temperature sensors that may collect data every 10 minutes, while soil nutrient sensors may collect data every hour. Time stamp alignment processing first needs to unify the time standard and convert the time stamps of all data to the same time zone standard time format. Then, according to the actual application requirements, a unified time interval is set, such as every 30 minutes as a time unit, and the purified growth environment data is resampled. For cases where there is no corresponding collected data at the set time point, linear interpolation or spline interpolation methods are used to calculate the estimated value at that time point. For cases where there are multiple source data at the same time point, the average value is taken or the sensor priority is selected. Through such processing, the originally inconsistent time distribution data is arranged into growth environment data sequences arranged in a unified time interval, each data sequence corresponding to an environmental parameter, such as temperature data sequence, humidity data sequence, etc. These sequences together constitute a set of growth environment data sequences.
[0026] Each growth environment data sequence in the set is subjected to feature entity recognition and classification to obtain a list of growth feature entity candidates and non-feature data segments. Feature entity recognition is the process of identifying information units with specific meanings and possibly related to the growth status of the plant from the data sequence. This process requires the combination of the biological characteristics of the plant growth and the pre-setting of a series of feature recognition rules. For example, in the temperature data sequence, "the average temperature is higher than 25°C for 5 consecutive days" can be identified as a feature entity; in the soil moisture data sequence, "the single-day moisture drop amplitude exceeds 30%" can also be identified as a feature entity. In the identification process, the data sequence is traversed through a sliding window, and the data in the window is subjected to statistical analysis and pattern matching. When the pre-set rules are met, the corresponding feature entity is marked, and its start and end time, value range, etc. are recorded. For parts of the data sequence that cannot be identified as feature entities, such as random fluctuations in small value changes, scattered data points without obvious rules, etc., they are classified as non-feature data segments. All identified feature entities are summarized to form a list of growth feature entity candidates, while non-feature data segments are stored separately as context information for subsequent processing. Through such classification, the growth environment data is divided into feature entities with potential analysis value and non-feature data that helps understand the context, providing a structured data foundation for subsequent modeling and analysis.
[0027] Embodiment 2: Take the first to-be-modeled data entity as the association keyword, perform semantic association query in the cloud collaborative knowledge base to obtain a plurality of candidate data entity entries matched with the first to-be-modeled data entity.
[0028] The process of semantic association query starts with semantic analysis of the first to-be-modeled data entity. The first to-be-modeled data entity usually contains multiple semantic elements, for example, when the first to-be-modeled data entity is "Malus halliana seedling stage continuous 7-day soil pH value lower than 5.0", its semantic elements include "Malus halliana" "seedling stage" "continuous 7 days" "soil pH value" "lower than 5.0". The analysis process extracts these elements one by one through natural language processing technology and performs part-of-speech tagging and semantic classification to clearly define the role and meaning of each element in the data entity.
[0029] The cloud collaborative knowledge base contains multi-dimensional structured and unstructured data. The structured data includes environmental parameter threshold table of different growth stages of Malus halliana, historical growth record database, disease and pest occurrence and environmental factor correlation table, etc.; the unstructured data includes Malus halliana cultivation technology documents, expert experience notes, research contents on Malus halliana growth in academic papers, etc. The knowledge base adopts a distributed storage architecture, and different types of data are stored in corresponding sub-libraries, and cross-library association is established through a semantic association graph. The nodes in the semantic association graph represent various entities, such as “growth stage”, “soil parameter”, “climatic condition”, etc., and the edges represent the association between the nodes, such as “influence”, “contain”, “correspond”, etc.
[0030] The query process first constructs a query vector based on the parsed semantic elements. Each dimension of the query vector corresponds to a feature value of a semantic element, for example, “seedling stage” corresponds to a specific value of the growth stage dimension, and “pH value less than 5.0” corresponds to the range value of the soil acidity and alkalinity dimension. Subsequently, the query vector is used for similarity retrieval in the index system of the cloud collaborative knowledge base. The index system adopts an inverted index structure, and the data entities in the knowledge base are segmented and indexed according to semantic elements, so that the query can quickly locate the possibly relevant data entity entries.
[0031] In the retrieval process, not only the directly matched semantic elements are considered, but also the query range is expanded through the semantic association graph. For example, when the query involves “soil pH value”, the system will automatically associate related concepts such as “soil acidity and alkalinity” and “soil microbial activity”, and retrieve data entity entries containing these concepts. At the same time, the system will preliminarily screen the retrieved candidate data entity entries and exclude obviously irrelevant entries, such as excluding data entries applicable to adult trees from the query results for Malus halliana seedling stage.
[0032] The screened candidate data entity entries need to be sorted through semantic similarity calculation. The calculation process is based on a pre-trained word vector model, which converts the candidate entries and the query vector into vector representations in a high-dimensional space, and measures the semantic similarity by calculating the cosine distance between the vectors. The smaller the distance, the closer the semantic association between the candidate entry and the first data entity to be modeled. After sorting, the top N (N is a preset value) entries with the highest semantic similarity are taken as the multiple candidate data entity entries matched with the first data entity to be modeled.
[0033] In addition, the query process also considers the timeliness and reliability of the data entity entries. For historical data entries, the system marks their collection time and source, and gives priority to data entries collected recently and from authoritative sources; for expert experience entries, their applicability is evaluated in combination with practical application records. Through such processing, the obtained candidate data entity entries not only match the first data entity to be modeled in semantics, but also have high reference value.
[0034] The final multiple candidate data entity entries cover different types of information, such as may include the suitable soil pH value range for the seedling stage of Chinese flowering crabapple, the historical growth record when the soil pH value is too low, the description of the influence of soil acidification on seedlings in related research, etc. These entries collectively constitute the basis data for subsequent collaborative disambiguation processing, providing multiple reference information for accurately analyzing the meaning of the first data entity to be modeled.
[0035] Embodiment 3: Taking non-feature data segments as context information, performing semantic-level collaborative aggregation analysis on multiple candidate data entity entries to obtain disambiguation data entity query response semantic aggregation encoding features, the specific process is as follows: Taking non-feature data segments as context information, performing semantic understanding on the context information to obtain context information semantic encoding vectors of the data entity to be modeled. Non-feature data segments contain various information related to the growth environment of Chinese flowering crabapple but not recognized as feature entities, which may be scattered environmental parameter records, numerical fluctuations without obvious rules, etc. Semantic understanding of these context information needs to be converted into a form that can be processed by a computer. First, the context information is processed by word segmentation, which splits continuous text or numerical sequences into independent semantic units, such as "May 10, 2024" "breeze" "intermittent rainfall" etc. Then, through word embedding technology, each semantic unit is mapped to a low-dimensional vector space to form an initial vector representation. These vectors can capture the potential associations between semantic units, for example, "rainfall" and "humidity increase" will have a closer distance in the vector space.
[0036] Using the data entity to be modeled embedding matrix to perform semantic embedding encoding on the context information of the data entity to be modeled to obtain the context sequence of the context information semantic embedding encoding vector of the data entity to be modeled. The data entity to be modeled embedding matrix is a parameter matrix trained by a large amount of text, and its dimension is set according to the semantic complexity. Each semantic unit is multiplied by this matrix to obtain the corresponding semantic embedding encoding vector. Arranging these vectors in the original order of the context information forms the context sequence. For example, the vectors corresponding to "intermittent rainfall" and "breeze" are arranged in the order of appearance to form a sequence reflecting the temporal relationship of this context.
[0037] The context sequence of the semantic embedding coding vector of the context information of the data entity to be modeled is coded based on a bidirectional GRU model to obtain a semantic coding vector of the context information of the data entity to be modeled. The bidirectional GRU model includes two parts of a forward GRU and a backward GRU. The forward GRU starts from the starting position of the context sequence and sequentially processes each semantic embedding coding vector to capture the forward timing characteristics in the sequence; the backward GRU starts from the end of the sequence and reversely processes each vector to capture the reverse timing characteristics. In the processing, each GRU unit updates the current hidden state according to the current input vector and the hidden state at the last time, so as to integrate the historical information. Finally, the hidden states of the forward GRU and the backward GRU at the last time step are spliced to obtain a coding vector containing complete semantic information of the context information.
[0038] The candidate data entity item embedding matrix is used to respectively structure code each candidate data entity item in the plurality of candidate data entity items to obtain a plurality of candidate data entity item semantic embedding coding vectors. The candidate data entity item embedding matrix is similar to the data entity embedding matrix to be modeled, but is optimized for the characteristics of the candidate data entity item. After each candidate data entity item is processed by word segmentation and structuring, it is converted into a fixed-dimensional semantic embedding coding vector by multiplication with the matrix. These vectors can reflect the semantic characteristics of the candidate item itself, such as the vector corresponding to "Haitang flower suitable temperature 15-25℃" containing temperature range, plant species and other information.
[0039] The data entity semantic collaborative aggregation coding is performed on the semantic coding vector of the context information of the data entity to be modeled and the plurality of candidate data entity item semantic embedding coding vectors to obtain a disambiguation data entity query response semantic aggregation coding vector as a disambiguation data entity query response semantic aggregation coding feature. The core of the collaborative aggregation coding is to fuse the context information and the candidate item information to form a comprehensive semantic representation. First, the correlation degree between the semantic coding vector of the context information of the data entity to be modeled and each candidate data entity item semantic embedding coding vector is calculated, and the correlation degree is calculated by vector dot product, as follows:
[0040] wherein, represents the correlation degree of the i th candidate data entity item and the context information, represents the semantic coding vector of the context information of the data entity to be modeled, represents the i th candidate data entity item semantic embedding coding vector, and represents the dot product operation of the vector.
[0041] Based on these correlation degrees, the candidate data entity item semantic embedding coding vectors are weighted to make the candidate item vectors with high context correlation degrees occupy a larger proportion in the aggregation process. Subsequently, the weighted candidate item vectors are fused with the context semantic coding vectors, and a non-linear transformation is performed through a multi-layer neural network to finally obtain a disambiguated data entity query response semantic aggregation coding vector. The vector integrates the context information and the semantic features of multiple candidate data entity items, providing a rich information foundation for subsequent semantic decoding.
[0042] The process of performing data entity semantic collaborative aggregation coding on the context information semantic coding vector of the data entity to be modeled and the semantic embedding coding vectors of multiple candidate data entity items to obtain a disambiguated data entity query response semantic aggregation coding vector is as follows: The context information semantic coding vector of the data entity to be modeled is enhanced based on an attention mechanism to obtain a context information semantic reinforced coding vector of the data entity to be modeled. The attention mechanism assigns appropriate weights to different parts of the context information by analyzing the correlation of the parts with the current task. For example, when the context information includes "continuous rainy weather" and "soil drainage is good", and the data entity to be modeled is soil moisture, the weight of "continuous rainy weather" will be increased, and the weight of "soil drainage is good" will also be adjusted according to its influence on the moisture. In this way, the features in the semantic coding vector that are closely related to the data entity to be modeled are reinforced, while the features with weak correlation are weakened, and the resulting semantic reinforced coding vector highlights the key context information.
[0043] The context information semantic reinforced coding vector of the data entity to be modeled and each of the candidate data entity item semantic embedding coding vectors in the multiple candidate data entity item semantic embedding coding vectors are subjected to data entity monomer semantic query to obtain a set of to-be-modeled-candidate data entity monomer semantic query score coding vectors. The data entity monomer semantic query is implemented by constructing a specific neural network module, which receives the semantic reinforced coding vector and the single candidate data entity item semantic embedding coding vector as input, performs a non-linear transformation on the two vectors through a multi-layer perceptron, and outputs a score coding vector representing the semantic matching degree of the two vectors. For example, when the data entity to be modeled is "soil moisture of Malus huangshanensis seedling stage" and the candidate data entity item is "soil moisture of seedling stage is maintained at 60%-70%", the output score coding vector will reflect a high matching degree; if the candidate item is "soil moisture of adult tree is 50%-60%", the matching degree of the score coding vector is lower.
[0044] Based on the feature set distribution characteristics of the set of score encoding vectors of the to-be-modeled-candidate data entity monomer semantic query, the set of score encoding vectors of the to-be-modeled-candidate data entity monomer semantic query is adaptively aggregated and analyzed based on a gating modulation to obtain a disambiguated data entity query response semantic aggregation encoding vector. The feature set distribution characteristics are the distribution patterns of the score encoding vectors in the feature space, for example, some vectors are concentrated in the high matching degree area, and some vectors are dispersed in the low matching degree area. The gating modulation mechanism processes these score encoding vectors through a gating network, and the gating network dynamically generates a gating coefficient according to the distribution characteristics of each score encoding vector. The size of the gating coefficient determines the contribution degree of the score encoding vector in the aggregation process. For the score encoding vectors distributed in the high matching degree area, the gating network gives a larger gating coefficient; for the vectors distributed in the low matching degree area, a smaller gating coefficient is given. The adaptive aggregation analysis obtains a comprehensive aggregation encoding vector by multiplying each score encoding vector by its corresponding gating coefficient and then accumulating. For example, if the gating coefficients of three score encoding vectors are 0.8, 0.15 and 0.05 respectively, the aggregation process mainly depends on the information of the first score encoding vector, while the secondary information of the last two vectors is also considered, and the disambiguated data entity query response semantic aggregation encoding vector finally formed can comprehensively reflect the overall semantic association between multiple candidate data entity entries and the to-be-modeled data entity in the context information.
[0045] The to-be-modeled data entity context information semantic reinforcement encoding vector and each candidate data entity item semantic embedding encoding vector in the plurality of candidate data entity item semantic embedding encoding vectors are input into a data entity monomer semantic collaborative query unit to obtain a set of initial to-be-modeled-candidate data entity monomer semantic query score encoding vectors. The data entity monomer semantic collaborative query unit includes a plurality of parallel sub-networks, each sub-network processes a group of to-be-modeled vectors and candidate vectors, and the sub-networks inside capture local association features between the two vectors through an attention mechanism, while alleviating the gradient disappearance problem of deep networks through a residual connection, ensuring that the output initial score encoding vector can accurately reflect the local semantic matching situation.
[0046] The set of initial to-be-modeled-candidate data entity monomer semantic query score encoding vectors is dynamically optimized and mapped to obtain a set of to-be-modeled-candidate data entity monomer semantic query score encoding vectors. The dynamic optimization and mapping is implemented through a dynamic adjustment network. The network adaptively adjusts the feature dimensions of the vectors according to the overall distribution of the initial score encoding vectors, for example, compresses redundant dimensions, expands key dimensions, and at the same time, through batch normalization processing, makes the distribution of the vectors more stable, and reduces the influence of outliers on the subsequent aggregation process. The score encoding vectors after the optimization and mapping have better feature expression ability while retaining the core semantic information, which can improve the accuracy of subsequent aggregation analysis.
[0047] In embodiment 5, based on the feature set self-distribution characteristics of the set of to-be-modeled-candidate data entity monomer semantic query score encoding vectors, the set of to-be-modeled-candidate data entity monomer semantic query score encoding vectors is adaptively aggregated and analyzed based on the gating modulation to obtain a disambiguated data entity query response semantic aggregation encoding vector. The process is as follows: Based on the feature set self-distribution characteristics of the set of to-be-modeled-candidate data entity monomer semantic query score encoding vectors, the monomer semantic matching degree of each to-be-modeled-candidate data entity monomer semantic query score encoding vector in the set of to-be-modeled-candidate data entity monomer semantic query score encoding vectors is determined to obtain a set of to-be-modeled-candidate data entity monomer semantic matching degrees. The feature set self-distribution characteristics are reflected in the distribution density, dispersion degree and clustering pattern of the score encoding vectors in the high-dimensional feature space. For example, some score encoding vectors may be clustered in a certain specific area, forming a high-density cluster, indicating that the candidate data entity entries corresponding to these vectors have similar semantic matching levels with the to-be-modeled data entity; while other vectors may be scattered in different areas, showing lower consistency. By analyzing these distribution characteristics, the average distance of each score encoding vector from the same type of vector, the density in the feature space and other indicators are calculated, and these indicators are converted into monomer semantic matching degrees. The matching degree is represented by a numerical value, and the higher the numerical value, the better the semantic fit between the candidate data entity entry corresponding to the vector and the to-be-modeled data entity.
[0048] The set of single-body semantic matching degrees of the to-be-modeled-candidate data entity is input into a relationship gate agent module to obtain a set of self-attention weights of the single-body query semantic of the to-be-modeled-candidate data entity. The relationship gate agent module includes a multi-layer neural network structure and can capture the mutual relationship between the single-body semantic matching degrees. For example, when the single-body semantic matching degrees of two candidate data entity entries are high and involve the same growth parameter, the module identifies this association and cooperatively adjusts the self-attention weights when generating the self-attention weights. The module first performs a nonlinear transformation on the input single-body semantic matching degrees, maps them to a new feature space, then filters out matching degree information with significant association through a gating mechanism, and calculates the self-attention weights based on this information. The numerical size of the self-attention weight is related to the importance of the corresponding candidate data entity entry, and the higher the weight, the more attention should be given to the entry in the aggregation process.
[0049] The set of single-body semantic query score encoding vectors of the to-be-modeled-candidate data entity is aggregated based on the set of single-body query semantic self-attention weights of the to-be-modeled-candidate data entity to obtain a semantic aggregation encoding vector of the disambiguated data entity query response. The aggregation process adopts a weighted summation method, multiplies each score encoding vector with its corresponding self-attention weight, and then adds all the product results to form a comprehensive encoding vector. For example, if the self-attention weight of a certain candidate data entity entry is 0.3 and its score encoding vector is [0.2, 0.5, 0.8], the contribution of the vector in the aggregation is [0.06, 0.15, 0.24], and the final semantic aggregation encoding vector is obtained by adding the contributions of all candidate entries. The vector integrates the semantic information of multiple candidate data entity entries and reflects the relative importance of different entries through the weight.
[0050] The semantic aggregation encoding vector of the disambiguated data entity query response is decoded based on a Transformer model to obtain the disambiguated data entity of the first to-be-modeled data entity. The decoder of the Transformer model is composed of multiple stacked decoding layers, each of which includes a multi-head self-attention mechanism and a feedforward neural network. During decoding, the decoder first receives the semantic aggregation encoding vector as the initial input, focuses on the key semantic features in the vector through the multi-head self-attention mechanism, and captures the temporal relationship in the sequence by combining the position encoding information. The feedforward neural network then nonlinearly transforms the output of the attention mechanism to further extract deep semantic features. After multiple rounds of decoding, the model converts the abstract encoding vector into a specific and unambiguous data entity description. For example, if the original to-be-modeled data entity has a vague expression of "humidity anomaly", the decoded disambiguated data entity can be "soil humidity below 40% for 3 consecutive days".
[0051] Based on the disambiguated data entity list, the management strategy of the remote management of the Chinese flowering crabapple is generated, specifically: the disambiguated data entity list is input into a pre-trained management strategy generation model to obtain the management strategy of the remote management of the Chinese flowering crabapple. The pre-trained management strategy generation model is trained through a large amount of Chinese flowering crabapple planting data and management cases, and can understand various information in the disambiguated data entity list, such as specific parameters such as soil humidity, illumination time, temperature change, etc. In the processing process, the model compares these data with the built-in Chinese flowering crabapple growth model, analyzes the differences between the current growth environment and the ideal state, and then generates the corresponding management strategy according to the preset decision logic. For example, when the disambiguated data entity in the list contains "insufficient illumination during flowering period", the model will generate the strategy of "turning on the light supplementing device for 3 hours per day"; when "low soil nitrogen content" is detected, the strategy of "applying nitrogen, phosphorus and potassium compound fertilizer, 50 kg per mu" is generated. The generated strategy covers irrigation, fertilization, light adjustment, pest prevention and other aspects, forming a complete remote management scheme.
[0052] It should be noted that in this paper, relationship terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between the entities or operations. Moreover, the terms "include", "contain" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or equipment including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or equipment.
[0053] Although the embodiments of the present application have been shown and described, it can be understood by those of ordinary skill in the art that various changes, modifications, replacements and variations can be made to the embodiments without departing from the principles and spirits of the present application, and the scope of the present application is defined by the appended claims and their equivalents.
Claims
1. A method for remote management and data modeling of crabapple flowers based on cloud collaborative architecture, characterized in that: include: S1: Receive the crabapple flower growth environment data collected by the terminal device; S2: Perform standardized preprocessing and multi-source data fusion analysis on the growth environment data to obtain a candidate list of growth feature entities and non-feature data fragments; S3: extracting a first data entity to be modeled from the growth feature entity candidate list; S4: extracting multiple candidate data entity entries matching the first data entity to be modeled from the cloud collaborative knowledge base; S5: Inputting multiple candidate data entity entries and non-feature data fragments matching the first data entity to be modeled into a collaborative disambiguation model to obtain a disambiguated data entity of the first data entity to be modeled, including: using the non-feature data fragments as context information, performing semantic-level collaborative aggregation analysis on the multiple candidate data entity entries to obtain semantic aggregation encoding features of disambiguated data entity query responses; and semantically decoding the semantic aggregation encoding features of the disambiguated data entity query responses to obtain a disambiguated data entity of the first data entity to be modeled; S6: Loop through S3 to S5 to obtain a list of disambiguated data entities; S7: Generate a remote management strategy for crabapple flowers based on the disambiguation data entity list.
2. The method for remote management and data modeling of crabapple flowers based on cloud collaborative architecture according to claim 1, characterized in that S2 include: Perform noise filtering and cleaning on the growth environment data to obtain purified growth environment data; Performing time stamp alignment processing on the purified growth environment data to obtain a set of growth environment data sequences; Feature entity recognition and classification are performed on each growth environment data sequence in the set of growth environment data sequences to obtain a growth feature entity candidate list and non-feature data segments.
3. The method for remote management and data modeling of crabapple flowers based on cloud collaborative architecture according to claim 2, characterized in that: S4 includes: using the first data entity to be modeled as an associated keyword, performing a semantic association query in a cloud collaborative knowledge base to obtain a plurality of candidate data entity entries that match the first data entity to be modeled.
4. The method for remote management and data modeling of crabapple flowers based on cloud collaborative architecture according to claim 3, characterized in that: Using non-feature data fragments as context information, a semantic-level collaborative aggregation analysis is performed on multiple candidate data entity entries to obtain semantic aggregation encoding features of disambiguated data entity query responses, including: Using non-feature data fragments as context information, the context information is semantically understood to obtain the semantic encoding vector of the context information of the data entity to be modeled; Using the candidate data entity entry embedding matrix, structurally encoding each candidate data entity entry in the plurality of candidate data entity entries is performed to obtain semantic embedding encoding vectors of the plurality of candidate data entity entries; The semantic encoding vector of the context information of the data entity to be modeled and the semantic embedding encoding vectors of multiple candidate data entity entries are subjected to data entity semantic collaborative aggregation encoding to obtain the disambiguated data entity query response semantic aggregation encoding vector as the disambiguated data entity query response semantic aggregation encoding feature.
5. The method for remote management and data modeling of crabapple flowers based on cloud collaborative architecture according to claim 4 is characterized in that: Using non-feature data fragments as context information, the context information is semantically understood to obtain the semantic encoding vector of the context information of the data entity to be modeled, including: Use the embedding matrix of the data entity to be modeled to perform semantic embedding encoding on the context information of the data entity to be modeled, so as to obtain a context sequence of the semantic embedding encoding vector of the context information of the data entity to be modeled; The context sequence of the semantic embedding encoding vector of the context information of the data entity to be modeled is subjected to context semantic encoding based on the bidirectional GRU model to obtain the semantic encoding vector of the context information of the data entity to be modeled.
6. The method for remote management and data modeling of crabapple flowers based on cloud collaborative architecture according to claim 5, characterized in that: Performing data entity semantic collaborative aggregation encoding on the semantic encoding vector of the context information of the data entity to be modeled and the semantic embedding encoding vectors of multiple candidate data entity entries to obtain a disambiguated data entity query response semantic aggregation encoding vector, including: Perform feature enhancement based on the attention mechanism on the semantic encoding vector of the context information of the data entity to be modeled to obtain the semantic enhanced encoding vector of the context information of the data entity to be modeled; Performing a data entity monomer semantic query on the semantic enhancement encoding vector of the context information of the data entity to be modeled and the semantic embedding encoding vectors of each candidate data entity entry in the multiple candidate data entity entry semantic embedding encoding vectors to obtain a set of semantic query score encoding vectors of the candidate data entity to be modeled; Based on the self-distribution characteristics of the feature set of the set of semantic query score encoding vectors of the entity to be modeled and the candidate data, an adaptive aggregation analysis based on gated modulation is performed on the set of semantic query score encoding vectors of the entity to be modeled and the candidate data to obtain the semantic aggregation encoding vector of the disambiguated data entity query response.
7. The method for remote management and data modeling of crabapple flowers based on cloud collaborative architecture according to claim 6, characterized in that: The semantic enhancement encoding vector of the context information of the data entity to be modeled and each of the semantic embedding encoding vectors of the candidate data entity entries in the semantic embedding encoding vectors of the multiple candidate data entity entries are used to perform a data entity monomer semantic query to obtain a set of semantic query score encoding vectors of the candidate data entity to be modeled, including: The context information semantic enhancement encoding vector of the data entity to be modeled and the semantic embedding encoding vectors of each candidate data entity entry in the multiple candidate data entity entry semantic embedding encoding vectors are respectively input into the data entity monomer semantic collaborative query unit to obtain a set of initial semantic query score encoding vectors of the candidate data entity to be modeled; The set of initial to-be-modeled-candidate data entity monomer semantic query score encoding vectors is subjected to dynamic feature optimization mapping to obtain a set of to-be-modeled-candidate data entity monomer semantic query score encoding vectors.
8. The method for remote management and data modeling of crabapple flowers based on cloud collaborative architecture according to claim 7, characterized in that: Based on the self-distribution characteristics of the feature set of the set of semantic query score encoding vectors of the entity to be modeled and the candidate data, an adaptive aggregation analysis based on gated modulation is performed on the set of semantic query score encoding vectors of the entity to be modeled and the candidate data to obtain a semantic aggregate encoding vector of the disambiguated data entity query response, including: Based on the self-distribution characteristics of the feature set of the set of single semantic query score encoding vectors of the entity to be modeled-candidate data, the single semantic matching degree of each single semantic query score encoding vector of the entity to be modeled-candidate data in the set of single semantic query score encoding vectors of the entity to be modeled-candidate data is determined to obtain a set of single semantic matching degrees of the entity to be modeled-candidate data; The set of semantic matching degrees between the entity to be modeled and the candidate data entity is input into the relational gating agent module to obtain the set of query semantic self-attention weights between the entity to be modeled and the candidate data entity; Based on the set of semantic self-attention weights of the to-be-modeled-candidate data entity monomer query, the set of to-be-modeled-candidate data entity monomer semantic query score encoding vectors is aggregated to obtain the disambiguated data entity query response semantic aggregate encoding vector.
9. The method for remote management and data modeling of crabapple flowers based on cloud collaborative architecture according to claim 8, characterized in that: Semantically decoding the semantic aggregation encoding features of the disambiguation data entity query response to obtain a disambiguation data entity of the first data entity to be modeled, including: The semantic aggregation encoding vector of the disambiguation data entity query response is semantically decoded based on the Transformer model to obtain the disambiguation data entity of the first data entity to be modeled.
10. The method for remote management and data modeling of crabapple flowers based on cloud collaborative architecture according to claim 8, characterized in that: Based on the disambiguation data entity list, a remote management strategy for Begonia flowers is generated, including: The disambiguated data entity list is input into the pre-trained management policy generation model to obtain the remote management policy for crabapple flowers.
Citation Information
Cited By
Method, system and equipment for analyzing abnormal phenomenon of semiconductor structure
CN121542487A