Sample data labeling method and system applied to water ecological laboratory

By constructing environmental factor response curves and time-varying features of biological attributes, a nonlinearly coupled specimen feature spectrum is generated, which solves the problem that the relationship between environmental parameters and biological attributes in aquatic ecological laboratories has not been deeply explored. This achieves close integration of specimen feature data and semantically rich annotation results, supporting ecological evolution analysis.

CN121637148BActive Publication Date: 2026-08-04BEIJING NORMAL UNIVERSITY +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING NORMAL UNIVERSITY
Filing Date
2025-11-14
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

In the data annotation methods of the aquatic ecology laboratory, the relationship between environmental parameters and biological attributes has not been deeply explored, resulting in weak data correlation, one-sided semantic information expression, and difficulty in supporting ecological evolution analysis.

Method used

By acquiring specimen samples and multi-source environmental metadata, we construct environmental factor response curves and time-varying features of biological attributes, generate an initial specimen feature spectrum containing nonlinear coupling relationships, and perform feature space mapping, multimodal semantic verification, and entity relationship reasoning to generate context-enhanced annotation vectors and track the semantic change process of specimen features.

Benefits of technology

Dynamically capturing the intrinsic relationship between environmental parameters and biological attributes enhances the fusion tightness and semantic accuracy of specimen feature data, enriches the knowledge content of annotation results, and fully records the ecological evolution process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121637148B_ABST
    Figure CN121637148B_ABST
Patent Text Reader

Abstract

The application provides a specimen data labeling method and system applied to an aquatic ecological laboratory, through acquiring specimen samples collected by the aquatic ecological laboratory and corresponding multi-source environmental metadata, environmental factor response curves are constructed and biological attribute time-varying characteristics are extracted for environmental parameter sequences and biological attribute observation data, multi-source environmental metadata and specimen sample attributes are coupled, and initial specimen characteristic spectra are generated; context-enhanced labeling vectors are generated through feature space mapping, multi-modal semantic verification and entity relationship reasoning; intermediate evolution label sets are generated through history trajectory similarity comparison and multi-dimensional evolution trend prediction; and labeling labels with ecological evolution semantics are generated through multi-dimensional semantic consistency verification and evolution path optimization integration of the intermediate evolution label sets. Through the method, the recording ability of the labeling result to the ecological evolution process of the specimen can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing, and more specifically, to a method and system for labeling specimen data in an aquatic ecology laboratory. Background Technology

[0002] In the field of aquatic ecology research, specimen data annotation is the process of describing the characteristics and labeling information of aquatic biological specimens collected in the laboratory. The annotation results provide fundamental data support for research such as species identification and ecological evolution analysis. Currently, specimen data annotation in aquatic ecology laboratories typically relies on manual or semi-automatic methods to label the morphological characteristics or basic attributes of specimens in isolation. During the annotation process, environmental parameters and biological attributes are often processed using linear superposition or independent recording methods, lacking in-depth exploration of the dynamic relationship between the two. Furthermore, the annotation results are mostly presented in the form of static labels, making it difficult to reflect the evolutionary process of specimen characteristics over time. This traditional annotation method results in weak data correlation and one-sided semantic information expression, making it difficult to support the systematic analysis of the ecological evolution process of specimens and affecting the accuracy and depth of data application in subsequent aquatic ecology research. Summary of the Invention

[0003] This invention provides a method and system for labeling specimen data in aquatic ecological laboratories.

[0004] In a first aspect, embodiments of the present invention provide a method for annotating specimen data in an aquatic ecological laboratory. The method includes: acquiring specimen samples collected in the aquatic ecological laboratory and corresponding multi-source environmental metadata, wherein the multi-source environmental metadata includes environmental parameter sequences and biological attribute observation data during the specimen collection process; constructing environmental factor response curves and extracting time-varying features of biological attributes from the environmental parameter sequences and biological attribute observation data; coupling multi-source environmental metadata and specimen sample attributes based on the environmental factor response curves and time-varying features of biological attributes to generate an initial specimen feature spectrum containing a nonlinear coupling relationship between environmental parameter trajectories and biological attribute response features; performing feature space mapping, multimodal semantic verification, and entity relationship reasoning on the initial specimen feature spectrum to generate a context-enhanced annotation vector containing the semantic association strength and evolutionary feature weights of the specimen samples in a multi-level knowledge system; performing historical trajectory similarity comparison and multi-dimensional evolutionary trend prediction on the context-enhanced annotation vector to generate an intermediate evolutionary label set recording the semantic change process of specimen features at different time scales; and integrating the intermediate evolutionary label set through multi-dimensional semantic consistency verification and evolutionary path optimization to generate annotation labels with ecological evolutionary semantics.

[0005] Secondly, embodiments of the present invention provide a computer system, including: a memory storing a computer program; and a processor for loading the computer program to implement the specimen data annotation method described above for use in an aquatic ecology laboratory.

[0006] The specimen data annotation method provided by this invention, applied to aquatic ecological laboratories, constructs environmental factor response curves and extracts time-varying features of biological attributes. It couples multi-source environmental metadata with specimen sample attributes to generate an initial specimen feature spectrum containing nonlinear coupling relationships. This dynamically captures the intrinsic correlation between environmental parameter trajectories and biological attribute response features, avoiding the loose data association problem caused by the independent processing of environmental data and biological attributes in traditional annotation methods, thus improving the fusion tightness of specimen feature data. Furthermore, by performing feature space mapping, multimodal semantic verification, and entity relationship reasoning on the initial specimen feature spectrum, it generates context enhancement containing semantic association strength and evolutionary feature weights across a multi-level knowledge system. The context-enhanced annotation vectors can deeply bind specimen features with the aquatic ecological knowledge system. Through multi-dimensional semantic verification, the consistency of annotation semantics is ensured, avoiding the semantic bias caused by single-dimensional annotation, and improving the semantic accuracy and knowledge richness of the annotation results. By comparing the historical trajectory similarity and predicting the multi-dimensional evolution trend of the context-enhanced annotation vectors, and combining multi-dimensional semantic consistency verification and evolution path optimization to integrate the intermediate evolution tag set, annotation tags with ecological evolution semantics are generated. This can track the semantic change process of specimen features at different time scales, avoid the limitation of static annotation that only records the instantaneous state, and improve the ability of the annotation results to fully record the ecological evolution process of the specimen. Attached Figure Description

[0007] Figure 1 This is a flowchart of a specimen data annotation method for use in an aquatic ecology laboratory, provided by an embodiment of the present invention.

[0008] Figure 2 This is a schematic diagram of the composition of a computer system provided in an embodiment of the present invention. Detailed Implementation

[0009] Please see Figure 1 The flowchart below illustrates a method for labeling specimen data in an aquatic ecology laboratory, as provided in an embodiment of the present invention. This method can be executed by a computer system and may include the following steps: Step S100: Obtain the specimen samples collected by the aquatic ecology laboratory and the corresponding multi-source environmental metadata, which includes environmental parameter sequences and biological attribute observation data during the specimen collection process.

[0010] Specimen samples are biological samples collected from aquatic ecosystems. Multi-source environmental metadata is a collection of various types of data related to specimen collection. Among these, environmental parameter sequences record the changes in environmental factors over time during specimen collection, such as the specific values ​​of environmental indicators like water temperature, pH, and dissolved oxygen concentration at different time points, reflecting the dynamic characteristics of the specimen's living environment. Biological attribute observation data are the observation and recording of the biological characteristics of the specimen itself, such as individual size, reproductive capacity, and growth rate, reflecting the organism's growth and development status in its environment.

[0011] Step S200: Construct environmental factor response curves and extract time-varying features of biological attributes from environmental parameter sequences and biological attribute observation data. Based on the environmental factor response curves and time-varying features of biological attributes, couple multi-source environmental metadata and specimen sample attributes to generate an initial specimen feature spectrum containing the nonlinear coupling relationship between environmental parameter trajectories and biological attribute response features.

[0012] Environmental factor response curves describe the relationship between biological attributes and environmental factors, reflecting the response patterns of organisms under different environmental factor conditions. Time-varying characteristics of biological attributes are the changes in biological attributes over time, reflecting the dynamic adaptation process of organisms during environmental changes. Coupling involves associating and integrating multi-source environmental metadata with specimen sample attributes, making the relationship between the two closer and clearer, thereby comprehensively reflecting the characteristics of the specimen in the corresponding environment.

[0013] In one implementation, step S200 may include the following steps S210-S260: Step S210: Perform multi-scale periodic decomposition on the environmental parameter sequence to separate environmental factor components with different oscillation periods, calculate the energy proportion of each environmental factor component, and construct a periodic contribution matrix with environmental factor type and oscillation period as dimensions. Based on the periodic contribution matrix, screen the dominant environmental factor components.

[0014] Environmental factor components with different oscillation periods represent the changes in environmental parameters at different time scales, with each component containing environmental change information for its corresponding period. Energy percentage represents the proportion of energy of each environmental factor component in the total energy, reflecting its contribution to the entire environmental parameter sequence. The periodic contribution matrix is ​​constructed with environmental factor type as rows and oscillation period as columns. Elements in the matrix represent the energy percentage of the corresponding environmental factor in its corresponding periodic component. This matrix provides a clear understanding of the contributions of different environmental factors across different periods. Dominant environmental factor components are those with a large energy percentage in the periodic contribution matrix; these components may have a more significant impact on biological properties and are the focus of subsequent analysis. Multi-scale periodic decomposition can be achieved using methods such as Empirical Mode Decomposition (EMD) and wavelet decomposition.

[0015] In one implementation, step S210 may include the following steps S211-S216: Step S211: Obtain the initial mode number range of the pre-set multi-scale periodic decomposition, optimize the mode number with the goal of maximizing the sum of the kurtosis values ​​of each component after decomposition, and determine the optimal mode decomposition parameters.

[0016] The initial range of modes in multi-scale periodic decomposition is a range of modes pre-defined by researchers based on experience and data characteristics. Kurtosis is a statistical measure describing the distribution pattern of data, reflecting the degree of peaks in the data distribution. In multi-scale periodic decomposition, the optimization of the modes aims to maximize the sum of the kurtosis values ​​of the decomposed components because a larger sum of kurtosis values ​​indicates that the decomposed components better highlight the periodic characteristics of the data and can more accurately reflect the inherent structure of the environmental parameter sequence. The optimal mode decomposition parameters are the modes that maximize the sum of kurtosis values ​​obtained during the optimization process. Using these parameters for decomposition yields environmental factor components that better reflect the actual situation.

[0017] Step S212: Perform multi-scale periodic decomposition on the environmental parameter sequence based on the optimal mode decomposition parameters to obtain multiple independent environmental factor components. Each environmental factor component contains time series data with the corresponding oscillation period.

[0018] Once the optimal mode decomposition parameters are determined, they can be used to perform multi-scale periodic decomposition on the environmental parameter sequence. The purpose of multi-scale periodic decomposition is to decompose the environmental parameter sequence into multiple independent components, each representing an oscillation period. These environmental factor components contain time series data of the corresponding oscillation periods. By analyzing these components, a deeper understanding of the periodic variation characteristics of environmental parameters can be obtained.

[0019] Step S213: Calculate the energy value of each environmental factor component. The energy value is obtained by integrating the sum of squares of the component time series. Normalize the energy values ​​of each component to obtain the energy proportion distribution.

[0020] The energy value reflects the proportion of energy a component occupies in the entire environmental parameter series. When calculating the energy value, for the time series data of each environmental factor component, each data point is first squared, and then the squared series is integrated. Integration can be performed using numerical integration methods, such as the trapezoidal rule or Simpson's rule. After obtaining the energy value of each component, the sum of all component energy values ​​is calculated. Finally, the energy value of each component is divided by the total energy value to obtain the energy proportion of that component.

[0021] Step S214: Construct a periodic contribution matrix with environmental factor type as the row and oscillation period as the column. The elements of the periodic contribution matrix are the energy proportion of the corresponding environmental factor in the corresponding periodic component.

[0022] The periodic contribution matrix is ​​a two-dimensional matrix with environmental factor types as rows and oscillation periods as columns. Each element in the matrix represents the energy proportion of the corresponding environmental factor in its respective period component. By constructing the periodic contribution matrix, the contribution of different environmental factors under different oscillation periods can be visually displayed, which helps in analyzing the relationship between environmental factors and biological properties. When constructing the periodic contribution matrix, the specific values ​​of the environmental factor types and oscillation periods are first determined. Environmental factor types can include common environmental indicators such as water temperature, light intensity, and dissolved oxygen concentration; the oscillation period is determined based on the different components obtained from multi-scale periodic decomposition. Then, the energy proportion of each environmental factor in different oscillation period components is filled into the corresponding positions in the matrix.

[0023] Step S215: Obtain the preset contribution threshold and filter environmental factor components whose element values ​​in the periodic contribution matrix exceed the contribution threshold as candidate dominant components.

[0024] Candidate dominant components are environmental factor components whose element values ​​exceed a contribution threshold in the periodic contribution matrix. These components are likely to have a more significant impact on biological properties and are therefore selected for further analysis. When determining the preset contribution threshold, the research focus, data characteristics, and previous research experience can be considered. For example, if the research focuses more on environmental factors with a greater impact on biological properties, the contribution threshold can be set higher; if it is desired to cover more potentially influential environmental factors, the threshold can be set lower. Then, each element in the periodic contribution matrix is ​​iterated through, and environmental factor components whose element values ​​exceed the contribution threshold are selected as candidate dominant components.

[0025] Step S216: Perform redundancy analysis on the candidate dominant components, calculate the mutual information value between the candidate dominant components, and remove the candidate dominant components whose mutual information value exceeds the set value.

[0026] Redundancy analysis aims to remove redundant information from candidate dominant components, avoiding redundant analysis and interference. Mutual information, a measure of the correlation between two random variables, is used in this step to measure the correlation between candidate dominant components. If the mutual information value between two candidate dominant components exceeds a set value, it indicates a strong correlation, potentially containing redundant information. Therefore, one of the components needs to be removed to improve the efficiency and accuracy of subsequent analysis.

[0027] When performing redundancy analysis, the mutual information value between candidate dominant components is first calculated. An entropy-based method can be used to calculate this value. For example, the entropy of each candidate dominant component is calculated first; entropy represents the uncertainty of a random variable. Then, the joint entropy of the two candidate dominant components is calculated. Finally, according to the definition of mutual information, mutual information equals the sum of the entropies of the two components minus their joint entropy. After obtaining the mutual information value, it is compared with a set value. If the mutual information value exceeds the set value, one of the candidate dominant components is removed.

[0028] Step S220: Perform time series alignment on the biological attribute observation data, extract the biological attribute response sequences that match the timestamps of the dominant environmental factor components, perform cross-correlation analysis on different environmental factor components and biological attribute response sequences to calculate the lag response coefficients, and construct a lag response coefficient matrix with the dominant environmental factor components and biological attribute response sequences as the dimensions.

[0029] Time series alignment unifies biological attribute observation data and dominant environmental factor components along the time dimension, ensuring their timestamps correspond for subsequent analysis. Biological attribute response sequences are biological attribute observation data matched to the timestamps of the dominant environmental factor components, reflecting the organism's response to changes in the corresponding environmental factors. Cross-correlation analysis can calculate the lag response coefficients of the biological attribute response sequences relative to the dominant environmental factor components; these coefficients represent the delay time in the biological attribute's response to changes in the environmental factors. The lag response coefficient matrix is ​​constructed with the dominant environmental factor components as rows and the biological attribute response sequences as columns. Matrix elements represent the lag response coefficients of corresponding combinations. This matrix provides a visual understanding of the lag response relationships between different environmental factors and biological attributes.

[0030] In one implementation, step S220 may include the following steps S221-S225: Step S221: Extract the timestamp set of the environmental parameter sequence, and perform linear interpolation on the timestamps of the biological attribute observation data to keep the biological attribute observation data and the environmental parameter sequence synchronized in the time dimension.

[0031] The timestamp set of the environmental parameter sequence records the specific time points when the environmental parameter data was collected, reflecting the temporal order of environmental parameter changes. Linear interpolation estimates the data values ​​at unknown time points by performing linear fitting between known data points. Linear interpolation is applied to the timestamps of the biological attribute observation data to ensure that the biological attribute observation data and the environmental parameter sequence are synchronized in the time dimension. When extracting the timestamp set of the environmental parameter sequence, the timestamp corresponding to each data point is directly obtained from the data records of the environmental parameter sequence. Then, for the timestamps of the biological attribute observation data, when there are missing or mismatched cases, linear interpolation is used to handle them.

[0032] Step S222: Extract the biological attribute index sequence from the aligned biological attribute observation data, and extract the biological attribute index sequence within the corresponding time period as the biological attribute response sequence based on the time interval of the dominant environmental factor component.

[0033] The biological attribute index sequence is a series of biological attribute data extracted from aligned biological attribute observation data, reflecting the attribute changes of organisms within a specific time range. The time interval of the dominant environmental factor component is the time range corresponding to that environmental factor component, representing a specific period of time for environmental factor changes.

[0034] When extracting biological attribute indicator sequences from aligned biological attribute observation data, the corresponding indicator data are selected from the data based on the type of biological attribute the study focuses on. For example, if the study focuses on the growth rate of organisms, then relevant growth rate data are extracted from the aligned data as the biological attribute indicator sequence. Then, based on the time interval of the dominant environmental factor components, the start and end times for the extraction are determined, and data within the corresponding time period is extracted from the biological attribute indicator sequence to obtain the biological attribute response sequence.

[0035] Step S223: Perform sliding window cross-correlation analysis on each dominant environmental factor component and the biological attribute response sequence. The window length is dynamically adjusted according to the oscillation period of the environmental factor component, and the correlation coefficient under different lag times is calculated.

[0036] Sliding window cross-correlation analysis is used to analyze the correlation between two time series. It calculates the correlation between the two series within a fixed-length window by sliding the window across the time series. The window length is dynamically adjusted according to the oscillation period of the environmental factor component to better capture changes in the correlation between environmental factors and biological attributes. The correlation coefficients at different lag times reflect the degree of correlation between the biological attribute response sequence and the dominant environmental factor component at different delay times.

[0037] When performing sliding window cross-correlation analysis, the window length is first determined based on the oscillation period of the dominant environmental factor component. For example, if the oscillation period is short, the window length can be set relatively small; if the oscillation period is long, the window length needs to be set larger. Then, the window is simultaneously slid across the dominant environmental factor component and the biological attribute response sequence, and the correlation coefficient between the two sequences is calculated at each window position. For each window position, the lag time of the biological attribute response sequence relative to the dominant environmental factor component is also varied, and the correlation coefficient at different lag times is calculated.

[0038] Step S224: Record the lag time and correlation coefficient value corresponding to the maximum correlation coefficient, and determine the correlation coefficient value as the lag response coefficient.

[0039] The lag time corresponding to the maximum correlation coefficient represents the optimal response delay time of the biological attribute response sequence relative to the dominant environmental factor component, reflecting the delayed response characteristics of biological attributes to changes in environmental factors. The maximum correlation coefficient value is determined as the lag response coefficient because this coefficient can most effectively reflect the correlation strength between biological attributes and environmental factors.

[0040] After obtaining the correlation coefficients at different lag times through sliding window cross-correlation analysis, these correlation coefficients are compared to find the maximum value. The lag time and correlation coefficient value corresponding to the maximum correlation coefficient are recorded. Then, this maximum correlation coefficient value is used as the lagged response coefficient for subsequent analysis and modeling to describe the lagged response relationship between biological attributes and environmental factors.

[0041] Step S225: Construct a lag response coefficient matrix with environmental factor components as rows and biological attribute response sequences as columns. The matrix elements are the lag response coefficients of the corresponding combinations. The negative coefficients in the lag response coefficient matrix are converted to absolute values ​​to ensure that the lag response coefficient matrix remains non-negative.

[0042] The lag response coefficient matrix is ​​a two-dimensional matrix with environmental factor components as rows and biological attribute response sequences as columns. Each element in the matrix represents the lag response coefficient of the corresponding combination of environmental factor component and biological attribute response sequence. When constructing the lag response coefficient matrix, the lag response coefficient of each combination of environmental factor component and biological attribute response sequence is filled into the corresponding position in the matrix. If a negative value appears in the calculation of the lag response coefficient, it means that there may be a negative correlation between the biological attribute and the environmental factor; however, to keep the matrix non-negative, the absolute value of the negative coefficient is taken.

[0043] Step S230: Using the environmental factor components and biological attribute indicators corresponding to the non-zero elements in the hysteresis response coefficient matrix as network nodes and the absolute value of the hysteresis response coefficient as the edge weight, construct an environmental factor-biological attribute cross-influence network, and perform topology analysis on the environmental factor-biological attribute cross-influence network to identify key influence paths.

[0044] In one implementation, step S230 may include the following steps S231-S236: Step S231: Determine the environmental factor components and biological attribute indicators corresponding to the non-zero elements in the hysteresis response coefficient matrix as network nodes. The node attributes include factor type and time scale characteristics.

[0045] Network nodes are the basic units constituting the cross-influence network of environmental factors and biological attributes, representing environmental factor components or biological attribute indicators. The environmental factor components and biological attribute indicators corresponding to the non-zero elements in the lagged response coefficient matrix are identified as network nodes because these elements indicate a certain correlation between environmental factors and biological attributes. Node attributes include factor type and time scale characteristics. Factor type distinguishes whether a node is an environmental factor component or a biological attribute indicator, while time scale characteristics reflect the time range of change in the corresponding environmental factor or biological attribute.

[0046] When determining network nodes, the lag response coefficient matrix is ​​traversed to identify the environmental factor components and biological attribute indicators corresponding to the non-zero elements. For each node, a corresponding factor type attribute is assigned, such as environmental factor type or biological attribute type. Simultaneously, its time-scale characteristic attribute is determined based on the oscillation period of the environmental factor components or the time range of the biological attribute indicators.

[0047] Step S232: Construct a directed weighted network with the absolute value of the hysteresis response coefficient as the edge weight, and the direction of the directed edge is from the environmental factor component to the biological attribute index.

[0048] In a directed weighted network, connections between nodes have direction and weight. Using the absolute value of the hysteresis response coefficient as the edge weight intuitively represents the strength of the influence between environmental factors and biological attributes. The direction of the directed edge points from the environmental factor component to the biological attribute indicator, indicating that the environmental factor has an impact on the biological attribute. When constructing a directed weighted network, for a given network node, directed edges are established between the corresponding environmental factor component node and the biological attribute indicator node based on the non-zero elements in the hysteresis response coefficient matrix. The weight of the edge is the absolute value of the hysteresis response coefficient, and the direction of the edge points from the environmental factor component node to the biological attribute indicator node.

[0049] Step S233: Calculate the degree centrality, betweenness centrality, and proximity centrality of each node in the network. Degree centrality represents the number of direct connections of a node, betweenness centrality represents the mediating role of a node in a path, and proximity centrality represents the average distance from a node to other nodes.

[0050] Degree centrality, betweenness centrality, and proximity centrality are metrics used to describe the importance of nodes and the topology of a network. Degree centrality is the number of direct connections a node has, reflecting its activity and influence within the network. Betweenness centrality is the frequency with which a node appears in all shortest paths in the network, reflecting its mediating role in information propagation and path connection. Proximity centrality is the reciprocal of the average shortest path length from a node to all other nodes, reflecting the node's reachability and propagation efficiency within the network.

[0051] Step S234: Obtain a pre-set centrality threshold, select nodes whose centrality index exceeds the centrality threshold as key nodes, and extract the directed edges between key nodes to form an initial influence path set.

[0052] When determining the pre-set centrality threshold, previous research experience and the structural characteristics of the current network can be considered. For example, if the goal is to select very important nodes, the centrality threshold can be set higher; if the goal is to cover more potentially influential nodes, the threshold can be set lower. Then, the degree centrality, betweenness centrality, and proximity centrality of each node in the network are compared with the centrality threshold, and nodes with centrality indicators exceeding the threshold are selected as key nodes. Finally, directed edges between key nodes are extracted, and these directed edges are connected to form an initial set of influence paths.

[0053] Step S235: Perform path length analysis on the initial set of affected paths, retain short paths whose path lengths meet the set conditions, and calculate the total weight of each path. The total weight is the product of the weights of the edges on the path.

[0054] Path length analysis is a method for filtering paths in the initial set of influence paths. By retaining short paths whose lengths meet set conditions, redundant paths can be reduced, improving the efficiency and accuracy of the analysis. The total weight of a path is the product of the weights of all edges in the path, reflecting the comprehensive influence of environmental factors on biological attributes along that path.

[0055] When performing path length analysis on the initial set of influencing paths, the conditions for setting the path length can be determined first. For example, a maximum path length threshold can be set, retaining only paths with a length less than this threshold. Then, for the retained paths, their total weight is calculated. Specifically, the weight of each edge in the path is multiplied to obtain the total weight of the path.

[0056] Step S236: Sort the paths in descending order of total weight, and select the top K paths as key impact paths, where K is greater than 0, and the key impact paths cover the relationship between major environmental factors and biological attributes.

[0057] Arranging the paths in descending order of total weight clearly demonstrates the order of importance of different paths. Selecting the top K paths as the key influencing paths, which have a larger total weight, indicates that their impact on biological attributes is more significant.

[0058] When sorting paths in descending order of total weight, a sorting algorithm can be used. For example, quicksort or heapsort can be used to sort the paths from largest to smallest based on their total weight. Then, based on a predetermined K value, the top K paths in the sorted list are selected as the key influencing paths.

[0059] Step S240: Optimize the nonlinear fitting model of the environmental factor response curve with the topological parameters of the key influence path, jointly model the environmental parameter sequence and the biological attribute response sequence to generate the dynamic response surface of the environmental factor, and perform time-frequency transformation on the biological attribute response sequence to extract the instantaneous frequency and instantaneous amplitude features to construct the dynamic descriptor of the time-varying features of the biological attribute.

[0060] In one implementation, step S240 may include the following steps S241-S246: Step S241: Extract the topological parameters of the key influencing paths. The topological parameters include path length, node degree and edge weight distribution. Use the topological parameters as prior knowledge for the nonlinear fitting model of the environmental factor response curve.

[0061] The topological parameters of a critical impact path describe its structural characteristics, including path length, node degree, and edge weight distribution. Path length reflects the indirectness of the influence between environmental factors and biological attributes; node degree represents the number of connections a node has, reflecting its importance in the network; and edge weight distribution represents the intensity distribution of the influence between environmental factors and biological attributes.

[0062] When extracting topological parameters of key impact paths, the paths themselves are first analyzed. For path length, the number of edges in the path is directly counted. For node degree, the number of direct connecting edges to each node is counted. For edge weight distribution, the range of edge weight values ​​and distribution patterns are analyzed. Then, these topological parameters are incorporated as prior knowledge into the nonlinear fitting model of the environmental factor response curve.

[0063] Step S242: Construct a nonlinear fitting model that includes a combination of kernel functions. The combination of kernel functions is a weighted combination of radial basis functions and periodic kernel functions, and the weights are dynamically adjusted based on the path topology parameters.

[0064] When constructing a nonlinear fitting model that includes a combination of kernel functions, radial basis functions (RBFs) and periodic kernel functions can be selected as the basic kernel functions. Then, the RBFs and periodic kernel functions are weighted and combined to obtain the kernel function combination. The weight coefficients are dynamically adjusted based on path topology parameters, such as path length, node degree, and edge weight distribution, using a specific algorithm to calculate the weight coefficients. The weight of the periodic kernel function can be increased when the path length is long, as a longer path may indicate periodic influences; the weight of the RBF can be increased when the node degree is large, as a large node degree may indicate strong local influences.

[0065] Step S243: Use the environmental parameter sequence as the model input and the biological attribute response sequence as the model output to train the nonlinear fitting model, and optimize the kernel function parameters through maximum likelihood estimation.

[0066] Using environmental parameter sequences as model input and biological attribute response sequences as model output allows the nonlinear fitting model to learn the relationship between environmental factors and biological attributes. Maximum likelihood estimation is a statistical method used to estimate model parameters; it finds the kernel function parameters that best match the model output to the actual biological attribute response sequences.

[0067] When training a nonlinear fitting model, the environmental parameter sequences and biological attribute response sequences are first preprocessed, such as through normalization, to ensure the data have the same scale. Then, the environmental parameter sequences are input into the nonlinear fitting model, which calculates the output based on the current kernel function parameters. The model output is then compared with the actual biological attribute response sequences to calculate the likelihood function. The likelihood function represents the probability of observing the actual biological attribute response sequence given the model parameters. Maximum likelihood estimation is used to adjust the kernel function parameters until the likelihood function reaches its maximum value.

[0068] Step S244: Based on the trained model, predict the response of the environmental parameter sequence, generate the predicted sequence of biological attribute response, and calculate the root mean square error between the predicted sequence and the actual observed sequence.

[0069] Predicting responses to environmental parameter sequences based on a trained model is intended to test the model's predictive power and accuracy. After generating predicted sequences of biological attribute responses, the root mean square error (RMSE) between the predicted and observed sequences can be used to quantitatively evaluate the model's fit. The RMSE reflects the average deviation between the predicted and actual values.

[0070] When predicting responses to environmental parameter sequences based on a trained model, the environmental parameter sequences are input into a trained nonlinear fitting model. The model calculates the output based on the learned relationships, obtaining the predicted sequence of biological attribute responses. To calculate the root mean square error between the predicted sequence and the actual observed sequence, the square of the difference between each corresponding data point in the predicted sequence and the actual observed sequence is first calculated. Then, the average of these squared values ​​is calculated, and finally, the square root of the average is taken.

[0071] Step S245: Adjust the kernel function weights based on error feedback, iteratively optimize the model until the error is lower than the set threshold, and obtain the optimized environmental factor response curve model.

[0072] Iteratively optimizing the model until the error falls below a set threshold ensures its accuracy and stability. The optimized environmental factor response curve model can more accurately describe the relationship between environmental factors and biological attributes.

[0073] When adjusting kernel function weights based on error feedback, the root mean square error (RMSE) between the calculated predicted sequence and the actual observed sequence is used to analyze the causes of the error. If the error is large, it indicates a poor model fit, requiring adjustment of the kernel function weights. Specific adjustment strategies can be used to adjust the kernel function weights based on the magnitude and direction of the error. For example, if the radial basis function has a large fitting error, its weight can be appropriately reduced; if the periodic kernel function has a large fitting error, its period parameter can be adjusted or its weight reduced. After each adjustment of the kernel function weights, the model is retrained and used for prediction, and a new RMSE is calculated. This process is repeated until the error falls below a set threshold.

[0074] Step S246: Simulate the response of different combinations of environmental parameters using the optimized environmental factor response curve model to generate a dynamic response surface of environmental factors that includes the response relationship between multi-dimensional environmental parameters and biological attributes.

[0075] By simulating responses to different combinations of environmental parameters using optimized environmental factor response curve models, a comprehensive understanding of the relationship between environmental factors and biological attributes can be achieved. The dynamic response surface of environmental factors is a three-dimensional or multi-dimensional model that demonstrates the relationship between multi-dimensional environmental parameters and biological attribute responses, intuitively reflecting the changes in biological attributes under different values ​​of environmental factors.

[0076] When conducting response simulations, the range of environmental parameter combinations to be simulated is first determined. For example, if the study involves three environmental factors—water temperature, light intensity, and dissolved oxygen concentration—the value range for each environmental factor can be determined. Then, a series of different combinations of environmental parameters are generated within this range. These combinations of environmental parameters are input into the optimized environmental factor response curve model, and the model calculates the biological attribute response value corresponding to each combination based on the learned relationships. These combinations of environmental parameters and their corresponding biological attribute response values ​​are then visualized to generate a dynamic response surface for the environmental factors. This surface can be displayed using 3D plotting tools or multidimensional data visualization methods.

[0077] Step S250: Couple the dynamic descriptor with the dynamic response surface of environmental factors, perform nonlinear dimensionality reduction on the coupled features, and retain the cross-correlation information between environmental parameters and biological attributes.

[0078] Dynamic descriptors contain information on the time-varying characteristics of biological attributes, such as instantaneous frequency and amplitude, reflecting how these attributes change over time. Dynamic response surfaces of environmental factors illustrate the relationship between multi-dimensional environmental parameters and biological attribute responses. Coupled with dynamic descriptors and dynamic response surfaces of environmental factors, the time-varying characteristics of biological attributes can be combined with the influence of environmental factors, providing a more comprehensive description of the relationship between biological attributes and environmental parameters.

[0079] When coupling dynamic descriptors with the dynamic response surfaces of environmental factors, the features of the dynamic descriptors are first associated with the coordinates and values ​​of the dynamic response surfaces. For example, the instantaneous frequency and amplitude features in the dynamic descriptors can be combined with the environmental parameters and biological attribute response values ​​in the dynamic response surfaces to form high-dimensional feature vectors. When performing nonlinear dimensionality reduction on the coupled features, various nonlinear dimensionality reduction algorithms can be used, such as nonlinear extensions of principal component analysis (PCA), such as kernel principal component analysis (KPCA), or methods like locally linear embedding (LLE) and isomap.

[0080] Step S260: Generate an initial specimen feature spectrum based on the dimensionality-reduced features, which includes the nonlinear coupling relationship between environmental parameter trajectories and biological attribute response features.

[0081] The dimensionality-reduced features have removed redundant information while retaining important cross-correlation information between environmental parameters and biological attributes. Based on these dimensionality-reduced features, an initial specimen feature spectrum is generated, which can represent the nonlinear coupling relationship between environmental parameter trajectories and biological attribute response characteristics in a concise and effective manner. The initial specimen feature spectrum is a comprehensive feature set that can provide important basis for subsequent specimen classification, comparison, and analysis.

[0082] When generating the initial specimen feature spectrum, the dimensionality-reduced features are first organized and processed. These features can be arranged according to certain rules, such as grouping them by environmental parameter type or biological attribute category. Then, based on the relationships between these features, a feature matrix or feature vector is constructed; this matrix or vector constitutes the initial specimen feature spectrum. During the construction process, the nonlinear coupling relationship between the environmental parameter trajectory and the biological attribute response features must be fully considered. This coupling relationship can be represented through feature multiplication, nonlinear transformations, or other methods.

[0083] Step S300: Perform feature space mapping, multimodal semantic verification, and entity relationship reasoning on the initial specimen feature spectrum to generate a context-enhanced annotation vector containing the semantic association strength and evolutionary feature weights of the specimen samples in a multi-level knowledge system.

[0084] Feature space mapping transforms the initial specimen feature spectrum from the original feature space to a new feature space that better reflects the semantic relationships between specimen features. Multimodal semantic verification performs semantic checks on specimen features from multiple dimensions to ensure semantic consistency and accuracy. Entity relationship reasoning analyzes the relationships between specimen features and entities in the knowledge system to infer the strength of their associations and the weights of their evolved features. Context-enhanced annotation vectors are vectors containing the semantic association strength and evolved feature weights of specimen samples within a multi-level knowledge system, providing richer contextual information for specimen annotation and understanding.

[0085] In one implementation, step S300 may include the following steps S310-S350: Step S310: Construct a specimen feature semantic space based on the hierarchical structure of the water ecology knowledge system. The spatial dimensions correspond to the hierarchical structure of the knowledge system, and the coordinate values ​​of each dimension represent the degree of association between the specimen features and the semantic concepts at that level.

[0086] The hierarchical structure of the aquatic ecological knowledge system encompasses a multi-level knowledge structure, ranging from macro to micro and from holistic to local, including levels such as ecosystem, species, and biological attributes. The specimen feature semantic space is a space used to represent the relationship between specimen features and semantic concepts within the knowledge system; its dimensions correspond to the hierarchical structure of the knowledge system. The coordinate values ​​of each dimension represent the degree of association between the specimen feature and the semantic concept at that level; a higher degree of association indicates a better match between the specimen feature and the semantic concept.

[0087] In one implementation, step S310 may include the following steps S311-S316: Step S311: Obtain the hierarchical structure of the aquatic ecological knowledge system, constructing a six-level semantic hierarchy from phylum, class, order, family, genus to species, with each semantic hierarchy containing multiple semantic concept nodes.

[0088] The hierarchical structure of the aquatic ecological knowledge system is a framework for classifying and organizing biological and environmental information within aquatic ecosystems. Constructing a six-level semantic hierarchy from phylum, class, order, family, genus to species is a commonly used classification method in biology, capable of systematically describing the taxonomic relationships among organisms. Each semantic level contains multiple semantic concept nodes, representing different taxonomic units or characteristic concepts within that level.

[0089] When acquiring the hierarchical structure of aquatic ecological knowledge systems, one can refer to biological taxonomic standards and relevant aquatic ecological research literature. For example, at the phylum level, it may include protozoa, arthropoda, etc.; at the class level, for protozoa, it may include flagellates, sarcodactyls, etc.; and further subdivided at the order, family, genus, and species levels. Each level of semantic concept node has its specific definition and characteristic description.

[0090] Step S312: Construct a corresponding feature dimension for each semantic level to obtain a multidimensional semantic space framework, wherein the size of the feature dimension is equal to the number of semantic concept nodes of that semantic level.

[0091] Constructing corresponding feature dimensions for each semantic level is to associate specimen features with the hierarchical structure of the aquatic ecological knowledge system. The multidimensional semantic space framework is a spatial structure used to represent the relationship between specimen features and semantic concepts in the knowledge system. Its dimensions correspond to the semantic levels, and the size of each dimension is equal to the number of semantic concept nodes at that semantic level.

[0092] When constructing feature dimensions, the size of the feature dimension is determined based on the number of semantic concept nodes at each semantic level. For example, at the gate level, if there are 5 semantic concept nodes (such as 5 different gates), the feature dimension size corresponding to the gate level is 5. Combining these feature dimensions yields the multidimensional semantic space framework.

[0093] Step S313: Collect standard feature descriptions of semantic concept nodes at each level, and perform vector transformation on the standard feature descriptions to obtain concept feature vectors.

[0094] The standard feature descriptions of semantic concept nodes at each level are detailed definitions and characteristic descriptions of the semantic concept, containing its essential features and key information distinguishing it from other concepts. Vector transformation of these standard feature descriptions yields concept feature vectors, converting semantic information into numerical vectors that can be processed by computers. Standard feature descriptions of semantic concept nodes at each level can be obtained from biological literature, specialized databases, and other sources. For example, the standard feature description for a species might include information on its morphological characteristics, physiological characteristics, and ecological habits. Word embedding techniques, such as Word2Vec and GloVe, can be used when performing vector transformation on the standard feature descriptions.

[0095] Step S314: Calculate the similarity between the initial specimen feature spectrum and the feature vectors of concepts at each semantic level, and normalize the similarity values ​​to use as the coordinate values ​​of the corresponding dimensions of the semantic space.

[0096] Calculating the similarity between the initial specimen feature spectrum and the feature vectors of concepts at each semantic level is to measure the degree of matching between specimen features and semantic concepts in the knowledge system. Normalizing the similarity values ​​and using them as coordinate values ​​for the corresponding dimensions of the semantic space makes the coordinate values ​​comparable and accurately represents the relationship between specimen features and semantic concepts in the semantic space.

[0097] When calculating similarity, various similarity metrics can be used, such as cosine similarity and Euclidean distance. After calculating the similarity value, it is normalized to the [0,1] interval. Normalization can be done using linear normalization methods, such as subtracting the minimum value from the similarity value and then dividing by the difference between the maximum and minimum values. The normalized similarity value is then used as the coordinate value of the corresponding dimension in the semantic space.

[0098] Step S315: Construct a distance metric function for the semantic space. The similarity of feature vectors of different specimens in the semantic space is measured based on Mahalanobis distance. The distance metric function integrates the semantic association weights between semantic levels.

[0099] Distance metrics in semantic space are used to measure the similarity of feature vectors from different specimens within the semantic space. These distance metrics integrate semantic association weights between semantic levels to account for the importance and relationships between different semantic levels when calculating distances.

[0100] When constructing the distance metric function for the semantic space, the covariance matrix of the data in the semantic space is first calculated. The covariance matrix reflects the correlation between the dimensions of the data. Then, according to the definition of Mahalanobis distance, the Mahalanobis distance is equal to the transpose of the vector difference multiplied by the inverse of the covariance matrix, and then multiplied by the vector difference. During the calculation, semantic association weights between semantic levels are introduced. These semantic association weights can be determined based on expert knowledge or statistical data analysis; for example, if a certain semantic level is more important for the classification and understanding of the specimens, its weight can be set higher. These weights are then incorporated into the Mahalanobis distance calculation to obtain a distance metric function that integrates semantic association weights.

[0101] Step S316: Perform principal component analysis to reduce the dimensionality of the high-dimensional semantic space, retaining principal components whose cumulative contribution rate exceeds a set proportion, to ensure the discriminative power and computational efficiency of the semantic space.

[0102] Principal component analysis (PCA) is used for dimensionality reduction in high-dimensional semantic spaces to reduce the dimensionality of the semantic space while preserving the main information of the data. The process begins by calculating the covariance matrix of the high-dimensional semantic space data. Then, eigenvalue decomposition is performed on the covariance matrix to obtain eigenvalues ​​and eigenvectors. Eigenvalues ​​represent the magnitude of the variance of the principal components, and eigenvectors represent the direction of the principal components. The eigenvalues ​​are sorted from largest to smallest, and the cumulative contribution rate is calculated. The cumulative contribution rate is the proportion of the sum of the top k eigenvalues ​​to the sum of all eigenvalues. Based on this proportion, the top k eigenvectors with a cumulative contribution rate exceeding this proportion are selected as principal components. The original data is then projected onto these principal components to obtain the dimensionality-reduced semantic space.

[0103] Step S320: Perform deep metric learning space mapping on the initial specimen feature spectrum, project the initial specimen feature spectrum onto the specimen feature semantic space, and generate the initial semantic coordinate vector.

[0104] In one implementation, step S320 may include the following steps S321-S326: Step S321: Construct a deep metric learning network. The deep metric learning network includes a feature extraction layer and a mapping layer. The feature extraction layer adopts a residual network structure, and the mapping layer adopts a fully connected network structure.

[0105] Deep metric learning networks are key models for mapping initial specimen feature spectra to the specimen feature semantic space. The feature extraction layer extracts more representative and discriminative features from the initial specimen feature spectra. Residual network structures possess strong feature extraction capabilities, effectively addressing the vanishing and exploding gradient problems in deep neural networks, allowing the network to be trained deeper and extract more complex features. The mapping layer maps the features extracted by the feature extraction layer to the specimen feature semantic space. Fully connected network structures can perform nonlinear transformations on the features, resulting in a suitable representation in the semantic space.

[0106] Step S322: Standardize the initial specimen feature spectrum to make the data of each dimension of the initial specimen feature spectrum conform to the zero mean unit variance distribution, so as to serve as the input data of the deep metric learning network.

[0107] Standardizing the initial specimen feature spectrum beforehand is crucial for ensuring data comparability and stability, and preventing scale differences across different dimensions from negatively impacting the training of deep metric learning networks. The zero-mean unit variance distribution is a standardized distribution that adjusts the mean of the data to 0 and the standard deviation to 1. Transforming the initial specimen feature spectrum into this distribution improves the network's training efficiency and convergence speed.

[0108] During the standardization preprocessing, the mean and standard deviation of each dimension of the initial specimen feature spectrum are first calculated. Then, for each dimension, the mean of that dimension is subtracted, and the result is divided by the standard deviation of that dimension to obtain the standardized data. After this processing, the mean of each dimension of the initial specimen feature spectrum is 0, and the standard deviation is 1.

[0109] Step S323: Obtain the preset triplet loss function. The triplet loss function includes the distance constraint between the anchor sample, the positive sample and the negative sample. The positive sample is the feature spectrum of the specimen of the same semantic category, and the negative sample is the feature spectrum of the specimen of different semantic categories.

[0110] The triplet loss function is a common loss function in deep metric learning. By constraining the distance relationship between anchor samples, positive samples, and negative samples, the network learns the semantic similarity between samples. Anchor samples are benchmark samples used for comparison, positive samples are the feature spectra of samples belonging to the same semantic category as anchor samples, and negative samples are the feature spectra of samples belonging to different semantic categories than anchor samples.

[0111] The predefined triplet loss function is in the form: L = max(d(a,p) - d(a,n) + α,0), where a represents the anchor sample, p represents the positive sample, n represents the negative sample, d represents the distance metric (such as Euclidean distance), and α is a positive boundary value. The purpose of this loss function is to minimize the distance d(a,p) between the anchor sample and the positive sample, maximize the distance d(a,n) between the anchor sample and the negative sample, and ensure that the difference between the two is greater than the boundary value α.

[0112] Step S324: Optimize the parameters of the deep metric learning network using the backpropagation algorithm, minimize the spatial distance between anchor samples and positive samples, maximize the spatial distance between anchor samples and negative samples, and iterate the training until the triplet loss function converges.

[0113] Backpropagation is an algorithm used in deep learning to optimize network parameters. It calculates the gradient of the loss function with respect to the network parameters and then updates the network parameters based on this gradient. In deep metric learning, backpropagation is used to optimize the parameters of the deep metric learning network. The goal is to minimize the spatial distance between anchor samples and positive samples, and maximize the spatial distance between anchor samples and negative samples, thereby enabling the network to learn the semantic similarity between samples. Iterative training continues until the triplet loss function converges, indicating that the network parameters have been adjusted to a good state, accurately representing the semantic relationships between samples in the sample feature semantic space.

[0114] Step S325: Fix the trained mapping layer parameters, input the preprocessed initial specimen feature spectrum for forward propagation, and obtain the semantic space coordinate vector output by the network.

[0115] Fixing the parameters of the trained mapping layer ensures the stability and consistency of the network when using the trained deep metric learning network for prediction. Forward propagation of the preprocessed initial specimen feature spectrum transforms the specimen features through the network, obtaining their coordinate representation in the semantic space. The semantic space coordinate vector is the network output, containing the location information of the specimen features in the semantic space.

[0116] After fixing the parameters of the trained mapping layer, the initial specimen feature spectrum, which has undergone normalization and preprocessing, is input into the feature extraction layer of the deep metric learning network. The feature extraction layer extracts features from the input specimen features to obtain more representative features. Then, the extracted features are input into the mapping layer, which maps the features into the specimen feature semantic space according to the fixed parameters. After forward propagation, the semantic space coordinate vector output by the network is finally obtained.

[0117] Step S326: Perform L2 normalization on the semantic space coordinate vector to make the coordinate values ​​of each dimension fall within the set range, and generate the initial semantic coordinate vector.

[0118] L2 normalization of semantic space coordinate vectors aims to make the dimensions of the vectors comparable and to restrict the coordinate values ​​to a defined interval. L2 normalization divides each element of the vector by its L2 norm, which is the square root of the sum of the squares of the vector's elements. Through L2 normalization, the magnitude of the coordinate vector is made 1, and the coordinate values ​​of each dimension fall within the interval [-1, 1]. This results in initial semantic coordinate vectors that have a better representation in the semantic space.

[0119] Step S330: Calculate the similarity between the initial semantic coordinate vector and the standard semantic template in the knowledge system, adjust the spatial mapping parameters according to the similarity distribution, and optimize the spatial distribution of the semantic coordinate vector.

[0120] Calculating the similarity between the initial semantic coordinate vector and the standard semantic template in the knowledge system is to evaluate the degree of matching between the representation of the specimen features in the semantic space and the standard semantics in the knowledge system. The standard semantic template is the standard representation of different semantic concepts in the knowledge system. Adjusting the spatial mapping parameters according to the similarity distribution is to make the position of the specimen features in the semantic space more reasonable, optimize the spatial distribution of the semantic coordinate vector, and thus improve the accuracy of the semantic representation of the specimen features.

[0121] When calculating similarity, methods such as cosine similarity and Euclidean distance can be used. After calculating the similarity between the initial semantic coordinate vector and each standard semantic template, the distribution of similarity is analyzed. If it is found that the semantic coordinate vectors of certain specimen features have generally low similarity to the standard semantic templates, it indicates that there may be a problem with the spatial mapping, and the spatial mapping parameters need to be adjusted. The spatial mapping parameters can be the mapping layer parameters of the deep metric learning network. By fine-tuning these parameters, the position of the specimen features in the semantic space can be made closer to the standard semantic templates.

[0122] Step S340: Construct a multimodal semantic verification rule base using semantic verification rules from multiple dimensions of morphology, physiology and ecology. Perform multimodal consistency verification on the optimized semantic coordinate vector, and correct the conflicting components in the semantic coordinate vector based on the verification results. Perform entity relationship reasoning on the implicit association between specimen features and conceptual entities in the knowledge system, and calculate the association strength.

[0123] The multimodal semantic verification rule base is a collection of semantic verification rules encompassing multiple dimensions, including morphology, physiology, and ecology. Morphological rules check whether the morphological features of a specimen conform to the characteristics of the corresponding species; physiological rules check whether the physiological attributes of the organism are normal; and ecological rules check whether the relationship between environmental parameters and biological attributes conforms to ecological principles. Multimodal consistency verification of the optimized semantic coordinate vector ensures that specimen features are consistent across multiple semantic dimensions. Correcting conflicting components in the semantic coordinate vector based on the verification results makes the semantic coordinate vector more accurately represent the specimen features. Entity relationship reasoning is performed on the implicit associations between specimen features and conceptual entities in the knowledge system to uncover potential relationships between specimen features and entities in the knowledge system and to calculate the association strength.

[0124] Step S350: Integrate the association strength and evolutionary feature weights into the semantic coordinate vector to generate a context-enhanced annotation vector containing the semantic association strength and evolutionary feature weights of the specimen samples in the multi-level knowledge system.

[0125] Integrating association strength and evolutionary feature weights into the semantic coordinate vector aims to include more contextual information, thereby more comprehensively representing the characteristics of the specimen sample within a multi-level knowledge system. Association strength represents the degree of association between the specimen features and conceptual entities in the knowledge system, while evolutionary feature weights reflect the importance of the specimen features during the evolutionary process. The context-enhanced annotation vector is the integrated result, providing richer information for specimen annotation.

[0126] Step S400: Perform historical trajectory similarity comparison and multi-dimensional evolution trend prediction on the context-enhanced annotation vector to generate an intermediate evolution label set of the semantic change process of the recorded specimen features at different time scales.

[0127] Historical trajectory similarity comparison of context-enhanced annotation vectors aims to identify specimens with similar historical evolutionary trajectories to the current specimen's features, thereby leveraging historical experience to analyze the current specimen's evolutionary trend. Multidimensional evolutionary trend prediction forecasts the future evolution of specimen features from multiple perspectives (such as morphology, physiology, and ecology). Intermediate evolutionary label sets are collections of labels recording the semantic changes of specimen features at different time scales, containing feature descriptions and evolutionary direction indicators at different points in time.

[0128] In one implementation, step S400 may include the following steps S410-S460: Step S410: Retrieve historical specimen annotation data from the laboratory specimen database that belong to the same genus as the current specimen sample, and extract the context-enhanced annotation vectors of the historical specimens as a reference vector set. The reference vector set contains annotation vector sequences from different collection time points.

[0129] Retrieving historical specimen annotation data from the laboratory specimen database that belongs to the same genus as the current specimen aims to obtain information on historical specimens with similar classification attributes, facilitating the comparison and analysis of historical trajectories. The context-enhanced annotation vectors of historical specimens contain information such as the semantic association strength and evolutionary feature weights of historical specimens within a multi-level knowledge system. Using these vectors as a reference vector set can provide insights into the evolutionary trend analysis of the current specimen. The reference vector set contains annotation vector sequences from different collection points, reflecting the characteristic changes of historical specimens at different times.

[0130] In one implementation, step S410 may include the following steps S411-S416: Step S411: Determine the biological taxonomic unit to which the current specimen belongs based on its taxonomic information, and construct a database retrieval keyword combination using the scientific name and characteristic description of the biological taxonomic unit. The keyword combination includes the taxonomic name and the characteristics of the collection environment.

[0131] Determining the biological taxonomic unit to which a current specimen belongs based on its taxonomic information is crucial for accurately locating historical specimens with similar taxonomic attributes within a laboratory specimen database. A biological taxonomic unit is the basic unit for classifying organisms in biology, such as phylum, class, order, family, genus, and species. Constructing a database search keyword combination using the scientific name and characteristic description of the biological taxonomic unit can improve the accuracy and relevance of the search. The keyword combination includes the taxonomic name and the characteristics of the collection environment. The taxonomic name is used to determine the classification range of the organism, while the collection environment characteristics are used to further filter historical specimens with similar collection environments to the current specimen.

[0132] Step S412: Query the metadata index table of the laboratory specimen database by searching for keyword combinations, and obtain a list of all associated historical specimen record IDs by index matching. The list of historical specimen record IDs is arranged in reverse order of collection time.

[0133] The metadata index table of the laboratory specimen database is an index structure that records basic information about all specimen records in the database, including specimen taxonomic information, collection time, collection location, and other metadata. Searching the metadata index table by combining keywords matches the keywords with the metadata to find historical specimen records related to those keywords. Index matching improves query efficiency and quickly locates records that meet the criteria. The list of historical specimen record IDs is sorted in reverse chronological order by collection time, allowing for priority access to more recent historical specimen records, facilitating subsequent analysis and comparison.

[0134] Step S413: Retrieve the corresponding historical specimen annotation data file from the database storage module based on the record ID list, parse the structured data fields of the file, and extract the annotated context-enhanced annotation vectors.

[0135] Retrieving the corresponding historical specimen annotation data file from the database storage module based on the record ID list is to obtain detailed annotation information for the historical specimens. The database storage module stores the specimen annotation data files, each containing relevant information for a historical specimen. Parsing the structured data fields of the file involves analyzing the data according to a specific structure to extract the required information. The labeled context-enhanced annotation vectors are crucial information contained in the file, recording the semantic association strength and evolutionary feature weights of the historical specimens within a multi-level knowledge system. When retrieving the data file, the corresponding file is located in the database storage module based on each ID in the record ID list. These files may be stored in specific formats, such as XML or JSON.

[0136] Step S414: Perform outlier detection on the extracted context-enhanced annotation vectors, calculate the mean and standard deviation of each vector dimension, and mark vectors that deviate from the mean by more than a set multiple of the standard deviation as outliers and remove them.

[0137] Outlier detection on the extracted context-enhanced labeled vectors aims to remove abnormal data and improve data quality and reliability. Outliers may arise from data collection errors, labeling mistakes, or other reasons, and can interfere with subsequent analysis and prediction. Calculating the mean and standard deviation of each vector dimension helps determine the normal range of the data.

[0138] During outlier detection, for each dimension of the extracted context-enhanced labeled vectors, the mean and standard deviation of all vector values ​​in that dimension are calculated. A fold threshold is set, for example, three standard deviations. For each dimension's value of each vector, if the value deviates from the mean of that dimension by more than three standard deviations, the vector is marked as an outlier. Vectors marked as outliers are then removed from the vector set.

[0139] Step S415: Perform data integrity verification on the remaining vectors, check whether there are missing values ​​in each dimension of the vectors, and fill the missing dimensions with the median of the vectors in the same batch.

[0140] Performing data integrity checks on the remaining vectors ensures data completeness and prevents missing data from affecting subsequent analysis and predictions. Checking for missing values ​​in each dimension of the vector involves iterating through each dimension and checking for any null or undefined values. During data integrity checks, for each remaining context-enhanced annotation vector, the values ​​of each dimension are checked. If a missing value is found in a dimension, the vector and the missing dimension's information are recorded. For the missing dimension, the median value of that dimension among vectors in the same batch is calculated. This median is then used as padding to fill in the missing dimension's position.

[0141] Step S416: Arrange the verified vectors in ascending order of the collection timestamp to construct a reference vector set containing time series characteristics. The vector set retains the association information between the specimen collection environment parameters and biological attribute observation data corresponding to each vector.

[0142] Arranging the validated vectors in ascending order of collection timestamps is to give the reference vector set time-series characteristics, reflecting the changes in historical specimen characteristics over time. Constructing a reference vector set with time-series characteristics facilitates subsequent time-series analysis and evolutionary trend prediction. The vector set retains the correlation information between the specimen collection environment parameters and biological attribute observation data corresponding to each vector, so that changes in environmental factors and biological attributes can be comprehensively considered during the analysis process, leading to a more comprehensive understanding of the specimen's evolutionary process.

[0143] During the sorting process, the vectors are arranged in ascending order based on the specimen collection timestamps corresponding to their respective samples using a sorting algorithm. Algorithms such as quicksort and mergesort can be used. After the sorting is complete, a reference vector set is constructed. This vector set not only contains the context-enhanced annotation vectors but also retains the correlation information between the specimen collection environment parameters (such as water temperature and light intensity) and biological attribute observation data (such as growth rate and reproduction rate) corresponding to each vector.

[0144] Step S420: Calculate the similarity between the current context-enhanced annotation vector and each vector in the reference vector set, and filter historical reference vectors whose similarity meets the set conditions.

[0145] Calculating the similarity between the current context-enhanced annotation vector and each vector in the reference vector set aims to identify historical specimens with similar features to the current specimen, allowing us to leverage historical experience to analyze the evolutionary trend of the current specimen. Filtering historical reference vectors that meet set similarity criteria selects the most relevant historical specimen information from the reference vector set, improving the accuracy and effectiveness of the analysis.

[0146] In one implementation, step S420 may include the following steps S421-S426: Step S421: Convert the current context-enhanced annotation vector and each vector in the reference vector set into column vectors of the same dimension, and normalize each dimension of the vector to ensure that the vector magnitude is consistent.

[0147] To accurately calculate the similarity between the current context-enhanced annotation vector and each vector in the reference vector set, the vectors are first normalized. They are converted into column vectors of the same dimension. This is because a unified vector form ensures consistency and accuracy in mathematical operations and comparisons. Vectors of different dimensions are difficult to compare and calculate similarity directly; unifying them into column vectors makes the vectors structurally comparable.

[0148] Step S422: Perform a dot product operation on the current vector and each reference vector, divide by the vector magnitude product to obtain the similarity value, and generate a distribution sequence containing the similarity values ​​of all reference vectors.

[0149] After normalizing the vectors, we can begin calculating the similarity between the current vector and each reference vector. The result of the dot product operation is affected by the vector magnitudes. To eliminate this effect, the result of the dot product operation needs to be divided by the product of the vector magnitudes. After this processing, the obtained similarity value can more accurately reflect the true degree of similarity between the two vectors.

[0150] Step S423: Perform statistical characteristic analysis on the similarity distribution sequence, calculate the median and interquartile range of the sequence, and dynamically determine the similarity threshold based on the median and interquartile range.

[0151] After obtaining the similarity distribution sequence, a suitable similarity threshold needs to be determined in order to filter out reference vectors with high similarity to the current vector. Statistical characteristic analysis of the similarity distribution sequence is an effective method for determining the threshold.

[0152] The median is the value in the middle when a sequence is arranged in ascending order; it reflects the middle level of the sequence. The interquartile range (ICM) is the difference between the upper and lower quartiles, measuring the dispersion of the middle 50% of the data in the sequence. By calculating the median and ICM, we can understand the central tendency and dispersion of a similarity distribution sequence.

[0153] The reason for dynamically determining the similarity threshold based on the median and interquartile range is that this method can adaptively adjust according to the actual situation of the similarity distribution sequence. Different datasets may have different distribution characteristics, and if a fixed threshold is used, it may not be able to accurately select suitable reference vectors. However, the threshold determined based on the median and interquartile range can better adapt to the distribution of the data and improve the accuracy of the selection.

[0154] Step S424: Select reference vectors with similarity values ​​greater than the dynamic threshold as the initial candidate set. If the number of candidate vectors exceeds the preset upper limit, then truncate the first preset number of vectors in descending order of similarity value.

[0155] Based on a dynamically determined similarity threshold, reference vectors with similarity values ​​greater than the threshold are selected from the similarity distribution sequence. These reference vectors have a high similarity to the current vector and are selected as the preliminary candidate set. The preliminary candidate set contains historical specimen information that may be of significant reference value for analyzing the current specimen's evolutionary trend. However, in some cases, the number of preliminary candidate vectors may exceed a preset upper limit. The preset upper limit is a quantity restriction set according to research needs and actual circumstances. If the number of candidate vectors is too large, it may increase the complexity and computational load of subsequent analysis, and may also introduce some irrelevant or redundant information. Therefore, when the number of candidate vectors exceeds the preset upper limit, the candidate sets need to be sorted in descending order of similarity value, and the first preset number of vectors are truncated.

[0156] Step S425: Perform a time distribution uniformity check on the vectors in the preliminary candidate set, calculate the interval distribution of vector timestamps, and if time clustering exists, perform equal-interval resampling.

[0157] After selecting the initial candidate set, the temporal distribution of the reference vectors also needs to be considered. The uniformity of the temporal distribution is crucial for accurately analyzing the evolutionary trend of the specimen. If the reference vectors are too concentrated in time, it may lead to a one-sided understanding of the specimen's evolutionary process and fail to fully reflect the characteristic changes of the specimen in different time periods.

[0158] The initial candidate set of vectors undergoes a temporal distribution uniformity check, primarily by calculating the interval distribution of vector timestamps. Timestamps record the time information of sample collection; analyzing the interval distribution of timestamps reveals whether the reference vectors are evenly distributed over time. If temporal clustering is observed—that is, an excessive number of reference vectors in some time periods and an insufficient number in others—equal-interval resampling is necessary.

[0159] Step S426: Perform semantic diversity assessment on the resampled vectors, calculate the average pairwise similarity between vectors, and remove redundant vectors if the average similarity exceeds a set value, while retaining historical reference vectors with significant semantic differences.

[0160] After performing equally spaced resampling, semantic diversity evaluation is required for the resampled vectors. The purpose of semantic diversity evaluation is to ensure that the retained reference vectors have sufficient diversity to provide richer information and avoid excessive redundant information between reference vectors. By calculating the similarity between each pair of resampled vectors and averaging these similarities, the average pairwise similarity between vectors can be obtained. If the average pairwise similarity exceeds a set value, it indicates that the similarity between vectors is high and there is a lot of redundant information.

[0161] Step S430: Extract the specimen feature evolution trajectory data corresponding to the filtered historical reference vectors, and perform segmentation processing on the trajectory data based on the timestamp information to obtain a set of evolutionary segments containing different time spans.

[0162] After selecting suitable historical reference vectors, the next step is to extract the specimen feature evolution trajectory data corresponding to these reference vectors. This trajectory data records the feature changes of historical specimens at different time points and is an important basis for analyzing specimen evolution trends.

[0163] Different time spans may correspond to different evolutionary stages and influencing factors. For example, a shorter time span may reflect the specimen's adaptation to environmental changes in a short period, while a longer time span may reflect the specimen's characteristic changes in a long-term evolutionary process. By obtaining a set of evolutionary segments containing different time spans, we can study the evolutionary trends of the specimen in more detail and analyze the factors affecting the changes in specimen characteristics in different time periods.

[0164] Step S440: Construct a time series prediction model that integrates long short-term memory network and attention mechanism. Use the current context-enhanced annotation vector and the set of evolutionary segments as input to the time series prediction model, and output a multi-timescale feature evolution trend prediction sequence.

[0165] To accurately predict the evolutionary trends of specimen features, a suitable time series prediction model needs to be constructed. A model integrating Long Short-Term Memory (LSTM) networks and attention mechanisms is an effective choice. LSTM networks can handle long-term dependencies in time series data. In predicting the evolutionary trends of specimen features, changes in specimen features may be influenced by multiple past time points. LSTM networks, through their memory units and gating mechanisms, can effectively capture these long-term dependencies, thus better predicting future feature changes. Attention mechanisms, on the other hand, help the model focus more on important parts of the time series data. During the evolution of specimen features, the importance of feature information at different time points for future prediction may vary. Attention mechanisms can automatically allocate different attention weights based on the characteristics of the data, allowing the model to focus more on time points and feature information that have a greater impact on the prediction results, thereby improving prediction accuracy.

[0166] The current context-enhanced annotation vector and the evolutionary fragment set are used as inputs to the time series prediction model. The current context-enhanced annotation vector contains information such as the semantic association strength and evolutionary feature weights of the current specimen in the multi-level knowledge system, while the evolutionary fragment set records the feature evolution of historical specimens at different time spans. By learning from these input data, the model outputs a multi-time-scale feature evolution trend prediction sequence, which reflects the feature change trends of the specimen at different time scales.

[0167] Step S450: Perform weighted fusion of the predicted sequence and historical evolution trajectory data, wherein the weights are dynamically allocated based on the importance of the time scale to generate a complete feature transition sequence containing the predicted information.

[0168] After obtaining the predicted sequence of feature evolution trends at multiple time scales, it is necessary to fuse it with historical evolution trajectory data. Historical evolution trajectory data records the actual changes in the specimen's past features, while the predicted sequence is a prediction of future feature changes. Fusing the two can make comprehensive use of historical and predictive information to obtain a more comprehensive and accurate sequence of feature changes.

[0169] Dynamically assigning weights based on the importance of time scales is a crucial step in the fusion process. Different time scales may have different importance in analyzing specimen evolution trends. For example, recent time scales may be more important for predicting short-term feature changes, while long-term time scales may be more critical for grasping the overall evolutionary trend of the specimen. By dynamically assigning weights according to the importance of time scales, the fused feature change sequence can more reasonably integrate information from different time scales.

[0170] During weighted fusion, the predicted sequence and historical evolutionary trajectory data are summed according to dynamically assigned weights to obtain a complete feature change sequence containing predictive information. This sequence not only includes past feature changes of the specimen but also predictive information about future feature changes, providing a rich information foundation for the subsequent generation of intermediate evolutionary label sets.

[0171] Step S460: Based on semantic feature extraction and timestamp marking of feature transition sequences, generate an intermediate evolution tag set of the semantic transition process of the recorded specimen features at different time scales. The intermediate evolution tag set contains feature descriptions and evolution direction indicators for each time node.

[0172] After obtaining the complete feature transition sequence containing predictive information, an intermediate evolutionary label set can be generated based on this sequence. Semantic feature extraction is a crucial step in generating this intermediate evolutionary label set. By performing semantic analysis on the feature transition sequence, representative semantic features are extracted from the sequence. These semantic features reflect the characteristic features and changing trends of the specimen at different time points. Based on semantic feature extraction and timestamp marking, an intermediate evolutionary label set recording the semantic change process of the specimen's features at different time scales is generated. This label set contains a feature description and an evolutionary direction indicator for each time point. The feature description details the characteristic state of the specimen at that time point, while the evolutionary direction indicator points out the direction of change of the specimen's features from one time point to the next, such as growth, decrease, or stabilization.

[0173] Step S500: Integrate the intermediate evolutionary tag set through multi-dimensional semantic consistency verification and evolutionary path optimization to generate annotation tags with ecological evolution semantics.

[0174] While intermediate evolutionary label sets record the semantic changes of specimen features at different time scales, they may contain semantic inconsistencies or unreasonable evolutionary paths. To generate high-quality annotation labels, multi-dimensional semantic consistency verification and evolutionary path optimization are required for the intermediate evolutionary label sets.

[0175] In one implementation, step S500 may include the following steps S510-S560: Step S510: Construct a multi-dimensional semantic consistency verification matrix with the label items of the intermediate evolution label set as rows and the semantic verification dimension as columns. The semantic verification dimension includes morphological semantics, physiological semantics and ecological semantics, and the matrix elements are the semantic descriptions of the label items in the corresponding dimensions.

[0176] To perform multi-dimensional semantic consistency verification, a multi-dimensional semantic consistency verification matrix must first be constructed. This matrix is ​​structured with rows representing the labels in the intermediate evolutionary label set, where each label represents a feature description and evolutionary direction indication of the specimen at a certain time point. Semantic verification dimensions are represented as columns, encompassing morphological semantics, physiological semantics, and ecological semantics. These dimensions perform semantic checks on the labels from different perspectives.

[0177] The matrix elements are semantic descriptions of the tags in their corresponding dimensions. By organizing and classifying the semantic information of the tags according to different dimensions and storing it in the matrix, it is convenient to perform multi-dimensional semantic analysis and verification of the tags in the future.

[0178] Step S520: Perform rule matching for each tag item across all semantic dimensions, compare the tag description with the standard rules in the multimodal semantic verification rule base, and calculate the consistency score based on the number and weight of the matching rules.

[0179] The multimodal semantic validation rule base stores a series of standard rules. These rules are formulated based on knowledge and research findings in fields such as biology, physiology, and ecology, representing semantically correct standards. The semantic description of each label item in the matrix in its corresponding semantic dimension is compared with the standard rules in the multimodal semantic validation rule base to check whether the label description conforms to the standard rules.

[0180] A consistency score is calculated based on the number and weight of matching rules. Different rules may have different importance, so each rule is assigned a corresponding weight. If a tag matches a large number of rules on a certain semantic dimension, and these rules have high weights, then the tag will have a higher consistency score on that semantic dimension.

[0181] Step S530: Identify semantic conflict labels based on the score matrix. Conflict labels are labels that score below a set threshold in any semantic dimension. Perform semantic correction on the conflict labels, based on the semantic representation of the high-scoring labels in the same dimension.

[0182] After calculating the consistency score for each label across all semantic dimensions, semantically conflicting labels can be identified based on the score matrix. A threshold is set; if a label's score on any semantic dimension falls below this threshold, it is identified as a semantically conflicting label. Semantically conflicting labels may arise due to data collection errors, labeling mistakes, or inaccurate understanding of specimen characteristics, and require correction.

[0183] When semantically correcting conflicting tags, the semantic representations of high-scoring tags in the same dimension are referenced. High-scoring tags in the same dimension indicate a higher degree of conformity to standard rules in that semantic dimension, and their semantic representations have higher reliability and accuracy. By referencing the semantic representations of high-scoring tags in the same dimension to correct conflicting tags, the conflicting tags can be made more semantically consistent with standard rules, improving the overall semantic consistency of the intermediate evolved tag set.

[0184] Step S540: Construct an evolution path directed graph with the corrected label items as nodes and the temporal order between labels as directed edges. The edge weight is the semantic similarity between adjacent label items, and the node attributes include the timestamp and semantic feature vector of the label.

[0185] After correcting the semantically conflicting labels, a directed graph of the evolutionary path needs to be constructed to further optimize the evolutionary path. The corrected labels serve as nodes, each representing the characteristic state of the specimen at a certain time point. The temporal relationship between labels is represented by directed edges, with the direction of the edges pointing from earlier label nodes to later label nodes, thus forming a directed graph structure that reflects the specimen's evolutionary process.

[0186] In one implementation, step S540 may specifically include the following steps: Step S541: Treat each corrected label item as an independent node in the directed graph of the evolution path. The node attribute fields include the timestamp information, semantic feature vector, and feature description text corresponding to the label. The semantic feature vector is generated by word embedding transformation of the label text.

[0187] When constructing the directed graph of the evolutionary path, each corrected label item is first transformed into an independent node in the graph. Each node represents the characteristic state of the specimen at the corresponding time. To comprehensively describe the information of the nodes, node attribute fields are set.

[0188] The node attribute field contains the timestamp information corresponding to the label. The timestamp clearly identifies the time node to which the label item corresponds, enabling a clear temporal sequence of specimen feature changes in the graph. The semantic feature vector is generated by word embedding transformation of the label text.

[0189] Step S542: Determine the direction of the directed edges between nodes based on the timestamp order of the label items, pointing from the earlier label node to the later label node, forming a time series directed connection.

[0190] After converting the corrected labels into nodes, it is necessary to determine the connections between the nodes. It is reasonable to determine the direction of directed edges between nodes based on the timestamp order of the labels, because the evolution of the specimen proceeds in chronological order, gradually evolving from an earlier state to a later state.

[0191] By setting the direction of the directed edges from the earlier label node to the later label node, a time-series directed connection is formed. This directed connection can intuitively show the evolution of specimen features over time, making the feature change trajectory of the specimen from one time point to the next clearly visible in the directed graph of the evolution path.

[0192] Step S543: Perform cosine similarity calculation on the semantic description text of adjacent label items, normalize the similarity value and use it as the weight value of the corresponding directed edge, and map the weight value range to the set interval.

[0193] Step S544: Add self-loop edges and jump connection edges to the directed graph. The weight of the self-loop edge is set to the self-similarity of the node's semantic feature vector. The jump connection edge connects non-contiguous label nodes that have a semantic correlation exceeding a set value.

[0194] To more comprehensively represent the evolutionary process of the specimen, in addition to basic adjacent node connections, self-loop edges and jump connections need to be added to the directed graph of the evolutionary path. A self-loop edge is an edge that points from one node to itself, with its weight set to the self-similarity of the node's semantic feature vector. Self-similarity reflects the consistency and stability of a node's own semantics. During the specimen's evolution, some features may remain relatively stable for a period of time; self-loop edges can be used to represent the persistent state of such features, and the weight of the self-loop edge reflects the degree of this stability.

[0195] Skip edges connect discontinuous label nodes whose semantic relevance exceeds a set threshold. During specimen evolution, some features may not change continuously but rather abruptly at certain points in time, yet these changes still maintain semantic connections. By setting skip edges, these discontinuous but meaningful semantic connections can be captured, allowing the directed graph of the evolutionary path to more comprehensively reflect the specimen's evolutionary process. A threshold for semantic relevance is set; skip edges are only added when the semantic relevance between discontinuous nodes exceeds this threshold. This avoids adding too many unnecessary connections, ensuring a clear and effective graph structure.

[0196] Step S545: Construct a node attribute index table, which includes node ID, timestamp range, and dimension index of semantic feature vector.

[0197] A node ID is a unique identifier for each node, allowing for quick and accurate location of a specific node in the graph. The timestamp range records the time interval corresponding to each node, which is extremely useful for querying nodes based on time criteria.

[0198] Step S546: Store the topology of the directed graph through an adjacency matrix. The matrix rows and columns are node IDs, and the matrix elements are the weight values ​​of the corresponding edges. Zero values ​​indicate no direct connection, thus generating an evolutionary path directed graph.

[0199] In an adjacency matrix, rows and columns correspond to node IDs, and each element represents the weight of the edge between the corresponding nodes. If two nodes are directly connected, the value of the matrix element is the weight of the directed edge; if there is no direct connection between the two nodes, the value of the matrix element is zero. In this way, the adjacency matrix can clearly represent the connection relationships between nodes in a directed graph and the weight of the edges.

[0200] The topological structure of the directed graph of evolutionary paths is stored in an adjacency matrix, which facilitates subsequent path search and analysis. During path search, the connection information between nodes and the weights of edges can be directly obtained from the adjacency matrix, allowing for the calculation of the cumulative weights of different paths. By storing the topological structure of the directed graph in an adjacency matrix, a complete directed graph of evolutionary paths is ultimately generated.

[0201] Step S550: Perform path search on the directed graph of the evolution path, calculate the cumulative weight of all possible paths, and select the path with the largest cumulative weight as the optimal evolution path.

[0202] Specifically, it may include the following steps: Step S551: Sort the nodes of the directed graph of the evolution path in ascending order of timestamps, construct the node access sequence, and determine the key time nodes by timestamp density clustering. The cluster centers correspond to the feature mutation points in the time series.

[0203] Before performing a path search, the nodes in the directed evolutionary path graph need to be sorted to ensure a more orderly traversal. Sort the nodes in the directed evolutionary path graph in ascending order of timestamps. This ensures that the order in which nodes are visited corresponds to the chronological order of the specimen's evolution. A node visit sequence is constructed, recording the order in which nodes are visited. During the path search, visiting nodes sequentially according to this sequence improves the efficiency and accuracy of the search.

[0204] Step S552: Perform path search on the directed graph. Define the state transition equation as the cumulative weight of the current node equals the sum of the cumulative weight of the previous node and the edge weight. Set the initial state to the cumulative weight of the earliest node as the weight of its self-loop edge.

[0205] After constructing the node visit sequence and identifying key time nodes, path searching can begin on the directed graph. To calculate the cumulative weight of each path, a state transition equation needs to be defined. This equation specifies how the cumulative weight is calculated when moving from one node to another. In this case, the cumulative weight of the current node is equal to the sum of the cumulative weight of its predecessor node and the weight of the edge connecting them. This calculation method aligns with the logic of path cumulative weights; by progressively accumulating edge weights, the cumulative weight from the starting node to the current node can be obtained. The initial state is set to the cumulative weight of the earliest time node, which is its self-loop edge weight. The earliest time node is the starting point of the evolutionary path; its cumulative weight is set to a self-loop edge weight because initially, the path only contains this node itself, and the self-loop edge weight reflects the semantic stability of that node. Starting from this initial state, the cumulative weights of subsequent nodes are calculated progressively according to the state transition equation, thus completing the calculation of the cumulative weights for all paths in the directed graph.

[0206] Step S553: ​​Maintain multiple candidate paths for each node. The number of candidate paths is dynamically adjusted according to the node's out-degree. A set proportion of paths with the highest cumulative weight are retained to avoid combinatorial explosion. The path storage format includes the node sequence and the total cumulative weight.

[0207] In the path search process, to find the optimal evolutionary path, it is necessary to maintain multiple candidate paths for each node. Each node may have multiple outgoing edges, and there may be different path choices starting from that node. Maintaining multiple candidate paths can take into account all possible path situations.

[0208] The number of candidate paths is dynamically adjusted based on the out-degree of a node. The out-degree of a node represents the number of edges originating from that node; the larger the out-degree, the more possible path choices there are. Dynamically adjusting the number of candidate paths based on the out-degree ensures comprehensive search while avoiding maintaining too many unnecessary paths.

[0209] Step S554: Introduce key node constraints during the path search process, forcing all key time nodes to be included in the path. If a candidate path is missing a key node, it is filled by skipping connecting edges, and the weight of the filled edge is reduced according to the semantic relevance.

[0210] To ensure that the searched paths reflect the key stages of specimen evolution, key node constraints are introduced during the path search process. If a candidate path is found to be missing key nodes during the search, the path is completed using jump edges. Jump edges connect non-contiguous label nodes that have a semantic relevance exceeding a set value, ensuring that the path includes key nodes.

[0211] The weight of the completed edge is reduced based on semantic relevance. Since jump edges connect non-contiguous nodes, the semantic relevance between them may not be as strong as that between adjacent nodes. Therefore, when calculating the cumulative weight of the path, the weight of the completed edge needs to be reduced. The degree of reduction is determined by the semantic relevance; the lower the semantic relevance, the greater the weight reduction of the completed edge.

[0212] Step S555: Calculate the cumulative weight of all complete paths. The cumulative weight is a weighted combination of the sum of the weights of all edges in the path and the self-similarity of the semantic feature vectors of the nodes. Sort the paths in descending order of cumulative weight.

[0213] After completing the path search and filling in all key nodes, the cumulative weight of all complete paths is calculated. The cumulative weight is calculated as a weighted combination of the sum of the weights of all edges in the path and the self-similarity of the semantic feature vectors of the nodes. The edge weights reflect the semantic similarity between adjacent nodes, while the self-similarity of the node's semantic feature vectors reflects the semantic stability of the node itself. This calculation method comprehensively considers the connection strength of the edges and the stability of the nodes in the path, and more comprehensively evaluates the merits of each path.

[0214] After calculating the cumulative weights of all complete paths, the paths are sorted in descending order of cumulative weight. This descending order places the path with the highest cumulative weight first, facilitating the direct selection of the optimal path later. This sorting method allows for a clear comparison of the cumulative weights of different paths, providing an intuitive basis for selecting the optimal evolutionary path.

[0215] Step S556: Select the path with the largest cumulative weight as the optimal evolution path, and verify whether the path covers all key time nodes and semantic coherence. If there are missing nodes, backtrack and adjust the path search parameters and recalculate.

[0216] After sorting the paths in descending order of cumulative weight, the path with the highest cumulative weight is selected as the optimal evolutionary path. The optimal evolutionary path represents the most reasonable and coherent path in the evolutionary process of the specimen, and can most accurately reflect the evolutionary trend of the specimen's characteristics and the semantics of ecological evolution.

[0217] Step S560: Integrate the tags on the optimal evolutionary path, arrange the semantic descriptions and evolutionary direction indicators of the tags in chronological order, and generate two-dimensional structured annotation tags containing time axis and semantic axis. The two-dimensional structured annotation tags have a complete expression of ecological evolution semantics.

[0218] Once the optimal evolutionary path is determined, labels with ecological evolutionary semantics can be generated based on this path. Specifically, this involves integrating the label entries along the optimal evolutionary path.

[0219] Arranging the labels along the optimal evolution path in chronological order clearly demonstrates the changes in specimen features at different time points. Each label includes a semantic description and an evolutionary direction indicator. The semantic description details the specimen's feature state at that time point, while the evolutionary direction indicator points to the direction of change of the specimen's features from one time point to the next. Arranging the semantic descriptions and evolutionary direction indicators in chronological order generates a two-dimensional structured label system containing a time axis and a semantic axis. The time axis represents the temporal sequence of specimen evolution, and the semantic axis represents the semantic state and direction of change of the specimen's features.

[0220] The various algorithms involved in the above descriptions of the embodiments of the present invention can all be obtained from relevant content in the prior art. To save space, they will not be elaborated on in the embodiments of the present invention. In addition, those skilled in the art can supplement the details based on common knowledge in the art when implementing the solutions of the present invention. For example, they can use normalization to eliminate dimensional conflicts before feature fusion, use interpolation to eliminate dimensional differences, reasonably set thresholds based on historical data, experience or business scenario requirements, train the model based on a general model training method, set the number of layers in the model structure based on actual needs, select activation functions, etc. The present invention will not provide redundant descriptions of overly detailed implementation processes.

[0221] Please see Figure 2 , Figure 2 This is a schematic diagram of a computer system provided in an embodiment of the present invention. The computer system includes at least a processor 101, a communication interface 102, and a memory 103. The processor 101, communication interface 102, and memory 103 can be connected via a bus or other means. The processor 101 (or Central Processing Unit, CPU) is the computing and control core of the computer system, capable of parsing various instructions and processing various data within the computer system. The communication interface 102 may optionally include a standard wired interface or a wireless interface (such as Wi-Fi, mobile communication interface, etc.), and can be used to send and receive data under the control of the processor 101; the communication interface 102 can also be used for data transmission and interaction within the computer system. The memory 103 is a storage device in the computer system used to store programs and data. It is understood that the memory 103 here can include the computer system's built-in memory, or it can include extended memory supported by the computer system. The memory 103 provides storage space, which stores the computer system's operating system; this invention does not limit this storage space.

[0222] In one embodiment, the processor 101 executes the specimen data annotation method for aquatic ecological laboratories provided above in the embodiments of the present invention by running a computer program in the memory 103.

Claims

1. A method for labeling specimen data in an aquatic ecology laboratory, characterized in that, The method includes: Acquire specimen samples collected by the aquatic ecology laboratory and corresponding multi-source environmental metadata, wherein the multi-source environmental metadata includes environmental parameter sequences and biological attribute observation data during the specimen collection process; Environmental factor response curves are constructed and time-varying features of biological attributes are extracted from environmental parameter sequences and biological attribute observation data. Based on the environmental factor response curves and biological attribute time-varying features, multi-source environmental metadata and specimen sample attributes are coupled to generate an initial specimen feature spectrum containing the nonlinear coupling relationship between environmental parameter trajectories and biological attribute response features. Specifically, this includes: performing multi-scale periodic decomposition on the environmental parameter sequences to separate environmental factor components with different oscillation periods; calculating the energy proportion of each environmental factor component and constructing a periodic contribution matrix with environmental factor type and oscillation period as dimensions; and selecting the dominant environmental factor component based on the periodic contribution matrix. Time series alignment is performed on the biological attribute observation data to extract biological attribute response sequences that match the timestamps of the dominant environmental factor components; cross-correlation analysis is performed on different environmental factor components and biological attribute response sequences to calculate lag response coefficients; and a lag response coefficient matrix is ​​constructed with the dominant environmental factor component and biological attribute response sequence as dimensions. A network of cross-influences between environmental factors and biological attributes is constructed, with environmental factor components and biological attribute indicators corresponding to non-zero elements in the hysteresis response coefficient matrix as network nodes and the absolute value of the hysteresis response coefficient as the edge weight. Topological analysis is performed on this network to identify key influence paths. A nonlinear fitting model for the environmental factor response curve is optimized using the topological parameters of the key influence paths. A joint modeling process is performed on the environmental parameter sequence and the biological attribute response sequence to generate a dynamic response surface for environmental factors. Simultaneously, a time-frequency transformation is performed on the biological attribute response sequence to extract instantaneous frequency and amplitude features, constructing a dynamic descriptor for the time-varying characteristics of biological attributes. The dynamic descriptor is coupled with the dynamic response surface of the environmental factors, and the coupled features are nonlinearly dimensionality-reduced to retain the cross-correlation information between environmental parameters and biological attributes. Based on the dimensionality-reduced features, an initial sample feature spectrum containing the nonlinear coupling relationship between the environmental parameter trajectory and the biological attribute response features is generated. The initial specimen feature spectrum is subjected to feature space mapping, multimodal semantic verification, and entity relationship reasoning to generate a context-enhanced annotation vector containing the semantic association strength and evolution feature weights of the specimen samples in a multi-level knowledge system. Historical trajectory similarity comparison and multi-dimensional evolution trend prediction are performed on the context-enhanced annotation vector to generate an intermediate evolution label set that records the semantic change process of specimen features at different time scales; By integrating intermediate evolutionary tag sets through multi-dimensional semantic consistency verification and evolutionary path optimization, annotation tags with ecological evolutionary semantics are generated.

2. The method according to claim 1, characterized in that, The process involves performing multi-scale periodic decomposition on the environmental parameter sequence to separate environmental factor components with different oscillation periods, calculating the energy proportion of each environmental factor component, constructing a periodic contribution matrix based on environmental factor type and oscillation period, and selecting dominant environmental factor components based on the periodic contribution matrix, including: Obtain the initial mode number range of the pre-set multi-scale periodic decomposition, optimize the mode number with the goal of maximizing the sum of the kurtosis values ​​of each component after decomposition, and determine the optimal mode decomposition parameters; Based on the optimal mode decomposition parameters, the environmental parameter sequence is decomposed into multiple scale periods to obtain multiple independent environmental factor components. Each environmental factor component contains time series data with a corresponding oscillation period. The energy value of each environmental factor component is calculated. The energy value is obtained by integrating the sum of squares of the component's time series. The energy values ​​of each component are then normalized to obtain the energy proportion distribution. A periodic contribution matrix is ​​constructed with environmental factor type as the row and oscillation period as the column. The elements of the periodic contribution matrix are the energy proportion of the corresponding environmental factor in the corresponding period component. Obtain a preset contribution threshold and filter environmental factor components whose element values ​​in the periodic contribution matrix exceed the contribution threshold as candidate dominant components; Redundancy analysis is performed on the candidate dominant components, the mutual information value between the candidate dominant components is calculated, and candidate dominant components with mutual information values ​​exceeding a set value are removed.

3. The method according to claim 2, characterized in that, The process involves time-series alignment of biological attribute observation data, extraction of biological attribute response sequences matching the timestamps of the dominant environmental factor components, cross-correlation analysis of different environmental factor components and biological attribute response sequences to calculate lag response coefficients, and construction of a lag response coefficient matrix with the dominant environmental factor components and biological attribute response sequences as dimensions, including: Extract the timestamp set of the environmental parameter sequence, and perform linear interpolation on the timestamps of the biological attribute observation data to keep the biological attribute observation data synchronized with the environmental parameter sequence in the time dimension; Biological attribute index sequences are extracted from aligned biological attribute observation data. Based on the time interval of the dominant environmental factor components, biological attribute index sequences within the corresponding time period are extracted as biological attribute response sequences. A sliding window cross-correlation analysis was performed on each dominant environmental factor component and the biological attribute response sequence. The window length was dynamically adjusted according to the oscillation period of the environmental factor component, and the correlation coefficient under different lag times was calculated. Record the lag time and correlation coefficient value corresponding to the maximum correlation coefficient, and determine the correlation coefficient value as the lag response coefficient. A hysteresis response coefficient matrix is ​​constructed with environmental factor components as rows and biological attribute response sequences as columns, where the matrix elements are the hysteresis response coefficients of the corresponding combinations.

4. The method according to claim 3, characterized in that, The process involves constructing an environmental factor-biological attribute cross-influence network using the environmental factor components and biological attribute indicators corresponding to the non-zero elements in the hysteresis response coefficient matrix as network nodes and the absolute value of the hysteresis response coefficient as the edge weight. Topology analysis is then performed on this network to identify key influence paths, including: The environmental factor components and biological attribute indicators corresponding to the non-zero elements in the hysteresis response coefficient matrix are determined as network nodes, and the node attributes include factor type and time scale characteristics. Using the absolute value of the hysteresis response coefficient as the edge weight, a directed weighted network is constructed, with the direction of the directed edges pointing from the environmental factor components to the biological attribute indicators. Degree centrality, betweenness centrality, and proximity centrality are calculated for each node in the network. Degree centrality represents the number of direct connections of a node, betweenness centrality represents the mediating role of a node in a path, and proximity centrality represents the average distance from a node to other nodes. Obtain a pre-set centrality threshold, filter nodes whose centrality index exceeds the centrality threshold as key nodes, and extract the directed edges between key nodes to form an initial set of influence paths; Perform path length analysis on the initial set of impact paths, retain short paths whose path lengths meet the set conditions, and calculate the total weight of each path. The total weight is the product of the weights of the edges on the path. The paths are sorted in descending order of total weight, and the top K paths are selected as the key impact paths, where K is greater than 0.

5. The method according to claim 1, characterized in that, The nonlinear fitting model for optimizing the environmental factor response curve using the topological parameters of the key influence path jointly models the environmental parameter sequence and the biological attribute response sequence to generate a dynamic response surface for the environmental factor, including: Extract the topological parameters of key influencing paths, including path length, node degree, and edge weight distribution. Use these topological parameters as prior knowledge for the nonlinear fitting model of environmental factor response curves. A nonlinear fitting model is constructed that includes a combination of kernel functions, wherein the combination of kernel functions is a weighted combination of radial basis functions and periodic kernel functions, and the weights are dynamically adjusted based on path topology parameters; The environmental parameter sequence is used as the model input, and the biological attribute response sequence is used as the model output. The nonlinear fitting model is trained, and the kernel function parameters are optimized by maximum likelihood estimation. Based on the trained model, the response to the environmental parameter sequence is predicted, a predicted sequence of biological attribute response is generated, and the root mean square error between the predicted sequence and the actual observed sequence is calculated. The kernel function weights are adjusted based on error feedback, and the model is iteratively optimized until the error is lower than a set threshold, thus obtaining the optimized environmental factor response curve model. The optimized environmental factor response curve model is used to simulate the response of different combinations of environmental parameters, and generate dynamic response surfaces of environmental factors that include the relationship between multi-dimensional environmental parameters and biological attributes.

6. The method according to claim 1, characterized in that, The process of performing feature space mapping, multimodal semantic verification, and entity relation reasoning on the initial specimen feature spectrum to generate a context-enhanced annotation vector containing the semantic association strength and evolutionary feature weights of the specimen samples in a multi-level knowledge system includes: A semantic space for specimen features is constructed based on the hierarchical structure of the aquatic ecological knowledge system. The dimensions of the space correspond to the hierarchical structure of the knowledge system, and the coordinate values ​​of each dimension represent the degree of association between the specimen features and the semantic concepts at that level. The initial specimen feature spectrum is subjected to a deep metric learning space mapping, and the initial specimen feature spectrum is projected onto the specimen feature semantic space to generate an initial semantic coordinate vector. Calculate the similarity between the initial semantic coordinate vector and the standard semantic template in the knowledge system, adjust the spatial mapping parameters according to the similarity distribution, and optimize the spatial distribution of the semantic coordinate vector; A multimodal semantic verification rule base is constructed using semantic verification rules from multiple dimensions including morphology, physiology, and ecology. Multimodal consistency verification is performed on the optimized semantic coordinate vectors, and conflicting components in the semantic coordinate vectors are corrected based on the verification results. Entity relationship reasoning is performed on the implicit association between specimen features and conceptual entities in the knowledge system, and the association strength is calculated. By integrating association strength and evolutionary feature weights into semantic coordinate vectors, a context-enhanced annotation vector is generated that contains the semantic association strength and evolutionary feature weights of specimen samples in a multi-level knowledge system.

7. The method according to claim 6, characterized in that, The specimen feature semantic space is constructed using the hierarchical structure of the aquatic ecological knowledge system. The spatial dimensions correspond to the hierarchical structure of the knowledge system, and the coordinate values ​​of each dimension represent the degree of association between the specimen features and the semantic concepts at that level, including: To acquire the hierarchical structure of the aquatic ecological knowledge system, a six-level semantic hierarchy is constructed from phylum, class, order, family, genus to species, with each semantic hierarchy containing multiple semantic concept nodes; A corresponding feature dimension is constructed for each semantic level to obtain a multidimensional semantic space framework, wherein the size of the feature dimension is equal to the number of semantic concept nodes at that semantic level; Collect standard feature descriptions of semantic concept nodes at each level, and perform vector transformation on the standard feature descriptions to obtain concept feature vectors; Calculate the similarity between the initial specimen feature spectrum and the feature vectors of concepts at each semantic level, and use the normalized similarity values ​​as the coordinate values ​​of the corresponding dimensions in the semantic space; A distance metric function is constructed for the semantic space, which measures the similarity of feature vectors of different specimens in the semantic space based on Mahalanobis distance. The distance metric function integrates the semantic association weights between semantic levels.

8. The method according to claim 7, characterized in that, The step of performing a deep metric learning space mapping on the initial specimen feature spectrum, projecting the initial specimen feature spectrum onto the specimen feature semantic space, and generating an initial semantic coordinate vector includes: A deep metric learning network is constructed, which includes a feature extraction layer and a mapping layer. The feature extraction layer adopts a residual network structure, and the mapping layer adopts a fully connected network structure. The initial specimen feature spectrum is standardized and preprocessed to make the data of each dimension of the initial specimen feature spectrum conform to a zero-mean unit variance distribution, so as to serve as the input data of the deep metric learning network. Obtain a preset triplet loss function, which includes the distance constraint between anchor samples, positive samples, and negative samples. The positive samples are the feature spectra of specimens of the same semantic category, and the negative samples are the feature spectra of specimens of different semantic categories. The parameters of the deep metric learning network are optimized by backpropagation algorithm, minimizing the spatial distance between anchor samples and positive samples, maximizing the spatial distance between anchor samples and negative samples, and iteratively training until the triplet loss function converges. The trained mapping layer parameters are fixed, and the preprocessed initial specimen feature spectrum is input for forward propagation to obtain the semantic space coordinate vector output by the network. The semantic space coordinate vector is subjected to L2 normalization to ensure that the coordinate values ​​of each dimension are within a set range, thereby generating an initial semantic coordinate vector.

9. A computer system, characterized in that, include: A memory, wherein a computer program is stored; A processor for loading the computer program to implement the specimen data annotation method for use in an aquatic ecology laboratory as described in any one of claims 1-8.