Specimen data labeling method and system applied to water ecology laboratory

By constructing environmental factor response curves and time-varying features of biological attributes, a specimen feature spectrum with nonlinear coupling relationship is generated. This solves the problem that the correlation between environmental parameters and biological attributes in aquatic ecological laboratories has not been deeply explored, and achieves close integration of specimen feature data and semantically rich annotation results, supporting ecological evolution analysis.

CN121637148AActive Publication Date: 2026-03-10BEIJING NORMAL UNIVERSITY +1
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-14
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

In the data annotation methods of the aquatic ecology laboratory, the relationship between environmental parameters and biological attributes has not been deeply explored, resulting in weak data correlation, one-sided semantic information expression, and difficulty in supporting ecological evolution analysis.

Method used

By acquiring specimen samples and multi-source environmental metadata, we construct environmental factor response curves and time-varying features of biological attributes, generate an initial specimen feature spectrum of nonlinear coupling relationships, and perform feature space mapping, multimodal semantic verification and entity relationship reasoning to generate context-enhanced annotation vectors and track the semantic change process of specimen features.

Benefits of technology

Dynamically capturing the intrinsic relationship between environmental parameters and biological attributes enhances the fusion tightness and semantic accuracy of specimen feature data, enriches the knowledge content of annotation results, and fully records the ecological evolution process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121637148A_ABST
    Figure CN121637148A_ABST
Patent Text Reader

Abstract

The invention provides a specimen data labeling method and system applied to a water ecology laboratory, and the method comprises the steps: obtaining a specimen sample collected by the water ecology laboratory and corresponding multi-source environment metadata, and carrying out the construction of an environmental factor response curve and the extraction of a biological attribute time-varying feature of an environmental parameter sequence and biological attribute observation data; coupling the multi-source environment metadata with the specimen sample attributes to generate an initial specimen characteristic spectrum; performing feature space mapping, multi-modal semantic verification and entity relationship reasoning on the context-enhanced annotation vector to generate a context-enhanced annotation vector; performing historical trajectory similarity comparison and multi-dimensional evolution trend prediction on the data to generate an intermediate evolution label set; and the intermediate evolution label set is optimized and integrated through multi-dimensional semantic consistency verification and an evolution path, and a labeling label with ecological evolution semantics is generated. Through the method provided by the invention, the recording capability of the labeling result on the ecological evolution process of the specimen can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of data processing, in particular to a specimen data labeling method and system applied to a water ecological laboratory. BACKGROUND

[0002] In the field of water ecological research, specimen data labeling is a process of describing the characteristics and marking the information of water biological specimens collected in the laboratory. The labeling results can provide basic data support for species identification, ecological evolution analysis and other researches. At present, the specimen data labeling of the water ecological laboratory usually relies on manual or semi-automatic methods to label the morphological characteristics or basic attributes of the specimens in isolation. In the labeling process, the environmental parameters and biological attributes are often processed in a linear superposition or independent recording manner, lacking deep mining of the dynamic correlation between the two. At the same time, the labeling results are mostly presented in the form of static labels, which are difficult to reflect the evolution process of specimen characteristics with time scale. This traditional labeling method leads to weak data correlation, one-sided semantic information expression, and difficulty in supporting systematic analysis of the ecological evolution process of specimens, affecting the accuracy and depth of data application in subsequent water ecological research. SUMMARY

[0003] The present application provides a specimen data labeling method and system applied to a water ecological laboratory.

[0004] In a first aspect, the present application provides a specimen data labeling method applied to a water ecological laboratory, which comprises: acquiring specimen samples collected by a water ecological laboratory and corresponding multi-source environmental metadata, wherein the multi-source environmental metadata includes environmental parameter sequences and biological attribute observation data in the specimen collection process; constructing an environmental factor response curve and extracting a biological attribute time-varying feature for the environmental parameter sequences and the biological attribute observation data, coupling the multi-source environmental metadata with the specimen sample attributes based on the environmental factor response curve and the biological attribute time-varying feature, and generating an initial specimen feature spectrum containing a nonlinear coupling relationship between environmental parameter trajectories and biological attribute response features; performing feature space mapping, multi-modal semantic verification and entity relationship reasoning on the initial specimen feature spectrum to generate a context-enhanced labeling vector containing the semantic correlation strength and evolution feature weight of the specimen samples in a multi-level knowledge system; performing historical trajectory similarity comparison and multi-dimensional evolution trend prediction on the context-enhanced labeling vector to generate an intermediate evolution label set recording the semantic change process of specimen characteristics at different time scales; and integrating the intermediate evolution label set through multi-dimensional semantic consistency verification and evolution path optimization to generate a labeling label with ecological evolution semantics.

[0005] In a second aspect, an embodiment of the present application provides a computer system, comprising: a memory, wherein a computer program is stored in the memory; and a processor configured to load the computer program to implement the specimen data labeling method applied to an aquatic ecological laboratory as described above.

[0006] The specimen data labeling method applied to an aquatic ecological laboratory provided by the present application can dynamically capture the internal correlation between the environmental parameter trajectory and the biological attribute response characteristics by constructing an environmental factor response curve and extracting time-varying characteristics of biological attributes, coupling multi-source environmental metadata and specimen sample attributes to generate an initial specimen characteristic spectrum containing a nonlinear coupling relationship, avoiding the problem of loose data correlation caused by independent processing of environmental data and biological attributes in traditional labeling methods, and improving the fusion tightness of specimen characteristic data; by performing feature space mapping, multi-modal semantic verification and entity relationship reasoning on the initial specimen characteristic spectrum, a context-enhanced labeling vector containing multi-level knowledge system semantic correlation strength and evolution characteristic weight is generated, which can deeply bind the specimen characteristics and the aquatic ecological knowledge system, ensure the consistency of the labeling semantics through multi-dimensional semantic verification, avoid the semantic one-sidedness caused by single-dimensional labeling, and improve the semantic accuracy and knowledge richness of the labeling result; by comparing the historical trajectory similarity and predicting the multi-dimensional evolution trend of the context-enhanced labeling vector, integrating the multi-dimensional semantic consistency verification and the evolution path optimization, and generating labeling labels with ecological evolution semantics, the semantic change process of the specimen characteristics under different time scales can be tracked, the limitation of static labeling only recording the instantaneous state can be avoided, and the complete recording ability of the labeling result on the ecological evolution process of the specimen can be improved. BRIEF DESCRIPTION OF DRAWINGS

[0007] Figure 1 is a flowchart of a specimen data labeling method applied to an aquatic ecological laboratory provided by an embodiment of the present application.

[0008] Figure 2 is a composition schematic diagram of a computer system provided by an embodiment of the present application. DETAILED DESCRIPTION

[0009] Please refer to Figure 1 is a flowchart of a specimen data labeling method applied to an aquatic ecological laboratory provided by an embodiment of the present application. The method can be executed by a computer system, and the method can include the following steps: Step S100: acquiring specimen samples collected by an aquatic ecological laboratory and corresponding multi-source environmental metadata, wherein the multi-source environmental metadata includes environmental parameter sequences and biological attribute observation data in the specimen collection process.

[0010] The specimen sample is a biological sample collected from a water ecological environment, and the multi-source environmental metadata is a collection of multiple types of data related to specimen collection. Among them, the environmental parameter sequence is a record of the change of environmental factors over time during the specimen collection process, such as the specific values of environmental indicators such as water temperature, pH, and dissolved oxygen concentration at different time points, reflecting the dynamic characteristics of the specimen's living environment. The biological attribute observation data is the observation and record of the biological characteristics of the specimen itself, such as individual size, reproductive capacity, and growth rate, reflecting the growth and development state of the organism in the environment.

[0011] Step S200: Environmental factor response curve construction and biological attribute time-varying feature extraction are performed on the environmental parameter sequence and the biological attribute observation data. Based on the environmental factor response curve and the biological attribute time-varying feature, the multi-source environmental metadata and the specimen sample attribute are coupled to generate an initial specimen feature spectrum containing the nonlinear coupling relationship between the environmental parameter trajectory and the biological attribute response feature.

[0012] The environmental factor response curve is a curve describing the relationship between biological attributes and environmental factors, reflecting the response law of organisms under different environmental factor conditions. The biological attribute time-varying feature is the change of biological attributes over time, which can reflect the dynamic adaptation process of organisms in the environmental change process. Coupling is to associate and integrate multi-source environmental metadata and specimen sample attributes, making the relationship between them more close and clear, so as to comprehensively reflect the characteristics of the specimen under the corresponding environment.

[0013] As an implementation, step S200 can include steps S210-S260 as follows: Step S210: Multi-scale periodic decomposition is performed on the environmental parameter sequence to separate environmental factor components of different oscillation periods, calculate the energy proportion of each environmental factor component, and construct a periodic contribution degree matrix with environmental factor type and oscillation period as dimensions. Based on the periodic contribution degree matrix, the dominant environmental factor component is selected.

[0014] The environmental factor components of different oscillation periods represent the changes of environmental parameters at different time scales, and each component contains environmental change information of the corresponding period. The energy proportion is the proportion of the energy of each environmental factor component in the total energy, reflecting the contribution degree of the component to the entire environmental parameter sequence. The periodic contribution degree matrix is a matrix constructed with environmental factor type as row and oscillation period as column. The elements in the matrix represent the energy proportion of the corresponding environmental factor in the corresponding period component. Through the matrix, the contribution of different environmental factors in different periods can be intuitively understood. The dominant environmental factor component is the environmental factor component with a larger energy proportion in the periodic contribution degree matrix. The influence of these components on biological attributes may be more significant, and they are the key objects for subsequent analysis. When performing multi-scale periodic decomposition, methods such as empirical mode decomposition (EMD) and wavelet decomposition can be used.

[0015] As an implementation, step S210 can include steps S211-S216 as follows: Step S211: Obtain the initial modal number range of the multi-scale periodic decomposition, and optimize the modal number to maximize the sum of kurtosis values of the decomposed components, to determine the optimal modal decomposition parameter.

[0016] The initial modal number range of the multi-scale periodic decomposition is a modal number value interval that is pre-set by researchers according to experience and data characteristics. The kurtosis value is a statistical quantity describing the shape of data distribution, reflecting the degree of data distribution peak. In the multi-scale periodic decomposition, the modal number is optimized to maximize the sum of kurtosis values of the decomposed components, because the greater the sum of kurtosis values, the more the decomposed components can highlight the periodic characteristics of the data, and can more accurately reflect the internal structure of the environmental parameter sequence. The optimal modal decomposition parameter is the modal number that maximizes the sum of kurtosis values obtained in the optimization process, and using this parameter for decomposition can obtain environmental factor components that are more in line with the actual situation.

[0017] Step S212: Based on the optimal modal decomposition parameter, the environmental parameter sequence is subjected to multi-scale periodic decomposition to obtain a plurality of independent environmental factor components, each of which contains time series data of a corresponding oscillation period.

[0018] After the optimal modal decomposition parameter is determined, the parameter can be used to perform multi-scale periodic decomposition on the environmental parameter sequence. The purpose of multi-scale periodic decomposition is to decompose the environmental parameter sequence into a plurality of independent components, each of which represents an oscillation period. These environmental factor components contain time series data of the corresponding oscillation period, and through analysis of these components, the periodic variation characteristics of the environmental parameters can be better understood.

[0019] Step S213: Calculate the energy value of each environmental factor component, the energy value is obtained by integrating the square sum of the component time series, and the energy value of each component is normalized to obtain the energy proportion distribution.

[0020] The energy value reflects the energy proportion of the component in the entire environmental parameter sequence. In calculating the energy value, for the time series data of each environmental factor component, each data point is squared, and then the squared sequence is integrated. The integral can be calculated by numerical integration methods such as trapezoidal integration and Simpson integration. After obtaining the energy value of each component, the sum of all component energy values is calculated. Finally, the energy value of each component is divided by the total energy value to obtain the energy proportion of the component.

[0021] Step S214: Construct a period contribution matrix with environmental factor types as rows and oscillation periods as columns, and the elements of the period contribution matrix are the energy proportions of the corresponding environmental factors in the corresponding period components.

[0022] The period contribution matrix is a two-dimensional matrix with environmental factor types as rows and oscillation periods as columns. Each element in the matrix represents the energy proportion of the corresponding environmental factor in the corresponding period component. By constructing the period contribution matrix, the contribution of different environmental factors under different oscillation periods can be intuitively displayed, which helps to analyze the relationship between environmental factors and biological attributes. When constructing the period contribution matrix, first, the specific values of environmental factor types and oscillation periods are determined. Environmental factor types can include common environmental indicators such as water temperature, light intensity, and dissolved oxygen concentration; oscillation periods are determined according to different components obtained by multi-scale period decomposition. Then, fill the energy proportion of each environmental factor in different oscillation period components into the corresponding position of the matrix.

[0023] Step S215: Obtain a preset contribution threshold, and select environmental factor components with element values exceeding the contribution threshold in the period contribution matrix as candidate dominant components.

[0024] The candidate dominant component is an environmental factor component with an element value exceeding the contribution threshold in the period contribution matrix. These components may have a more significant impact on biological attributes, so they are selected as objects for further analysis. When obtaining the preset contribution threshold, the focus of the study, the characteristics of the data, and previous research experience can be considered comprehensively. For example, if the study focuses more on environmental factors that have a greater impact on biological attributes, set the contribution threshold higher; if you want to cover more environmental factors that may have an impact, set the threshold lower. Then, traverse each element in the period contribution matrix, and select environmental factor components with element values exceeding the contribution threshold as candidate dominant components.

[0025] Step S216: Perform redundancy analysis on the candidate dominant components, calculate the mutual information values between the candidate dominant components, and remove candidate dominant components with mutual information values exceeding the set value.

[0026] Redundancy analysis is to remove redundant information in candidate dominant components to avoid repeated analysis and interference. Mutual information value is an index to measure the correlation between two random variables, which is used to measure the correlation between candidate dominant components in this step. If the mutual information value between two candidate dominant components exceeds the set value, it means that there is a strong correlation between them, and they may contain redundant information. Therefore, one of the components needs to be removed to improve the efficiency and accuracy of subsequent analysis.

[0027] In the redundancy analysis, mutual information values between candidate principal components are calculated first. Entropy-based methods can be used to calculate the mutual information values. For example, the entropy of each candidate principal component is calculated first, which represents the uncertainty of a random variable. Then the joint entropy of two candidate principal components is calculated. Finally, according to the definition of mutual information, the mutual information is equal to the sum of the entropies of the two components minus their joint entropy. After obtaining the mutual information values, they are compared with a set value. If the mutual information value exceeds the set value, one of the candidate principal components is removed.

[0028] Step S220: Time series alignment of biological attribute observation data, extraction of biological attribute response sequences matching the timestamps of dominant environmental factor components, cross-correlation analysis of different environmental factor components and biological attribute response sequences to calculate lag response coefficients, and construction of lag response coefficient matrix with dominant environmental factor components and biological attribute response sequences as dimensions.

[0029] Time series alignment is to unify biological attribute observation data and dominant environmental factor components in the time dimension, so that their timestamps can be matched for subsequent analysis. Biological attribute response sequences are biological attribute observation data matching the timestamps of dominant environmental factor components, reflecting the response of the organism to the corresponding environmental factor changes. Through cross-correlation analysis, the lag response coefficient of the biological attribute response sequence relative to the dominant environmental factor component can be calculated, which represents the response delay time of the biological attribute to the environmental factor change. The lag response coefficient matrix is a matrix constructed with dominant environmental factor components as rows and biological attribute response sequences as columns. The matrix elements represent the lag response coefficients of the corresponding combinations. Through the matrix, the lag response relationship between different environmental factors and biological attributes can be intuitively understood.

[0030] As an implementation, step S220 can include steps S221-S225 as follows: Step S221: Extract the timestamp set of the environmental parameter sequence, and perform linear interpolation processing on the timestamps of the biological attribute observation data to keep the biological attribute observation data and the environmental parameter sequence synchronized in the time dimension.

[0031] The timestamp set of the environmental parameter sequence records the specific time points of environmental parameter data collection, reflecting the time sequence of environmental parameter changes. Linear interpolation processing estimates the data values at unknown time points by linear fitting between known data points. Linear interpolation processing is performed on the timestamps of the biological attribute observation data to keep the biological attribute observation data and the environmental parameter sequence synchronized in the time dimension. When extracting the timestamp set of the environmental parameter sequence, the timestamp corresponding to each data point in the environmental parameter sequence is directly obtained from the data record of the environmental parameter sequence. Then, for the timestamps of the biological attribute observation data, when there are missing or mismatching cases, linear interpolation method is used for processing.

[0032] Step S222: Extract the biological attribute index sequence from the aligned biological attribute observation data, and extract the biological attribute index sequence within the corresponding time period as the biological attribute response sequence based on the time interval of the dominant environmental factor component.

[0033] The biological attribute index sequence is a series of biological attribute data extracted from aligned biological attribute observation data, reflecting the attribute changes of organisms within a specific time range. The time interval of the dominant environmental factor component is the time range corresponding to that environmental factor component, representing a specific period of time for environmental factor changes.

[0034] When extracting biological attribute indicator sequences from aligned biological attribute observation data, the corresponding indicator data are selected from the data based on the type of biological attribute the study focuses on. For example, if the study focuses on the growth rate of organisms, then relevant growth rate data are extracted from the aligned data as the biological attribute indicator sequence. Then, based on the time interval of the dominant environmental factor components, the start and end times for the extraction are determined, and data within the corresponding time period is extracted from the biological attribute indicator sequence to obtain the biological attribute response sequence.

[0035] Step S223: Perform sliding window cross-correlation analysis on each dominant environmental factor component and the biological attribute response sequence. The window length is dynamically adjusted according to the oscillation period of the environmental factor component, and the correlation coefficient under different lag times is calculated.

[0036] Sliding window cross-correlation analysis is used to analyze the correlation between two time series. It calculates the correlation between the two series within a fixed-length window by sliding the window across the time series. The window length is dynamically adjusted according to the oscillation period of the environmental factor component to better capture changes in the correlation between environmental factors and biological attributes. The correlation coefficients at different lag times reflect the degree of correlation between the biological attribute response sequence and the dominant environmental factor component at different delay times.

[0037] When performing sliding window cross-correlation analysis, the window length is first determined based on the oscillation period of the dominant environmental factor component. For example, if the oscillation period is short, the window length can be set relatively small; if the oscillation period is long, the window length needs to be set larger. Then, the window is simultaneously slid across the dominant environmental factor component and the biological attribute response sequence, and the correlation coefficient between the two sequences is calculated at each window position. For each window position, the lag time of the biological attribute response sequence relative to the dominant environmental factor component is also varied, and the correlation coefficient at different lag times is calculated.

[0038] Step S224: record the lag time corresponding to the maximum correlation coefficient and the correlation coefficient value, and determine the correlation coefficient value as the lag response coefficient.

[0039] The lag time corresponding to the maximum correlation coefficient represents the optimal response delay time of the biological attribute response sequence relative to the dominant environmental factor component, reflecting the response delay characteristics of the biological attribute to the change of the environmental factor. The maximum correlation coefficient value is determined as the lag response coefficient because this coefficient can most effectively reflect the correlation strength between the biological attribute and the environmental factor.

[0040] After obtaining the correlation coefficients at different lag times through sliding window cross-correlation analysis, the correlation coefficients are compared to find the maximum value. The lag time corresponding to the maximum correlation coefficient and the correlation coefficient value are recorded. Then, the maximum correlation coefficient value is taken as the lag response coefficient for subsequent analysis and modeling to describe the lag response relationship between the biological attribute and the environmental factor.

[0041] Step S225: construct a lag response coefficient matrix with environmental factor components as rows and biological attribute response sequences as columns, and the matrix elements are the lag response coefficients of the corresponding combinations. The negative coefficients in the lag response coefficient matrix are converted to absolute values to make the lag response coefficient matrix non-negative.

[0042] The lag response coefficient matrix is a two-dimensional matrix with environmental factor components as rows and biological attribute response sequences as columns. Each element in the matrix represents the lag response coefficient of the corresponding combination of environmental factor components and biological attribute response sequences. When constructing the lag response coefficient matrix, the lag response coefficient of each combination of environmental factor components and biological attribute response sequences is filled into the corresponding position of the matrix. If a negative value appears in the calculation of the lag response coefficient, it means that there may be a negative correlation between the biological attribute and the environmental factor. However, in order to make the matrix non-negative, the negative coefficients are taken as absolute values.

[0043] Step S230: construct an environmental factor-biological attribute cross-influence network with the environmental factor components and biological attribute indicators corresponding to the non-zero elements in the lag response coefficient matrix as network nodes and the absolute values of the lag response coefficients as edge weights, and perform topological analysis on the environmental factor-biological attribute cross-influence network to identify key influence paths.

[0044] As an implementation manner, step S230 can include steps S231-S236 as follows: Step S231: determine the environmental factor components and biological attribute indicators corresponding to the non-zero elements in the lag response coefficient matrix as network nodes, and the node attributes include factor type and time scale characteristics.

[0045] The network node is a basic unit of the environmental factor-biological attribute cross-influence network, and represents an environmental factor component or a biological attribute index. The environmental factor component and the biological attribute index corresponding to the non-zero elements in the lag response coefficient matrix are determined as the network nodes, because these elements represent a certain correlation between the environmental factor and the biological attribute. The node attribute includes a factor type and a time scale feature. The factor type is used to distinguish whether the node is an environmental factor component or a biological attribute index, and the time scale feature reflects the time range of the change of the environmental factor or the biological attribute corresponding to the node.

[0046] In determining the network nodes, the lag response coefficient matrix is traversed to find the environmental factor components and the biological attribute indexes corresponding to the non-zero elements. For each node, a corresponding factor type attribute, such as an environmental factor type or a biological attribute type, is assigned. At the same time, according to the oscillation period of the environmental factor component or the time range of the biological attribute index, the time scale feature attribute is determined.

[0047] Step S232: A directed weighted network is constructed with the absolute value of the lag response coefficient as the edge weight, and the direction of the directed edge is that the environmental factor component points to the biological attribute index.

[0048] In the directed weighted network, the connection between nodes has a direction and a weight. The absolute value of the lag response coefficient is used as the edge weight, which can directly represent the influence strength between the environmental factor and the biological attribute. The direction of the directed edge is that the environmental factor component points to the biological attribute index, indicating that the environmental factor has an influence on the biological attribute. In constructing the directed weighted network, for the determined network nodes, according to the non-zero elements in the lag response coefficient matrix, a directed edge is established between the corresponding environmental factor component node and the biological attribute index node. The weight of the edge is the absolute value of the lag response coefficient, and the direction of the edge is from the environmental factor component node to the biological attribute index node.

[0049] Step S233: Degree centrality, betweenness centrality, and closeness centrality of each node in the network are calculated, wherein the degree centrality represents the number of direct connections of the node, the betweenness centrality represents the intermediary role of the node in the path, and the closeness centrality represents the average distance of the node to other nodes.

[0050] The degree centrality, betweenness centrality, and closeness centrality are indexes for describing the importance and topological structure of the network nodes. The degree centrality is the number of direct connections of the node, reflecting the activity and influence of the node in the network. The betweenness centrality is the frequency of the node appearing in all shortest paths in the network, embodying the intermediary role of the node in information propagation and path connection. The closeness centrality is the reciprocal of the average shortest path length of the node to all other nodes, reflecting the accessibility and propagation efficiency of the node in the network.

[0051] Step S234: Obtain a pre-set centrality threshold, filter nodes with centrality indicators exceeding the centrality threshold as key nodes, and extract directed edges between key nodes to form an initial influence path set.

[0052] In obtaining the pre-set centrality threshold, reference can be made to previous research experience and current network structure characteristics for setting. For example, if very important nodes are desired to be filtered, the centrality threshold can be set higher; if more possible influential nodes are desired to be covered, the threshold can be set lower. Then, the degree centrality, betweenness centrality and closeness centrality of each node in the network are compared with the centrality threshold, and nodes with centrality indicators exceeding the threshold are filtered as key nodes. Finally, directed edges between key nodes are extracted, and these directed edges are connected to form an initial influence path set.

[0053] Step S235: Perform path length analysis on the initial influence path set, retain short paths with path lengths meeting set conditions, and calculate total weights of each path, which is the product of edge weights on the path.

[0054] Path length analysis is a method of filtering paths in the initial influence path set. By retaining short paths with path lengths meeting set conditions, redundancy can be reduced, and efficiency and accuracy of analysis can be improved. The total weight of a path is the product of all edge weights on the path, reflecting the comprehensive influence strength of environmental factors on biological attributes on the path.

[0055] In performing path length analysis on the initial influence path set, the set conditions for path length can be determined first. For example, a maximum path length threshold can be set, and only paths with path lengths less than the threshold are retained. Then, for the retained paths, their total weights are calculated. Specifically, the weights of each edge in the path are multiplied to obtain the total weight of the path.

[0056] Step S236: Sort paths in descending order of total weight, and select the top K paths as key influence paths, where K is greater than 0, and the key influence paths cover the main association between environmental factors and biological attributes.

[0057] Sorting paths in descending order of total weight can clearly show the importance order of different paths. Selecting the top K paths as key influence paths, these paths have larger total weights, indicating that they have more significant influence on biological attributes.

[0058] In sorting paths in descending order of total weight, sorting algorithms can be used to sort paths. For example, quicksort algorithm or heap sort algorithm can be used to sort paths from large to small according to their total weights. Then, according to the pre-determined K value, the top K paths are selected as key influence paths.

[0059] Step S240: Optimize the nonlinear fitting model of the environmental factor response curve with the topological parameters of the key influence path, jointly model the environmental parameter sequence and the biological attribute response sequence to generate an environmental factor dynamic response surface, and perform time-frequency transformation on the biological attribute response sequence to extract the instantaneous frequency and instantaneous amplitude characteristics to construct a dynamic descriptor of the biological attribute time-varying characteristics.

[0060] As an implementation, step S240 can include steps S241-S246: Step S241: Extract the topological parameters of the key influence path, including path length, node degree, and edge weight distribution, and use the topological parameters as prior knowledge of the nonlinear fitting model of the environmental factor response curve.

[0061] The topological parameters of the key influence path are parameters that describe the structural characteristics of the key influence path, including path length, node degree, and edge weight distribution. Path length reflects the degree of indirect influence between environmental factors and biological attributes; node degree represents the number of connections of a node, reflecting the importance of the node in the network; and edge weight distribution represents the strength distribution of the influence between environmental factors and biological attributes.

[0062] When extracting the topological parameters of the key influence path, the key influence path is first analyzed. For path length, the number of edges contained in the path is directly counted. For node degree, the number of directly connected edges of each node is counted. For edge weight distribution, the value range, distribution pattern, and other characteristics of the edge weight are analyzed. Then, these topological parameters are integrated into the nonlinear fitting model of the environmental factor response curve as prior knowledge.

[0063] Step S242: Construct a nonlinear fitting model containing a kernel function combination, which is a weighted combination of radial basis functions and periodic kernel functions, with the weights dynamically adjusted based on path topological parameters.

[0064] In constructing a nonlinear fitting model containing a kernel function combination, radial basis functions and periodic kernel functions can be selected as basic kernel functions. Then, the radial basis functions and periodic kernel functions are combined to obtain a kernel function combination. The weight coefficients are dynamically adjusted based on the path topological parameters, for example, using a certain algorithm to calculate the weight coefficients according to the path length, node degree, and edge weight distribution. The weight of the periodic kernel function can be increased when the path length is longer, as a longer path may indicate a periodic influence; the weight of the radial basis function can be increased when the node degree is larger, as a larger node degree may indicate a stronger local influence.

[0065] Step S243: Train the nonlinear fitting model with the sequence of environmental parameters as input and the sequence of biological attribute responses as output, and optimize the kernel function parameters by maximum likelihood estimation.

[0066] The sequence of environmental parameters as input and the sequence of biological attribute responses as output is to let the nonlinear fitting model learn the relationship between environmental factors and biological attributes. Maximum likelihood estimation is a statistical method for estimating model parameters, and through maximum likelihood estimation, the kernel function parameters that best match the model output and the actual biological attribute response sequence can be found.

[0067] When training the nonlinear fitting model, the sequence of environmental parameters and the sequence of biological attribute responses are first preprocessed, such as normalization, to make the data have the same scale. Then, the sequence of environmental parameters is input into the nonlinear fitting model, and the model calculates the output result according to the current kernel function parameters. The model output result is compared with the actual biological attribute response sequence, and the likelihood function is calculated. The likelihood function represents the probability of observing the actual biological attribute response sequence given the model parameters. Through maximum likelihood estimation, the kernel function parameters are adjusted to maximize the likelihood function.

[0068] Step S244: Based on the trained model, the response of the sequence of environmental parameters is predicted to generate the predicted sequence of biological attribute responses, and the root mean square error of the predicted sequence and the actual observed sequence is calculated.

[0069] Based on the trained model, the response of the sequence of environmental parameters is predicted to test the prediction ability and accuracy of the model. After generating the predicted sequence of biological attribute responses, the root mean square error of the predicted sequence and the actual observed sequence is calculated to quantitatively evaluate the fitting effect of the model. The root mean square error reflects the average deviation between the predicted value and the actual value.

[0070] When predicting the response of the sequence of environmental parameters based on the trained model, the sequence of environmental parameters is input into the trained nonlinear fitting model, and the model calculates the output result according to the learned relationship to obtain the predicted sequence of biological attribute responses. When calculating the root mean square error of the predicted sequence and the actual observed sequence, first calculate the square of the difference between each corresponding data point of the predicted sequence and the actual observed sequence, then calculate the average of these square values, and finally take the square root of the average value.

[0071] Step S245: Adjust the kernel function weight based on the error feedback, iteratively optimize the model until the error is lower than the set threshold, and obtain the optimized environmental factor response curve model.

[0072] The iterative optimization model ensures the accuracy and stability of the model until the error is below the set threshold. The optimized environmental factor response curve model can more accurately describe the relationship between environmental factors and biological attributes.

[0073] When adjusting the kernel function weight based on error feedback, the root mean square error between the calculated prediction sequence and the actual observation sequence is analyzed to find the cause of the error. If the error is large, it indicates that the model fitting effect is not good, and the weight of the kernel function needs to be adjusted. Based on the size and direction of the error, certain adjustment strategies can be used to adjust the weight of the kernel function. For example, if the fitting error of the radial basis function is large, its weight can be appropriately reduced; if the fitting error of the periodic kernel function is large, its period parameter can be adjusted or its weight can be reduced. After adjusting the weight of the kernel function each time, the model is retrained and predicted, and the new root mean square error is calculated. This process is repeated until the error is below the set threshold.

[0074] Step S246: Response simulation is performed on different environmental parameter combinations through the optimized environmental factor response curve model to generate an environmental factor dynamic response surface containing the response relationship between multi-dimensional environmental parameters and biological attributes.

[0075] Response simulation on different environmental parameter combinations through the optimized environmental factor response curve model can comprehensively understand the relationship between environmental factors and biological attributes. The environmental factor dynamic response surface is a three-dimensional or multi-dimensional model that shows the relationship between multi-dimensional environmental parameters and biological attribute responses, and can intuitively reflect the changes in biological attributes under different values of environmental factors.

[0076] When performing response simulation, the range of environmental parameter combinations to be simulated is first determined. For example, if the study involves three environmental factors: water temperature, light intensity, and dissolved oxygen concentration, the value range of each environmental factor can be determined. Then, a series of different environmental parameter combinations are generated within this value range. These environmental parameter combinations are input into the optimized environmental factor response curve model, and the model calculates the corresponding biological attribute response values for each combination based on the learned relationship. The environmental parameter combinations and corresponding biological attribute response values are visualized to generate an environmental factor dynamic response surface. Three-dimensional plotting tools or multi-dimensional data visualization methods can be used to display the surface.

[0077] Step S250: The dynamic descriptor is coupled with the environmental factor dynamic response surface, and the cross-correlation information between environmental parameters and biological attributes is preserved through nonlinear dimensionality reduction of the coupled features.

[0078] The dynamic descriptor contains information of time-varying characteristics of biological attributes, such as instantaneous frequency and instantaneous amplitude characteristics, reflecting the changes of biological attributes over time. The dynamic response surface of environmental factors shows the relationship between multi-dimensional environmental parameters and biological attribute responses. Coupling the dynamic descriptor with the dynamic response surface of environmental factors can combine the time-varying characteristics of biological attributes with the influence of environmental factors, and more comprehensively describe the relationship between biological attributes and environmental parameters.

[0079] When coupling the dynamic descriptor with the dynamic response surface of environmental factors, the characteristics of the dynamic descriptor are first associated with the coordinates and values of the dynamic response surface of environmental factors. For example, the instantaneous frequency and instantaneous amplitude characteristics in the dynamic descriptor can be combined with the environmental parameters and biological attribute response values in the dynamic response surface of environmental factors to form a high-dimensional feature vector. When performing nonlinear dimensionality reduction on the coupled features, various nonlinear dimensionality reduction algorithms can be used, such as nonlinear extensions of principal component analysis (PCA), such as kernel principal component analysis (KPCA), or local linear embedding (LLE), isometric mapping (Isomap), etc.

[0080] Step S260: Based on the reduced features, an initial specimen feature spectrum containing the nonlinear coupling relationship between the environmental parameter trajectory and the biological attribute response characteristics is generated.

[0081] The reduced features have removed redundant information while retaining important cross-correlation information between environmental parameters and biological attributes. Generating an initial specimen feature spectrum based on these reduced features can represent the nonlinear coupling relationship between the environmental parameter trajectory and the biological attribute response characteristics in a concise and effective manner. The initial specimen feature spectrum is a comprehensive feature set that can provide an important basis for subsequent specimen classification, comparison, and analysis.

[0082] In generating the initial specimen feature spectrum, the reduced features are first arranged and organized. The reduced features can be arranged according to certain rules, such as grouping by type of environmental parameter or category of biological attribute. Then, according to the relationship between these features, a feature matrix or feature vector is constructed, which is the initial specimen feature spectrum. In the construction process, the nonlinear coupling relationship between the environmental parameter trajectory and the biological attribute response characteristics should be fully considered, such as through the product of features, nonlinear transformation, etc. to reflect this coupling relationship.

[0083] Step S300: Perform feature space mapping, multi-modal semantic verification, and entity relationship reasoning on the initial specimen feature spectrum to generate a context-enhanced annotation vector containing the semantic association strength and evolution characteristic weight of the specimen sample in the multi-level knowledge system.

[0084] Feature space mapping is the process of transforming the initial specimen feature spectrum from the original feature space to a new feature space that better reflects the semantic relationships between specimen features. Multi-modal semantic verification is a multi-dimensional semantic check of specimen features to ensure semantic consistency and accuracy. Entity relationship reasoning is the analysis of the relationship between specimen features and entities in the knowledge system to infer their association strength and evolution feature weight. Context-enhanced annotation vector is a vector that contains the semantic association strength and evolution feature weight of the specimen sample in the multi-level knowledge system, providing more rich context information for specimen annotation and understanding.

[0085] As an implementation, step S300 can include steps S310-S350 as follows: Step S310: Construct a specimen feature semantic space with the hierarchical structure of the water ecological knowledge system, the spatial dimensions correspond to the hierarchical structure of the knowledge system, and the coordinate value of each dimension represents the association degree of the specimen feature and the hierarchical semantic concept.

[0086] The hierarchical structure of the water ecological knowledge system is a multi-level knowledge structure that includes macro to micro and whole to part, such as ecosystem level, species level, and biological attribute level. The specimen feature semantic space is a space used to represent the association relationship between specimen features and semantic concepts in the knowledge system, and its dimensions correspond to the hierarchical structure of the knowledge system. The coordinate value of each dimension represents the association degree of the specimen feature and the hierarchical semantic concept. The higher the association degree, the better the matching degree of the specimen feature and the semantic concept.

[0087] As an implementation, step S310 can include steps S311-S316 as follows: Step S311: Obtain the hierarchical structure of the water ecological knowledge system, and construct six-level semantic hierarchy from door, class, order, family, genus to species, each semantic hierarchy containing multiple semantic concept nodes.

[0088] The hierarchical structure of the water ecological knowledge system is a framework for classifying and organizing biological and environmental information in the water ecosystem. The construction of six-level semantic hierarchy from door, class, order, family, genus to species is a common classification method in biology, which can systematically describe the classification relationship of organisms. Each semantic hierarchy contains multiple semantic concept nodes, which represent different classification units or feature concepts at this level.

[0089] In obtaining the hierarchical structure of the water ecological knowledge system, the biological taxonomy standard and related water ecological research literature can be referred to. For example, at the level of phylum, it can include Protozoa, Arthropoda, etc.; at the level of class, for Protozoa, it can include Mastigophora, Sarcodina, etc.; at the levels of order, family, genus, and species, it is successively subdivided. Each semantic concept node at each level has its specific definition and characteristic description.

[0090] Step S312: constructing a corresponding feature dimension for each semantic level to obtain a multi-dimensional semantic space framework, wherein the size of the feature dimension is equal to the number of semantic concept nodes of the semantic level.

[0091] The construction of the corresponding feature dimension for each semantic level is to associate the specimen characteristics with the hierarchical structure of the water ecological knowledge system. The multi-dimensional semantic space framework is a spatial structure for representing the relationship between the specimen characteristics and the semantic concepts in the knowledge system, and the dimensions correspond to the semantic levels, and the size of each dimension is equal to the number of semantic concept nodes of the semantic level.

[0092] In constructing the feature dimension, for each semantic level, the size of the feature dimension is determined according to the number of semantic concept nodes of the level. For example, at the level of phylum, if there are 5 semantic concept nodes (such as 5 different phyla), the size of the feature dimension corresponding to the phylum level is 5. Combining these feature dimensions, the multi-dimensional semantic space framework is obtained.

[0093] Step S313: collecting standard feature descriptions of semantic concept nodes at each level, and performing vector conversion on the standard feature descriptions to obtain concept feature vectors.

[0094] The standard feature description of the semantic concept node at each level is a detailed definition and characteristic description of the semantic concept, which contains the essential characteristics of the semantic concept and the key information that distinguishes it from other concepts. The vector conversion of the standard feature description to obtain the concept feature vector is to convert the semantic information into a numerical vector that can be processed by a computer. In collecting the standard feature descriptions of the semantic concept nodes at each level, they can be obtained from biological literature, professional databases, etc. For example, the standard feature description of a species can include its morphological characteristics, physiological characteristics, ecological habits, etc. In vector conversion of the standard feature description, word embedding techniques such as Word2Vec, GloVe, etc. can be used.

[0095] Step S314: calculating the similarity between the initial specimen feature spectrum and the concept feature vectors of each semantic level, and normalizing the similarity value as the coordinate value of the corresponding dimension of the semantic space.

[0096] The similarity between the initial specimen feature spectrum and the concept feature vector of each semantic level is calculated to measure the matching degree of the specimen feature and the semantic concept in the knowledge system. The normalized similarity value is used as the coordinate value of the corresponding dimension in the semantic space, which can make the coordinate value comparable and accurately represent the association between the specimen feature and the semantic concept in the semantic space.

[0097] In the calculation of the similarity, various similarity measurement methods can be used, such as cosine similarity, Euclidean distance, etc. After calculating the similarity value, normalization processing is performed to convert it to the [0, 1] interval. Linear normalization method can be used for normalization, such as subtracting the minimum value and then dividing by the difference between the maximum value and the minimum value. The normalized similarity value is used as the coordinate value of the corresponding dimension in the semantic space.

[0098] Step S315: Construct the distance measurement function of the semantic space, measure the similarity of different specimen feature vectors in the semantic space based on Mahalanobis distance, and integrate the semantic association weights between semantic levels in the distance measurement function.

[0099] The distance measurement function of the semantic space is used to measure the similarity of different specimen feature vectors in the semantic space. The distance measurement function integrates the semantic association weights between semantic levels to consider the importance and association relationship between different semantic levels when calculating the distance.

[0100] In constructing the distance measurement function of the semantic space, the covariance matrix of the data in the semantic space is first calculated. The covariance matrix reflects the correlation between the dimensions of the data. Then, according to the definition of Mahalanobis distance, Mahalanobis distance is equal to the transpose of the vector difference multiplied by the inverse matrix of the covariance matrix and then multiplied by the vector difference. In the calculation process, the semantic association weights between semantic levels are introduced. The semantic association weights can be determined according to expert knowledge or data statistical analysis, for example, if a certain semantic level is more important for the classification and understanding of the specimen, its weight can be set higher. These weights are integrated into the calculation of Mahalanobis distance to obtain the distance measurement function integrated with the semantic association weights.

[0101] Step S316: Perform principal component analysis dimension reduction optimization on the high-dimensional semantic space, retain the principal components with cumulative contribution rate exceeding the set proportion, and ensure the discrimination and calculation efficiency of the semantic space.

[0102] The principal component analysis dimension reduction optimization of the high-dimensional semantic space is to reduce the dimension of the semantic space while retaining the main information of the data. When performing the principal component analysis dimension reduction optimization, first, the covariance matrix of the high-dimensional semantic space data is calculated. Then, the covariance matrix is subjected to eigenvalue decomposition to obtain eigenvalues and eigenvectors. The eigenvalues represent the variance size of the principal components, and the eigenvectors represent the direction of the principal components. The eigenvalues are sorted from large to small, and the cumulative contribution rate is calculated. The cumulative contribution rate is the proportion of the sum of the first k eigenvalues to the sum of all eigenvalues. According to the set proportion, the first k eigenvectors with a cumulative contribution rate exceeding the proportion are selected as the principal components. The original data is projected onto these principal components to obtain the reduced semantic space.

[0103] Step S320: mapping the initial specimen feature spectrum to the specimen feature semantic space by deep metric learning space mapping to generate an initial semantic coordinate vector.

[0104] As an implementation, step S320 can include steps S321-S326 as follows: Step S321: constructing a deep metric learning network, the deep metric learning network including a feature extraction layer and a mapping layer, the feature extraction layer adopting a residual network structure, and the mapping layer adopting a fully connected network structure.

[0105] The deep metric learning network is a key model for realizing the mapping of the initial specimen feature spectrum to the specimen feature semantic space. The function of the feature extraction layer is to extract more representative and discriminative features from the initial specimen feature spectrum. The residual network structure has strong feature extraction capability and can effectively solve the problems of gradient disappearance and gradient explosion in deep neural networks, so that the network can be trained deeper and more complex features can be extracted. The function of the mapping layer is to map the features extracted by the feature extraction layer to the specimen feature semantic space. The fully connected network structure can perform nonlinear transformation on the features so that they can be appropriately represented in the semantic space.

[0106] Step S322: standardizing the initial specimen feature spectrum for pretreatment, so that the data of each dimension of the initial specimen feature spectrum conforms to the zero mean unit variance distribution, to serve as the input data of the deep metric learning network.

[0107] The standardization pretreatment of the initial specimen feature spectrum is to make the data comparable and stable, and to avoid the influence of the scale difference of different dimensional data on the training of the deep metric learning network. The zero mean unit variance distribution is a standardization distribution that adjusts the mean of the data to 0 and the standard deviation to 1. After converting the data of each dimension of the initial specimen feature spectrum into this distribution, the training efficiency and convergence speed of the network can be improved.

[0108] In the standardization preprocessing, first calculate the mean and standard deviation of each dimension of the initial sample feature spectrum. Then, for each dimension of data, subtract the mean of the dimension and divide by the standard deviation of the dimension to get the standardized data. After processing, the mean of each dimension of the initial sample feature spectrum is 0 and the standard deviation is 1.

[0109] Step S323: Obtain a preset triplet loss function, which includes distance constraints between anchor samples, positive samples and negative samples, wherein the positive samples are sample feature spectra of the same semantic category, and the negative samples are sample feature spectra of different semantic categories.

[0110] The triplet loss function is a common loss function in deep metric learning, which constrains the distance relationship between anchor samples, positive samples and negative samples, so that the network learns the semantic similarity between samples. The anchor sample is a reference sample for comparison, the positive sample is a sample feature spectrum belonging to the same semantic category as the anchor sample, and the negative sample is a sample feature spectrum belonging to a different semantic category from the anchor sample.

[0111] The form of the preset triplet loss function is: L=max(d(a,p)-d(a,n)+α,0), where a represents the anchor sample, p represents the positive sample, n represents the negative sample, d represents the distance metric function (such as Euclidean distance), and α is a positive boundary value. The purpose of this loss function is to make the distance d(a,p) between the anchor sample and the positive sample as small as possible, and the distance d(a,n) between the anchor sample and the negative sample as large as possible, and the difference between the two is greater than the boundary value α.

[0112] Step S324: Optimize the parameters of the deep metric learning network through the back propagation algorithm, minimize the spatial distance between the anchor sample and the positive sample, maximize the spatial distance between the anchor sample and the negative sample, and iterate until the triplet loss function converges.

[0113] The back propagation algorithm is an algorithm used to optimize network parameters in deep learning, which calculates the gradient of the loss function with respect to the network parameters, and then updates the network parameters according to the gradient. In deep metric learning, the back propagation algorithm is used to optimize the parameters of the deep metric learning network, the purpose is to minimize the spatial distance between the anchor sample and the positive sample, and maximize the spatial distance between the anchor sample and the negative sample, so that the network can learn the semantic similarity between samples. Iterative training until the triplet loss function converges, which means that the parameters of the network have been adjusted to a better state, and can accurately represent the semantic relationship between samples in the sample feature semantic space.

[0114] Step S325: Fix the trained mapping layer parameters, input the preprocessed initial sample feature spectrum for forward propagation, and get the semantic space coordinate vector output by the network.

[0115] The fixed mapping layer parameters are used to maintain the stability and consistency of the network when using the trained deep metric learning network for prediction. The pre-processed initial specimen feature spectrum is input for forward propagation, which converts the specimen features through the network to obtain the coordinate representation in the semantic space of the specimen features. The semantic space coordinate vector is the result of the network output, which contains the position information of the specimen features in the semantic space.

[0116] After fixing the trained mapping layer parameters, the standardized pre-processed initial specimen feature spectrum is input into the feature extraction layer of the deep metric learning network. The feature extraction layer extracts features from the input specimen features to obtain more representative features. Then, the extracted features are input into the mapping layer, which maps the features to the semantic space of the specimen features according to the fixed parameters. After forward propagation, the semantic space coordinate vector output by the network is obtained.

[0117] Step S326: L2 normalization processing is performed on the semantic space coordinate vector to make the coordinate values of each dimension within a set interval, generating an initial semantic coordinate vector.

[0118] L2 normalization processing is performed on the semantic space coordinate vector to make the coordinate values of each dimension within a set interval, generating an initial semantic coordinate vector.

[0119] Step S330: Calculate the similarity between the initial semantic coordinate vector and the standard semantic template in the knowledge system, adjust the space mapping parameters according to the similarity distribution, and optimize the spatial distribution of the semantic coordinate vector.

[0120] The similarity between the initial semantic coordinate vector and the standard semantic template in the knowledge system is calculated to evaluate the matching degree of the specimen feature representation in the semantic space with the standard semantics in the knowledge system. The standard semantic template is the standard representation of different semantic concepts in the knowledge system. The space mapping parameters are adjusted according to the similarity distribution to make the position of the specimen feature in the semantic space more reasonable, optimize the spatial distribution of the semantic coordinate vector, and improve the representation accuracy of the specimen feature in the semantic level.

[0121] In calculating the similarity, cosine similarity, Euclidean distance, etc. methods can be used. After calculating the similarity between the initial semantic coordinate vector and each standard semantic template, the distribution of the similarity is analyzed. If it is found that the similarity between the semantic coordinate vector of some specimen characteristics and the standard semantic template is generally low, it indicates that there may be a problem with the space mapping, and the space mapping parameters need to be adjusted. The space mapping parameters can be the mapping layer parameters of the deep metric learning network, and by fine-tuning these parameters, the position of the specimen characteristics in the semantic space is closer to the standard semantic template.

[0122] Step S340: Construct a multi-modal semantic verification rule library with morphological, physiological, and ecological multi-dimensional semantic verification rules, perform multi-modal consistency verification on the optimized semantic coordinate vector, correct the conflict components in the semantic coordinate vector based on the verification result, perform entity relationship reasoning on the implicit association between the specimen characteristics and the concept entities in the knowledge system, and calculate the association strength.

[0123] The multi-modal semantic verification rule library is a collection of semantic verification rules in multiple dimensions such as morphology, physiology, and ecology. The morphological rules are used to check whether the morphological characteristics of the specimen conform to the corresponding species characteristics; the physiological rules are used to check whether the physiological attributes of the organism are normal; and the ecological rules are used to check whether the relationship between the environmental parameters and the biological attributes conforms to the ecological laws. The multi-modal consistency verification on the optimized semantic coordinate vector is to ensure that the specimen characteristics have consistency in multiple dimensions of semantics. The correction of the conflict components in the semantic coordinate vector based on the verification result is to make the semantic coordinate vector more accurately represent the specimen characteristics. The entity relationship reasoning on the implicit association between the specimen characteristics and the concept entities in the knowledge system is to mine the potential relationship between the specimen characteristics and the entities in the knowledge system, and to calculate the association strength.

[0124] Step S350: Integrate the association strength and the evolutionary characteristic weight into the semantic coordinate vector to generate a context-enhanced annotation vector containing the semantic association strength and the evolutionary characteristic weight of the specimen sample in the multi-level knowledge system.

[0125] Integrating the association strength and the evolutionary characteristic weight into the semantic coordinate vector is to make the semantic coordinate vector contain more context information, so as to more comprehensively represent the characteristics of the specimen sample in the multi-level knowledge system. The association strength represents the degree of association between the specimen characteristics and the concept entities in the knowledge system, and the evolutionary characteristic weight reflects the importance of the specimen characteristics in the evolutionary process. The context-enhanced annotation vector is the result of integration, which can provide more abundant information for the annotation of the specimen.

[0126] Step S400: Perform history trajectory similarity comparison and multi-dimensional evolutionary trend prediction on the context-enhanced annotation vector to generate an intermediate evolution label set recording the semantic change process of the specimen characteristics at different time scales.

[0127] The history trajectory similarity comparison of the context-enhanced annotation vector is to find out the specimens with similar history evolution trajectory as the current specimen, so as to analyze the evolution trend of the current specimen by learning from the history experience. The multi-dimensional evolution trend prediction is to predict the future evolution of the specimen characteristics from multiple angles (such as morphology, physiology, ecology, etc.). The intermediate evolution label set is a label set recording the semantic change process of the specimen characteristics at different time scales, which contains the feature description and evolution direction indication of the specimen at different time points.

[0128] As an implementation, step S400 can include steps S410-S460 as follows: Step S410: retrieving the historical specimen annotation data of the same genus as the current specimen sample from the laboratory specimen database, extracting the context-enhanced annotation vector of the historical specimen as the reference vector set, and the reference vector set contains the annotation vector sequence at different collection time points.

[0129] Retrieving the historical specimen annotation data of the same genus as the current specimen sample from the laboratory specimen database is to obtain the historical specimen information with similar classification attributes, so as to perform the history trajectory comparison and analysis. The context-enhanced annotation vector of the historical specimen contains information such as semantic association strength and evolution feature weight in the multi-level knowledge system, and taking these vectors as the reference vector set can provide reference for the evolution trend analysis of the current specimen. The reference vector set contains the annotation vector sequence at different collection time points, and these sequences reflect the feature change of the historical specimen at different times.

[0130] As an implementation, step S410 can include steps S411-S416 as follows: Step S411: determining the biological classification unit to which the current specimen sample belongs based on the taxonomic information of the current specimen sample, constructing a database retrieval keyword combination with the scientific name and feature description of the biological classification unit, and the keyword combination contains the taxonomic name and collection environment features.

[0131] Determining the biological classification unit to which the current specimen sample belongs based on the taxonomic information of the current specimen sample is to accurately locate the historical specimen with similar classification attributes in the laboratory specimen database. The biological classification unit is the basic unit of biological classification in biology, such as phylum, class, order, family, genus, and species. Constructing a database retrieval keyword combination with the scientific name and feature description of the biological classification unit can improve the accuracy and pertinence of the retrieval. The keyword combination contains the taxonomic name and collection environment features, the taxonomic name is used to determine the classification range of the organism, and the collection environment features are used to further filter the historical specimens similar to the collection environment of the current specimen.

[0132] Step S412: Query the metadata index table of the laboratory specimen database by searching for the keyword combination, obtain the list of all associated historical specimen record IDs through index matching, and arrange the list of historical specimen record IDs in descending order of collection time.

[0133] The metadata index table of the laboratory specimen database is an index structure that records the basic information of all specimen records in the database, including taxonomic information, collection time, collection location, and other metadata. By searching for the keyword combination to query the metadata index table, the keyword is matched with the metadata to find the historical specimen records related to the keyword. Index matching can improve query efficiency and quickly locate records that meet the conditions. The list of historical specimen record IDs obtained is arranged in descending order of collection time, which can prioritize newer historical specimen records for subsequent analysis and comparison.

[0134] Step S413: According to the record ID list, retrieve the corresponding historical specimen annotation data file from the database storage module, parse the structured data fields of the file, and extract the annotated context-enhanced annotation vector.

[0135] According to the record ID list, retrieve the corresponding historical specimen annotation data file from the database storage module, in order to obtain detailed annotation information of historical specimens. The database storage module is where specimen annotation data files are stored, each file contains information related to a historical specimen. Parsing the structured data fields of the file is to parse the data in the file according to a certain structure and extract the required information. The annotated context-enhanced annotation vector is important information contained in the file, which records the semantic association strength and evolution feature weight of historical specimens in the multi-level knowledge system. When retrieving data files, according to each ID in the record ID list, find the corresponding file from the database storage module. These files may be stored in a specific format, such as XML, JSON, etc.

[0136] Step S414: Perform outlier detection on the extracted context-enhanced annotation vector, calculate the mean and standard deviation of each vector dimension, and mark and remove vectors that deviate from the mean by more than a certain multiple of the standard deviation as outliers.

[0137] Performing outlier detection on the extracted context-enhanced annotation vector is to remove outliers in the data and improve the quality and reliability of the data. Outliers may be caused by data collection errors, annotation errors, etc., and can interfere with subsequent analysis and prediction. Calculating the mean and standard deviation of each vector dimension is to determine the normal range of the data.

[0138] During outlier detection, for each dimension of the extracted context-enhanced labeled vectors, the mean and standard deviation of all vector values ​​in that dimension are calculated. A fold threshold is set, for example, three standard deviations. For each dimension's value of each vector, if the value deviates from the mean of that dimension by more than three standard deviations, the vector is marked as an outlier. Vectors marked as outliers are then removed from the vector set.

[0139] Step S415: Perform data integrity verification on the remaining vectors, check whether there are missing values ​​in each dimension of the vectors, and fill the missing dimensions with the median of the vectors in the same batch.

[0140] Performing data integrity checks on the remaining vectors ensures data completeness and prevents missing data from affecting subsequent analysis and predictions. Checking for missing values ​​in each dimension of the vector involves iterating through each dimension and checking for any null or undefined values. During data integrity checks, for each remaining context-enhanced annotation vector, the values ​​of each dimension are checked. If a missing value is found in a dimension, the vector and the missing dimension's information are recorded. For the missing dimension, the median value of that dimension among vectors in the same batch is calculated. This median is then used as padding to fill in the missing dimension's position.

[0141] Step S416: Arrange the verified vectors in ascending order of the collection timestamp to construct a reference vector set containing time series characteristics. The vector set retains the association information between the specimen collection environment parameters and biological attribute observation data corresponding to each vector.

[0142] Arranging the validated vectors in ascending order of collection timestamps is to give the reference vector set time-series characteristics, reflecting the changes in historical specimen characteristics over time. Constructing a reference vector set with time-series characteristics facilitates subsequent time-series analysis and evolutionary trend prediction. The vector set retains the correlation information between the specimen collection environment parameters and biological attribute observation data corresponding to each vector, so that changes in environmental factors and biological attributes can be comprehensively considered during the analysis process, leading to a more comprehensive understanding of the specimen's evolutionary process.

[0143] During the sorting process, the vectors are arranged in ascending order based on the specimen collection timestamps corresponding to their respective samples using a sorting algorithm. Algorithms such as quicksort and mergesort can be used. After the sorting is complete, a reference vector set is constructed. This vector set not only contains the context-enhanced annotation vectors but also retains the correlation information between the specimen collection environment parameters (such as water temperature and light intensity) and biological attribute observation data (such as growth rate and reproduction rate) corresponding to each vector.

[0144] Step S420: Calculate the similarity between the current context-enhanced annotation vector and each vector in the reference vector set, and filter historical reference vectors whose similarity meets the set conditions.

[0145] Calculating the similarity between the current context-enhanced annotation vector and each vector in the reference vector set aims to identify historical specimens with similar features to the current specimen, allowing us to leverage historical experience to analyze the evolutionary trend of the current specimen. Filtering historical reference vectors that meet set similarity criteria selects the most relevant historical specimen information from the reference vector set, improving the accuracy and effectiveness of the analysis.

[0146] In one implementation, step S420 may include the following steps S421-S426: Step S421: Convert the current context-enhanced annotation vector and each vector in the reference vector set into column vectors of the same dimension, and normalize each dimension of the vector to ensure that the vector magnitude is consistent.

[0147] To accurately calculate the similarity between the current context-enhanced annotation vector and each vector in the reference vector set, the vectors are first normalized. They are converted into column vectors of the same dimension. This is because a unified vector form ensures consistency and accuracy in mathematical operations and comparisons. Vectors of different dimensions are difficult to compare and calculate similarity directly; unifying them into column vectors makes the vectors structurally comparable.

[0148] Step S422: Perform a dot product operation on the current vector and each reference vector, divide by the vector magnitude product to obtain the similarity value, and generate a distribution sequence containing the similarity values ​​of all reference vectors.

[0149] After normalizing the vectors, we can begin calculating the similarity between the current vector and each reference vector. The result of the dot product operation is affected by the vector magnitudes. To eliminate this effect, the result of the dot product operation needs to be divided by the product of the vector magnitudes. After this processing, the obtained similarity value can more accurately reflect the true degree of similarity between the two vectors.

[0150] Step S423: Perform statistical characteristic analysis on the similarity distribution sequence, calculate the median and interquartile range of the sequence, and dynamically determine the similarity threshold based on the median and interquartile range.

[0151] After obtaining the similarity distribution sequence, a suitable similarity threshold needs to be determined in order to filter out reference vectors with high similarity to the current vector. Statistical characteristic analysis of the similarity distribution sequence is an effective method for determining the threshold.

[0152] The median is the value in the middle when a sequence is arranged in ascending order; it reflects the middle level of the sequence. The interquartile range (ICM) is the difference between the upper and lower quartiles, measuring the dispersion of the middle 50% of the data in the sequence. By calculating the median and ICM, we can understand the central tendency and dispersion of a similarity distribution sequence.

[0153] The reason for dynamically determining the similarity threshold based on the median and interquartile range is that this method can adaptively adjust according to the actual situation of the similarity distribution sequence. Different datasets may have different distribution characteristics, and if a fixed threshold is used, it may not be able to accurately select suitable reference vectors. However, the threshold determined based on the median and interquartile range can better adapt to the distribution of the data and improve the accuracy of the selection.

[0154] Step S424: Select reference vectors with similarity values ​​greater than the dynamic threshold as the initial candidate set. If the number of candidate vectors exceeds the preset upper limit, then truncate the first preset number of vectors in descending order of similarity value.

[0155] Based on a dynamically determined similarity threshold, reference vectors with similarity values ​​greater than the threshold are selected from the similarity distribution sequence. These reference vectors have a high similarity to the current vector and are selected as the preliminary candidate set. The preliminary candidate set contains historical specimen information that may be of significant reference value for analyzing the current specimen's evolutionary trend. However, in some cases, the number of preliminary candidate vectors may exceed a preset upper limit. The preset upper limit is a quantity restriction set according to research needs and actual circumstances. If the number of candidate vectors is too large, it may increase the complexity and computational load of subsequent analysis, and may also introduce some irrelevant or redundant information. Therefore, when the number of candidate vectors exceeds the preset upper limit, the candidate sets need to be sorted in descending order of similarity value, and the first preset number of vectors are truncated.

[0156] Step S425: Perform a time distribution uniformity check on the vectors in the preliminary candidate set, calculate the interval distribution of vector timestamps, and if time clustering exists, perform equal-interval resampling.

[0157] After selecting the initial candidate set, the temporal distribution of the reference vectors also needs to be considered. The uniformity of the temporal distribution is crucial for accurately analyzing the evolutionary trend of the specimen. If the reference vectors are too concentrated in time, it may lead to a one-sided understanding of the specimen's evolutionary process and fail to fully reflect the characteristic changes of the specimen in different time periods.

[0158] The initial candidate set of vectors undergoes a temporal distribution uniformity check, primarily by calculating the interval distribution of vector timestamps. Timestamps record the time information of sample collection; analyzing the interval distribution of timestamps reveals whether the reference vectors are evenly distributed over time. If temporal clustering is observed—that is, an excessive number of reference vectors in some time periods and an insufficient number in others—equal-interval resampling is necessary.

[0159] Step S426: Perform semantic diversity assessment on the resampled vectors, calculate the average pairwise similarity between vectors, and remove redundant vectors if the average similarity exceeds a set value, while retaining historical reference vectors with significant semantic differences.

[0160] After performing equally spaced resampling, semantic diversity evaluation is required for the resampled vectors. The purpose of semantic diversity evaluation is to ensure that the retained reference vectors have sufficient diversity to provide richer information and avoid excessive redundant information between reference vectors. By calculating the similarity between each pair of resampled vectors and averaging these similarities, the average pairwise similarity between vectors can be obtained. If the average pairwise similarity exceeds a set value, it indicates that the similarity between vectors is high and there is a lot of redundant information.

[0161] Step S430: Extract the specimen feature evolution trajectory data corresponding to the filtered historical reference vectors, and perform segmentation processing on the trajectory data based on the timestamp information to obtain a set of evolutionary segments containing different time spans.

[0162] After selecting suitable historical reference vectors, the next step is to extract the specimen feature evolution trajectory data corresponding to these reference vectors. This trajectory data records the feature changes of historical specimens at different time points and is an important basis for analyzing specimen evolution trends.

[0163] Different time spans may correspond to different evolutionary stages and influencing factors. For example, a shorter time span may reflect the specimen's adaptation to environmental changes in a short period, while a longer time span may reflect the specimen's characteristic changes in a long-term evolutionary process. By obtaining a set of evolutionary segments containing different time spans, we can study the evolutionary trends of the specimen in more detail and analyze the factors affecting the changes in specimen characteristics in different time periods.

[0164] Step S440: Construct a time series prediction model that integrates long short-term memory network and attention mechanism. Use the current context-enhanced annotation vector and the set of evolutionary segments as input to the time series prediction model, and output a multi-timescale feature evolution trend prediction sequence.

[0165] To accurately predict the evolutionary trends of specimen features, a suitable time series prediction model needs to be constructed. A model integrating Long Short-Term Memory (LSTM) networks and attention mechanisms is an effective choice. LSTM networks can handle long-term dependencies in time series data. In predicting the evolutionary trends of specimen features, changes in specimen features may be influenced by multiple past time points. LSTM networks, through their memory units and gating mechanisms, can effectively capture these long-term dependencies, thus better predicting future feature changes. Attention mechanisms, on the other hand, help the model focus more on important parts of the time series data. During the evolution of specimen features, the importance of feature information at different time points for future prediction may vary. Attention mechanisms can automatically allocate different attention weights based on the characteristics of the data, allowing the model to focus more on time points and feature information that have a greater impact on the prediction results, thereby improving prediction accuracy.

[0166] The current context-enhanced annotation vector and the evolutionary fragment set are used as inputs to the time series prediction model. The current context-enhanced annotation vector contains information such as the semantic association strength and evolutionary feature weights of the current specimen in the multi-level knowledge system, while the evolutionary fragment set records the feature evolution of historical specimens at different time spans. By learning from these input data, the model outputs a multi-time-scale feature evolution trend prediction sequence, which reflects the feature change trends of the specimen at different time scales.

[0167] Step S450: Perform weighted fusion of the predicted sequence and historical evolution trajectory data, wherein the weights are dynamically allocated based on the importance of the time scale to generate a complete feature transition sequence containing the predicted information.

[0168] After obtaining the predicted sequence of feature evolution trends at multiple time scales, it is necessary to fuse it with historical evolution trajectory data. Historical evolution trajectory data records the actual changes in the specimen's past features, while the predicted sequence is a prediction of future feature changes. Fusing the two can make comprehensive use of historical and predictive information to obtain a more comprehensive and accurate sequence of feature changes.

[0169] Dynamically assigning weights based on the importance of time scales is a crucial step in the fusion process. Different time scales may have different importance in analyzing specimen evolution trends. For example, recent time scales may be more important for predicting short-term feature changes, while long-term time scales may be more critical for grasping the overall evolutionary trend of the specimen. By dynamically assigning weights according to the importance of time scales, the fused feature change sequence can more reasonably integrate information from different time scales.

[0170] During weighted fusion, the predicted sequence and historical evolutionary trajectory data are summed according to dynamically assigned weights to obtain a complete feature change sequence containing predictive information. This sequence not only includes past feature changes of the specimen but also predictive information about future feature changes, providing a rich information foundation for the subsequent generation of intermediate evolutionary label sets.

[0171] Step S460: Based on semantic feature extraction and timestamp marking of feature transition sequences, generate an intermediate evolution tag set of the semantic transition process of the recorded specimen features at different time scales. The intermediate evolution tag set contains feature descriptions and evolution direction indicators for each time node.

[0172] After obtaining the complete feature transition sequence containing predictive information, an intermediate evolutionary label set can be generated based on this sequence. Semantic feature extraction is a crucial step in generating this intermediate evolutionary label set. By performing semantic analysis on the feature transition sequence, representative semantic features are extracted from the sequence. These semantic features reflect the characteristic features and changing trends of the specimen at different time points. Based on semantic feature extraction and timestamp marking, an intermediate evolutionary label set recording the semantic change process of the specimen's features at different time scales is generated. This label set contains a feature description and an evolutionary direction indicator for each time point. The feature description details the characteristic state of the specimen at that time point, while the evolutionary direction indicator points out the direction of change of the specimen's features from one time point to the next, such as growth, decrease, or stabilization.

[0173] Step S500: Integrate intermediate evolutionary tag sets through multi-dimensional semantic consistency verification and evolutionary path optimization to generate annotation tags with ecological evolution semantics.

[0174] While intermediate evolutionary label sets record the semantic changes of specimen features at different time scales, they may contain semantic inconsistencies or unreasonable evolutionary paths. To generate high-quality annotation labels, multi-dimensional semantic consistency verification and evolutionary path optimization are required for the intermediate evolutionary label sets.

[0175] In one implementation, step S500 may include the following steps S510-S560: Step S510: Construct a multi-dimensional semantic consistency verification matrix with the label items of the intermediate evolution label set as rows and the semantic verification dimension as columns. The semantic verification dimension includes morphological semantics, physiological semantics and ecological semantics, and the matrix elements are the semantic descriptions of the label items in the corresponding dimensions.

[0176] To perform multi-dimensional semantic consistency verification, a multi-dimensional semantic consistency verification matrix must first be constructed. This matrix is ​​structured with rows representing the labels in the intermediate evolutionary label set, where each label represents a feature description and evolutionary direction indication of the specimen at a certain time point. Semantic verification dimensions are represented as columns, encompassing morphological semantics, physiological semantics, and ecological semantics. These dimensions perform semantic checks on the labels from different perspectives.

[0177] The matrix elements are semantic descriptions of the tags in their corresponding dimensions. By organizing and classifying the semantic information of the tags according to different dimensions and storing it in the matrix, it is convenient to perform multi-dimensional semantic analysis and verification of the tags in the future.

[0178] Step S520: Perform rule matching for each tag item across all semantic dimensions, compare the tag description with the standard rules in the multimodal semantic verification rule base, and calculate the consistency score based on the number and weight of the matching rules.

[0179] The multimodal semantic validation rule base stores a series of standard rules. These rules are formulated based on knowledge and research findings in fields such as biology, physiology, and ecology, representing semantically correct standards. The semantic description of each label item in the matrix in its corresponding semantic dimension is compared with the standard rules in the multimodal semantic validation rule base to check whether the label description conforms to the standard rules.

[0180] A consistency score is calculated based on the number and weight of matching rules. Different rules may have different importance, so each rule is assigned a corresponding weight. If a tag matches a large number of rules on a certain semantic dimension, and these rules have high weights, then the tag will have a higher consistency score on that semantic dimension.

[0181] Step S530: Identify semantic conflict labels based on the score matrix. Conflict labels are labels that score below a set threshold in any semantic dimension. Perform semantic correction on the conflict labels, based on the semantic representation of the high-scoring labels in the same dimension.

[0182] After calculating the consistency score for each label across all semantic dimensions, semantically conflicting labels can be identified based on the score matrix. A threshold is set; if a label's score on any semantic dimension falls below this threshold, it is identified as a semantically conflicting label. Semantically conflicting labels may arise due to data collection errors, labeling mistakes, or inaccurate understanding of specimen characteristics, and require correction.

[0183] When semantically correcting conflicting tags, the semantic representations of high-scoring tags in the same dimension are referenced. High-scoring tags in the same dimension indicate a higher degree of conformity to standard rules in that semantic dimension, and their semantic representations have higher reliability and accuracy. By referencing the semantic representations of high-scoring tags in the same dimension to correct conflicting tags, the conflicting tags can be made more semantically consistent with standard rules, improving the overall semantic consistency of the intermediate evolved tag set.

[0184] Step S540: Construct an evolution path directed graph with the corrected label items as nodes and the temporal order between labels as directed edges. The edge weight is the semantic similarity between adjacent label items, and the node attributes include the timestamp and semantic feature vector of the label.

[0185] After correcting the semantically conflicting labels, a directed graph of the evolutionary path needs to be constructed to further optimize the evolutionary path. The corrected labels serve as nodes, each representing the characteristic state of the specimen at a certain time point. The temporal relationship between labels is represented by directed edges, with the direction of the edges pointing from earlier label nodes to later label nodes, thus forming a directed graph structure that reflects the specimen's evolutionary process.

[0186] In one implementation, step S540 may specifically include the following steps: Step S541: Treat each corrected label item as an independent node in the directed graph of the evolution path. The node attribute fields include the timestamp information, semantic feature vector, and feature description text corresponding to the label. The semantic feature vector is generated by word embedding transformation of the label text.

[0187] When constructing the directed graph of the evolutionary path, each corrected label item is first transformed into an independent node in the graph. Each node represents the characteristic state of the specimen at the corresponding time. To comprehensively describe the information of the nodes, node attribute fields are set.

[0188] The node attribute field contains the timestamp information corresponding to the label. The timestamp clearly identifies the time node to which the label item corresponds, enabling a clear temporal sequence of specimen feature changes in the graph. The semantic feature vector is generated by word embedding transformation of the label text.

[0189] Step S542: Determine the direction of the directed edges between nodes based on the timestamp order of the label items, pointing from the earlier label node to the later label node, forming a time series directed connection.

[0190] After converting the corrected labels into nodes, it is necessary to determine the connections between the nodes. It is reasonable to determine the direction of directed edges between nodes based on the timestamp order of the labels, because the evolution of the specimen proceeds in chronological order, gradually evolving from an earlier state to a later state.

[0191] By setting the direction of the directed edges from the earlier label node to the later label node, a time-series directed connection is formed. This directed connection can intuitively show the evolution of specimen features over time, making the feature change trajectory of the specimen from one time point to the next clearly visible in the directed graph of the evolution path.

[0192] Step S543: Perform cosine similarity calculation on the semantic description text of adjacent label items, normalize the similarity value and use it as the weight value of the corresponding directed edge, and map the weight value range to the set interval.

[0193] Step S544: Add self-loop edges and jump connection edges to the directed graph. The weight of the self-loop edge is set to the self-similarity of the node's semantic feature vector. The jump connection edge connects non-contiguous label nodes that have a semantic correlation exceeding a set value.

[0194] To more comprehensively represent the evolutionary process of the specimen, in addition to basic adjacent node connections, self-loop edges and jump connections need to be added to the directed graph of the evolutionary path. A self-loop edge is an edge that points from one node to itself, with its weight set to the self-similarity of the node's semantic feature vector. Self-similarity reflects the consistency and stability of a node's own semantics. During the specimen's evolution, some features may remain relatively stable for a period of time; self-loop edges can be used to represent the persistent state of such features, and the weight of the self-loop edge reflects the degree of this stability.

[0195] Skip edges connect discontinuous label nodes whose semantic relevance exceeds a set threshold. During specimen evolution, some features may not change continuously but rather abruptly at certain points in time, yet these changes still maintain semantic connections. By setting skip edges, these discontinuous but meaningful semantic connections can be captured, allowing the directed graph of the evolutionary path to more comprehensively reflect the specimen's evolutionary process. A threshold for semantic relevance is set; skip edges are only added when the semantic relevance between discontinuous nodes exceeds this threshold. This avoids adding too many unnecessary connections, ensuring a clear and effective graph structure.

[0196] Step S545: Construct a node attribute index table, which includes node ID, timestamp range, and dimension index of semantic feature vector.

[0197] A node ID is a unique identifier for each node, allowing for quick and accurate location of a specific node in the graph. The timestamp range records the time interval corresponding to each node, which is extremely useful for querying nodes based on time criteria.

[0198] Step S546: Store the topology of the directed graph through an adjacency matrix. The matrix rows and columns are node IDs, and the matrix elements are the weight values ​​of the corresponding edges. Zero values ​​indicate no direct connection, thus generating an evolutionary path directed graph.

[0199] In an adjacency matrix, rows and columns correspond to node IDs, and each element represents the weight of the edge between the corresponding nodes. If two nodes are directly connected, the value of the matrix element is the weight of the directed edge; if there is no direct connection between the two nodes, the value of the matrix element is zero. In this way, the adjacency matrix can clearly represent the connection relationships between nodes in a directed graph and the weight of the edges.

[0200] The topological structure of the directed graph of evolutionary paths is stored in an adjacency matrix, which facilitates subsequent path search and analysis. During path search, the connection information between nodes and the weights of edges can be directly obtained from the adjacency matrix, allowing for the calculation of the cumulative weights of different paths. By storing the topological structure of the directed graph in an adjacency matrix, a complete directed graph of evolutionary paths is ultimately generated.

[0201] Step S550: Perform path search on the directed graph of the evolution path, calculate the cumulative weight of all possible paths, and select the path with the largest cumulative weight as the optimal evolution path.

[0202] Specifically, it may include the following steps: Step S551: Sort the nodes of the directed graph of the evolution path in ascending order of timestamps, construct the node access sequence, and determine the key time nodes by timestamp density clustering. The cluster centers correspond to the feature mutation points in the time series.

[0203] Before performing a path search, the nodes in the directed evolutionary path graph need to be sorted to ensure a more orderly traversal. Sort the nodes in the directed evolutionary path graph in ascending order of timestamps. This ensures that the order in which nodes are visited corresponds to the chronological order of the specimen's evolution. A node visit sequence is constructed, recording the order in which nodes are visited. During the path search, visiting nodes sequentially according to this sequence improves the efficiency and accuracy of the search.

[0204] Step S552: Perform path search on the directed graph. Define the state transition equation as the cumulative weight of the current node equals the sum of the cumulative weight of the previous node and the edge weight. Set the initial state to the cumulative weight of the earliest node as the weight of its self-loop edge.

[0205] After constructing the node visit sequence and identifying key time nodes, path searching can begin on the directed graph. To calculate the cumulative weight of each path, a state transition equation needs to be defined. This equation specifies how the cumulative weight is calculated when moving from one node to another. In this case, the cumulative weight of the current node is equal to the sum of the cumulative weight of its predecessor node and the weight of the edge connecting them. This calculation method aligns with the logic of path cumulative weights; by progressively accumulating edge weights, the cumulative weight from the starting node to the current node can be obtained. The initial state is set to the cumulative weight of the earliest time node, which is its self-loop edge weight. The earliest time node is the starting point of the evolutionary path; its cumulative weight is set to a self-loop edge weight because initially, the path only contains this node itself, and the self-loop edge weight reflects the semantic stability of that node. Starting from this initial state, the cumulative weights of subsequent nodes are calculated progressively according to the state transition equation, thus completing the calculation of the cumulative weights for all paths in the directed graph.

[0206] Step S553: ​​Maintain multiple candidate paths for each node. The number of candidate paths is dynamically adjusted according to the node's out-degree. A set proportion of paths with the highest cumulative weight are retained to avoid combinatorial explosion. The path storage format includes the node sequence and the total cumulative weight.

[0207] In the path search process, to find the optimal evolutionary path, it is necessary to maintain multiple candidate paths for each node. Each node may have multiple outgoing edges, and there may be different path choices starting from that node. Maintaining multiple candidate paths can take into account all possible path situations.

[0208] The number of candidate paths is dynamically adjusted based on the out-degree of a node. The out-degree of a node represents the number of edges originating from that node; the larger the out-degree, the more possible path choices there are. Dynamically adjusting the number of candidate paths based on the out-degree ensures comprehensive search while avoiding maintaining too many unnecessary paths.

[0209] Step S554: Introduce key node constraints during the path search process, forcing all key time nodes to be included in the path. If a candidate path is missing a key node, it is filled by skipping connecting edges, and the weight of the filled edge is reduced according to the semantic relevance.

[0210] To ensure that the searched paths reflect the key stages of specimen evolution, key node constraints are introduced during the path search process. If a candidate path is found to be missing key nodes during the search, the path is completed using jump edges. Jump edges connect non-contiguous label nodes that have a semantic relevance exceeding a set value, ensuring that the path includes key nodes.

[0211] The weight of the completed edge is reduced based on semantic relevance. Since jump edges connect non-contiguous nodes, the semantic relevance between them may not be as strong as that between adjacent nodes. Therefore, when calculating the cumulative weight of the path, the weight of the completed edge needs to be reduced. The degree of reduction is determined by the semantic relevance; the lower the semantic relevance, the greater the weight reduction of the completed edge.

[0212] Step S555: Calculate the cumulative weight of all complete paths. The cumulative weight is a weighted combination of the sum of the weights of all edges in the path and the self-similarity of the semantic feature vectors of the nodes. Sort the paths in descending order of cumulative weight.

[0213] After completing the path search and filling in all key nodes, the cumulative weight of all complete paths is calculated. The cumulative weight is calculated as a weighted combination of the sum of the weights of all edges in the path and the self-similarity of the semantic feature vectors of the nodes. The edge weights reflect the semantic similarity between adjacent nodes, while the self-similarity of the node's semantic feature vectors reflects the semantic stability of the node itself. This calculation method comprehensively considers the connection strength of the edges and the stability of the nodes in the path, and more comprehensively evaluates the merits of each path.

[0214] After calculating the cumulative weights of all complete paths, the paths are sorted in descending order of cumulative weight. This descending order places the path with the highest cumulative weight first, facilitating the direct selection of the optimal path later. This sorting method allows for a clear comparison of the cumulative weights of different paths, providing an intuitive basis for selecting the optimal evolutionary path.

[0215] Step S556: Select the path with the largest cumulative weight as the optimal evolution path, and verify whether the path covers all key time nodes and semantic coherence. If there are missing nodes, backtrack and adjust the path search parameters and recalculate.

[0216] After sorting the paths in descending order of cumulative weight, the path with the highest cumulative weight is selected as the optimal evolutionary path. The optimal evolutionary path represents the most reasonable and coherent path in the evolutionary process of the specimen, and can most accurately reflect the evolutionary trend of the specimen's characteristics and the semantics of ecological evolution.

[0217] Step S560: Integrate the tags on the optimal evolutionary path, arrange the semantic descriptions and evolutionary direction indicators of the tags in chronological order, and generate two-dimensional structured annotation tags containing time axis and semantic axis. The two-dimensional structured annotation tags have a complete expression of ecological evolution semantics.

[0218] Once the optimal evolutionary path is determined, labels with ecological evolutionary semantics can be generated based on this path. Specifically, this involves integrating the labels from the optimal evolutionary path.

[0219] Arranging the labels along the optimal evolution path in chronological order clearly demonstrates the changes in specimen features at different time points. Each label includes a semantic description and an evolutionary direction indicator. The semantic description details the specimen's feature state at that time point, while the evolutionary direction indicator points to the direction of change of the specimen's features from one time point to the next. Arranging the semantic descriptions and evolutionary direction indicators in chronological order generates a two-dimensional structured label system containing a time axis and a semantic axis. The time axis represents the temporal sequence of specimen evolution, and the semantic axis represents the semantic state and direction of change of the specimen's features.

[0220] The various algorithms involved in the above descriptions of the embodiments of the present invention can all be obtained from relevant content in the prior art. To save space, they will not be elaborated on in the embodiments of the present invention. In addition, those skilled in the art can supplement the details based on common knowledge in the art when implementing the solutions of the present invention. For example, they can use normalization to eliminate dimensional conflicts before feature fusion, use interpolation to eliminate dimensional differences, reasonably set thresholds based on historical data, experience or business scenario requirements, train the model based on a general model training method, set the number of layers in the model structure based on actual needs, select activation functions, etc. The present invention will not provide redundant descriptions of overly detailed implementation processes.

[0221] Please see Figure 2 , Figure 2 This is a schematic diagram of a computer system provided in an embodiment of the present invention. The computer system includes at least a processor 101, a communication interface 102, and a memory 103. The processor 101, communication interface 102, and memory 103 can be connected via a bus or other means. The processor 101 (or Central Processing Unit, CPU) is the computing and control core of the computer system, capable of parsing various instructions and processing various data within the computer system. The communication interface 102 may optionally include a standard wired interface or a wireless interface (such as Wi-Fi, mobile communication interface, etc.), and can be used to send and receive data under the control of the processor 101; the communication interface 102 can also be used for data transmission and interaction within the computer system. The memory 103 is a storage device in the computer system used to store programs and data. It is understood that the memory 103 here can include the computer system's built-in memory, or it can include extended memory supported by the computer system. The memory 103 provides storage space, which stores the computer system's operating system; this invention does not limit this storage space.

[0222] In one embodiment, the processor 101 executes the specimen data annotation method for aquatic ecological laboratories provided above in the embodiments of the present invention by running a computer program in the memory 103.

Claims

1. A specimen data labeling method applied to a water ecological laboratory, characterized in that, The method comprises: acquiring specimen samples collected by a water ecological laboratory and corresponding multi-source environmental metadata, the multi-source environmental metadata including environmental parameter sequences and biological attribute observation data in the specimen collection process; constructing an environmental factor response curve and extracting a biological attribute time-varying feature from the environmental parameter sequences and the biological attribute observation data, coupling the multi-source environmental metadata and the specimen sample attributes based on the environmental factor response curve and the biological attribute time-varying feature, and generating an initial specimen feature spectrum including a nonlinear coupling relationship between an environmental parameter trajectory and a biological attribute response feature; mapping the initial specimen feature spectrum to a feature space, verifying a multi-modal semantic, and reasoning an entity relationship to generate a context-enhanced labeling vector including a semantic correlation strength and an evolution feature weight of the specimen sample in a multi-level knowledge system; comparing the context-enhanced labeling vector with a historical trajectory similarity and predicting a multi-dimensional evolution trend to generate an intermediate evolution label set recording a semantic change process of a specimen feature at different time scales; integrating the intermediate evolution label set through multi-dimensional semantic consistency verification and evolution path optimization to generate a labeled label with ecological evolution semantics.

2. The method of claim 1, wherein, The method comprises: performing multi-scale periodic decomposition on the environmental parameter sequences to separate environmental factor components of different oscillation periods, calculating the energy proportion of each environmental factor component, and constructing a periodic contribution degree matrix with the environmental factor type and oscillation period as dimensions, and selecting dominant environmental factor components based on the periodic contribution degree matrix; aligning the biological attribute observation data in a time sequence, extracting biological attribute response sequences matching the timestamps of the dominant environmental factor components, and performing cross-correlation analysis on different environmental factor components and biological attribute response sequences to calculate lag response coefficients, and constructing a lag response coefficient matrix with the dominant environmental factor components and biological attribute response sequences as dimensions; constructing an environmental factor-biological attribute cross-influence network with the environmental factor components and biological attribute indicators corresponding to the non-zero elements in the lag response coefficient matrix as network nodes and the absolute values of the lag response coefficients as edge weights, and performing topological analysis on the environmental factor-biological attribute cross-influence network to identify key influence paths; optimizing a nonlinear fitting model of the environmental factor response curve with the topological parameters of the key influence paths, jointly modeling the environmental parameter sequences and the biological attribute response sequences to generate an environmental factor dynamic response surface, and simultaneously performing time-frequency transformation on the biological attribute response sequences to extract instantaneous frequency and instantaneous amplitude features and construct dynamic descriptors of the biological attribute time-varying features; coupling the dynamic descriptors with the environmental factor dynamic response surface, performing nonlinear dimensionality reduction on the coupled features, and retaining cross-correlation information between the environmental parameters and the biological attributes; An initial sample feature spectrum containing a nonlinear coupling relationship between an environmental parameter trajectory and a biological attribute response feature is generated based on the reduced features.

3. The method of claim 2, wherein, The environmental parameter sequence is subjected to multi-scale periodic decomposition, and environmental factor components of different oscillation periods are separated. The energy proportion of each environmental factor component is calculated, and a periodic contribution degree matrix is constructed with the environmental factor type and oscillation period as dimensions. Dominant environmental factor components are screened based on the periodic contribution degree matrix, including: An initial modal number range for multi-scale periodic decomposition is obtained. The modal number is optimized to maximize the sum of kurtosis values of each component after decomposition, and the optimal modal decomposition parameter is determined. The environmental parameter sequence is subjected to multi-scale periodic decomposition based on the optimal modal decomposition parameter, and a plurality of independent environmental factor components are obtained. Each environmental factor component contains time series data of a corresponding oscillation period. The energy value of each environmental factor component is calculated. The energy value is obtained by integrating the square sum of the component time series. The energy proportion distribution is obtained by normalizing the energy value of each component. A periodic contribution degree matrix is constructed with the environmental factor type as the row and the oscillation period as the column. The elements of the periodic contribution degree matrix are the energy proportions of the corresponding environmental factors in the corresponding periodic components. A preset contribution degree threshold is obtained, and environmental factor components with element values exceeding the contribution degree threshold in the periodic contribution degree matrix are selected as candidate dominant components. Redundancy analysis is performed on the candidate dominant components, the mutual information value between the candidate dominant components is calculated, and the candidate dominant components with mutual information values exceeding a set value are removed.

4. The method of claim 3, wherein, The biological attribute observation data is subjected to time series alignment, and the biological attribute response sequence matching the timestamp of the dominant environmental factor component is extracted. Cross-correlation analysis is performed on different environmental factor components and biological attribute response sequences to calculate the lag response coefficient. A lag response coefficient matrix is constructed with the dominant environmental factor component and the biological attribute response sequence as dimensions, including: The timestamp set of the environmental parameter sequence is extracted, and the timestamp of the biological attribute observation data is subjected to linear interpolation processing to synchronize the biological attribute observation data and the environmental parameter sequence in the time dimension. The biological attribute index sequence is extracted from the aligned biological attribute observation data. Based on the time interval of the dominant environmental factor component, the biological attribute index sequence in the corresponding time period is intercepted as the biological attribute response sequence. Sliding window cross-correlation analysis is performed on each dominant environmental factor component and the biological attribute response sequence. The window length is dynamically adjusted according to the oscillation period of the environmental factor component, and the correlation coefficient at different lag times is calculated. The lag time and correlation coefficient value corresponding to the maximum correlation coefficient are recorded, and the correlation coefficient value is determined as the lag response coefficient. A lag response coefficient matrix is constructed with the environmental factor component as the row and the biological attribute response sequence as the column. The matrix elements are the lag response coefficients of the corresponding combinations.

5. The method of claim 4, wherein, The non-zero element corresponding to the environmental factor component and the biological attribute index in the lag response coefficient matrix is determined as a network node, and the node attribute includes factor type and time scale characteristics. The absolute value of the lag response coefficient is used as the edge weight to construct a directed weighted network, and the direction of the directed edge is from the environmental factor component to the biological attribute index. Degree centrality, betweenness centrality and closeness centrality of each node in the network are calculated, wherein the degree centrality represents the number of direct connections of the node, the betweenness centrality represents the intermediary role of the node in the path, and the closeness centrality represents the average distance from the node to other nodes. A pre-set centrality threshold is obtained, and nodes with a centrality index exceeding the centrality threshold are selected as key nodes, and the directed edges between the key nodes are extracted to form an initial influence path set. The initial influence path set is analyzed for path length, and short paths that meet the set conditions are retained, and the total weight of each path is calculated, wherein the total weight is the product of the edge weights on the path. The paths are sorted in descending order of total weight, and the top K paths are selected as key influence paths, wherein K is greater than 0. The topological parameters of the key influence paths are extracted, including path length, node degree and edge weight distribution, which are used as prior knowledge of the nonlinear fitting model of the environmental factor response curve.

6. The method of claim 2, wherein, A nonlinear fitting model including a kernel function combination is constructed, wherein the kernel function combination is a weighted combination of a radial basis function and a periodic kernel function, and the weight is dynamically adjusted based on the path topological parameters. The environmental parameter sequence is used as the model input, and the biological attribute response sequence is used as the model output, and the nonlinear fitting model is trained, and the kernel function parameters are optimized by maximum likelihood estimation. Based on the trained model, the response of the environmental parameter sequence is predicted to generate a predicted sequence of biological attribute responses, and the root mean square error of the predicted sequence and the actual observed sequence is calculated. Based on the error feedback, the kernel function weight is adjusted, and the model is iteratively optimized until the error is below the set threshold to obtain the optimized environmental factor response curve model. Different environmental parameter combinations are simulated by the optimized environmental factor response curve model to generate an environmental factor dynamic response surface including the relationship between multi-dimensional environmental parameters and biological attributes. The initial specimen feature spectrum is mapped to a feature space, multi-modal semantic verification and entity relationship reasoning are performed, and a context-enhanced annotation vector including the semantic association strength and evolution feature weight of the specimen sample in a multi-level knowledge system is generated, including: ​ 7. The method of claim 1, wherein, ​ A specimen characteristic semantic space is constructed based on a hierarchical structure of an aquatic ecological knowledge system, dimensions of the space correspond to the hierarchical structure of the knowledge system, and coordinate values of each dimension represent an association degree of the specimen characteristic and the hierarchical semantic concept; An initial specimen characteristic spectrum is subjected to deep metric learning space mapping to project the initial specimen characteristic spectrum to the specimen characteristic semantic space to generate an initial semantic coordinate vector; Similarity of the initial semantic coordinate vector and a standard semantic template in the knowledge system is calculated, space mapping parameters are adjusted according to the similarity distribution, and space distribution of the semantic coordinate vector is optimized; A multi-modal semantic verification rule library is constructed based on morphological, physiological and ecological multi-dimensional semantic verification rules, multi-modal consistency verification is performed on the optimized semantic coordinate vector, conflicting components in the semantic coordinate vector are corrected based on a verification result, an entity relationship is reasoned for an implicit association of the specimen characteristic and a concept entity in the knowledge system, and an association strength is calculated; The association strength and an evolutionary characteristic weight are integrated into the semantic coordinate vector to generate a context-enhanced annotation vector containing a semantic association strength and an evolutionary characteristic weight of the specimen sample in the multi-level knowledge system.

8. The method of claim 7, wherein, The specimen characteristic semantic space is constructed based on a hierarchical structure of an aquatic ecological knowledge system, dimensions of the space correspond to the hierarchical structure of the knowledge system, and coordinate values of each dimension represent an association degree of the specimen characteristic and the hierarchical semantic concept, and the method comprises the following steps: A hierarchical structure of an aquatic ecological knowledge system is obtained, six semantic levels are constructed from door, class, order, family, genus to species, each semantic level contains a plurality of semantic concept nodes; A corresponding characteristic dimension is constructed for each semantic level to obtain a multi-dimensional semantic space framework, wherein the characteristic dimension size is equal to the number of semantic concept nodes of the semantic level; Standard characteristic descriptions of the semantic concept nodes of each level are collected, and the standard characteristic descriptions are subjected to vector conversion to obtain concept characteristic vectors; Similarity of the initial specimen characteristic spectrum and the concept characteristic vectors of each semantic level is calculated, and the similarity values are normalized to serve as coordinate values of corresponding dimensions of the semantic space; A distance measurement function of the semantic space is constructed, the similarity of different specimen characteristic vectors in the semantic space is measured based on Mahalanobis distance, and the distance measurement function integrates the semantic association weights between the semantic levels.

9. The method of claim 8, wherein, The initial specimen characteristic spectrum is subjected to deep metric learning space mapping to project the initial specimen characteristic spectrum to the specimen characteristic semantic space to generate an initial semantic coordinate vector, and the method comprises the following steps: A deep metric learning network is constructed, the deep metric learning network comprises a feature extraction layer and a mapping layer, the feature extraction layer adopts a residual network structure, and the mapping layer adopts a fully connected network structure; The initial specimen characteristic spectrum is subjected to standardization preprocessing, so that the dimension data of the initial specimen characteristic spectrum conforms to a zero mean unit variance distribution to serve as input data of the deep metric learning network; A preset triplet loss function is obtained, the triplet loss function comprises distance constraints between an anchor sample, a positive sample and a negative sample, wherein the positive sample is a specimen characteristic spectrum of the same semantic category, and the negative sample is a specimen characteristic spectrum of different semantic categories; parameters of the deep metric learning network are optimized by a back propagation algorithm to minimize spatial distance between the anchor sample and the positive sample, maximize spatial distance between the anchor sample and the negative sample, and iteratively train until the triplet loss function converges; the trained mapping layer parameters are fixed, pre-processed initial sample feature spectrum is input for forward propagation, and a semantic space coordinate vector output by the network is obtained; L2 normalization processing is performed on the semantic space coordinate vector to make each dimension coordinate value in a set interval, and an initial semantic coordinate vector is generated.

10. A computer system, characterized by comprising: a memory having a computer program stored therein; a processor configured to load the computer program to implement the specimen data labeling method for aquatic ecological laboratory according to any one of claims 1-9.

Citation Information

Patent Citations

  • Multi-mode-based training method and system for cervical pathology image classification model

    CN120808067A

  • Full-dimensional monitoring system of intelligent specimen cabinet

    CN120848338A

  • System and method of screening biological or biomedical specimens

    US20220351530A1