Big data analysis-based alien invasive species early warning method
By collecting and integrating multi-dimensional data, combining image morphology comparison and gene sequence analysis, and combining multi-feature decision trees and fluid dynamics models, the problem of data integration and assessment in the early warning of invasive alien species has been solved, and the prevention and control effect of early detection, early assessment, early warning and early treatment has been achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-05
- Publication Date
- 2026-03-27
AI Technical Summary
Existing technologies for early warning of invasive alien species suffer from deficiencies in data integration effectiveness, spread prediction accuracy, and risk assessment system, leading to delayed or misjudgments and failing to meet the needs for precise prevention and control in complex ecological scenarios.
The system employs multi-dimensional data acquisition and integrated processing, combining image morphology comparison and gene sequence homology analysis. It conducts risk assessment through a multi-feature decision tree integration mechanism, simulates diffusion paths using a fluid dynamics model, generates early warning information, and pushes differentiated disposal suggestions and records response status in real time.
It has enabled early detection, early assessment, early warning, and early response to invasive alien species, improving the accuracy of early warning and the effectiveness of prevention and control, and ensuring accurate identification and timely response in sensitive areas.
Smart Images

Figure CN121743968A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of biological invasion prevention and control, in particular to a foreign invasive species early warning method based on big data analysis. BACKGROUND
[0002] Foreign invasive species have become a major threat to global ecological security and economic development. After being introduced into a new ecosystem through ship transportation, trade logistics, and natural dispersion, they can reproduce rapidly due to the lack of natural enemies, destroy local species diversity, and erode ecological niches, posing a threat to agricultural planting, aquaculture, water conservancy facilities, and causing corresponding economic losses. Among them, waterborne invasive species are particularly difficult to prevent and control due to their rapid spread through water bodies.
[0003] To address these issues, existing technologies have gradually developed from traditional manual patrols to early warning combined with data technology. For example, some solutions collect environmental parameters (such as water temperature and salinity) to determine the survival adaptability of species, or analyze the transmission probability of invasive species based on ship transportation trajectories. Another technology attempts to establish a database of foreign species to identify species through morphological comparison or genetic sequencing. However, current early warning methods still have technical bottlenecks, especially in terms of data integration effectiveness, spread prediction accuracy, and risk assessment systematization, making it difficult to form a complete closed loop of identification, assessment, prediction, and response, resulting in delayed or incorrect early warnings, which cannot meet the precise prevention and control needs in complex ecological scenarios. SUMMARY
[0004] The technical problem to be solved by the present application is to provide a foreign invasive species early warning method based on big data analysis, which can achieve early detection, early assessment, early warning, and early disposal of foreign invasive species.
[0005] To solve the above technical problems, the technical solution of the present application is as follows: In a first aspect, a foreign invasive species early warning method based on big data analysis is provided, which comprises: Performing multi-dimensional data collection to obtain environmental parameter data, ship transportation data, and biological sample data; Based on multi-dimensional data, performing data integration processing by time aligning and formatting the environmental parameter data, ship transportation data, and biological sample data to form a standardized data set; Based on the standardized data set, performing foreign invasive species identification by comparing a pre-set global foreign invasive species database to obtain a matching species identification result; According to the matching species identification result, combining the environmental parameter data and ship transportation data, performing risk assessment by analyzing the species survival condition adaptability and transmission probability characteristics, and calculating the colonization risk level based on a multi-feature decision tree integration mechanism to obtain a risk level assessment result; The risk level assessment result is subjected to diffusion prediction analysis, and a species diffusion path is simulated according to meteorological data and through a fluid dynamics model; in the simulation process, a vertical diffusion component is determined by calculating a cross product of a water flow direction vector and a terrain gradient vector to correct a diffusion trajectory of the species, and finally, early warning information is formed; Based on the early warning information, a disposal suggestion and early warning content are formed by performing early warning response and information publishing; the disposal suggestion is delivered to a terminal of a relevant department, and the early warning content is published to a public platform; a delivery state and a publishing state of the early warning information are recorded in real time to update an early warning response log.
[0006] Further, multi-dimensional data collection is performed to obtain environmental parameter data, ship navigation data and biological sample data, including: The environmental parameter data of the port water area is collected, and the environmental parameter data includes water temperature, salinity, pH value and water flow speed; The navigation data related to the ship is obtained, and the navigation data includes ballast water discharge records, route trajectories and cargo types; The biological sample data is collected, and the biological sample data is converted into standardized image data and gene sequence data to obtain converted biological sample data; The environmental parameter data, the navigation data and the converted biological sample data are integrated to form a set of original data to be processed.
[0007] Further, based on the multi-dimensional data, data integration processing is performed, and the environmental parameter data, the ship navigation data and the biological sample data are time-aligned and unified in format to form a standardized data set, including: The original data set is subjected to timestamp alignment processing, and the environmental parameter data, the navigation data and the biological sample data are unified to the same time reference to obtain time-aligned data; The time-aligned data is subjected to format standardization processing, and the data from different sources is converted into a unified numerical format and data structure to obtain format-standardized environmental parameter data, navigation data and biological sample data; The format-standardized environmental parameter data, navigation data and biological sample data are subjected to data fusion processing to form a standardized data set with consistent space-time reference.
[0008] Further, based on the standardized data set, alien invasive species identification is performed, that is, a matching species identification result is obtained by comparing a preset global alien invasive species database, including: The biological sample data is extracted from the standardized data set, and the biological sample data includes image data and gene sequence data; The image data is compared with the morphological feature data in the preset global invasive alien species database for similarity comparison, to obtain a morphological feature comparison result; meanwhile, the genetic sequence data is compared with the genetic sequence data in the preset global invasive alien species database for homology analysis, to obtain a genetic sequence analysis result; Based on the morphological feature comparison result and the genetic sequence analysis result, a comprehensive matching degree is calculated by weighted fusion to determine the matching species and the confidence thereof; The matching species are filtered according to a preset confidence threshold, to obtain a matching species identification result, which contains the species name and matching confidence information.
[0009] Further, according to the matching species identification result, environmental parameter data and ship navigation data are combined for risk assessment, by analyzing the species survival condition adaptability and the import probability characteristics, and calculating the colonization risk level based on a multi-feature decision tree integration mechanism, a risk level evaluation result is obtained, including: The species name is extracted from the matching species identification result, and the corresponding species survival condition parameters are obtained from the preset global invasive alien species database based on the species name; The environmental parameter data in the standardized data set are analyzed for adaptability with the species survival condition parameters, to obtain an adaptability analysis result; Based on the shipping route trajectory in the navigation data, the minimum distance between the ship position point and the boundary of the high-risk area of the invasive alien species is calculated to quantify the spatial proximity; based on the spatial proximity, the import probability analysis is performed by combining the ballast water discharge record, to obtain an import probability analysis result; Based on the adaptability analysis result and the import probability analysis result, combined with the historical invasion case data sorted in advance, a plurality of basic risk assessment results are obtained by using a plurality of decision trees to process multiple feature dimensions in parallel; The plurality of basic risk assessment results are integrated and processed, and the final colonization risk level is determined by using a voting mechanism, and the risk level evaluation result containing low, medium and high levels is obtained based on the final colonization risk level.
[0010] Further, the risk level evaluation result is analyzed for diffusion prediction, and the species diffusion path is simulated according to the meteorological data and by using a fluid dynamics model; during the simulation process, the water flow direction vector is calculated, which is defined as a three-dimensional vector; the east direction is the x-axis, the north direction is the y-axis, and the vertical direction is the z-axis, the initial value of the vertical direction is set to 0 and is adjusted only with the terrain gradient; the cross product of the terrain gradient vector and the three-dimensional vector constructed based on the water area elevation data, containing x, y and z axis components, is used to determine the vertical diffusion component to correct the diffusion trajectory of the species, and finally the warning information is formed, including: extracting species information marked as a high risk level and location data thereof in the risk level assessment result; based on the species information marked as a high risk level and the location data thereof, combining real-time acquired meteorological data and hydrological data, simulating a diffusion process of the species in the water body through a fluid dynamics model; in the simulation process of the fluid dynamics model, determining a vertical diffusion component by calculating a cross product of a water flow direction vector and a terrain gradient vector, correcting a horizontal diffusion trajectory based on the diffusion component, and obtaining a corrected diffusion trajectory; spatially superimposing and analyzing the corrected diffusion trajectory and a preset sensitive area boundary to identify an agricultural breeding area and an ecological protection area that may be affected, and obtaining a spatial superimposition analysis result; based on the spatial superimposition analysis result, forming early warning information including a diffusion path diagram, an affected area, and a risk level.
[0011] Further, based on the early warning information, forming a disposal suggestion and early warning content by executing early warning response and information publishing; delivering the disposal suggestion to a terminal of a relevant department, and publishing the early warning content to a public platform; and recording a delivery state and a publishing state of the early warning information in real time to update an early warning response log, including: analyzing and processing the early warning information to obtain species characteristics, a risk level, affected area data, and diffusion path data; based on the risk level, automatically matching a corresponding disposal measure from a preset disposal scheme library to form a disposal suggestion including specific operation instructions; based on the affected area data, determining a terminal list of the relevant department that needs to be pushed the early warning, and automatically transmitting the disposal suggestion to the corresponding terminal through a data interface based on the terminal list; based on the species characteristics, the risk level, and the affected area data, generating early warning content for different user groups, and publishing and pushing the early warning content to the public platform through a preset publishing channel; real-time monitoring the delivery state of the disposal suggestion and the publishing state of the early warning content to update the early warning response log and complete early warning closed-loop management.
[0012] In a second aspect, an early warning system for alien invasive species based on big data analysis includes: an acquisition module for multi-dimensional data acquisition to obtain environmental parameter data, ship navigation data, and biological sample data; a processing module for data integration processing based on the multi-dimensional data, forming a standardized data set by time alignment and format unification of the environmental parameter data, the ship navigation data, and the biological sample data; The identification module is configured to identify the alien invasive species based on the standardized data set, i.e., by comparing a preset global alien invasive species database to obtain a matching species identification result. The evaluation module is configured to perform risk evaluation according to the matching species identification result, in combination with environmental parameter data and ship navigation data, by analyzing the adaptability of the survival conditions of the species and the transmission probability characteristics, and calculating the colonization risk level based on a multi-feature decision tree integration mechanism to obtain a risk level evaluation result. The simulation module is configured to perform diffusion prediction analysis on the risk level evaluation result, simulate the diffusion path of the species according to meteorological data and by a fluid dynamics model; in the simulation process, the diffusion component in the vertical direction is determined by calculating the cross product of the water flow direction vector and the terrain gradient vector to correct the diffusion trajectory of the species, and finally form the early warning information. The execution module is configured to form the disposal suggestion and the early warning content by performing the early warning response and the information release based on the early warning information; the disposal suggestion is delivered to the terminal of the relevant department, and the early warning content is released to the public platform; the delivery state and the release state of the early warning information are recorded in real time to update the early warning response log.
[0013] In a third aspect, a computing device includes: one or more processors; a memory device storing one or more programs, when the one or more programs are executed by the one or more processors, the one or more processors implement the method.
[0014] In a fourth aspect, a computer readable storage medium stores a program, when the program is executed by a processor, the method is implemented.
[0015] The above-mentioned scheme of the present application at least has the following beneficial effects: Because the multi-dimensional data collection of covering environmental parameters, ship navigation and biological sample data, time alignment and format unified data integration processing means are adopted, the technical problems of scattered data sources, incompatible formats and inability to collaborative analysis in the prior art are overcome, and the technical effect of providing consistent space-time benchmark data for species identification and risk assessment is achieved; because the dual identification and weighted fusion calculation means of combining image morphology comparison and gene sequence homology analysis are adopted, the technical problems of low precision and insufficient reliability of single identification method are overcome, and the technical effect of accurately determining the matched species and its real-time reliability is achieved; because the multi-feature decision tree integration mechanism combining species survival adaptability, ship transmission probability and historical cases and the voting evaluation means are adopted, the technical problems of single risk assessment dimension and strong subjectivity of results in the prior art are overcome, and the technical effect of objectively outputting low, medium and high three-level colonization risk levels is achieved; because the diffusion prediction means of calculating the vertical diffusion component by the cross product of the flow direction vector and the terrain gradient vector to correct the trajectory is adopted, the technical problems of ignoring the terrain influence and large diffusion path prediction deviation in the existing fluid dynamics simulation are overcome, and the technical effect of accurately identifying the affected range of sensitive areas such as agricultural breeding areas and ecological protection areas is achieved; because the closed-loop management means of automatic matching of a preset disposal scheme library, differentiated pushing of departments and the public and real-time updating of early warning response logs is adopted, the technical problems of lack of follow-up tracking, delayed response and no traceability in the existing early warning are overcome, and the technical effect of realizing the whole process of early detection, early assessment, early warning and early disposal of alien invasive species is achieved. BRIEF DESCRIPTION OF DRAWINGS
[0016] Figure 1 is a flowchart of a method for early warning of alien invasive species based on big data analysis provided by an embodiment of the present application.
[0017] Figure 2 is a schematic diagram of a system for early warning of alien invasive species based on big data analysis provided by an embodiment of the present application. DETAILED DESCRIPTION
[0018] Exemplary embodiments of the present disclosure will be described in greater detail below with reference to the accompanying drawings. Although exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. Rather, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be accurately conveyed to those skilled in the art.
[0019] As shown in Figure 1 An embodiment of the present application proposes a method for early warning of alien invasive species based on big data analysis, which comprises the following steps: Step 1, multi-dimensional data acquisition is performed to obtain environmental parameter data, ship navigation data and biological sample data; Step 2, based on the multi-dimensional data, data integration processing is performed, and by time alignment and format unification of the environmental parameter data, the ship navigation data and the biological sample data, a standardized data set is formed; Step 3, based on the standardized data set, alien invasive species identification is performed, that is, by comparing a preset global alien invasive species database, a matching species identification result is obtained; Step 4, according to the matching species identification result, combining the environmental parameter data and the ship navigation data, risk assessment is performed, by analyzing the species survival condition adaptability and the import probability characteristics, and based on a multi-feature decision tree integration mechanism, a colonization risk level is calculated, and a risk level evaluation result is obtained; Step 5, the risk level evaluation result is subjected to diffusion prediction analysis, and according to meteorological data and by a fluid dynamics model, a species diffusion path is simulated; in the simulation process, the cross product of the water flow direction vector and the terrain gradient vector is calculated to determine the diffusion component in the vertical direction, so as to correct the diffusion trajectory of the species, and finally form an early warning information; Step 6, based on the early warning information, by executing early warning response and information publishing, disposal suggestions and early warning contents are formed; the disposal suggestions are delivered to the terminal of the relevant department, and the early warning contents are published to the public platform; the delivery state and the publishing state of the early warning information are recorded in real time to update the early warning response log.
[0020] In the embodiment of the present application, multi-dimensional collection of environmental, ship navigation and biological sample data can comprehensively cover the core monitoring information, avoid the blind area caused by single data, and provide sufficient basis for subsequent analysis; time alignment and format unification processing of multi-source data can eliminate data differences and time deviations, form a standardized data set, and effectively improve the processing efficiency and accuracy of subsequent links; relying on the global alien invasive species database for identification can improve the matching accuracy by using extensive species resources, and reduce misjudgment and omission; combining species survival adaptability, import probability and multi-feature decision tree integration mechanism to assess the risk can objectively quantify the colonization risk level, make the low, medium and high risk division more scientific, and reduce subjective errors; based on meteorological data and fluid dynamics model to simulate the diffusion path, and then correcting the trajectory by vector cross product, can accurately predict the species diffusion range, clearly identify sensitive areas such as agricultural breeding areas and ecological protection areas, and point out the direction for key prevention and control; finally, by differentiating the push of disposal suggestions and early warning contents, and recording the delivery and publishing state in real time, not only can the relevant departments timely dispose and the public timely know, but also can realize the traceability of early warning response, form a complete prevention and control closed loop, and significantly improve the overall efficiency of alien invasive species early warning and prevention and control.
[0021] In a preferred embodiment of the present application, the above step 1 can include: Step 1.1, collecting environmental parameter data of the port water area, wherein the environmental parameter data includes water temperature, salinity, pH value and water flow speed, specifically including: arranging environmental parameter collection equipment in different key areas of the port water area, including the water area near the ship berthing berth, the port water inlet area and the ecological sensitive water area in the port, etc.; selecting sensors with real-time monitoring function to collect water temperature, salinity, pH value and water flow speed data, wherein the water temperature sensor adopts high-precision temperature probe, the salinity sensor measures the water conductivity by electrode method to convert salinity, the pH value sensor is calibrated regularly to ensure measurement accuracy, and the water flow speed sensor is installed on the underwater fixed support to capture the flow speed change of different water layers; the collection equipment automatically records data at a frequency of every 15 minutes, and uploads the data to the data receiving terminal in real time, while preliminarily screening the uploaded data to eliminate abnormal values caused by equipment failure, to ensure that the obtained environmental parameter data accurately reflects the actual environmental conditions of the port water area.
[0022] Step 1.2, obtaining ship-related shipping data, including ballast water discharge record, route trajectory and cargo type, specifically including: establishing data interaction channel with port ship management system to obtain ship-related shipping data; for ballast water discharge record, requiring the ship to submit information such as source sea area, discharge amount and discharge time of ballast water through ship declaration terminal before arriving at the port, automatically receiving these declaration data and matching and verifying with the actual arrival time of the ship, if there is a difference between the declaration data and the actual situation, feedback to the port management department for verification in time; for route trajectory data, real-time record the sailing path of the ship from the departure port to the destination port, including the coordinates of the sea area, the sailing speed and the stay time in each sea area, etc., and arrange these trajectory data in time sequence into continuous route record; for cargo type data, extract the detailed category information of the ship's carried cargo from the port cargo declaration platform, including cargo name, packaging method and whether it belongs to the cargo type of easy-to-carry alien species, etc., and supplement the cargo type related data according to the preliminary verification results of the port cargo inspection department, to ensure the completeness and authenticity of the shipping data.
[0023] Step 1.3, collecting biological sample data, converting the biological sample data into standardized image data and gene sequence data to obtain converted biological sample data, specifically comprising: professional sampling personnel using various methods to collect biological samples in the port area, the sampling range covering water samples in the port water area, attached samples on the deck and cargo hold surface of the ship, and residual samples on the packaging of the transported goods, etc.; for the collected biological samples, first classification and labeling, and marking the sampling time, sampling location and sample preliminary morphological characteristics on the sample container; then processing the biological samples to convert them into standardized image data and gene sequence data, wherein the image data is obtained by using a high-definition digital camera under uniform lighting conditions and shooting distance to take multiple-angle shots of the overall morphology and local features of the biological samples, and then the captured images are uniformly adjusted to a resolution of 1920x1080 and saved according to the naming rules of sampling time, sampling location and sample number; the gene sequence data is obtained by first extracting genomic DNA from the biological sample, ensuring efficiency and purity during the extraction process, then sequencing the extracted DNA, and after obtaining the original gene sequence data, arranging and labeling it according to the internationally accepted gene sequence format standard, removing redundant sequences generated during sequencing, and forming standardized gene sequence data.
[0024] Step 1.4, integrating environmental parameter data, shipping data and converted biological sample data to form a set of raw data to be processed, specifically comprising: In the embodiments of the present application, the environmental parameter data of water temperature, salinity, pH value and flow velocity of the port water area is collected, which can accurately obtain the key environmental conditions for the survival of alien invasive species, provide core basic data for analyzing the survival adaptability of the species, and avoid the deviation of adaptability judgment caused by the lack of environmental data; the shipping data of ballast water discharge records, shipping route and cargo type of the ship are obtained, which can fully master the key carrier information of the spread of alien species through ships, the ballast water discharge records are associated with the species introduction route, the shipping route assists in judging whether the ship passes through the high-incidence area, and the cargo type investigates the species attachment risk, which provides complete basis for analyzing the introduction probability; after collecting the biological sample data, it is converted into standardized image data and gene sequence data, which not only eliminates the format difference of biological data caused by different collection methods through standardization to ensure that the data can be uniformly compared, but also relies on image and gene dual dimensions to provide more comprehensive feature support for accurate species identification and reduce the limitations of single data type identification; integrating the environmental parameter data, shipping data and converted biological sample data to form a set of raw data to be processed can associate scattered environmental, shipping and biological data, avoid data fragmentation, lay a solid foundation for carrying out time alignment and format unified data integration processing, and ensure that the data source of the whole early warning process is complete and relevant.
[0025] In a preferred embodiment of the present application, step 2 above can include: Step 2.1, time stamp alignment processing is performed on the original data set, time-aligned data is obtained by unifying the environmental parameter data, shipping data and biological sample data to the same time reference, specifically including: determining a unified time reference, selecting coordinated universal time as the time standard for all data to ensure consistency of the time reference dimension of different source data; extracting the time stamp information of the environmental parameter data, shipping data and biological sample data from the original data set, wherein the time stamp of the environmental parameter data is the recording time automatically collected by the sensor, the time stamp of the shipping data is the time generated by the ship's Beidou positioning system, and the time stamp of the biological sample data is the sample collection time recorded by the sampling personnel; in view of the problem of different data collection frequencies, for shipping data with higher collection frequency, such as positioning data collected once every 1 minute, data aggregation is performed according to a time window of every 15 minutes, and key information in each time window is taken, such as average position, whether to discharge ballast water, etc.; for biological sample data with lower collection frequency, such as sample data collected once, the time stamp is corresponded to the nearest 15-minute time window; then the time stamps of all data are uniformly formatted, converted into a string format of year, month, day, hour, minute and second, and abnormal time stamps are checked, such as time records exceeding a reasonable time range or being repeated, the three types of data passing the verification are arranged in chronological order, and finally time-aligned data is obtained.
[0026] Step 2.2, the time-aligned data is processed for format standardization, by converting data from different sources into a unified numerical format and data structure, to obtain the format-standardized environmental parameter data, shipping data and biological sample data, which specifically includes: first, according to the data type, a unified format standard is formulated, wherein the numerical format specifies that the water temperature in the environmental parameter data retains one decimal place, the salinity retains two decimal places, the pH value retains one decimal place, and the water flow velocity retains two decimal places; the coordinate value in the shipping data retains six decimal places, and the sailing speed retains one decimal place; in terms of data structure, the structured data table form is adopted, and the time stamp, sampling location, water temperature, salinity, pH value, and water flow velocity fields are set for the environmental parameter data; the time stamp, ship number, route coordinate, ballast water discharge, and cargo type fields are set for the shipping data; the time stamp, sampling location, sample number, image data storage path, and gene sequence data identifier field are set for the biological sample data; according to the formulated format standard, the three types of time-aligned data are processed one by one, the data in the environmental parameter data that does not meet the numerical format requirements is rounded and adjusted, the coordinate data in the shipping data in different formats, such as degree-minute-second format, is converted to decimal format, and the dispersedly stored image and gene sequence data in the biological sample data are re-stored according to the unified path rule and the corresponding identifier is recorded; after the format conversion is completed, it is checked whether each piece of data meets the preset standard, and the entries that do not meet the standard are modified again, and finally the format-standardized environmental parameter data, shipping data and biological sample data are obtained.
[0027] Step 2.3, the environment parameter data, shipping data and biological sample data after format standardization are subjected to data fusion processing to form a standardized data set with consistent space-time reference, specifically including: first, a unified space-time reference is determined, the WGS84 coordinate system is adopted as the reference standard of all geographic position data to ensure the consistency of spatial information such as sampling position and ship route coordinates, and the coordinated universal time is used as the time reference; the three types of data after format standardization are associated and matched with the space-time reference as the core, specifically, the environment parameter data, shipping data and biological sample data in the same time stamp, the same space position, such as the same latitude and longitude range, are correspondingly associated, for example, the shipping data of a ship parked at a specific berth in a certain period of time is associated with the environment parameter data collected near the berth in the same period and the biological sample data collected in the area of the berth; if a certain type of data is missing in the same space-time during the association process, the data is supplemented by a data completion rule, for example, the average value of the same type of data in adjacent space-time is used for supplement; if data conflict occurs, for example, there are two different water temperature values in the environment parameter data in the same space-time, the data with a collection time closer to the space-time node is selected as the effective data; all the data after association processing are integrated into a unified database, an association index between the data is established to facilitate fast calling in the later stage, and finally a standardized data set with consistent space-time reference and containing complete environment, shipping and biological information is formed.
[0028] In the embodiment of the application, the time stamp alignment processing is performed on the original data set, the environment parameter data, shipping data and biological sample data are unified to the same time reference, different types of data can be accurately associated according to the time dimension, the analysis interference caused by time deviation is eliminated, and the judgment error caused by different data time is avoided; the data after time alignment is subjected to format standardization processing, different sources of data are converted into a unified numerical format and data structure, the problem of data format confusion can be solved, the additional format adjustment link in the subsequent processing is saved, and the data processing efficiency is effectively improved; the three types of data after format standardization are fused to form a standardized data set with consistent space-time reference, the core information of the environment, shipping and biological data can be fully integrated, a unified and complete data source is provided for species identification, risk assessment and other links, the accuracy and continuity of data application in the whole early warning process are ensured, and the problem of insufficient data integration effectiveness in the prior art is effectively improved.
[0029] In a preferred embodiment of the application, step 3 can include: Step 3.1, extracting biological sample data from the standardized data set, the biological sample data including image data and gene sequence data, specifically including: determining the target space-time range for which the biological sample data needs to be extracted based on the consistent space-time reference of the standardized data set, which is consistent with the key port area of the previous data collection, such as ship berthing, ecological sensitive water area and the like, and the corresponding time window; filtering out the entries containing the biological sample data identifier field from the standardized data set, and distinguishing and extracting the image data and gene sequence data according to the field, wherein the image data is obtained by calling the file corresponding to the data storage path, and the gene sequence data is obtained by reading the complete sequence record corresponding to the gene sequence data identifier; synchronously associating the space-time information corresponding to each biological sample data, i.e. the sampling time and the sampling position, during the extraction process, to ensure that the extracted image data and gene sequence data can form a space-time correspondence with the previous environmental parameter data and shipping data, and avoid disconnection of the biological sample data from other data.
[0030] Step 3.2, comparing the image data with the morphological feature data in the preset global alien invasive species database to obtain a morphological feature comparison result; at the same time, performing homology analysis on the gene sequence data and the gene sequence data in the preset global alien invasive species database to obtain a gene sequence analysis result, specifically including: for the image data, first removing noise interference in the image by using image preprocessing technology, such as background impurities and abnormal areas caused by light reflection, and then extracting key morphological features of the biological sample, including sample contour, surface texture, organ structure size and the like; comparing the extracted morphological features with the species morphological feature data in the preset global alien invasive species database one by one to calculate the feature similarity, such as contour coincidence degree and texture consistency; generating a morphological feature comparison result according to the calculation result, which contains multiple species names and corresponding similarity values closest to the sample morphology; for the gene sequence data, first extracting conserved fragments from the standardized gene sequence, i.e. gene fragments with significant differences between species and stable inheritance, and then comparing the conserved fragments with the gene sequence conserved fragments of each species in the preset global alien invasive species database, and counting the number and proportion of matching base pairs, to generate a gene sequence analysis result according to the base pair matching, which contains multiple species names and corresponding homology proportions with the highest homology to the sample gene sequence.
[0031] Step 3.3, based on the morphological feature comparison result and the gene sequence analysis result, the comprehensive matching degree is calculated by weighted fusion calculation to determine the matching species and its confidence, specifically including: first, according to the accuracy rate data of historical species identification cases, the weight distribution of morphological feature comparison result and gene sequence analysis result is determined, wherein the weight of gene sequence analysis result is set to 0.6 due to the stability of genetic characteristics, and the weight of morphological feature comparison result is set to 0.4 due to the possible deviation caused by the influence of environment, and the weight can be dynamically adjusted according to the identification characteristics of different species types; for the same candidate species, multiply its morphological similarity value by the morphological feature weight, and multiply its gene homology ratio by the gene sequence weight, and the sum of the two is the comprehensive matching degree of the species; after traversing all candidate species and calculating their respective comprehensive matching degrees, the species with the highest comprehensive matching degree is selected as the preliminary matching species, and the comprehensive matching degree is converted into confidence, such as comprehensive matching degree 0.9 corresponding to confidence 90%, the preliminary matching species and its corresponding confidence value are determined, ensuring that the matching result can comprehensively reflect the matching of morphological and genetic characteristics.
[0032] Step 3.4, according to the preset confidence threshold, the matching species is screened to obtain the matching species identification result, which contains species name and matching confidence information, specifically including: first, based on the misjudgment rate data of historical alien species identification, the confidence threshold is set, such as setting the threshold to 80%, this threshold can be adjusted according to the port prevention and control demand and the species harm degree, and the threshold setting needs to ensure that the low confidence matching result caused by the similarity of characteristics can be filtered, while avoiding missing the real invasive species; compare the confidence of the preliminary matching species with the preset confidence threshold, if the confidence is higher than or equal to the threshold, the species is retained as an effective matching species; if the confidence is lower than the threshold, the species is excluded, and the range of comparison is expanded or the biological sample features are supplemented again in step 3.2; for the retained effective matching species, the complete species name is arranged, including the Latin name and the common name, and the corresponding matching confidence value, that is, the calculated confidence, is labeled, and finally the matching species identification result containing the species name and the matching confidence information is formed.
[0033] In the embodiment of the present application, the image data and gene sequence data are extracted from the standardized data set, the core feature information of the biological sample can be accurately obtained, and a direct and key analysis basis for species identification is provided; the image data is compared with the morphological feature data of the global alien invasive species database, and the gene sequence data is compared with the gene sequence of the database for homology analysis, through the comparison in the morphological and gene dual dimensions, the species characteristics can be comprehensively captured, and the one-sidedness caused by single feature recognition can be avoided; the weighted fusion calculation comprehensive matching degree is carried out based on the two types of analysis results, the importance of different characteristics can be combined to obtain a more objective matching result, and the matching species and the confidence degree are accurately determined; the matching species is screened according to the pre-set confidence threshold, the matching result with low confidence can be filtered, it is ensured that the final identification result contains reliable species name and matching confidence information, and the accuracy and reliability of the alien invasive species identification are effectively improved.
[0034] In a preferred embodiment of the present application, the above step 4 can include: Step 4.1, extracting the species name from the matching species identification result, and obtaining the corresponding species survival condition parameters from the pre-set global alien invasive species database based on the species name, specifically including: first, extracting the complete species name from the matching species identification result, the name contains the Latin name and the common name of the species, and the spelling accuracy of the name needs to be checked during extraction to avoid database retrieval deviation caused by name errors; then, logging into the pre-set global alien invasive species database, inputting the extracted species name in the database, triggering the accurate retrieval function of the database; according to the species name, the complete survival condition parameters corresponding to the species are called, these parameters include the key information such as the suitable water temperature range, the salinity range, the pH value range and the adaptation interval of the water flow speed; after the retrieval is completed, the obtained species survival condition parameters are arranged according to the structure of the species name, the parameter type and the parameter range, so that each parameter can be clearly corresponded to the species.
[0035] Step 4.2, performing adaptability analysis on the environmental parameter data in the standardized data set and the species survival condition parameters to obtain adaptability analysis results, specifically including: extracting environmental parameter data corresponding to the port water area from the standardized data set, including the average value and fluctuation range of water temperature, the average value and fluctuation range of salinity, the average value and fluctuation range of pH value, and the common numerical interval of water flow speed within a period of time; comparing and analyzing the extracted environmental parameter data and the obtained species survival condition parameters one by one, for the water temperature parameter, judging whether the fluctuation range of the environmental water temperature is completely or partially within the species suitable water temperature range, and calculating the overlap proportion; for the parameters of salinity, pH value, water flow speed, etc., the same comparison method is adopted, and the overlap proportions of the environmental parameters and the species suitable parameters are calculated respectively; comprehensively judging according to the overlap proportions of all parameters, if the overlap proportions of all parameters are all more than 80%, it is determined as high adaptability; if the overlap proportion is between 50% and 80%, it is determined as medium adaptability; if the overlap proportion is less than 50%, it is determined as low adaptability, and finally the adaptability analysis results including the comparison of each parameter and the comprehensive adaptability level are formed.
[0036] Step 4.3, based on the route trajectory in the shipping data, the minimum distance between the ship position point and the preset high-risk sea area boundary of the alien invasive species is calculated to quantify the spatial proximity; based on the spatial proximity, the transmission probability analysis result is obtained by combining the ballast water discharge record, specifically including: first, the complete route trajectory of the ship is extracted from the shipping data, the trajectory data contains the latitude and longitude coordinates of the ship at different time points during the voyage, and these coordinates are arranged in chronological order into a continuous sequence of position points; load the boundary data of the preset high-risk sea area of the alien invasive species, the preset rule of the boundary data is: combined with the high-risk sea area record of the global alien invasive species database in the past 10 years of historical invasion cases, the ecological environment similarity (such as water temperature, salinity interval matching degree greater than or equal to 80%) between the target sea area and the known high-risk sea area and the monthly average stop frequency of the ship in the sea area in the past 3 years (greater than or equal to 5 times / month), the initial boundary is delineated by spatial overlay analysis; and every quarter, according to the newly added invasion case data of the global alien invasive species database and the sea area ecological monitoring data, such as annual change of water temperature, the boundary range is dynamically adjusted to ensure the timeliness of the boundary; the boundary data of the preset high-risk sea area of the alien invasive species is presented in the form of latitude and longitude coordinate string, which clearly divides the range of the high-risk sea area, and the straight-line distance from each ship position point in the route trajectory to the boundary of the high-risk sea area is calculated in turn, the minimum distance value is selected from all the calculation results, which is the minimum distance between the ship and the high-risk sea area boundary; the spatial proximity is quantified according to the minimum distance value, if the minimum distance is less than 10 nautical miles, it is determined as high spatial proximity, if it is between 10 nautical miles and 50 nautical miles, it is determined as medium spatial proximity, and if it is greater than 50 nautical miles, it is determined as low spatial proximity; then combined with the ballast water discharge record of the ship, if the ship has a ballast water discharge record in the high spatial proximity area and the discharge amount is large, it is determined as high transmission probability; if there is a discharge record in the medium spatial proximity area or the discharge amount is small, it is determined as medium transmission probability; if there is no discharge record or no discharge amount in the low spatial proximity area, it is determined as low transmission probability, and finally the transmission probability analysis result is formed.
[0037] Step 4.4, based on the results of the adaptation analysis and the results of the transmission probability analysis, combined with the pre-processed historical invasion case data, by using multiple decision trees to process multiple feature dimensions in parallel, multiple basic risk assessment results are obtained, including: first, the pre-collected historical invasion case data is processed, which contains the survival condition adaptation level, transmission probability level, and final colonization risk level of the invading species in past external invasion events, as well as auxiliary information such as the port environment and ship navigation characteristics of the event, and these data are standardized according to the structure of case number, adaptation level, transmission probability level, and colonization risk level; the results of the adaptation analysis and the results of the transmission probability analysis are used as input data for multiple decision trees together with the processed historical invasion case data; three decision trees with different structures are selected, the first decision tree focuses on processing the correlation between the adaptation analysis results and the adaptation-related data in the historical cases, the second decision tree focuses on processing the correlation between the transmission probability analysis results and the transmission probability-related data in the historical cases, and the third decision tree comprehensively processes the overall correlation between adaptation, transmission probability, and historical cases; the three decision trees run in parallel, each making logical judgments and inferences based on the input feature dimensions, and each decision tree outputs a corresponding colonization risk level recommendation, which are the multiple basic risk assessment results.
[0038] Step 4.5, multiple basic risk assessment results are integrated and processed, and the final colonization risk level is determined by using a voting mechanism, based on which a risk level assessment result containing low, medium, and high levels is obtained, including: collecting multiple basic risk assessment results and counting the colonization risk levels in these results, counting the number of times each level, i.e. low, medium, and high, appears; using a voting mechanism for integrated processing, if a certain level appears the most, then that level is determined as the final colonization risk level; if two levels appear the same number of times, refer to the colonization risk level of the historical invasion case data under the same combination of the current adaptation analysis results and transmission probability analysis results, and use the level appearing in the historical case as the final colonization risk level; if three levels appear the same number of times, a special case, then by default the level is determined as medium risk; after determining the final colonization risk level, combined with the previous adaptation analysis results and transmission probability analysis results, a risk level assessment result containing the final risk level and level determination basis is formed, which clearly indicates which level of the low, medium, and high levels the colonization risk of the current matching species belongs to.
[0039] In the embodiment of the present application, the species name is extracted from the matching species identification result, and the corresponding survival condition parameters are obtained from the database according to the species name, so that the survival demand of the species can be accurately locked, an accurate basis for judging whether the environment is suitable is provided, and analysis deviation caused by parameter mismatch is avoided; the adaptability analysis is performed on the standardized environment parameters and the species survival condition parameters, so that it can be clearly judged whether the current port water environment meets the species colonization demand, and a core environmental adaptation basis for risk assessment is provided; by quantifying the spatial proximity through calculating the minimum distance between the ship and the high-invasion sea area, and combining with the ballast water discharge record to analyze the import probability, the possibility of species import through the ship can be objectively evaluated, and the error caused by subjective judgment is avoided; based on the adaptability, the import probability and the historical case data, the multiple decision trees are used to process multiple feature dimensions in parallel, so that multiple key factors of risk assessment can be covered, the one-sidedness of single dimension analysis is avoided, and a more comprehensive basic risk result is obtained; the final colonization risk level is determined by integrating multiple basic risk results through the voting mechanism, so that the analysis deviation of different decision trees can be balanced, the division of low, medium and high levels is more objective and scientific, and the systematicness and accuracy of risk assessment are effectively improved.
[0040] In a preferred embodiment of the present application, the above step 5 can include: Step 5.1, extract the species information and its location data marked as high risk level in the risk level evaluation result, specifically including: first, call the database file storing the risk level evaluation result, which records the complete information of all matching species, including species name, matching credibility, colonization risk level and associated sampling information; filter out the species entries with high risk level clearly marked in the file, extract detailed species information from these entries, including the Latin name, common name, typical morphological characteristics of the species, such as body size, body color, special organ structure, and known ecological harm characteristics, such as the predation relationship with local aquatic organisms and the destruction mode of water ecological balance; at the same time, the location data corresponding to the high-risk species is extracted, which is the sampling point latitude and longitude coordinates recorded by the positioning device during the early biological sample collection, and the sampling time information needs to be extracted synchronously to ensure the timeliness of the location data; after extraction, arrange the species information and location data into a structured data table in the order of species Latin name, common name, morphological characteristics, ecological harm, sampling latitude and longitude, and sampling time, and then associate and verify the data table with the standardized data set formed in the early stage to confirm that the sampling location and sampling time are completely consistent with the environmental parameter data and the time and space range of shipping data in the standardized data set, so as to avoid data misplacement.
[0041] Step 5.2, based on the species information of high-risk level and its location data, combined with real-time acquisition of meteorological data and hydrological data, the diffusion process of the species in the water body is simulated by fluid dynamics model, specifically including: obtaining real-time meteorological data within 50 kilometers around the port from the authoritative meteorological data platform, including real-time wind speed updated every hour, its unit is meter per second; wind direction, angle value based on the north as the benchmark; air temperature, unit is degree Celsius, and precipitation probability in the next 24 hours, each data is marked with specific acquisition time; obtain real-time hydrological data from the port dedicated hydrological monitoring station, covering real-time water flow speed recorded every hour, unit is meter per second; water flow direction, angle value based on the north as the benchmark; water temperature, unit is degree Celsius, and water level change amplitude, unit is meter, ensure that the time of hydrological data is completely synchronized with the time of meteorological data; based on the information of high-risk species, the motion characteristic parameters of the species in the water body are determined by querying the preset global alien invasive species database, such as if the species is passive diffusion type, record its floating ability, like the volume and weight ratio of floating object; if the species has active movement ability, record its average swimming speed, unit is meter per second; real-time meteorological data, real-time hydrological data and species motion characteristic parameters are input into fluid dynamics model to simulate water flow state, and then combined with species motion characteristics to calculate the moving track of species in water body. The simulation process is divided into stages according to 12 hours as a time period, the moving direction and distance of the species in each stage is calculated according to the average meteorological conditions and average hydrological conditions in the period, and the position coordinates of the species at different time nodes, such as 12 hours, 24 hours, 36 hours after the simulation starts, are gradually generated, so as to completely present the diffusion process of the species in the water body.
[0042] Step 5.3, in the process of fluid dynamics model simulation, the vertical direction diffusion component is determined by calculating the cross product of the water flow direction vector and the terrain gradient vector, the horizontal diffusion trajectory is corrected based on the diffusion component, and the corrected diffusion trajectory is obtained, which specifically includes: in the process of fluid dynamics model simulation, first, the water flow direction information of the current time period is extracted from the real-time hydrological data, and the water flow direction information is converted into a water flow direction vector in a three-dimensional coordinate system, i.e. east as the x-axis, north as the y-axis, and vertical as the z-axis, the vertical initial value is set to 0 and is adjusted only with the terrain gradient; this coordinate system takes the sampling point as the origin, the east direction as the positive direction of the x-axis, and the north direction as the positive direction of the y-axis, the direction of the vector is consistent with the actual flow direction of the water flow, and the modulus of the vector is equal to the value of the water flow speed; the terrain data of the water area is obtained, including the elevation value of each coordinate point in the water area, the unit is meter, the terrain slope of different positions is calculated according to the elevation value, the unit is degree, and the slope direction is an angle value based on the north; the slope and slope information are converted into a terrain gradient vector, a three-dimensional vector based on the water area elevation data, including x-axis, y-axis and z-axis components, corresponding to horizontal east, horizontal north and vertical gradient changes respectively; the direction of the vector points to the direction of the fastest decrease of the terrain elevation, and the modulus of the vector is proportional to the slope value; then the cross product operation is performed on the water flow direction vector and the terrain gradient vector, the strength of the vertical diffusion is judged according to the numerical value of the operation result, and the greater the absolute value, the higher the strength; the direction of the vertical diffusion is judged according to the positive and negative signs of the operation result, the positive value indicates that the species has a tendency to diffuse to the upper layer of water, and the negative value indicates that the species has a tendency to diffuse to the lower layer of water; the horizontal diffusion trajectory obtained by previous simulation is adjusted according to the diffusion component in the vertical direction, for example, when the operation result is positive and the absolute value is large, the water depth level corresponding to the horizontal trajectory is appropriately increased, such as from 5 meters under water to 2 meters under water; when the operation result is negative and the absolute value is large, the water depth level corresponding to the horizontal trajectory is appropriately reduced, such as from 2 meters under water to 5 meters under water, and finally the corrected diffusion trajectory considering both the horizontal diffusion trend and the vertical diffusion trend is obtained.
[0043] Step 5.4, spatial overlay analysis of the corrected diffusion trajectory and the preset sensitive area boundary to identify the agricultural breeding areas and ecological protection areas that may be affected, to obtain a spatial overlay analysis result, specifically including: first, structuring the preset sensitive area boundary data, wherein the boundary data of the agricultural breeding areas is derived from the official breeding area planning documents published by the local agricultural and rural departments, and each breeding area data is converted into structured information containing the name, the administrative area to which it belongs, the coordinate string of longitude and latitude, and the main breeding varieties. The coordinate string is a coordinate sequence that connects multiple coordinate points in a clockwise order to form a closed area. The boundary data of the ecological protection areas is derived from the natural protection area directory published by the local ecological and environmental departments, and is also converted into structured information containing the name, the protection level, the coordinate string of longitude and latitude, and the main protection objects. The coordinate string is also a coordinate sequence that connects multiple coordinate points in a clockwise order to form a closed area. Then, the corrected diffusion trajectory data is processed and converted into a time series of coordinate sets, each element of which contains a simulation time node, such as 12 hours and 24 hours after the simulation starts, and the longitude and latitude coordinates of the species at the corresponding time. Subsequently, spatial overlay analysis is carried out, and the coordinate comparison method is used to determine whether each coordinate point in the trajectory data falls within the coordinate range of the sensitive area. First, the closed polygon formed by the coordinate string of the sensitive area is determined, and then the position relationship between the trajectory coordinate point and the polygon is calculated. If the coordinate point is inside or on the boundary of the polygon, it is determined that the corresponding sensitive area may be affected. For the sensitive areas determined to be affected, record their name, type, whether it is an agricultural breeding area or an ecological protection area, the expected arrival time of the species, the simulation time node in the corresponding trajectory, and calculate the intersection area of the total area of the sensitive area and the area where the trajectory coordinate point is located by the polygon area calculation formula. Divide the intersection area by the total area to get the impact area ratio, expressed in percentage. Finally, form a spatial overlay analysis result containing all the above information.
[0044] Step 5.5, based on the results of spatial superposition analysis, forming early warning information including diffusion path diagram, impact area and risk level, specifically including: based on the results of spatial superposition analysis, constructing a species diffusion path diagram. When constructing, use a continuous solid line to mark the corrected diffusion trajectory, and at every other simulation time node on the trajectory, such as every 12 hours, mark the corresponding specific time, such as 24 hours of simulation, and the latitude and longitude coordinates of the species at that time; use different colored closed figures to distinguish sensitive area types, mark agricultural breeding areas with blue closed figures, and mark ecological protection areas with green closed figures; for sensitive areas that are determined to be potentially affected, add diagonal lines inside the closed figure to highlight; then sort the impact area information into a text table according to the fixed structure of sensitive area type, sensitive area name, expected impact time, and impact area ratio; the table must clearly distinguish between agricultural breeding areas and ecological protection areas, and each entry must record the corresponding information completely to ensure that the information is clear and traceable; then integrate the diffusion path diagram, impact area information table, and high-risk species information into early warning information; first mark the complete name of the high-risk species, including the Latin name and common name, and the corresponding colonization risk level, that is, high risk, to ensure consistency with the results of the previous risk level assessment; finally, supplement the risk prompt content, reminding relevant regional management departments to prepare for prevention and control before the species is expected to arrive, such as setting up water interception facilities, increasing the frequency of regional monitoring, reminding the public to avoid engaging in activities that may promote the spread of species in affected areas, such as unauthorized discharge of breeding wastewater, and randomly discarding aquatic organism samples.
[0045] The specific construction and training process of the fluid dynamics model is as follows: When constructing the fluid dynamics model, first combine the requirements of the water-borne invasive species diffusion prediction in the present application, and determine that the core goal of the model is to accurately simulate the diffusion process of high-risk species in port water. The simulation range is within 50 kilometers of the port and surrounding water, covering key areas such as ship docking berths, agricultural breeding areas, and ecological protection areas, to ensure that the model output can directly serve the subsequent sensitive area impact analysis. At the same time, it is clear that the model needs to consider both horizontal and vertical diffusion simulation, which makes up for the limitations of traditional models that only focus on horizontal trajectories, which is compatible with the need to calculate the vertical diffusion component by vector cross product.
[0046] Based on the results of previous data collection and processing, the core parameters required by the model are screened and divided into three categories. The first category is the environmental driving parameter, including the real-time flow velocity, flow direction and water temperature of the port water area. These data come from real-time hydrological data, as well as wind speed and direction, which also come from real-time meteorological data. They directly affect the flow state of the water body. The second category is the terrain parameter, that is, the digital elevation model data of the port water area, which is used to calculate the terrain gradient vector and support the simulation of vertical diffusion components. The third category is the species characteristic parameter, which is determined according to the biological characteristics of high-risk species, such as whether they have the ability of active movement and floating ability, etc. For example, the floating volume to weight ratio of passive diffusion species and the average swimming speed of active diffusion species. These data come from the query results of the global alien invasive species database.
[0047] The model is divided into two core modules: water flow simulation module and species diffusion simulation module. The modules are linked through data interface. The water flow simulation module is responsible for calculating the flow trajectory and velocity field of the water body according to the environmental driving parameters and terrain parameters. It is realized by restoring the flow direction vector, while simulating the flow rate difference of different water layers. The species diffusion simulation module is based on the results of water flow simulation, combined with the species characteristic parameters, to calculate the moving trajectory of the species in the water body. It also integrates the vertical diffusion component calculation logic, that is, the vertical diffusion trend is determined by the cross product of the flow direction vector and the terrain gradient vector, to realize the coordinated simulation of horizontal and vertical directions. In addition, the module reserves a data input interface, which can directly interface with the extracted high-risk species position data as the diffusion starting point, ensuring seamless connection between the model and the previous data flow process.
[0048] The framework is built using hierarchical calculation logic. The first layer is the water state calculation layer, which converts water flow velocity, wind direction and other parameters into flow vectors of each coordinate point in the water body by solving the water motion equation. The second layer is the species motion calculation layer, which calculates the horizontal movement distance and direction of the species in unit time based on the water flow vector and species characteristic parameters, and introduces the vertical diffusion component calculation logic to adjust the distribution of the species in different water layers. The third layer is the trajectory output layer, which organizes the calculated species position coordinates in time sequence, such as recording the position once every 12 hours, to form a trajectory data format consistent with the requirements of the diffusion process simulation in step 5.2, which is convenient for subsequent superposition analysis with the sensitive area boundary.
[0049] Then the model training process is carried out, and the training data is derived from two types of key data. The first type is historical environmental data, including hourly hydrological data of the port water area in the past 5 years, such as flow speed, direction, water temperature, etc., and meteorological data, such as wind speed, direction, etc., which are obtained through port hydrological monitoring stations and meteorological platforms. The collection dimension of real-time data is consistent to ensure data compatibility. The second type is historical diffusion verification data, which extracts historical invasion case data similar to the ecological characteristics of the port water area in the global alien invasive species database, including the actual diffusion trajectory of the invasive species in the case, such as the location coordinates of the species at different time nodes, as well as species characteristic parameters, and corresponding environmental data. These data include both high-risk species cases that have occurred invasions, and species cases that have not successfully colonized, which can fully support model training.
[0050] Referring to the previous data processing logic of the application, such as time alignment, format standardization, the training data is preprocessed. First, clean up the outliers in the environmental data, such as instantaneous flow anomalies caused by equipment failure, and use the average value of adjacent period data to complete the missing values. Then convert the angle parameters such as flow direction and wind direction into two-dimensional vector format, which corresponds to the vector calculation requirement, and standardize the numerical parameters such as flow speed and wind speed to the interval of 0 to 1 to avoid affecting model calculation due to unit differences. Finally, the historical diffusion trajectory data is split by time sequence, with each 12 hours as a data segment, consistent with the model simulation period, to ensure that the training rhythm matches the model application rhythm. Based on the historical environmental data, the water flow simulation module is initialized, and the average hydrological and meteorological parameters of the past year are input. The initial water flow field is obtained by running the module, which is compared with the actual monitored water flow state at the same period, and the internal parameters such as water flow resistance coefficient and terrain influence weight in the module are adjusted until the error between the initial flow field and the actual monitoring value is less than 10%. Combined with the species characteristic parameters in the historical diffusion cases, the species diffusion simulation module is initialized. For example, for passive diffusion species, the initial floating coefficient is set to 0.6, which is based on historical data of similar species. For active diffusion species, the initial swimming speed is set to 0.2 meters per second to ensure that the initial state of the module conforms to the actual motion law of the species.
[0051] Subsequently, iterative training is performed in stages, 3-year historical environmental data is selected as the training set, and the training batches are divided according to quarters. After each batch of data is input into the water flow module, the deviation between the flow rate and flow direction output by the calculation module and the actual monitoring data is calculated, and the gradient descent method is used to adjust the parameters in the module, such as the calculation weight of the water flow direction vector and the influence coefficient of the terrain gradient on the flow rate. After completing the training of 1 quarter of data, the next quarter of data is used for verification. If the deviation exceeds 15%, return to the previous batch to adjust the parameters again, until the deviation of all training batches is stable within 10%, to ensure that the module can accurately restore the water flow state under different environmental conditions. The trained water flow module and the species diffusion module are linked, and the complete data of the historical diffusion case is input, including environmental data, species characteristic parameters and actual diffusion trajectory. First, fix the water flow module parameters, and only adjust the parameters in the species diffusion module, such as the calculation weight of the vertical diffusion component and the interaction coefficient between the species and the water, so that the coincidence degree of the diffusion trajectory output by the module and the actual trajectory reaches more than 80%. Then release the water flow module parameters, and jointly train the two modules as a whole to simulate the complete process of species diffusion under different environmental conditions, for example, when there is a high flow rate in the historical case. Strengthen the wind direction of the environment, focus on adjusting the adaptability of the module to extreme environments to ensure that the model can still output stably under complex conditions.
[0052] Finally, the model verification and precision optimization process is performed, 2-year historical diffusion case data not involved in training is selected as the verification set, environmental data and species characteristic parameters in the verification set are input into the trained model to obtain the simulated diffusion trajectory, which is compared with the actual diffusion trajectory in the case. Calculate the deviation of both at key time nodes, such as the time when the species enters the sensitive area. If the deviation is less than 24 hours and the position deviation is less than 5 kilometers, which meets the precision requirements of port water area prevention and control, the model verification is passed. If the deviation exceeds the threshold, analyze the reasons and adjust accordingly, for example, if the species characteristic parameter setting is not accurate, requery the global alien invasive species database to supplement the parameter details.
[0053] Based on the verification results, special optimization is performed in combination with the specific characteristics of the port water area in the present application, such as the distribution of sensitive areas, common high-risk species types, etc. For example, according to the characteristics of slow water flow around the agricultural breeding area in the port, the species diffusion coefficient under low-speed water flow in the model is adjusted. For the calculation of the vertical diffusion component, the weight distribution of the vector cross product operation is optimized to ensure that the vertical diffusion simulation is more consistent with the actual port terrain, such as the difference in vertical diffusion between shallow and deep water areas. After optimization, verification is performed again until the diffusion trajectory simulation precision of the model in the application scenario of the present application meets the requirements, and the model can be directly used for diffusion prediction of high-risk species.
[0054] In the embodiments of the present application, the information of high-risk grade species and its location data are extracted, the invasive species needing attention and the starting point of diffusion are accurately locked, the dispersion of prevention and control resources on low-risk species is avoided, the core object and initial position basis for subsequent diffusion simulation are provided, the diffusion process is simulated based on the high-risk species information, location data, real-time meteorological and hydrological data and fluid dynamics model, the influence of actual environmental factors on species diffusion is fully considered, the simulation process is more in line with the real movement of species in water, and the simulation deviation caused by deviation from the actual environment is reduced, the vertical diffusion component is determined by calculating the vector cross product during simulation, the horizontal trajectory is corrected, the limitations of traditional simulation only focusing on the horizontal direction are made up, the diffusion trend in the vertical direction is considered, and the prediction accuracy of the diffusion trajectory is effectively improved, the agricultural breeding areas and ecological protection areas that may be affected are clearly identified by spatial overlay analysis of the corrected trajectory and the boundary of the sensitive area, the key areas for prevention and control are determined, and the prevention and control omissions caused by the failure to locate the sensitive influence area are avoided, the warning information including the diffusion path diagram, the influence area and the risk grade is formed based on the overlay analysis result, the species diffusion risk can be intuitively and comprehensively presented, clear guidance is provided for the relevant departments to formulate prevention and control measures and for the public to understand the risk situation, and efficient risk response work is facilitated.
[0055] In a preferred embodiment of the present application, step 6 can include: Step 6.1, the warning information is analyzed and processed to obtain species characteristics, risk grade, influence area data and diffusion path data, specifically including: first, the complete warning information generated is obtained, which includes the diffusion path diagram, the influence area table, the risk grade and the species related content; the warning information is disassembled according to the preset data analysis rule, the species characteristics are extracted from the species related content, including the Latin name, the common name, the typical morphological identification points of the species, such as body size, body color difference, and the main ecological harm mode, such as whether to prey on local organisms or to block water conservancy facilities; the risk grade is extracted from the explicit label of the warning information, that is, the three levels of high, medium and low, and the determination basis summary corresponding to the risk grade is confirmed; the influence area data is extracted from the influence area table, covering the name of the affected area, the administrative range, the latitude and longitude range and the expected influence time. The diffusion path data is extracted from the diffusion path diagram, including the species position coordinates, the diffusion direction and the diffusion speed estimation value at different time nodes; all the extracted data are arranged into a standardized data table according to the structure of species characteristics, risk grade, influence area data and diffusion path data, the data in the table is subjected to completeness check, if some kind of data is missing, it is returned to the previous step to supplement and obtain, and the key data obtained by analysis is ensured to be comprehensive and accurate.
[0056] Step 6.2, based on the risk level, automatically match the corresponding treatment measures from the preset treatment scheme library to form a treatment proposal containing specific operation instructions, including: calling the preset treatment scheme library, which is divided into high, medium and low risk sub-libraries according to risk level, each sub-library stores standardized treatment measures corresponding to the risk level; measures in the high-risk sub-library include emergency interception facility layout, such as setting up a gillnet at the key node of the spread path; special monitoring team formation, clear monitoring frequency and monitoring range; affected area biological investigation, specify the investigation time and investigation method; measures in the medium-risk sub-library include increasing the monitoring frequency, such as adjusting from once a day to twice a day; setting up regional warning signs, specifying the sign location and content; starting the relevant department communication mechanism; measures in the low-risk sub-library include regular data tracking, such as weekly summary of spread data; public prevention knowledge propaganda, determine the propaganda content framework; according to the risk level obtained by analysis, automatically locate to the corresponding sub-library in the treatment scheme library, and call all suitable treatment measures from the sub-library; combined with the species characteristics and influence area data obtained by analysis, adjust the details of the called treatment measures, for example, optimize the mesh size of the interception facility for species good at drilling, adjust the investigation range for breeding area influence, finally form a treatment proposal containing operation subjects such as agricultural department and environmental protection department, specific operation steps, completion time limit and expected effect, to ensure that each operation instruction is clear and executable.
[0057] Step 6.3, according to the influence area data, determine the related department terminal list that needs to be pushed for early warning, based on the terminal list, automatically transmit the treatment proposal to the corresponding terminal through the data interface, including: first analyze the influence area data, determine the corresponding responsible department according to the type and range of the affected area; if the affected area includes agricultural breeding area, associate the agricultural and rural department; if it contains ecological protection area, associate the ecological environment department; if it involves water conservancy facilities, associate the water conservancy department; if it crosses multiple administrative regions, associate the same level of competent department of the corresponding region; call the terminal information of these responsible departments from the preset department terminal database, including the IP address, port number and terminal responsible person contact of the department exclusive data receiving terminal, and organize the related department terminal list that needs to be pushed for early warning; connect with each terminal in the terminal list through the preset data transmission interface, after successful connection, convert the treatment proposal into a data format suitable for the terminal, such as XML format, JSON format, after conversion, automatically transmit the treatment proposal to the corresponding terminal; real-time check the integrity of data transmission during transmission, if a terminal fails to transmit, immediately start a backup transmission channel to resend, and record the transmission failure reason and retry times, to ensure that all responsible department terminals can receive the treatment proposal.
[0058] Step 6.4, based on the species characteristics, risk level and impact area data, generate early warning content for different user groups, and push the early warning content to the public platform through the preset release channel, including: first, divide different user groups, mainly including affected area breeder groups, ordinary public groups, related research institution groups and community management groups; for breeder groups, generate early warning content combined with species characteristics and impact area data, focusing on the harmful ways of species to breeding varieties, breeding area prevention measures, such as breeding water filtration method, and suspected species discovery reporting channel; for ordinary public groups, generate easy-to-understand early warning content, including simple identification pictures of species, public contact or spread precautions, such as not randomly discarding aquatic pets, and risk area safety tips; for research institution groups, generate professional early warning content including species gene sequence key fragments, diffusion model parameters and impact area ecological background data; for community management groups, generate early warning content including community propaganda task list, regional patrol focus and emergency contact person information; determine the preset release channel according to the information acquisition habit of different user groups, breeder groups through the agricultural department industry platform, ordinary public groups through local government public number, social media platform, research institution groups through professional academic data platform, community management groups through government office system; release corresponding early warning content one by one according to the determined channel, and record the release time and release channel information after release.
[0059] Step 6.5, real-time monitoring of the delivery state of the treatment suggestion and the publishing state of the early warning content to update the early warning response log and complete the early warning closed-loop management, specifically including: starting a real-time monitoring program to monitor the delivery state of the treatment suggestion, mainly monitoring the receiving confirmation information of each department terminal, the opening viewing record of the treatment suggestion and the data integrity feedback; if no receiving confirmation is received from a certain terminal for more than 30 minutes, a reminder information is automatically sent to the person in charge of the terminal, and the name of the unconfirmed terminal and the number of reminders are recorded; the publishing state of the early warning content is monitored, including the content display state of each publishing channel, whether it is normally displayed, whether there is content missing, user click volume statistics and feedback message collection; if it is found that the content display of a certain channel is abnormal, the channel maintenance personnel are immediately contacted for repair, and the abnormal channel name, the abnormal type and the repair time are recorded; the monitored delivery state data such as the number of successfully received terminals and the number of transmission failures are recorded in the early warning response log in real time, together with the publishing state data such as the click volume of each channel and the abnormal repair situation, and the operation time, the operator and the key data summary of each link are also included in the log; the early warning response log is regularly summarized and analyzed, and if problems such as delay in delivery of treatment suggestions or insufficient coverage of early warning content publishing are found, optimization schemes such as increasing standby terminals and expanding publishing channels are proposed to ensure that the early warning is traceable from generation, pushing, response to recording, and finally complete the early warning closed-loop management.
[0060] In the embodiment of the present application, the early warning information is analyzed and processed to obtain species characteristics, risk level and other key data, which can accurately extract the core information of the early warning, provide clear basis for matching treatment measures and pushing early warning content, and avoid deviation in subsequent work caused by incomplete information extraction; based on the risk level, the treatment measures are automatically matched from the treatment scheme library and the operation instructions are formed, which can save the tedious process of manual scheme screening, make the treatment suggestion more suitable for the actual situation of the risk, and improve the efficiency and accuracy of the prevention and control measure making; according to the influence area data, the department terminal list is determined and the treatment suggestion is automatically transmitted, which can ensure that the relevant departments timely receive the corresponding prevention and control instructions, avoid information transmission lag, and solve the problem of untimely early warning response; the early warning content is generated for different user groups and published through the preset channel, which can enable the public to accurately obtain early warning information meeting their own needs, improve the public's awareness of the invasion risk, and help to form a situation of government-enterprise coordinated prevention and control; real-time monitoring of the delivery state of the treatment suggestion and the publishing state of the early warning content and updating the log can record the whole process of early warning response, realize the closed-loop management of early warning from generation to response, and effectively improve the deficiency that the existing technology cannot form a complete prevention and control closed loop.
[0061] As shown in Figure 2 The embodiment of the present application also provides an alien invasive species early warning system based on big data analysis, which comprises: The acquisition module is configured to acquire environmental parameter data, ship navigation data and biological sample data through multi-dimensional data acquisition. The processing module is configured to perform data integration processing based on the multi-dimensional data, and form a standardized data set by time alignment and format unification of the environmental parameter data, the ship navigation data and the biological sample data. The identification module is configured to identify alien invasive species based on the standardized data set, that is, to obtain a matching species identification result by comparing a preset global alien invasive species database. The evaluation module is configured to perform risk evaluation according to the matching species identification result, in combination with the environmental parameter data and the ship navigation data, to obtain a risk level evaluation result by analyzing species survival condition adaptability and import probability characteristics and calculating a colonization risk level based on a multi-feature decision tree integration mechanism. The simulation module is configured to perform diffusion prediction analysis on the risk level evaluation result, to simulate a species diffusion path according to meteorological data and through a fluid dynamics model; in the simulation process, a diffusion component in a vertical direction is determined by calculating a cross product of a water flow direction vector and a terrain gradient vector to correct a diffusion trajectory of the species, and finally to form early warning information. The execution module is configured to form disposal suggestions and early warning content by performing early warning response and information release based on the early warning information; the disposal suggestions are delivered to a terminal of a relevant department, and the early warning content is released to a public platform; and a delivery state and a release state of the early warning information are recorded in real time to update an early warning response log.
[0062] The above is a preferred embodiment of the present application. It should be noted that those skilled in the art can make some improvements and refinements without departing from the principles of the present application, and these improvements and refinements should also be considered within the scope of protection of the present application.
Claims
1. A method for early warning of invasive alien species based on big data analysis, characterized in that, The method includes: Multi-dimensional data collection was conducted to obtain environmental parameter data, shipping data, and biological sample data. Based on multi-dimensional data, data integration and processing are carried out. By aligning environmental parameter data, shipping data, and biological sample data in time and unifying their formats, a standardized dataset is formed. Based on standardized datasets, invasive alien species are identified, that is, by comparing with a pre-set global database of invasive alien species, the matching species identification results are obtained. Based on the matching species identification results, combined with environmental parameter data and shipping data, a risk assessment is conducted. By analyzing the species' adaptability to living conditions and the probability of introduction, and based on a multi-feature decision tree ensemble mechanism, the colonization risk level is calculated, and the risk level assessment result is obtained. The risk level assessment results are used to predict and analyze the spread of species. Based on meteorological data, the species spread path is simulated using a fluid dynamics model. During the simulation, the vertical spread component is determined by calculating the cross product of the water flow direction vector and the topographic gradient vector to correct the species spread trajectory and ultimately generate early warning information. Based on the early warning information, the system generates handling suggestions and early warning content by executing early warning responses and information dissemination; the handling suggestions are delivered to relevant department terminals, while the early warning content is published to the public platform; the delivery status and publication status of early warning information are recorded in real time to update the early warning response log.
2. The method for early warning of invasive alien species based on big data analysis according to claim 1, characterized in that, Multi-dimensional data collection was conducted to obtain environmental parameter data, shipping data, and biological sample data, including: Collect environmental parameter data of the port waters, including water temperature, salinity, pH value and water flow velocity; Acquire shipping data related to the vessel, including ballast water discharge records, route trajectory, and cargo type; Collect biological sample data, convert the biological sample data into standardized image data and gene sequence data, and obtain the converted biological sample data; Environmental parameter data, shipping data, and converted biological sample data are integrated to form a raw data set to be processed.
3. The method for early warning of invasive alien species based on big data analysis according to claim 2, characterized in that, Based on multi-dimensional data, data integration and processing are performed. By aligning environmental parameter data, shipping data, and biological sample data over time and standardizing their formats, a standardized dataset is formed, including: The original dataset is timestamped by unifying environmental parameter data, shipping data, and biological sample data to the same time base to obtain time-aligned data. The time-aligned data is format-standardized by converting data from different sources into a unified numerical format and data structure to obtain format-standardized environmental parameter data, shipping data, and biological sample data. The standardized environmental parameter data, shipping data, and biological sample data are fused together to form a standardized dataset with a consistent spatiotemporal benchmark.
4. The method for early warning of invasive alien species based on big data analysis according to claim 3, characterized in that, Based on standardized datasets, invasive alien species identification is performed, which involves comparing the results with a pre-defined global database of invasive alien species to obtain matching species identification results, including: Biological sample data is extracted from standardized datasets, including image data and gene sequence data. The image data is compared with the morphological feature data in the preset global invasive alien species database to obtain the morphological feature comparison results; at the same time, the gene sequence data is compared with the gene sequence data in the preset global invasive alien species database to obtain the gene sequence analysis results. Based on the results of morphological feature comparison and gene sequence analysis, a weighted fusion calculation is performed to determine the matching species and their confidence level. The matching species are filtered according to the preset confidence threshold to obtain the matching species identification results, which include the species name and matching confidence information.
5. The method for early warning of invasive alien species based on big data analysis according to claim 4, characterized in that, Based on the matched species identification results, combined with environmental parameter data and shipping data, a risk assessment is conducted. This involves analyzing the species' adaptability to living conditions and the probability of introduction, and calculating the colonization risk level based on a multi-feature decision tree ensemble mechanism. The resulting risk level assessment includes: Extract species names from the matched species identification results, and obtain corresponding species survival condition parameters from a pre-set global database of invasive alien species based on the species names; The environmental parameter data in the standardized dataset are subjected to a fit analysis with the species survival condition parameters to obtain the fit analysis results. Based on shipping route trajectories in shipping data, the minimum distance between the ship's position and the boundary of a pre-defined high-incidence sea area for invasive alien species is calculated to quantify spatial proximity; based on spatial proximity, the introduction probability analysis is performed by combining ballast water discharge records to obtain the introduction probability analysis results. Based on the results of adaptability analysis and ingress probability analysis, combined with pre-organized historical intrusion case data, multiple basic risk assessment results are obtained by using multiple decision trees to process multiple feature dimensions in parallel. Multiple basic risk assessment results are integrated and processed, and the final colonization risk level is determined by a voting mechanism. Based on the final colonization risk level, a risk level assessment result including low, medium and high levels is obtained.
6. The method for early warning of invasive alien species based on big data analysis according to claim 5, characterized in that, The risk level assessment results are used to conduct diffusion prediction analysis, and species diffusion paths are simulated based on meteorological data and fluid dynamics models. During the simulation, the vertical diffusion component is determined by calculating the cross product of the water flow direction vector and the terrain gradient vector to correct the species diffusion trajectory, ultimately generating early warning information, including: Extract information on species marked as high-risk in the risk level assessment results and their location data; Based on species information and location data at high risk levels, combined with real-time meteorological and hydrological data, the diffusion process of species in water bodies is simulated using a fluid dynamics model. In the process of fluid dynamics model simulation, the vertical diffusion component is determined by calculating the cross product of the water flow direction vector and the terrain gradient vector. The horizontal diffusion trajectory is then corrected based on the diffusion component to obtain the corrected diffusion trajectory. The corrected diffusion trajectory is spatially overlaid with the pre-defined sensitive area boundary to identify potentially affected agricultural and aquaculture areas and ecological protection zones, and the spatial overlay analysis results are obtained. Based on the spatial overlay analysis results, early warning information is generated that includes a diffusion path map, affected areas, and risk levels.
7. The method for early warning of invasive alien species based on big data analysis according to claim 6, characterized in that, Based on the early warning information, the system executes early warning responses and disseminates information to generate handling recommendations and early warning content. The handling recommendations are then delivered to relevant departmental terminals, while the early warning content is published on public platforms. The transmission and dissemination status of early warning information are recorded in real time to update the early warning response log, including: The early warning information is analyzed and processed to obtain data on species characteristics, risk level, affected area, and spread path. Based on the risk level, the system automatically matches the corresponding response measures from a pre-set response plan library to generate response suggestions that include specific operational instructions. Based on the data of the affected area, a list of relevant department terminals that need to be pushed with early warnings is determined, and handling suggestions are automatically transmitted to the corresponding terminals through the data interface based on the terminal list; Based on species characteristics, risk levels, and affected area data, early warning content is generated for different user groups, and the early warning content is released and pushed to the public platform through preset release channels. Real-time monitoring of the delivery status of handling suggestions and the release status of early warning content is used to update the early warning response log and complete the closed-loop management of early warnings.
8. A big data analysis-based early warning system for invasive alien species, wherein the system implements the method as described in any one of claims 1 to 7, characterized in that, include: The acquisition module is used to collect multi-dimensional data to obtain environmental parameter data, shipping data, and biological sample data. The processing module is used to perform data integration processing based on multi-dimensional data. It forms a standardized dataset by aligning environmental parameter data, shipping data, and biological sample data in time and unifying their formats. The identification module is used to identify invasive alien species based on a standardized dataset, that is, to obtain the matching species identification results by comparing with a pre-set global database of invasive alien species; The assessment module is used to conduct risk assessment based on the matching species identification results, combined with environmental parameter data and shipping data. It analyzes the species' adaptability to living conditions and the probability of introduction, and calculates the colonization risk level based on a multi-feature decision tree ensemble mechanism to obtain the risk level assessment result. The simulation module is used to perform diffusion prediction analysis based on the risk level assessment results, and simulates the species diffusion path based on meteorological data and a fluid dynamics model. During the simulation, the vertical diffusion component is determined by calculating the cross product of the water flow direction vector and the terrain gradient vector, so as to correct the species diffusion trajectory and ultimately form early warning information. The execution module is used to generate handling suggestions and warning content based on the warning information by executing the warning response and information dissemination. The recommendations for handling the situation will be sent to the relevant departments' terminals, and the warning information will be published on the public platform. The delivery status and publication status of the warning information will be recorded in real time to update the warning response log.
9. A computing device, characterized in that, include: One or more processors; A storage device for storing one or more programs, which, when executed by one or more processors, cause the one or more processors to implement the method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program that, when executed by a processor, implements the method as described in any one of claims 1 to 7.