Double-flow neural network based on composite attention mechanism and immersion prediction method and device
Patent Information
- Application Number
- CN202511009767.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-22
- Publication Date
- 2025-11-14
Smart Images

Figure CN120953869A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of neural network frame sequence understanding technology, specifically to a two-stream neural network based on a composite attention mechanism and a method and device for predicting ship hull damage and flooding. Background Technology
[0002] As the global economy develops, the demand for maritime trade continues to rise, thanks to the low cost, large cargo capacity, and long transport distances of both land and sea transportation. While development brings convenience, it also necessitates consideration of safety hazards. Accidents not only cause casualties but also result in substantial economic losses and secondary disasters such as environmental pollution. Research into the immersion process can better aid in emergency decision-making after a shipwreck and in improving ship structures.
[0003] Research on flooding processes continues to evolve to better aid emergency decision-making and structural improvements after shipwrecks. A mainstream approach is to statically simulate water flow within a ship using Computational Fluid Dynamics (CFD). This method, by incorporating more detail during computation, can simulate fluid motion with higher accuracy. However, CFD calculations are typically complex and time-consuming, making them unsuitable as a decision-making system. In recent years, with the rapid development of artificial intelligence, deep learning techniques have also been applied to flooding research. Existing studies use a hybrid CNN and LSTM network, employing a series of flooding images at fixed time intervals as a dataset to predict the flooding time of compartments. The recurrent neural network recursively memorizes and processes the features of each frame, ultimately forming the overall features of the image sequence.
[0004] The datasets used in deep learning methods come from ship model experiments, an experimental method that simulates and studies the navigation performance and behavior of actual ships in water using scaled-down ship models. Physical networks are also used to validate the simulation results of the computer software. This method can approximate the results of real-world ship experiments as closely as possible through data transformation, addressing the problem of limited data on real-world ship immersion events.
[0005] Existing CNN-LSTM networks can predict the immersion time of ship model flooding datasets with high accuracy, but there are still some issues with the adaptation of the network to the dataset. First, some images in the flooding dataset are quite complex, potentially containing elements such as the ship hull, cabin doors, and the spreading water surface within the same image. CNNs have poor feature capture capabilities for such complex images. Even disregarding the unique characteristics of ship model experiments, flooding images taken in actual ship cabins may increase image complexity due to ship tilting and collapsed items. Therefore, adapting to complex images remains a problem to be solved. Second, due to the backpropagation characteristics of recurrent neural networks, features from earlier stages are gradually lost in long sequences, with greater focus on later features. This leads to a bias in the modeling of the overall flooding features. Combined with the relatively small changes in water surface features in later stages, this results in a poor grasp of the flooding process characteristics. In the early stages of flooding, the water surface does not cover the entire bottom of the cabin, and the process of water spreading can be clearly seen, with significant differences between images.
[0006] In summary, existing technologies suffer from shortcomings such as insufficient ability to extract details from complex immersion images, easy imbalance in temporal feature modeling, and lack of effective modeling of optical flow motion information. Summary of the Invention
[0007] To address the shortcomings of existing technologies, such as insufficient ability to extract details from complex immersion images, easy imbalance in temporal feature modeling, and lack of effective modeling of optical flow motion information, the technical solution provided by this invention is as follows: A two-stream neural network based on a composite attention mechanism and an inundation prediction method, comprising: The steps include: acquiring videos of the breached immersion experiment and immersion time information for each compartment, constructing a frame sequence image by extracting frames of the video at equal intervals, performing fuzzy classification on the immersion time, and generating training data. The step of preprocessing the frame sequence images and converting them into tensors in a format acceptable to neural networks; The steps of inputting the preprocessed image tensor into the two-stream neural network model are as follows: the first branch extracts the spatial structure and temporal dynamic features of the image sequence, and the second branch extracts inter-frame optical flow information and motion features. The steps are to uniformly represent the features of the first and second branches and input them into the multi-compartment output head to output the immersion time category prediction results for each compartment; The output of each compartment is optimized using a composite cross-entropy loss function with learnable weights to complete the training of the neural network. The steps are to convert the predicted flooding time categories into corresponding actual time values and output the flooding time prediction results for each compartment.
[0008] Furthermore, a preferred implementation is provided in which frame sequence image preprocessing includes scaling, normalization, and tensor transformation using the torchvision library.
[0009] Furthermore, a preferred implementation is provided in which the first branch uses a Swing Transformer structure for image block segmentation and multi-layer displacement window attention calculation.
[0010] Furthermore, a preferred implementation is provided in which the second branch extracts multi-resolution features through a cross-scale image embedding module and extracts optical flow information and image displacement features by combining an inter-frame attention mechanism.
[0011] Furthermore, a preferred implementation method is provided, in which the output features of the two branches are fused into a unified representation vector by concatenating the channel dimensions and performing one-dimensional convolution dimensionality reduction.
[0012] A two-stream neural network based on a composite attention mechanism and an inundation prediction device are also provided, comprising: This module acquires videos of the breached immersion experiment and immersion time information for each compartment, constructs a frame sequence image by extracting frames from the video at equal intervals, performs fuzzy classification on the immersion time, and generates training data. The module preprocesses the frame sequence images and converts them into tensors in a format acceptable to neural networks; The preprocessed image tensor is input into the module of the two-stream neural network model. The first branch extracts the spatial structure and temporal dynamic features of the image sequence, and the second branch extracts the inter-frame optical flow information and motion features. The module unifies the features of the first and second branches and inputs them into the multi-compartment output head to output the immersion time category prediction results for each compartment. The module that performs gradient optimization on the output of each compartment using a composite cross-entropy loss function with learnable weights to complete the training of the neural network. This module converts the predicted flooding time categories into corresponding actual time values and outputs the predicted flooding time results for each compartment.
[0013] A two-stream neural network and flooding prediction system based on a composite attention mechanism are also provided to implement the method, including: The acquisition unit is used to acquire image sequences and immersion time data of each compartment during the breaching and immersion process; The preprocessing unit is used to scale and normalize the image sequence and classify and label the immersion time data; The dual-flow neural network unit includes a spatiotemporal feature extraction branch and an optical flow feature extraction branch, which respectively extract the spatiotemporal dynamic information and inter-frame motion information of the image; The fusion unit is used to concatenate and reduce the dimensionality of the features output from the two branches to form a unified representation; The output unit is used to output the immersion time category results of each compartment based on the fusion characteristics and convert them into the actual time range.
[0014] A computer storage medium is also provided for storing a computer program, which, when read by the computer, executes the method.
[0015] A computer is also provided, including a processor and a storage medium, wherein the computer executes the method when the processor reads a computer program stored in the storage medium.
[0016] A computer program product is also provided, which, when executed, implements the method described.
[0017] Compared with the prior art, the advantages of the technical solution provided by the present invention are as follows: This approach constructs a composite attention mechanism combining temporal and spatial attention. During the spatiotemporal information extraction process, it effectively integrates the spatial layout and inter-frame dynamic features of each frame in the image sequence, improving the overall modeling capability for water spread and hull attitude changes in complex immersion scenarios. Compared to traditional methods using only CNN or LSTM, this approach significantly enhances the image sequence's ability to perceive details of the initial immersion process, effectively preventing the dilution of early image information during long-term propagation.
[0018] This solution introduces a dual-stream neural network structure, where the main branch extracts spatiotemporal attention features, and the auxiliary branch extracts motion optical flow features based on inter-frame differences. Feature integration is achieved through an end-effector fusion module, effectively enhancing the ability to recognize implicit motion trends in images. Compared to a single-path network design, this structure is more sensitive in recognizing visual features closely related to immersion speed, such as hull rolling and water surface ripples, thereby improving the accuracy and stability of predicting the immersion time of multiple compartments.
[0019] This approach employs a multi-scale feature extraction and fusion structure in the optical flow branch. It extracts image motion features at different resolutions through a cross-scale image embedding module and combines this with an inter-frame attention mechanism to form a unified optical flow representation, thus improving the model's responsiveness to dynamic changes at different scales. Compared to single-scale optical flow estimation methods in existing research, this approach has a stronger ability to capture fine-grained motion features (such as water flow edge variations), significantly improving the resolution and accuracy of immersion process modeling.
[0020] This approach introduces a composite loss function with adaptive weights during network training, dynamically adjusting and optimizing the weights for the prediction results of multiple compartments, thereby solving the training bias problem that easily occurs in multi-output tasks. Compared with the traditional method of using static loss weighting in multi-task learning, this method is more robust, effectively ensuring the consistency of prediction performance across compartments and improving the generalization ability of the overall model.
[0021] This approach employs fuzzy labeling and preprocessing of simulated experimental data, combined with fixed-frame-count sampling and image tensor conversion, to enable the raw video data to quickly adapt to the input format of the neural network model. This ensures the balance of the network input data while controlling computational resource consumption. Compared to some studies that directly use complete videos for training, this approach significantly shortens the model inference time without sacrificing prediction accuracy, meeting the real-time requirements of ship emergency decision-making.
[0022] It is applicable to the rapid prediction and decision support of multi-compartment flooding time in emergency response to ship hull damage accidents. Attached Figure Description
[0023] Figure 1 This is a schematic diagram of the internal structure of a model for predicting ship hull damage and flooding time based on a composite attention mechanism.
[0024] Figure 2 This is an example of a flooding image collected in a ship hull breach flooding time prediction method based on a composite attention mechanism.
[0025] Figure 3 This is a schematic diagram of the dual-flow network structure in a ship hull breach flooding time prediction method based on a composite attention mechanism.
[0026] Figure 4 This is a schematic diagram of the spatiotemporal flow and motion optical flow in the dual-flow network structure of the ship hull breach flooding time prediction method based on the composite attention mechanism.
[0027] Figure 5 This is a graph showing the change in the average accuracy of a two-stream network.
[0028] Figure 6 This is a graph showing the change in loss during the training of a two-stream network.
[0029] Figure 7 This represents the error between the output of the dual-stream network and the actual immersion time when converted to the actual immersion time of the ship. Detailed Implementation
[0030] To make the advantages and benefits of the technical solution provided by the present invention clearer, the technical solution provided by the present invention will now be described in further detail with reference to the accompanying drawings, specifically: Implementation Method 1: This implementation method provides a two-stream neural network based on a composite attention mechanism and a flooding prediction method, including: The steps include: acquiring videos of the breached immersion experiment and immersion time information for each compartment, constructing a frame sequence image by extracting frames of the video at equal intervals, performing fuzzy classification on the immersion time, and generating training data. The step of preprocessing the frame sequence images and converting them into tensors in a format acceptable to neural networks; The steps of inputting the preprocessed image tensor into the two-stream neural network model are as follows: the first branch extracts the spatial structure and temporal dynamic features of the image sequence, and the second branch extracts inter-frame optical flow information and motion features. The steps are to uniformly represent the features of the first and second branches and input them into the multi-compartment output head to output the immersion time category prediction results for each compartment; The output of each compartment is optimized using a composite cross-entropy loss function with learnable weights to complete the training of the neural network. The steps are to convert the predicted flooding time categories into corresponding actual time values and output the flooding time prediction results for each compartment.
[0031] Frame sequence image preprocessing includes scaling, normalization, and tensor transformation using the torchvision library.
[0032] The first branch uses the Swing Transformer structure for image block partitioning and multi-layer displacement window attention calculation.
[0033] The second branch extracts multi-resolution features through a cross-scale image embedding module and combines it with an inter-frame attention mechanism to extract optical flow information and image displacement features.
[0034] The unified representation method uses channel dimension concatenation and one-dimensional convolution dimensionality reduction to fuse the output features of the two branches into a unified representation vector.
[0035] A two-stream neural network and flooding prediction system based on a composite attention mechanism are also provided to implement the method, including: The acquisition unit is used to acquire image sequences and immersion time data of each compartment during the breaching and immersion process; The preprocessing unit is used to scale and normalize the image sequence and classify and label the immersion time data; The dual-flow neural network unit includes a spatiotemporal feature extraction branch and an optical flow feature extraction branch, which respectively extract the spatiotemporal dynamic information and inter-frame motion information of the image; The fusion unit is used to concatenate and reduce the dimensionality of the features output from the two branches to form a unified representation; The output unit is used to output the immersion time category results of each compartment based on the fusion characteristics and convert them into the actual time range.
[0036] Implementation Method Two: This implementation method is a further detailed description of the technical solution provided in Implementation Method One, specifically: A method for predicting the flooding time of ship compartments based on a dual-stream neural network with a composite attention mechanism is presented. The overall process includes three main stages: data acquisition and preprocessing, construction and training of the dual-stream neural network, and output and transformation of prediction results. The first step involved conducting a immersion experiment on a model ship with breaches and collecting images and tagging data. The experiment used a 1:200 scale model of a real ship, with multiple breach locations set. In each immersion test, one breach was activated, while the others were sealed. Cameras recorded the entire water spread process. An immersion alarm was installed inside each compartment, automatically triggering when the water level reached the height of the hatch and recording the corresponding time, thus creating immersion time-tag data for the compartment.
[0037] The second step involves structuring the collected video and tag data. The 30-second video is uniformly extracted into a 30-frame image sequence, and tools such as TorchVision are used to perform scaling, standardization, and other processing operations on the images, converting them into a tensor format suitable for input into the neural network. Simultaneously, the immersion time tags are fuzzily classified into multiple time period categories based on a set interval to adapt to the multi-classification task of the neural network.
[0038] The third step is to construct and train a two-stream neural network model. The network consists of two information processing branches: a spatiotemporal attention extraction stream and a motion optical flow extraction stream. The former serves as the main branch, using a convolutional-attention mechanism to extract spatial structure and temporal dynamic features from the image sequence; the latter serves as an auxiliary branch, using inter-frame attention and cross-scale fusion modules to extract optical flow information and ship sway features between image frames.
[0039] The fourth step involves setting up a feature fusion module and a multi-output prediction head at the end of the neural network. Features from the two branches are concatenated and dimensionality reduced along the channel dimension in the fusion module to form a unified representation, which is then fed into a multi-head fully connected layer. Each output head corresponds to the prediction of the immersion time category of a compartment, adapting to the needs of multi-compartment immersion judgment.
[0040] The fifth step involves designing a composite loss function and training the network. Cross-entropy is used as the basic loss function, and learnable weight coefficients are introduced for each output head to adaptively adjust the gradient update contribution of different compartment outputs during training, effectively avoiding the problem of some outputs dominating the overall training. During training, a cosine annealing learning rate strategy and the AdamW optimizer are implemented to enhance convergence stability.
[0041] The sixth step involves network inference and outputting prediction results. The neural network receives image sequence input and outputs the immersion time category for each compartment. Combined with a pre-defined category-time mapping table, this is further converted into a minute-level estimate of the actual immersion time. If the category prediction is biased, the median value of adjacent categories is used as a fault-tolerant correction to improve robustness in actual use.
[0042] The seventh step involves using the output results for emergency decision-making regarding flooding of ship compartments. The predicted flooding time for each compartment provides data support for strategies such as personnel evacuation and compartment sealing, thereby improving the efficiency of ship emergency response and safety assurance capabilities.
[0043] Implementation Method 3: Combination Figure 1-7 This embodiment describes the technical solution provided above in further detail through specific examples. Specifically: Part 1: This embodiment provides a data collection and processing method for training ship hull breach and flooding time, including: S1. Conduct a breach immersion experiment using a network that simulates a real ship, and record the water surface spread in the breached compartment and the immersion time of all compartments during immersion. S2. Process the data collected in S1. For video data, capture the frame at certain time intervals and construct a frame sequence. For immersion time data, blur it. S3. Use S1 to S2 to predict the flooding time of damaged hulls, and complete the preparation of ship flooding time prediction data based on the composite attention mechanism.
[0044] Furthermore, the ship model used in S1 includes a step of setting the dimensional elements of the compartments, which include length, width, height, and bulkhead thickness. The breach dimensional elements include breach size and breach distribution.
[0045] In the preparation of the ship model in S1, the location of the breach is divided into 23 categories.
[0046] Furthermore, S2 processes the video data and immersion time data as follows: S2.1. Set the frame sampling time interval according to the duration of the immersion video data, ensuring that the frame sequence length is 30 frames, and that the time interval for each frame is equal. For a 30-second video duration, the sampling interval is 1 frame per second. S2.2. The obtained immersion time data is fuzzy processed, and the range of category classification is determined based on the distribution of immersion time obtained from the experiment. S2.3 After completing S2.1 to S2.2, use the transforms function in the torchvision library to further process the data, including scaling, normalization, etc.
[0047] Part Two: Constructing a two-stream neural network for predicting ship compartment flooding time based on a composite attention mechanism, including: S4. Pre-loaded module: ImageNet-1K and Vimeo90K datasets are used as pre-training datasets for the spatiotemporal attention module; S5, Spatiotemporal Attention Extraction Module, the first branch of the dual-stream network. It processes the immersion process data input to the neural network, extracting the temporal continuity information and spatial information of the immersion chamber contained within the immersion data. S6, Motion Optical Flow Information Extraction Module, the second branch of the dual-flow network. It processes the data input to the neural network, extracting the water surface change features and ship rolling features needed for flood prediction from the image. S7, End-point Attention Fusion Module, fuses the attention information obtained from the first and second branches at the end of the network, and the final information state is used to predict the flooding time of the ship's compartments. S8, Multi-task training module, used to train a neural network with multiple outputs using a composite loss function with adaptive weights; S9, the prediction module, based on S4-S7, completes the prediction of the ship's flooding time based on the attention mechanism. The time prediction is performed using a linear neural network.
[0048] Specifically: See Figures 1 to 7 Step 1: Ship model preparation: In this implementation method, a model that imitates the structure of a real ship is first selected. In order to simulate the movement of water in the hull after the compartment is breached, small holes of the same size are made at different positions on the ship wall. The height of the holes is roughly divided into three categories. Several small holes are evenly made on the compartment wall at each height. The size of the holes is restored to the actual scale to correspond to a breach with a radius of 1.2 meters. Figure 1 This diagram illustrates the interior of the network's cabins. The model ship contains several cabins of varying capacities, separated by walls. In the event of flooding, after a breach, water will first submerge the breached cabin. Once the water level reaches a certain height, water will seep through the cabin doors, spreading to all other cabins until the entire ship is flooded. To detect whether a cabin is flooded, a flood alarm system is required in the network. Table 1 shows the external dimensions of the model ship.
[0049]
[0050] Step 2: Immersion Test of the Ship Model A water immersion experiment was conducted based on the ship model from step one to obtain visual information on the spread of water under different breach conditions and water immersion time data for different compartments. Figure 2 Here is an example of a captured image of the submersion.
[0051] Each time the compartment was flooded, one small hole in one of the chambers was opened, while the others were sealed. Cameras positioned above the network recorded changes in the breached chambers to observe water movement; the collected flooding video of the breached chambers was then input into the network. Each chamber was equipped with a flood alarm, which was automatically triggered when the floodwaters reached a designated level. The time from the start of flooding to alarm triggering was recorded. The alarm trigger height was set to the same height as the chamber door, as flooding of the door would hinder escape, and the evacuation target time should be before the door is submerged. Water began to spread from the breached chambers; when all chamber doors were submerged, the cameras stopped recording the time. Flooding times for all four chambers were collected for each flood. Ship model tests can simulate the immersion scenario in a real ship to a certain extent. The time data obtained from the immersion test in this paper can be converted into the immersion time in the actual ship navigation scenario. Assuming that the location of the breach in the real ship is the same as the location of the breach in the network, according to Bernoulli's principle, the formula for the water flow velocity v at the breach can be obtained.
[0052]
[0053] in and These represent the heights from the sea surface and the breach to the seabed, respectively. The open flow coefficient is a constant. Let g be the acceleration due to gravity. By calculating the water flow velocity at the breach in both the network and the actual ship, the conversion coefficient between the network immersion time and the actual ship immersion time can be obtained, thus converting the experimentally simulated immersion time into the actual ship immersion time.
[0054] Step 3: Data Processing The data obtained from the immersion experiment were further processed. A total of 816 sets of experimental data were obtained. Different lighting conditions were set in the experiment to improve the generalization ability of the network. Considering that the prediction accuracy of the network was not yet sufficient to support regression prediction, the label data was further blurred during data processing. Statistically, all immersion times did not exceed 130 seconds, and the immersion times were divided into categories every 10 seconds. The network would then predict the category of the immersion time. In order to be input into the network, the immersion videos were also preprocessed into image sequences. The duration of the immersion videos was 30 seconds. The software extracted frames from the videos at the same intervals, resulting in a 30-frame immersion image sequence. Longer sequences allow the network to perceive more information, but inference will take longer and require expensive equipment to run the network. Considering that auxiliary decision-making usually requires a fast reaction time, the preprocessing limited the number of frames to an acceptable range. The training set and test set were split in an 8:2 ratio. In the training set and test set, the number of times each compartment was considered a breach compartment was kept as even as possible.
[0055] Step 4: Neural Network Training This step covers the construction and training of the basic neural network, dynamic training, and updating the dataset.
[0056] 4.1 Building the Basic Neural Network This implementation method first requires building a basic spatiotemporal attention neural network. Spatiotemporal attention is a mechanism for processing data with spatial and temporal dimensions. It combines the spatial feature distribution of the data of interest with the dynamic changes of the data of interest in the temporal dimension to improve the network's ability to model complex spatiotemporal relationships.
[0057] 4.2 Dual-stream network setup This implementation requires combining two different feature extraction structures into a complete dual-stream network to extract two different features, thereby increasing feature richness and enhancing the network's generalization ability. The network consists of a spatiotemporal attention information flow and a motion optical flow. The spatiotemporal flow constructed in section 4.1 serves as the main branch, using a spatiotemporal attention algorithm to extract the pattern of appearance changes over time. The motion optical flow serves as an auxiliary branch, using an inter-frame attention algorithm to extract the motion relationship between every two frames, and the optical flow information in the data is fused. The two features are fused through a feature fusion module and finally fed into a classification module to be converted into the probability of immersion time.
[0058] 4.3 After the multi-compartment training neural network is built, the network is trained using the prepared flooding dataset. Since the task to be predicted is the flooding time of multiple compartments, the network has multiple outputs in one prediction. During training, a composite loss function is needed to calculate the gradient. However, by adding adaptive weights to dynamically update the weights of each output, the training process can be optimized and over-optimization of a single output can be prevented.
[0059] 4.4 Conversion of Prediction Results The network's inference results are categorical data, which need to be converted into actual immersion time through numerical transformation.
[0060] The above steps specifically include the following: 4.1 Basic Neural Network Setup The first step is to build a basic spatiotemporal attention network, primarily consisting of two modules: a spatial feature extraction structure and a temporal feature extraction structure. These two structures are linearly connected. The overall network approach involves the spatial feature extraction structure converting pixel information from images into features that the network can understand. These features include information such as the angle of the hull and the area of the water surface. By combining this information from the image sequence, the amplitude of the rocking and the speed of water spread can be determined. The immersion time after the hull breach is related to these physical phenomena. The temporal feature extraction structure is responsible for modeling the features of the image sequence, and the results are then converted into probabilities for each category through a classification output module.
[0061] The entire network processing flow is as follows: First, the image is mapped into small image patches by convolutional operations. These image patches undergo multiple layers of shifted window attention to obtain attention between image patches. This attention is used to acquire the overall features of the image. After obtaining the feature information of all images in the sequence, these features are integrated into a global image feature vector. Next, the temporal feature extraction structure performs attention operations again, with the object being the obtained global image feature vector, and calculates the correlation between image features to model relationships. The linear layer contains several parallel and parameter-independent linear heads. The immersion time of each compartment is obtained by a separate linear head. The feature extraction result vector is projected into the immersion time category probabilities of multiple compartments. Table 2 shows the training parameters of the spatiotemporal attention.
[0062] The above description of several specific embodiments further details the technical solution provided by the present invention in order to highlight the advantages and benefits of the technical solution provided by the present invention. However, the above-described specific embodiments are not intended to limit the present invention. Any reasonable modifications and improvements to the present invention, combinations of embodiments, and equivalent substitutions based on the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A two-stream neural network based on a composite attention mechanism and a flooding prediction method, characterized in that, include: The steps include: acquiring videos of the breached immersion experiment and immersion time information for each compartment, constructing a frame sequence image by extracting frames of the video at equal intervals, performing fuzzy classification on the immersion time, and generating training data. The step of preprocessing the frame sequence images and converting them into tensors in a format acceptable to neural networks; The steps of inputting the preprocessed image tensor into the two-stream neural network model are as follows: the first branch extracts the spatial structure and temporal dynamic features of the image sequence, and the second branch extracts inter-frame optical flow information and motion features. The steps are to uniformly represent the features of the first and second branches and input them into the multi-compartment output head to output the immersion time category prediction results for each compartment; The output of each compartment is optimized using a composite cross-entropy loss function with learnable weights to complete the training of the neural network. The steps are to convert the predicted flooding time categories into corresponding actual time values and output the flooding time prediction results for each compartment.
2. The two-stream neural network based on a composite attention mechanism and the flooding prediction method according to claim 1, characterized in that, Frame sequence image preprocessing includes scaling, normalization, and tensor transformation using the torchvision library.
3. The two-stream neural network based on a composite attention mechanism and the flooding prediction method according to claim 1, characterized in that, The first branch uses the Swing Transformer structure for image block partitioning and multi-layer displacement window attention calculation.
4. The two-stream neural network based on a composite attention mechanism and the flooding prediction method according to claim 1, characterized in that, The second branch extracts multi-resolution features through a cross-scale image embedding module and combines it with an inter-frame attention mechanism to extract optical flow information and image displacement features.
5. The two-stream neural network based on a composite attention mechanism and the flooding prediction method according to claim 1, characterized in that, The unified representation method uses channel dimension concatenation and one-dimensional convolution dimensionality reduction to fuse the output features of the two branches into a unified representation vector.
6. A two-stream neural network based on a composite attention mechanism and an immersion prediction device, characterized in that, include: This module acquires videos of the breached immersion experiment and immersion time information for each compartment, constructs a frame sequence image by extracting frames from the video at equal intervals, performs fuzzy classification on the immersion time, and generates training data. The module preprocesses the frame sequence images and converts them into tensors in a format acceptable to neural networks; The preprocessed image tensor is input into the module of the two-stream neural network model. The first branch extracts the spatial structure and temporal dynamic features of the image sequence, and the second branch extracts the inter-frame optical flow information and motion features. The module unifies the features of the first and second branches and inputs them into the multi-compartment output head to output the immersion time category prediction results for each compartment. The module that performs gradient optimization on the output of each compartment using a composite cross-entropy loss function with learnable weights to complete the training of the neural network. This module converts the predicted flooding time categories into corresponding actual time values and outputs the predicted flooding time results for each compartment.
7. A two-stream neural network and flooding prediction system based on a composite attention mechanism, characterized in that, To implement the method of claim 1, the method comprises: The acquisition unit is used to acquire image sequences and immersion time data of each compartment during the breaching and immersion process; The preprocessing unit is used to scale and normalize the image sequence and classify and label the immersion time data; The dual-flow neural network unit includes a spatiotemporal feature extraction branch and an optical flow feature extraction branch, which respectively extract the spatiotemporal dynamic information and inter-frame motion information of the image; The fusion unit is used to concatenate and reduce the dimensionality of the features output from the two branches to form a unified representation; The output unit is used to output the immersion time category results of each compartment based on the fusion characteristics and convert them into the actual time range.
8. A computer storage medium for storing computer programs, characterized in that, When the computer program is read by the computer, the computer executes the method of claim 1.
9. A computer, comprising a processor and a storage medium, characterized in that, When the processor reads the computer program stored in the storage medium, the computer executes the method of claim 1.
10. A computer program product, as a computer program, is characterized by: When the computer program is executed, it implements the method of claim 1.