Multidimensional information parallel targeting extraction method based on multiphase flow droplet / bubble spatial characteristics
Through multi-dimensional backbone lightweight neural network and related technical means, problems such as large amount of calculation and time-consuming training in complex bubble/droplet image recognition are solved, and efficient feature extraction and recognition effects are achieved.
Patent Information
- Application Number
- CN202510039717.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-10
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-01-10
AI Technical Summary
In the process of complex bubble/drop image recognition, there is a problem of large calculations, time-consuming training, disappearance or explosion of gradients, difficulty in extracting feature details, multiple deformations and inconsistent scales.
A multi-dimensional backbone lightweight neural network is used to extract image features layered, and capture local texture information from shallow to deep aggregate global semantic features. Combining deep-separable convolution, pruning and quantization techniques to reduce computational costs. The residual structure ResNet and dynamic feature fusion mechanism are introduced to alleviate the problem of gradient vanishing.
The resolution capability of the model is improved, the balance between fine-grained feature extraction and computational efficiency in complex scenarios is achieved, and an efficient solution is provided for complex image recognition.
Smart Images

Figure CN120126115A_ABST
Abstract
Description
Technical Field The present invention relates to a method for extracting multi-dimensional information of multiphase flow liquid, and particularly to a method for parallel targeted extraction of multi-dimensional information based on the spatial characteristics of multiphase flow droplets / bubbles. Background Art Gas-liquid / liquid-liquid two-phase flows widely exist in industrial applications such as petrochemical industry, nuclear power generation, wastewater treatment, mass transfer equipment, and pipelines. Gas-liquid / liquid-liquid two-phase flows have important practical significance in the petrochemical process, especially in fields such as petroleum cracking, natural gas treatment, and chemical synthesis. The petroleum cracking process is a key step in the petrochemical industry. Through the mass transfer and heat transfer characteristics of gas-liquid / liquid-liquid two-phase flows, the cracking efficiency of hydrocarbon molecules can be effectively improved. By using the efficient contact between the catalyst and the gas-liquid mixture, the reaction rate of hydrocarbon compounds in a high-temperature and high-pressure environment is enhanced, thereby producing more high-value-added light oil products. The optimized application of gas-liquid / liquid-liquid two-phase flows in these processes can not only improve the utilization rate of crude oil, but also reduce energy consumption and production costs, meeting the current demand of the petrochemical industry for high efficiency, greenness, and low carbon. In a gas-liquid / liquid-liquid two-phase flow system, bubble / droplet parameters play a crucial role in characterizing the system characteristics. For example, the shape, surface area, and size of bubbles / droplets are all affected by the flow pattern. Therefore, high-precision bubble / droplet recognition can more accurately evaluate a gas-liquid / liquid-liquid two-phase flow system. A convolutional neural network with high universality and robustness can accurately identify bubbles / droplets, obtain bubble / droplet parameters, and track the process of bubble / droplet coalescence and breakup in a complex gas-liquid / liquid-liquid system. Therefore, the parameters obtained by the neural network play an important role in optimizing the mixer design and obtaining the stability of the gas-liquid / liquid-liquid system.
[0004] The features of complex images may require a network with a relatively large depth for processing, and deeper networks have a large amount of computation, and training and inference are more time-consuming. Training large-scale CNNs requires high-performance hardware support, especially GPUs or TPUs with sufficient video memory. The problem of gradient disappearance or explosion in deep networks will affect the training stability and convergence speed, especially when dealing with complex tasks. In complex images, details may be hidden in the background, and the model needs to have stronger resolution ability. Bubble / droplet images have various postures, perspectives, and shape deformations, increasing the difficulty of recognition. Moreover, bubbles / droplets appear at different scales, and traditional CNNs may be difficult to adapt. Therefore, a neural network design method that can extract complex multi-dimensional data and has a more lightweight model is needed. Summary of the Invention The object of the present invention is to provide a multi-dimensional information parallel targeting extraction method based on the spatial characteristics of multiphase flow droplets / bubbles, which effectively solves the problems existing in the recognition process of complex bubble / droplet images, such as large computational amount, long training time, gradient disappearance or explosion, difficulty in extracting feature details, various deformations and inconsistent scales. This method is optimized by using a multi-dimensional backbone lightweight neural network. The multi-dimensional backbone network extracts image features layer by layer, capturing local texture information from the shallow layer to aggregating global semantic features in the deep layer, and can effectively process targets of different scales and deformations. At the same time, the lightweight design uses depthwise separable convolution, pruning and quantization techniques to significantly reduce the computational cost, reduce the dependence on hardware, and improve the inference efficiency. In addition, combining the residual structure ResNet and the dynamic feature fusion mechanism can alleviate the problem of gradient disappearance and improve the stability and convergence speed of the model. While improving the model's resolution ability, this method takes into account the extraction of fine-grained features and computational efficiency in complex scenarios, providing an efficient solution for complex image recognition. The technical solution adopted by the present invention is as follows: A multi-dimensional information parallel targeting extraction method based on the spatial characteristics of multiphase flow droplets / bubbles, the method includes the target dynamic detection, prediction of multi-dimensional information and the overall lightweight of the model: the specific method steps are as follows: Step 1: Use bubble / droplet images under different working conditions as the data set, where 70% - 80% is used for the training set and 30% - 20% is used for the validation set; Step 2: Calibrate the bubbles / droplets in the image, and then perform operations such as automatic orientation, size scaling, grayscale conversion, cropping and isolating objects on the image; When extracting bubble / droplet information in three-dimensional space, it is first necessary to clarify the types of bubble / droplet information, such as the size, shape, distribution position, number density, and movement trajectory of bubbles / droplets; different types of bubble / droplet information usually correspond to different research needs. For example, the research on dynamic characteristics may focus on movement trajectory and speed information, while the research on structural characteristics may pay more attention to the size and shape of bubbles / droplets; after determining the types of bubble / droplet information to be extracted, it is necessary to design the number of backbones according to the complexity and data dimension of the information; different backbones can be used to collect one or more different types of information respectively, so as to achieve a comprehensive coverage and efficient analysis of bubble / droplet characteristics; when designing the backbone, it is necessary to balance the acquisition accuracy, processing speed and system complexity; for some simple research needs, optimize the system structure by reducing the number of backbones and combining functions; while for complex research objectives, increase the number of backbones to ensure the independent acquisition and efficient processing of multiple information. Step 4. In the design of the head and tail of each level of the backbone, embed the C2DAT module respectively to optimize information extraction and computing resource allocation; improve the information processing efficiency of the backbone, and ensure the pertinence and accuracy of data processing; embed the C2DAT module at the head of the backbone, and its core function is to focus on and strengthen the specific information extracted by the backbone; by extracting and filtering the specific features of the input data, this module can accurately capture the key information related to the target task, while suppressing the interference factors irrelevant to the backbone task; the introduction of this module enables the backbone to effectively reduce the redundancy of input information at the initial stage, thereby optimizing the data flow efficiency; at the tail of the backbone, also embed the C2DAT module to further focus on the data output directly related to the target information; at this stage, the main function of this module is to screen and streamline the intermediate data generated after being processed by the middle layer of the backbone; by eliminating redundant or irrelevant calculation results, the tail module can reduce the occupancy of computing resources by redundant information, thereby releasing computing power so that more computing power can be invested in key tasks; in addition, the tail module can also enhance the expressiveness and interpretability of the target information, and provide a highly focused data output for subsequent tasks; this dual-module design optimizes the input and output data streams at the head and tail of the backbone respectively; at the head, the C2DAT module can directionally capture specific features when the data flows in; while at the tail, the module realizes the enhanced expression of the target information through further focusing and redundancy elimination; through this structure, the backbone can not only achieve more efficient information processing, but also improve the computing power utilization efficiency, ensuring that the allocation of computing resources is more reasonable and scientific; Step 5. The training adopts compound iterative training, including: feature extraction stage, optimization stage and iterative adjustment; Step 6. Use the neural network model to identify bubbles / droplets to obtain bubble / droplet information, track the movement state of bubbles / droplets, and reconstruct the shape of bubbles / droplets in three-dimensional space; accurately separate bubbles / droplets from the background, extract geometric features and dynamic information, and combine multi-view data to achieve high-precision three-dimensional reconstruction of the shape of bubbles / droplets. The described multi-dimensional information parallel targeted extraction method based on the spatial characteristics of multiphase flow droplets / bubbles, this method is based on an optical image containing multiple types of information, and uses Figure 2 the multi-backbone neural network shown to extract various feature information in the image, and the process is as follows: Step 1. Preprocess the image to set the resolution and then input it into the Input layer of the neural network; Step 2. Send the dataset after simple convolution into the multi-dimensional backbone network composed of multiple backbones to deeply mine the required information. Obtain various positions, sizes, shapes, and movement states of the target; Step 3: Embed the C2DAT dynamic attention module at the head and tail of each backbone channel to focus on target information, remove redundant calculations, and release computing power; Step 4: Concatenate and combine the information obtained from the multi-dimensional backbone and participate in subsequent convolutional pooling, and complete target recognition, classification, tracking, and prediction operations through functions such as NMS and Softmax. In the described multi-dimensional information parallel targeted extraction method based on the spatial characteristics of multiphase flow droplets / bubbles, in Step 1, a dataset is constructed by collecting multi-dimensional complex bubble / droplet images with high deformation degree, large overlap degree, and large offset angle caused by different operating conditions, strong disturbances, and high turbulence levels; among them, the dataset is divided into a training set and a validation set according to a certain proportion, with 70% - 80% used for the training set and 30% - 20% used for the validation set to ensure the generalization ability and robustness of the model in complex environments. In the described neural network design method for multi-dimensional information parallel targeted extraction and model lightweighting, in Step 2, the collected images are labeled to extract key information, and the images are uniformly processed, including data preprocessing steps such as automatic orientation, size scaling, grayscale conversion, cropping, and object isolation to standardize the data format. In the described multi-dimensional information parallel targeted extraction method based on the spatial characteristics of multiphase flow droplets / bubbles, the complex optical images obtained in Step 3 have various types of information including: semantic information, structural information, temporal information, context information, spatial information, feature interaction information, and scale invariance. Semantic information: The multi-dimensional backbone model uses semantic information to identify bubbles / droplets in the image and the differences between bubbles / droplets and the fluid, helping the model understand the key elements in the image and their meanings; Structural information: Structural information covers the external contour features and internal details of bubbles / droplets; by reading the structural information, the multi-dimensional backbone model can accurately capture the morphological features of bubbles / droplets and perform efficient characterization and visualization on them; Temporal information: When analyzing video data, the multi-dimensional backbone model can track the changes of bubbles / droplets over time through temporal information, such as movement speed, direction, merging or splitting behaviors; Context information: The multi-dimensional backbone model analyzes the context relationships of bubbles / droplets in the entire flow field by identifying the relative positions and interactions between bubbles / droplets; Spatial information: Spatial information is very important for analyzing the distribution, density, and overall flow pattern of bubbles / droplets in the fluid. The multi-dimensional backbone model can accurately extract the spatial features in the image, including the precise positions of bubbles / droplets in the image and the spatial relationships between bubbles / droplets; Feature interaction information: Feature interaction information helps the multi-dimensional backbone model learn and understand the complex interactions between bubble / droplet features (such as size and shape), which contributes to the analysis of the characteristics of gas-liquid flow at a higher level; Scale invariance: The multi-dimensional backbone model identifies similar features that appear at different scales through various convolutional and pooling layers; this means that the model can effectively identify and classify bubbles / droplets regardless of their size changes; In step three, a multi-level backbone is composed of a main branch, a reversible auxiliary branch, ResNet, GELAN, and MobileNet; the main branch only obtains partial surface information, and the original image data is used to obtain global feature information through preliminary convolution and then directly enters the tail of the backbone along the main branch for feature splicing; the reversible auxiliary branch provides backpropagation support for the main branch by generating reliable gradient information; and multiple backbones extracted from various specific information extract different information; ResNet introduces a residual module to solve the problem of gradient disappearance in deep networks and capture deeper features; GELAN is used to identify the size changes and shape diversity of bubbles / droplets, and at the same time segment the image to obtain complete spatial information; MobileNet is used for real-time auxiliary calculation of other branches to ensure sufficient computing power for each level of the backbone. In the described multi-dimensional information parallel targeted extraction method based on the spatial features of droplets / bubbles in multiphase flow, the C2DAT module is embedded in step four to enhance the computing performance and contribute to the overall lightweight of the model; the C2DAT module is introduced into the head and tail of the backbone network to efficiently extract bubble / droplet features; in step four, the feature maps from the previous layer are input into the head C2DAT module, and these feature maps have been processed by the previous convolutional layers and contain preliminary image information; the C2DAT in step four includes standard convolution, depth convolution, pointwise convolution, and the combined formula is as follows: Y c = W * X in + z (1) Y de = W de * X in (2) Y p = W p * Y de (3) Y m = [Y c , Y de , Y a (4) where Y is the output feature map, X in is the input feature map, W is the convolution kernel, * is the convolution calculation, and z is the bias; In Step 4, C2DAT solves the problems of high memory and computational costs caused by the input layer receiving a wide range of features in traditional methods. At the same time, it can also effectively reduce the impact of irrelevant parts in the image. The C2DAT network structure is as Figure 3 shown. C2DAT can dynamically select the key points and value pairs of attention according to the needs of the data, enabling the model to focus on relevant regions and thus extract more informative features. The deformable attention mechanism introduced by C2DAT only focuses on a small part of the image. In the deformable attention mechanism, C2DAT dynamically selects sampling points instead of processing the entire image fixedly. This dynamic selection mechanism enables the model to concentrate on processing the regions that are most important for the current task. A multi-dimensional information parallel targeted extraction method based on the spatial features of multiphase flow droplets / bubbles. In Step 5, the training adopts compound iterative training, including: feature extraction stage, optimization stage, and iterative adjustment; In the feature extraction stage, the model extracts the features of the input data through forward propagation; these feature representations are the basis for future classification or prediction tasks; In the optimization stage, the loss function is used to optimize the parameters of the model; this step usually includes calculating the loss and then adjusting the network weights through the backpropagation algorithm to reduce the prediction error; In the iterative adjustment stage, after the first iteration, the feature extractor is adjusted according to the performance of the model, and the network structure or parameters are modified, and then feature extraction and optimization are performed again; multiple iterations are carried out to gradually improve the model's understanding and prediction ability of the data. A multi-dimensional information parallel targeted extraction method based on the spatial features of multiphase flow droplets / bubbles. In Step 6, the three-dimensional size of the characteristic bubbles / droplets is estimated using the planar projection of the bubbles / droplets; the major axis a, minor axis c, and inclination angle β of the bubbles / droplets are obtained from the recognition image. Therefore, it is also necessary to obtain the major axis b of the elliptical bubbles / droplets perpendicular to the image plane to estimate the volume and cross-sectional area of the bubbles / droplets; a spatial coordinate system is established using the equatorial radius and polar radius of the ellipsoid; to obtain the horizontal cross-section of the ellipsoid at an arbitrary height h, it is necessary to determine β and the inclination angle θ in the planar coordinate system, as follows: The elliptical cross-section equation of the ellipsoid at an arbitrary height h is as follows: In the formula, a 1 is the major axis of the elliptical cross-section. b 1 is the minor axis of the elliptical cross-section. The formula for the cross-sectional area of the ellipse is as follows: A s = πa 1 b 1 (9) The bubble / droplet volume formula is: A multi-dimensional information parallel targeted extraction method based on the spatial characteristics of multiphase flow droplets / bubbles, which improves the indicators of the neural network model; provides a reference for multi-dimensional information acquisition and the overall lightweighting of the model. The beneficial effects of the present invention are: Based on the design of a multi-backbone neural network, multi-dimensional information in complex images is systematically and comprehensively extracted. At the same time, the C2DAT module is introduced, which not only significantly improves the computing performance but also contributes to the overall lightweighting of the model. Through parallel design, the multi-backbone neural network can make full use of the characteristics of the main branch, auxiliary branch, and lightweight network to comprehensively mine local and global features in the image from different levels and angles. This structure ensures that the model can capture global semantic information while also taking into account local details and dynamic changes. In the multi-dimensional backbone network, the information extracted by each branch is fused through feature splicing and combination to achieve the comprehensive acquisition of global information. This fusion strategy organically integrates the features from different network branches, enabling the model to more accurately model the feature distribution and interrelationships in complex scenarios. The embedding of the C2DAT module further improves the efficiency and effect of feature processing. Through the dynamic attention mechanism, C2DAT can focus on the most relevant target regions in the image while ignoring unnecessary background information and noise, thereby optimizing the allocation of the model's computing resources. Its deformable structure makes the attention mechanism more flexible, capable of adapting to various scenarios and feature distributions, and enhancing the model's ability to capture target features. The introduction of the C2DAT module significantly contributes to the lightweight design of the model. By reducing the processing and calculation of irrelevant regions, this module effectively reduces the overall computational amount of the model, making the inference process for complex scenarios more efficient. On this basis, the multi-backbone neural network and the C2DAT module work together to achieve a balance between performance and efficiency, ensuring both high-precision feature extraction of the model and meeting the requirements of real-time calculation and lightweighting. This design provides an efficient, flexible, and highly performant solution for complex image recognition tasks. BRIEF DESCRIPTION OF THE DRAWINGS Figure 1 It is the overall architecture diagram of the neural network of the present invention, which includes information of each layer of the neural network; Figure 2 It is the overall flowchart of image recognition of the present invention, revealing the entire process of image recognition and target detection; Figure 3 It is the structural diagram of the C2DAT module of the present invention; Figure 4 These are the various performance index diagrams after the training of the present invention is completed. Specific Embodiments The following will illustrate the implementation manners of the present invention through specific examples. Those skilled in the art can easily understand other advantages and effects of the present invention according to the content disclosed in this specification. In addition, the present invention can also be implemented or applied through various other specific implementation manners. The details in this specification can also be modified or adjusted based on different perspectives and uses without departing from the spirit of the present invention. It should be specifically noted that, without conflict, the following embodiments and the features therein can be flexibly combined to meet different requirements and application scenarios. The following further describes the specific implementation manners of the present invention in detail. For data processing, it includes steps such as normalization, data augmentation, and size adjustment. The core of normalization is to adjust the pixel values of the image to a specific range, such as [0, 1] or [-1, 1]. This can eliminate the dimensional differences in the data, make the data distribution more uniform, thereby improving the stability of neural network training and accelerating the model convergence process. In terms of data augmentation, by operations such as rotation, cropping, scaling, flipping, or adjusting brightness, the diversity of training data can be artificially increased. This not only helps to improve the generalization ability of the model but also effectively reduces the overfitting phenomenon. For size adjustment, usually the images need to be unified to a fixed size to adapt to the input requirements of the network, which can be achieved through scaling. If it is necessary to maintain the original ratio of the image, zero-padding can also be added to avoid information loss caused by deformation. These processing steps complement each other, building a solid data foundation for deep learning tasks and improving the performance and training efficiency of the model. For normalization: where μ is the pixel mean and σ is the standard deviation. The backbone part adopts a multi-dimensional backbone network architecture composed of a main branch, a reversible auxiliary branch, ResNet, GELAN, and MobileNet. The backbone network is responsible for extracting multi-scale features. Among them, the main operations include convolution, batch normalization, activation function, and feature fusion, which are respectively formulas 12 - 15: where X(i, j) is the input feature map, W(m, n) is the convolution kernel, b is the bias, and Y(i, j) is the output feature map. where μ and σ 2 are the mean and variance. γ, β are learnable parameters. Feature fusion is: Z(i,j) = X 1 (i,j) + X 2 (i,j) (15) Z(i,j) = Concat(X 1 (i,j), X 2 (i,j)) (16) The Neck layer consists of a Feature Pyramid Network (FPN) and a Path Aggregation Network (PANet). The FPN fuses high-level features (rich in semantic information) and low-level features (rich in detail information) from top to bottom, while the PANet aggregates low-level features (containing spatial information) from bottom to top to further enhance the detection ability of small objects. The features of different layers are merged through a multi-scale feature fusion method. Among them, upsampling interpolates or deconvolves low-resolution features: Y up (i,j) = Interpolate(X(i,j)) (17) Given a k×k window, the max pooling operation selects the maximum value within the window for output: y i,j = max{x i+k-1,j+k-1} (18) Average pooling is: where x is the input feature map, y is the pooled feature map, i,j represents the starting position of the pooling operation, and k is the size of the pooling window. Calculating the average value of the entire input feature map and outputting a single value, global average pooling is: where H and W are the height and width of the feature map respectively. In addition to the pooling operation, the convolution operation itself can also achieve downsampling. By setting the stride greater than 1, the convolution operation can be directly changed to downsampling. When the stride S is greater than 1, not all adjacent input signals will be sampled, and the size of the output signal will be reduced. The detection head predicts the object classification and bounding box parameters for each grid cell. The Softmax function predicts the object class: The bounding box regression outputs the bounding box parameters and then converts them into actual coordinates: where (c x , c y ) is the center coordinate of the grid cell, pw and p h are the width and height of the prior box Use a composite loss function to comprehensively optimize classification, regression, and object detection errors. Use IoU optimization to improve the bounding box matching degree and FocalLoss to reduce the weights of easily classified samples. FocalLoss is as follows FL(p t ) = -α(1 - p t ) γ log(p t ) (25) where (1 - p t ) γ is a "modulation term". When pt is large (i.e., the model's prediction is relatively accurate), (1 - p t ) γ will become smaller, thus reducing the loss for easily classified samples. The existence of this term makes FocalLoss pay more attention to those difficult-to-classify samples. α is a balancing factor, usually used in the case of class imbalance to adjust the weights of positive and negative classes. For the case of imbalance between positive and negative classes, a suitable α is usually selected to balance their effects. The γ parameter controls the degree of attention to difficult-to-classify samples. When γ = 0, FocalLoss degenerates into the standard cross-entropy loss; as γ increases, the loss of the model for easily classified samples will be smaller, enhancing the attention to difficult-to-classify samples. The total loss formula is L = L classification + L box regression + L objectness (26) Remove overlapping bounding boxes through NMS (Non-Maximum Suppression) and only retain the box with the highest confidence, thus ensuring that the output detection results are clear and accurate. Given two boxes B i and B j , their intersection over union (IoU) is calculated by the following formula where ∣B i ∩B j ∣ represents the area of the intersection region of the two boxes, and ∣B i ∪B j ∣ represents the area of the union region of the two boxes. Quantization and Knowledge Distillation are used in the model to reduce the computational complexity and memory requirements of the model while maintaining good performance. The following are their basic principles and formulas. The goal of quantization is to convert parameters, activation values, etc. in a high-precision (e.g., 32-bit floating-point) model into a lower-precision representation (such as 8-bit integers), thereby reducing the storage and computational costs of the model. Assume the original parameter w is a floating-point number, and after quantization, the parameter becomes an integer. The quantization process can be expressed by the following formula: round() represents the rounding operation. min(w) is the smallest value in w, which is usually used to shift the data to the non-negative range. Δ is the quantization step size, and its calculation formula is: where Q is the number of quantization bits. The quantization process realizes compression in storage and calculation by mapping continuous floating-point values to discrete integer values. Knowledge distillation is a method of transferring knowledge from a complex large model (teacher model) to a small model (student model). The student model is trained by mimicking the output probability distribution of the teacher model, thereby obtaining a performance similar to that of the teacher model, but with fewer parameters and faster inference speed. Assume the probability distribution output by the teacher model is p teacher , and the probability distribution output by the student model is p student . The loss function of knowledge distillation usually consists of two parts: one is the standard classification loss, and the other is the distillation loss, which guides the student model to learn by minimizing the difference between the output distributions of the student and teacher models: L KD = α·l CE (p student ,y)+(1 - α)·T 2 ·L KL (p student ,p teacher ) (31) L CE (p teacher ,y) is the cross-entropy loss of the student model, and y is the true label. L KL (p student ,p teacher ) is the Kullback-Leibler divergence between the student model and the teacher model, which measures the difference between the two probability distributions. T is the temperature parameter, which is used to smooth the output probability distribution of the teacher model so that the student can learn more detailed distribution information. α is the weight coefficient of the two loss parts, and α controls the balance between the label true loss and the distillation loss of the student model. Among them Finally, visualize the detection results. Map the bounding box back to the original image: Convert the network output of a fixed size to the original image coordinates. Figure 1 This is the overall architecture diagram of the neural network of the present invention, which contains information of each layer of the neural network. Figure 2 This is the overall flowchart of image recognition of the present invention, revealing the entire process of image recognition and target detection. Figure 3 This is the structure diagram of the C2DAT module of the present invention. Figure 4 These are the various performance index diagrams after the training of the present invention is completed. Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications to these embodiments once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed as including the preferred embodiments as well as all changes and modifications falling within the scope of the present invention. The above embodiments are only used to illustrate the present invention and are not intended to limit the present invention. Those of ordinary skill in the relevant technical field can make various modifications and improvements to the present invention without departing from the spirit and scope of the present invention. Therefore, all equivalent technical solutions also belong to the scope of the present invention, and the patent protection scope of the present invention shall be defined by the claims.
Claims
1. A multi-dimensional information parallel targeted extraction method based on the spatial characteristics of multiphase flow droplets / bubbles, characterized in that: The method includes target dynamic detection and prediction of multi-dimensional information and overall lightweighting of the model. The specific method steps are as follows: Step 1: Use bubble / droplet images under different working conditions as data sets, of which 70% to 80% are used for training sets and 30% to 20% are used for validation sets; Step 2: calibrate the bubbles / droplets in the image, and then automatically orient, resize, convert grayscale, crop, and isolate the object; Step 3: When extracting bubble / droplet information in three-dimensional space, it is first necessary to clarify the type of bubble / droplet information, such as the size, shape, distribution position, number density, and motion trajectory of the bubble / droplet; different types of bubble / droplet information usually correspond to different research needs. For example, the study of dynamic characteristics may focus on motion trajectory and speed information, while the study of structural characteristics may pay more attention to the size and shape of the bubble / droplet; after determining the type of bubble / droplet information to be extracted, the number of trunks needs to be designed according to the complexity of the information and the data dimension; different trunks can be used to collect one or more different types of information respectively, so as to achieve comprehensive coverage and efficient analysis of bubble / droplet characteristics; when designing the trunk, it is necessary to weigh the acquisition accuracy, processing speed and system complexity; for some simple research needs, the system structure is optimized by reducing the number of trunks and merging functions; and for complex research goals, the number of trunks is increased to ensure the independent collection and efficient processing of multiple information; Step 4. In the design of the head and tail of each level of the backbone, C2DAT modules are respectively embedded to optimize information extraction and computing resource allocation; improve the information processing efficiency of the backbone and ensure the pertinence and accuracy of data processing; embed the C2DAT module in the head of the backbone, and its core function is to focus on and strengthen the specific information that the backbone is responsible for extracting; by extracting and filtering specific features of the input data, the module can accurately capture key information related to the target task, while suppressing interference factors unrelated to the backbone task; the introduction of this module enables the backbone to effectively reduce the redundancy of input information in the initial stage, thereby optimizing data flow efficiency; at the tail of the backbone, the C2DAT module is also embedded to further focus on the data output directly related to the target information; at this stage, the main function of this module is to The intermediate data generated after processing is screened and streamlined; by eliminating redundant or irrelevant calculation results, the tail module can reduce the occupation of computing resources by redundant information, thereby releasing computing power so that more computing power can be invested in key tasks; in addition, the tail module can also enhance the expressiveness and interpretability of target information, and provide highly focused data output for subsequent tasks; this dual-module design optimizes the input and output data streams at the head and tail of the trunk respectively; at the head, the C2DAT module can capture specific features in a targeted manner when data flows in; and at the tail, the module can achieve enhanced expression of target information through further focusing and redundant elimination; through this structure, the trunk can not only achieve more efficient information processing, but also improve the efficiency of computing power utilization, ensuring a more rational and scientific allocation of computing resources; Step 5: Training adopts compound iterative training including: feature extraction stage, optimization stage and iterative adjustment; Step 6. Use a neural network model to identify bubbles / droplets to obtain bubble / droplet information, track bubble / droplet motion status, and reconstruct the bubble / droplet morphology in three-dimensional space; accurately separate bubbles / droplets from the background, extract geometric features and dynamic information, and combine multi-view data to achieve high-precision three-dimensional reconstruction of bubble / droplet morphology.
2. The multi-dimensional information parallel targeted extraction method based on multiphase flow droplet / bubble spatial characteristics according to claim 1 is characterized in that: The method is based on an optical image containing multiple information, and uses a multi-trunk neural network as shown in FIG2 to extract multiple feature information from the image. The process is as follows: Step 1: Preprocess the image to set the resolution and then input it into the neural network input layer; Step 2: Send the simple convolution data set into a multi-dimensional backbone network composed of multiple backbones to deeply mine the required information. Obtain the target's various positions, sizes, shapes, and motion states; Step 3: Embed the C2DAT dynamic attention module at the head and tail of each trunk channel to focus on the target information, remove redundant calculations and release computing power; Step 4: Concatenate and combine the information obtained from the multi-dimensional backbone and participate in subsequent convolution pooling, as well as complete target recognition, classification, tracking and prediction operations through functions such as NMS and Softmax.
3. The multi-dimensional information parallel targeted extraction method based on multiphase flow droplet / bubble spatial characteristics according to claim 1 is characterized in that: In the step 1, a data set is constructed by collecting multidimensional complex bubble / droplet images with high deformation, large overlap, and large offset angle caused by different operating conditions, strong disturbances, and high turbulence levels; wherein the data set is divided into a training set and a validation set in proportion, 70% to 80% for the training set and 30% to 20% for the validation set, to ensure the generalization ability and robustness of the model in complex environments.
4. The neural network design method for parallel targeted extraction of multi-dimensional information and lightweight model according to claim 1, characterized in that: In the step 2, the collected images are annotated to extract key information therein, and the images are uniformly processed, including automatic orientation, size scaling, grayscale conversion, cropping, and object isolation data preprocessing steps to standardize the data format.
5. The multi-dimensional information parallel targeted extraction method based on multiphase flow droplet / bubble spatial characteristics according to claim 1 is characterized in that: The complex optical image obtained in step three has a variety of information including: semantic information, structural information, temporal information, contextual information, spatial information, feature interaction information and scale invariance. Semantic information: The multi-dimensional backbone model uses semantic information to identify bubbles / droplets in the image and the difference between bubbles / droplets and fluids, helping the model understand the key elements in the image and their meaning; Structural information: Structural information covers the external contour features and internal details of bubbles / droplets. By reading the structural information, the multidimensional backbone model can accurately capture the morphological characteristics of bubbles / droplets and efficiently characterize and visualize them. Time series information: When analyzing video data, the multidimensional backbone model can track the changes of bubbles / droplets over time through time series information, such as movement speed, direction, merging or splitting behavior; Contextual information: The multidimensional backbone model analyzes the contextual relationship between bubbles / droplets in the entire flow field by identifying their relative positions and interactions; Spatial information: Spatial information is very important for analyzing the distribution, density, and overall flow pattern of bubbles / droplets in the fluid. The multidimensional backbone model can accurately extract the spatial features in the image, including the exact location of bubbles / droplets in the image and the spatial relationship between bubbles / droplets; Feature interaction information: Feature interaction information can help the multidimensional backbone model learn and understand the complex interactions between bubble / droplet features (such as size and shape), which helps to analyze the characteristics of gas-liquid flow at a higher level; Scale invariance: The multi-dimensional backbone model recognizes similar features that appear at different scales through various convolutional and pooling layers; this means that the model can effectively identify and classify bubbles / droplets regardless of how their sizes vary; In step three, a multi-level backbone is composed of a main branch, a reversible auxiliary branch, ResNet, GELAN and MobileNet. The main branch only obtains part of the surface information, and obtains global feature information through preliminary convolution of the original image data, and then directly enters the tail of the backbone along the main branch for feature splicing. The reversible auxiliary branch provides backpropagation support for the main branch by generating reliable gradient information. Multiple backbones extracted from a variety of unique information extract different information. ResNet introduces a residual module to solve the gradient vanishing problem of deep networks and capture deeper features. GELAN is used to identify the size changes and shape diversity of bubbles / droplets, and at the same time segment the image to obtain complete spatial information. MobileNet is used to assist other branches in real-time calculations to ensure sufficient computing power at all levels of the backbone.
6. The multi-dimensional information parallel targeted extraction method based on multiphase flow droplet / bubble spatial characteristics according to claim 1 is characterized in that: In step 4, the C2DAT module is embedded to enhance the computing performance and help the overall lightweight of the model; the C2DAT module is introduced into the head and tail of the backbone network to efficiently extract bubble / droplet features; in step 4, the feature maps from the previous layer are input to the head C2DAT module, which have been processed by the previous convolution layer and contain preliminary image information; in step 4, C2DAT includes standard convolution, depth convolution, point-by-point convolution and merging formulas as shown below: Y c =W*X in +z (1) Y de =W de *X in (2)Y p =W p *Y de (3)Y m =[Y c ,Y de ,Y a ] (4) Where Y is the output feature map, X in is the input feature map, W is the convolution kernel, * is the convolution calculation, and z is the bias; In step 4, C2DAT solves the high memory and computational cost problems caused by the input layer receiving a wide range of features in traditional methods, and can also effectively reduce the impact of irrelevant parts of the image. The C2DAT network structure is shown in Figure 3. C2DAT can dynamically select attention keypoints and value pairs according to the needs of the data, so that the model can focus on the relevant areas and extract more informative features. The deformable attention mechanism introduced by C2DAT only focuses on a small part of the image. In the deformable attention mechanism, C2DAT dynamically selects sampling points instead of fixed processing of the entire image. This dynamic selection mechanism allows the model to focus on the most important areas for the current task.
7. The multi-dimensional information parallel targeted extraction method based on multiphase flow droplet / bubble spatial characteristics according to claim 1 is characterized in that: The training in step 5 adopts compound iterative training including: feature extraction stage, optimization stage and iterative adjustment; In the feature extraction stage, the model extracts features of the input data through forward propagation; these feature representations are the basis for future classification or prediction tasks; The optimization phase uses the loss function to optimize the model's parameters; this step usually involves calculating the loss and then adjusting the network weights through the back-propagation algorithm to reduce the prediction error; In the iterative adjustment phase, after the first iteration, the feature extractor is adjusted and the network structure or parameters are modified according to the performance of the model, and then feature extraction and optimization are performed again; multiple iterations are performed to gradually improve the model's understanding and prediction capabilities of the data.
8. The method for parallel targeted extraction of multidimensional information based on multiphase flow droplet / bubble spatial characteristics according to claim 1, characterized in that: In step 6, the plane projection of the bubble / droplet is used to estimate the three-dimensional size of the characteristic bubble / droplet; the major axis a, minor axis c and inclination angle β of the bubble / droplet are obtained from the recognition image. Therefore, it is also necessary to obtain the major axis b of the elliptical bubble / droplet perpendicular to the image plane to estimate the volume and cross-sectional area of the bubble / droplet; the equatorial radius and polar radius of the ellipsoid are used to establish a spatial coordinate system; in order to obtain the horizontal section of the ellipsoid at any height h, it is necessary to determine β and the inclination angle θ in the plane coordinate system, as shown below: The equation of the elliptical section at any height h of the ellipsoid is as follows: Where a1 is the major axis of the elliptical cross section. b1 is the minor axis of the elliptical cross section. The formula for the cross-sectional area of an ellipse is as follows: A s =πa1b1 (9) The bubble / drop volume formula is:
9. The method for parallel targeted extraction of multidimensional information based on multiphase flow droplet / bubble spatial characteristics according to claim 1, characterized in that: The method improves the indicators of the neural network model and provides a reference for multi-dimensional information acquisition and overall lightweighting of the model.
Citation Information
Patent Citations
Target detection algorithm based on improved YOLOv8s
CN117456167A
Gas-liquid two-phase flow phase interface segmentation method
CN117593321A
Intravascular ultrasound image segmentation and reconstruction method based on multi-scale prior coordination mechanism
CN118644498A
Transform-based global positioning method for automatic driving commercial vehicle on structured road
CN118691779A
Gas-liquid two-phase flow bubble identification method based on improved Mask R-CNN
CN119006980A
Cited By
Method for predicting pressure characteristics of multiphase flow in pipeline based on bubble attention mechanism
CN120611639A
Method for parallel extraction of multi-scale features from microbubbles to large bubbles based on multi-branch neural architecture
CN121788857A